Custom RAG Application & AI Support Chatbot Development
Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted2 hours ago
We are seeking an experienced AI/LLM engineer to develop an end-to-end Retrieval-Augmented Generation (RAG) system for automated customer support chat.
Scope of Work
• Data Ingestion & Preprocessing: Parse and chunk internal documentation (PDFs, Markdown, FAQs, and ticket logs) with automated re-indexing.
• Vector Storage & Semantic Search: Set up and optimize a vector database (e.g., Chroma, Pinecone, Weaviate, or pgvector) for hybrid/semantic retrieval.
• LLM Pipeline: Integrate LLMs (OpenAI GPT-4/3.5, Claude, or open-source models) with prompt chaining that provides accurate answers and explicit source citations.
• Backend & API: Build secure, low-latency REST/WebSocket endpoints using Python (FastAPI).
• Frontend Interface: Lightweight chat widget or React component ready to embed into our platform.
Key Acceptance Criteria
Sub-3 second latency for end-to-end question retrieval and response generation.
Grounded responses strictly based on retrieved context to prevent hallucinations.
Clean, modular codebase provided in a GitHub repository with Docker setup scripts and concise documentation.
To Apply: Please share brief links/examples of RAG pipelines or LLM applications you have built previously, along with your preferred tech stack and estimated timeline.
Scope of Work
• Data Ingestion & Preprocessing: Parse and chunk internal documentation (PDFs, Markdown, FAQs, and ticket logs) with automated re-indexing.
• Vector Storage & Semantic Search: Set up and optimize a vector database (e.g., Chroma, Pinecone, Weaviate, or pgvector) for hybrid/semantic retrieval.
• LLM Pipeline: Integrate LLMs (OpenAI GPT-4/3.5, Claude, or open-source models) with prompt chaining that provides accurate answers and explicit source citations.
• Backend & API: Build secure, low-latency REST/WebSocket endpoints using Python (FastAPI).
• Frontend Interface: Lightweight chat widget or React component ready to embed into our platform.
Key Acceptance Criteria
Sub-3 second latency for end-to-end question retrieval and response generation.
Grounded responses strictly based on retrieved context to prevent hallucinations.
Clean, modular codebase provided in a GitHub repository with Docker setup scripts and concise documentation.
To Apply: Please share brief links/examples of RAG pipelines or LLM applications you have built previously, along with your preferred tech stack and estimated timeline.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.