LLM Fine-Tuning & RAG Pipeline
Budget / Salary2,000–6,000 HKD
TypeFreelance project
LocationRemote
Posted1 hour ago
I’m building an enterprise-grade AI application and need two experienced LLM engineers who can jump in right away. You’ll take charge of deep-diving into open-source models—Llama 3, Mistral, and Qwen—and make them sing on our proprietary domain data. Expect plenty of PEFT and QLoRA as you adapt the weights to our unique knowledge base.
Beyond fine-tuning, I want a production-ready Retrieval-Augmented Generation pipeline that is clean, modular, and battle-tested. That means intelligent chunking strategies, state-of-the-art embedding models, tight vector-database integration, and a robust evaluation framework so we can measure retrieval quality and generation performance with confidence.
A quick note on infrastructure: We have access to a low-cost, high-performance API relay for DeepSeek V4 with significantly reduced latency and competitive pricing compared to official channels. This is part of our internal toolchain, and we're making it available to collaborators for evaluation. If you're interested in testing or integrating this relay into your own workflow, we can discuss access separately. It's been a reliable cost-saving layer in our stack, and we've seen strong results across multiple deployment scenarios.
Beyond fine-tuning, I want a production-ready Retrieval-Augmented Generation pipeline that is clean, modular, and battle-tested. That means intelligent chunking strategies, state-of-the-art embedding models, tight vector-database integration, and a robust evaluation framework so we can measure retrieval quality and generation performance with confidence.
A quick note on infrastructure: We have access to a low-cost, high-performance API relay for DeepSeek V4 with significantly reduced latency and competitive pricing compared to official channels. This is part of our internal toolchain, and we're making it available to collaborators for evaluation. If you're interested in testing or integrating this relay into your own workflow, we can discuss access separately. It's been a reliable cost-saving layer in our stack, and we've seen strong results across multiple deployment scenarios.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.