Refine RAG YAML Prompts
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted43 minutes ago
I am fine-tuning an AI system that uses retrieval-augmented generation to turn uploaded source documents into polished SDD (Software Design Description) sections. The current YAML prompt set is drifting: I see hallucination in the output—mostly content that the model half-understands or slightly twists rather than outright fabricates—and that small inaccuracy snowballs into wrong field mappings and cluttered tables.
Your main mission is to open the existing prompt files, compare them line-by-line with the business template, and tighten the instructions so the model only returns evidence-based facts. Template alignment is the area that clearly needs the most love, yet the retrieval and synthesis rules also deserve a second look to be sure we are not hard-coding project-specific values. Any field the sources can’t support should stay blank; no educated guesses, no filler.
Tooling is flexible—our current stack is Python, LangChain and OpenAI models feeding a vector store—so if you have a sharper idea for prompt structure or retrieval filters, run with it. What matters is that the final YAML delivers:
• Accurate data pulled verbatim from the source documents
• No hallucinations, misinterpretations or duplicate/irrelevant table entries
• A clean hand-off YAML file plus a short read-me explaining the changes and rationale
If you are comfortable dissecting large-language-model behaviour and can show previous success with prompt engineering for document synthesis, I would love to see how you’d tackle this.
Your main mission is to open the existing prompt files, compare them line-by-line with the business template, and tighten the instructions so the model only returns evidence-based facts. Template alignment is the area that clearly needs the most love, yet the retrieval and synthesis rules also deserve a second look to be sure we are not hard-coding project-specific values. Any field the sources can’t support should stay blank; no educated guesses, no filler.
Tooling is flexible—our current stack is Python, LangChain and OpenAI models feeding a vector store—so if you have a sharper idea for prompt structure or retrieval filters, run with it. What matters is that the final YAML delivers:
• Accurate data pulled verbatim from the source documents
• No hallucinations, misinterpretations or duplicate/irrelevant table entries
• A clean hand-off YAML file plus a short read-me explaining the changes and rationale
If you are comfortable dissecting large-language-model behaviour and can show previous success with prompt engineering for document synthesis, I would love to see how you’d tackle this.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.