Natural AI Voice Call Optimisation
Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted2 hours ago
I already have a full STT → LLM → TTS → telephony stack in production, yet callers still hear obvious lag, awkward pauses and a robotic tone. I’m not looking for a rebuild; I need an expert who can dive into the existing codebase and tune it for truly human-sounding conversations.
Your main target area is the Text-to-Speech component, though every tweak must be validated end-to-end. The agent must speak Indian English naturally and switch to Hindi when a caller does. Support for other accents is a nice bonus but not essential.
Key TTS fixes I expect:
• Pronunciation accuracy
• Voice naturalness
• Response speed
Scope of work
– Profile the complete pipeline to locate bottlenecks (latency, buffering, audio encoding).
– Implement low-latency TTS strategies or swap voices if needed while preserving the existing architecture.
– Optimise turn-taking so interruptions are handled smoothly and the agent never gets “stuck.”
– Fine-tune prosody, pacing and intonation until the disclosure “This is an AI-powered call” is the only clue it’s synthetic.
– Regression-test with live calls and supply before/after recordings that clearly demonstrate the improvements.
Acceptance criteria
1. Average round-trip response time in live calls ≤ 800 ms.
2. No audible stutter or clipped words in a 10-minute stress test.
3. Pronunciation errors reduced by at least 90 % compared to baseline sample.
4. Deliverables: optimised code/configs, test logs, and the comparative audio files.
If you have proven, real-time voice-agent experience, let’s make this voice sound human.
Your main target area is the Text-to-Speech component, though every tweak must be validated end-to-end. The agent must speak Indian English naturally and switch to Hindi when a caller does. Support for other accents is a nice bonus but not essential.
Key TTS fixes I expect:
• Pronunciation accuracy
• Voice naturalness
• Response speed
Scope of work
– Profile the complete pipeline to locate bottlenecks (latency, buffering, audio encoding).
– Implement low-latency TTS strategies or swap voices if needed while preserving the existing architecture.
– Optimise turn-taking so interruptions are handled smoothly and the agent never gets “stuck.”
– Fine-tune prosody, pacing and intonation until the disclosure “This is an AI-powered call” is the only clue it’s synthetic.
– Regression-test with live calls and supply before/after recordings that clearly demonstrate the improvements.
Acceptance criteria
1. Average round-trip response time in live calls ≤ 800 ms.
2. No audible stutter or clipped words in a 10-minute stress test.
3. Pronunciation errors reduced by at least 90 % compared to baseline sample.
4. Deliverables: optimised code/configs, test logs, and the comparative audio files.
If you have proven, real-time voice-agent experience, let’s make this voice sound human.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.