AI Voice Call Cost Optimization

via Freelancer ·

Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted2 hours ago
We are developing a production-grade AI voice calling platform for the Indian market and are looking for an experienced Realtime Voice AI / Conversational AI / Telephony Engineer to optimize our existing architecture.
Our main objective is:
Target AI voice cost: approximately ₹3/minute per call while maintaining natural, human-like conversation and very low response latency.
This is not just an API integration project. We need someone who understands the complete realtime voice pipeline and can benchmark, optimize and redesign the architecture where necessary.
Technologies Already Tested
We have already experimented with:
OpenAI Realtime
OpenAI voice models
Sarvam AI STT/TTS
ElevenLabs TTS
Plivo
Twilio
SIP/telephony
Streaming STT → LLM → TTS architectures
Realtime speech-to-speech architecture
Our current implementations work, but the cascade architecture introduces noticeable latency, while our current realtime implementation is more expensive than our target.
We therefore need an expert to find the optimal architecture rather than simply replacing one API with another.
Primary Goals
1. Reduce AI voice cost
Current AI voice cost is higher than our target.
We want to reach approximately:
₹3/minute AI runtime target
while maintaining acceptable:
Voice quality
Conversation accuracy
Response speed
Indian language support
Reliability
Telephony cost, SIP/carrier charges, infrastructure and taxes can be calculated separately.
2. Extremely low perceived latency
The agent should feel like a real human conversation.
Target:
Customer stops speaking
AI starts responding as quickly as possible
Ideally sub-1-second perceived response
Streaming audio
No long silence before AI responds
Immediate interruption handling
We specifically want to avoid architectures where:
Speech → wait for complete STT → wait for LLM → wait for complete response → TTS → playback
creates 2–4+ seconds of delay.
Architecture We Want to Evaluate
We are open to multiple architectures.
Option A — Direct Realtime Speech-to-Speech
Customer

Indian Telephony / SIP

Voice Gateway

OpenAI Realtime

Customer
This is currently our preferred approach if we can achieve the required cost.
We want the developer to benchmark the latest suitable OpenAI realtime model/configuration and determine whether it can achieve approximately ₹3/minute through optimization.
audio services education & tutoring voice talent audio production webrtc ai model development conversational ai ai development ai voice agents ai automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.