Ai contain

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I need an audio-based AI that takes plain text and returns a natural-sounding English voice output. The end goal is a production-ready text-to-speech (TTS) engine I can integrate into web and mobile apps.

Here is what I expect as the finished work:
• A trained TTS model (neural or parametric) optimised for clear, human-like English speech
• An inference script or lightweight API endpoint (Python/Flask or Node preferred) that accepts text and streams back audio (MP3 or WAV) in real time
• Simple usage documentation plus a quick demo project showing how to call the service

Acceptance criteria
• Latency under two seconds for a 200-character string on a standard GPU instance
• No noticeable robotic artefacts at common sampling rates (22 kHz or higher)
• The demo reproduces punctuation, numbers and common abbreviations accurately

Use any modern framework—TensorFlow, PyTorch, or specialised TTS libraries such as Tacotron 2, FastSpeech, or VITS—as long as the final voice quality meets the criteria above. If you have pre-trained checkpoints or proprietary techniques that accelerate training, feel free to suggest them. Source code must be included.
python android software architecture arduino tensorflow ai text-to-speech ai model development ai development
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.