Speaker Recognition in Noisy Environments

via Freelancer ·

Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted3 hours ago
I need a robust solution that can pick out individual human voices even when the recording is full of background clatter—think cafés, factory floors, or busy streets. The sole objective is to identify who is speaking; transcription and emotion analysis are outside the scope.

I’m open to any modern approach—deep-learning architectures such as ECAPA-TDNN, x-vector, or a custom CNN-RNN hybrid—so long as the final model remains reliable when the signal-to-noise ratio drops. Python is preferred for the pipeline, and frameworks like PyTorch or TensorFlow are perfectly acceptable. Please work with publicly licensable datasets or clearly state any proprietary material you intend to use, and describe your noise-augmentation strategy up front so I can vet it.

Deliverables
• A trained speaker-recognition model capable of handling noisy audio
• A lightweight API or CLI demo that accepts a wav/mp3 file and returns the speaker ID plus a confidence score
• A short technical report covering data sources, preprocessing, training procedure, and validation results, including accuracy on a withheld noisy test set (≥90 % target)

Share your proposed workflow, expected timeline, and any questions you need answered before kicking off. I’m ready to move quickly once I’m confident the approach can handle real-world noise.
python audio services software architecture machine learning (ml) c++ programming audio processing deep learning speech recognition
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.