Off-the-Shelf Dual-Channel Speech Corpora
Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted50 minutes ago
I need ready-made, professionally recorded conversational speech that I can license immediately for AI training. All material must be dual channel (each speaker isolated), feature spontaneous human dialogue, and come with accurate time-aligned transcriptions plus basic speaker metadata.
Languages and expected hours per language are as follows:
• Spanish – US (es-US): 220h
• French – Canada (fr-CA): 220h
• Arabic – UAE (ar-AE): 220h
• Italian (it-IT): 220h
• Thai (th-TH): 195h
• Turkish (tr-TR): 220h
• Hebrew (he-IL): 220h
• Dutch (nl-NL): 220h
• Greek (el-GR): 220h
• Vietnamese (vi-VN): 220h
The recordings should cover a healthy mix of customer service interactions, casual conversations, and technical support calls. Within those, I specifically want sessions that touch on Healthcare, Finance, Travel & Tourism, and everyday general topics so the dataset remains domain-diverse.
Key requirements
• Dual-channel WAV or FLAC with 8 kHz, 16-bit PCM (or higher)
• Human-generated speech only; no synthetic voices or read scripts
• Natural background noise acceptable as long as speech is clear
• Turn-level, verbatim transcription in UTF-8 with speaker labels and timestamps
• Documentation outlining collection methodology, consent, and demographics (age, gender, accent, mic type)
Acceptance criteria
1. Sample package (5-10 minutes) passes a quick audio quality and transcription accuracy check.
2. Full corpus delivery via secure download or shipped drive.
3. Metadata and licenses confirm worldwide, perpetual research & commercial use.
If you already have all or part of this inventory, let me know which languages, hours, and domains you can supply and your proposed delivery schedule.
Languages and expected hours per language are as follows:
• Spanish – US (es-US): 220h
• French – Canada (fr-CA): 220h
• Arabic – UAE (ar-AE): 220h
• Italian (it-IT): 220h
• Thai (th-TH): 195h
• Turkish (tr-TR): 220h
• Hebrew (he-IL): 220h
• Dutch (nl-NL): 220h
• Greek (el-GR): 220h
• Vietnamese (vi-VN): 220h
The recordings should cover a healthy mix of customer service interactions, casual conversations, and technical support calls. Within those, I specifically want sessions that touch on Healthcare, Finance, Travel & Tourism, and everyday general topics so the dataset remains domain-diverse.
Key requirements
• Dual-channel WAV or FLAC with 8 kHz, 16-bit PCM (or higher)
• Human-generated speech only; no synthetic voices or read scripts
• Natural background noise acceptable as long as speech is clear
• Turn-level, verbatim transcription in UTF-8 with speaker labels and timestamps
• Documentation outlining collection methodology, consent, and demographics (age, gender, accent, mic type)
Acceptance criteria
1. Sample package (5-10 minutes) passes a quick audio quality and transcription accuracy check.
2. Full corpus delivery via secure download or shipped drive.
3. Metadata and licenses confirm worldwide, perpetual research & commercial use.
If you already have all or part of this inventory, let me know which languages, hours, and domains you can supply and your proposed delivery schedule.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.