Multilingual Speech Data – Existing/Ready-Made Datasets Required
Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted2 hours ago
We are looking for agencies, call centres, data providers, research organizations, or individuals who already have existing/ready-made audio datasets for a multilingual speech data project.
Important: This is not primarily a new recording project. We are specifically looking for previously recorded, legally transferable and redistributable audio data that meets the required language, accent, and dialect specifications.
Languages Required
English
Japanese
Spanish
Arabic
Korean
Thai
Vietnamese
Indonesian
Malay
Portuguese
Data Requirements
Existing/ready-made audio data is strongly preferred.
The data must be legally transferable and redistributable to the client.
Data may include:
Single-speaker recordings
Two-person conversations/dialogues
Multi-person conversations
Multi-person conversational data with 3 or more speakers is highly preferred.
The data may cover a wide range of domains, including:
Technology, healthcare, education, finance, transportation, manufacturing, retail, automotive, gaming, sports, customer service, and other relevant domains.
For standard-accent data, recordings should correspond to the relevant language and region.
For accent and dialect data, there is no specific restriction on the recording format or domain, subject to the client's evaluation and acceptance.
Who We Are Looking For
We welcome proposals from:
Speech/audio data providers
Call centres with existing recordings
AI/ML training data companies
Language-data agencies
Research and data-collection organizations
Companies with legally reusable speech datasets
Individuals or organizations possessing suitable transferable datasets
Please Include the Following in Your Proposal
If you have relevant existing datasets, please provide:
Language and dialect/accent available
Approximate total duration of the audio (hours/minutes)
Number of speakers
Whether the data is:
Single-speaker
Two-person conversation
Multi-person conversation
Number of recordings/files, if available
Recording format (WAV, MP3, etc.)
Audio quality/specifications, if available
General domain/content type
Confirmation that the data is legally transferable and redistributable
Whether you can provide a sample dataset for evaluation
Expected price/rate
Location/company details, if applicable
Important
Please do not submit generic proposals.
We are specifically interested in individuals, companies, agencies, call centres, or organizations that already possess suitable recorded speech data or can provide access to an existing dataset.
The final project volume, budget, and exact requirements will be confirmed after the client evaluates the available resources and samples.
If you have relevant existing datasets matching the above requirements, please submit your proposal with the requested details.
We look forward to working with suitable data providers for this project.
Important: This is not primarily a new recording project. We are specifically looking for previously recorded, legally transferable and redistributable audio data that meets the required language, accent, and dialect specifications.
Languages Required
English
Japanese
Spanish
Arabic
Korean
Thai
Vietnamese
Indonesian
Malay
Portuguese
Data Requirements
Existing/ready-made audio data is strongly preferred.
The data must be legally transferable and redistributable to the client.
Data may include:
Single-speaker recordings
Two-person conversations/dialogues
Multi-person conversations
Multi-person conversational data with 3 or more speakers is highly preferred.
The data may cover a wide range of domains, including:
Technology, healthcare, education, finance, transportation, manufacturing, retail, automotive, gaming, sports, customer service, and other relevant domains.
For standard-accent data, recordings should correspond to the relevant language and region.
For accent and dialect data, there is no specific restriction on the recording format or domain, subject to the client's evaluation and acceptance.
Who We Are Looking For
We welcome proposals from:
Speech/audio data providers
Call centres with existing recordings
AI/ML training data companies
Language-data agencies
Research and data-collection organizations
Companies with legally reusable speech datasets
Individuals or organizations possessing suitable transferable datasets
Please Include the Following in Your Proposal
If you have relevant existing datasets, please provide:
Language and dialect/accent available
Approximate total duration of the audio (hours/minutes)
Number of speakers
Whether the data is:
Single-speaker
Two-person conversation
Multi-person conversation
Number of recordings/files, if available
Recording format (WAV, MP3, etc.)
Audio quality/specifications, if available
General domain/content type
Confirmation that the data is legally transferable and redistributable
Whether you can provide a sample dataset for evaluation
Expected price/rate
Location/company details, if applicable
Important
Please do not submit generic proposals.
We are specifically interested in individuals, companies, agencies, call centres, or organizations that already possess suitable recorded speech data or can provide access to an existing dataset.
The final project volume, budget, and exact requirements will be confirmed after the client evaluates the available resources and samples.
If you have relevant existing datasets matching the above requirements, please submit your proposal with the requested details.
We look forward to working with suitable data providers for this project.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.