Precise Audio Transcribers Wanted – AI Project

via Freelancer ·

Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted2 hours ago
Audio Transcriptionist – AI Training Data (Ermis Transcription Project)

Job Type: Contract / Freelance (Remote) Category: Data Annotation / AI Training Data / Linguistics Compensation: [Insert rate — e.g., per-audio-file or hourly]

About the Role

We're looking for detail-oriented Transcriptionists to join the Ermis Transcription project, supporting the development of a text-to-speech AI model. In this role, you'll listen to short audio clips of human speech and produce precise, spoken-form transcriptions that capture not just the words spoken, but every filler, hesitation, repetition, and non-verbal vocal cue exactly as it occurs. Your work directly shapes how an AI model learns to understand and generate natural human speech.

This is detail-heavy, linguistically rigorous work. If you have a strong ear for spoken language, enjoy following precise style guides, and take pride in getting the small things right, this role is for you.

What You'll Do
Listen to short audio recordings of human speech and produce accurate, spoken-form transcriptions
Capture disfluencies faithfully — filled pauses ("um," "uh"), backchannels ("mm-hm," "uh-huh"), stutters, self-corrections, and repetitions — exactly as spoken, without cleaning up grammar or "fixing" the speaker's language
Apply standardized formatting rules for numbers, acronyms, letter sequences, alphanumeric codes, proper nouns, catalog entities (songs, movies, brands, etc.), URLs, and usernames
Use a defined tagging system to mark features like multiple speakers, overlapping speech, cross talk, laughter, non-lexical sounds, background speech, audio artifacts, garbled audio, and sensitive content
Research unfamiliar names, brands, and catalog entities to confirm correct spelling using an approved dictionary/resource hierarchy
Identify and label speaker gender (masculine-sounding, feminine-sounding, unknown) and nativity (native, non-native, unknown, TTS) for each speaker in a file
Distinguish between genuinely unintelligible audio and clearly intelligible speech, using best-guess transcription where appropriate
Follow a strict "transcribe what you hear, not what you expect" standard — no deletions, no hallucinated words, no substitutions
Submit completed transcriptions through the designated task tool
What We're Looking For
Excellent English listening comprehension and a strong ear for accents, dialects, and informal/colloquial speech
High attention to detail and comfort following a detailed, rule-based style guide (50+ pages of specific conventions)
Strong written English with clean spelling and formatting habits
Comfort doing quick online research to verify names, spellings, and catalog entities (music, film, TV, brands, etc.)
Patience to replay and closely listen to short audio segments, sometimes multiple times, to ensure accuracy
Prior experience in transcription, captioning, linguistics, audio QA, or data annotation is a plus but not required
Reliable internet connection and a quiet listening environment with headphones
Important Notes
Use of Automatic Speech Recognition (ASR) tools to generate transcriptions is strictly prohibited. All transcriptions must be produced manually by the transcriber.
This work involves listening to a wide variety of real, informal speech, which may occasionally include sensitive or explicit content; annotators use their judgment and a defined tag to flag such content.
Training and detailed style-guide documentation will be provided; a comprehension check or short paid assessment may be part of onboarding.
Why This Work Matters

Every transcription becomes a training signal for a speech AI model — teaching it what real human speech (including all its natural imperfections) sounds like. Accuracy and consistency in this role have a direct, measurable impact on the quality and reliability of the resulting AI product.
audio services transcription copy typing audio production linguistics natural language processing data annotation speech recognition
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.