Python Bot for Audiobook TTS & Localization

via Freelancer ·

Budget / Salary₹600–1,500
TypeFreelance project
LocationRemote
Posted1 hour ago
Job Post Title:
Python / Automation Expert Needed for Pocket FM-Style Multi-Voice TTS Audiobook Bot (Dialogue Aware)

Job Description:
Hi, I run a novel-publishing platform and create high-volume audiobooks daily. I am looking for an expert Python developer to build a dedicated Text-to-Speech (TTS) Automation, Translation & Content Editing Tool.

Since I process large files daily, I need a robust, permanent solution. If you cannot deliver the core features within a free/stable architecture without crashing, please do not bid.

Core Project Requirements (Triple Input Mode):
1. Multi-Functional Modes:
* Mode A (Full Audiobook Pipeline): The bot takes an English TXT file, translates it to Hindi, applies Pocket FM-style localization, and exports to a Multi-Voice MP3.
* Mode B (Direct Hindi TTS): If I upload a file that is already in Hindi, the bot must skip the translation step entirely and directly process the Hindi text for character/voice detection and MP3 export.
* Mode C (Plain Text Translation): The bot must also function as a simple translator. If selected, it should take a 10 MB English TXT file and directly translate it into plain Hindi text WITHOUT any modifications, edits, or name alterations, and export it back as a plain .txt file (no audio generation).

2. Pocket FM-Style Localization & Editing (Bulk Replace): The tool must automatically Indianize the entire text without altering the core story or plot. It must accept a secondary configuration/mapping file (Excel or text) to bulk-replace content. For example:
* Names: Xiao Chen Rohan, Lin Fan Lakhan
* Currency/Terms: Yuan Rupees, Sect/Clan Ashram/Gharana
* Locations: Chinese cities Indian cities (e.g., Delhi, Mumbai)

3. Dialogue-Aware Multi-Voice Output: The system must automatically scan the novel text file. Anything inside quotation marks ("...") must be rendered in a Female/Character Voice, and the standard text outside quotes must be rendered in a Male/Narrator Voice (using free Microsoft Edge-TTS or similar clear, natural Indian voices like Madhur/Swara).

4. No Code Formatting Needed From Me: I will NOT manually format text files with tags like "Narrator:" or "Rahul:". The code must be smart enough to parse native novel structures dynamically in both English and Hindi inputs.

5. Large File Handling (Up to 10 MB): The pipeline must easily handle text files up to 10 MB at once without freezing, hanging, or throwing memory/timeout errors. It should split the book into chapters/parts automatically and merge them cleanly into separate exported MP3s or text files.

6. No Hidden API Costs / 100% Free Execution: The system must utilize completely free, high-limit models or engines (e.g., Edge-TTS or zero-cost high-limit routing endpoints). I will not pay monthly API/hosting subscriptions.

7. Robust Deployment Interface: I prefer a stable, responsive web interface that features custom speed control (1.0x to 1.5x), volume boost, password protection, and a visual progress tracker (Part 1/120 processing...).

8. Dedicated "Generate Audio" Button (Separate Trigger): The system must NOT start generating audio automatically when a file is uploaded. After the file is successfully uploaded and processed by the system, the Multi-Voice audio conversion pipeline must ONLY start executing when I explicitly click a separate, dedicated "Generate Audio" action button on the interface.

Strict Terms & Conditions (Zero Excuse & Non-Dispute Policy):
* Cancellation & No-Dispute Agreement: By bidding on this project, the freelancer explicitly agrees that this is an "All-or-Nothing" delivery. If the tool fails to process a 10 MB file, crashes, or fails at multi-voice detection, the project will be cancelled. The freelancer strictly agrees NOT to file any dispute or claim partial payments upon cancellation of an incomplete/broken product.
* 100% Code Handover & Ownership: Upon completion, the freelancer must hand over the complete, clean source code (.py files), requirements.txt, and a full deployment guide. I will have 100% full ownership of the data and code.
* No Server/Environment Excuses: The developer must ensure the tool runs flawlessly on the final hosted platform. I will NOT accept post-award excuses regarding free cloud hosting sleep modes, Termux compatibility issues, mobile OS limitations, or continuous script time-outs. Finding a stable, working environment is fully the developer's responsibility.
* Single Milestone Payment Policy: The project budget will be kept in a single milestone locked until final delivery. No partial or mid-project payments will be released.
* Pre-Award Working Sample: Before I award you this project, you must provide a short 2-minute MP3 audio sample generated from a sample text snippet I provide, demonstrating both Pocket FM-style word replacements and automated dialogue transitions.
python translation machine learning (ml) software development web development audio processing automation api development natural language processing speech synthesis
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.