AI Script-to-Video Agent using API
Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted2 hours ago
I want a self-contained AI agent that takes a written script as input and returns a fully rendered, realistic-looking MP4. The flow I have in mind is simple for the user—drop in the text, pick a resolution, press run—yet sophisticated behind the scenes:
• The agent parses the script, breaks it into scenes and shots, and generates matching realistic visuals (you may lean on services like Runway Gen-2, Pika Labs, Stable Video Diffusion, or any alternative you trust).
• Dialogue is converted to lifelike voice-over, timed to the generated footage, then mixed with ambient sound or music.
• Footage is stitched, color-matched, and exported as a single MP4 ready to share.
You are free to choose the stack—Python, Node, or a no-code/low-code orchestration—so long as it can run locally on a decent GPU or call external APIs with minimal setup. A lightweight web interface or CLI helper that hides complexity from the end user is essential.
Deliverables
1. Source code and packaged agent (Docker or virtual-env preferred)
2. A sample MP4 produced from a test script I will provide
3. Step-by-step setup and usage guide
4. Short video walkthrough of the workflow in action
Acceptance criteria
• The output video matches the scene descriptions in a visibly realistic style.
• Length, frame rate, and audio are synced and free from obvious artefacts.
• One-click (or single command) operation from raw script to finished MP4.
• All libraries, APIs, and model dependencies are documented.
If you have already built something similar, feel free to show a reel or repo link in your bid. Let’s bring words to life on screen.
• The agent parses the script, breaks it into scenes and shots, and generates matching realistic visuals (you may lean on services like Runway Gen-2, Pika Labs, Stable Video Diffusion, or any alternative you trust).
• Dialogue is converted to lifelike voice-over, timed to the generated footage, then mixed with ambient sound or music.
• Footage is stitched, color-matched, and exported as a single MP4 ready to share.
You are free to choose the stack—Python, Node, or a no-code/low-code orchestration—so long as it can run locally on a decent GPU or call external APIs with minimal setup. A lightweight web interface or CLI helper that hides complexity from the end user is essential.
Deliverables
1. Source code and packaged agent (Docker or virtual-env preferred)
2. A sample MP4 produced from a test script I will provide
3. Step-by-step setup and usage guide
4. Short video walkthrough of the workflow in action
Acceptance criteria
• The output video matches the scene descriptions in a visibly realistic style.
• Length, frame rate, and audio are synced and free from obvious artefacts.
• One-click (or single command) operation from raw script to finished MP4.
• All libraries, APIs, and model dependencies are documented.
If you have already built something similar, feel free to show a reel or repo link in your bid. Let’s bring words to life on screen.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.