Consistent AI Avatar Workflow
Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted3 hours ago
I’m ready to level-up my production stack with a rock-solid LTX 2.3 + ComfyUI pipeline that keeps both the face and the voice perfectly locked from the first frame to the last. The end goal is a reusable workflow I can drop into:
• UGC-style ads that feel authentic yet never drift
• Short narrative pieces
• Full-length films
My priority is absolute character and audio consistency—no jitter, morphing, or timbre shifts even in multi-minute takes. Platform-wise, I want the solution to stay flexible enough to deploy on web, mobile, and desktop front-ends as the projects demand.
Key expectations
1. A repeatable LTX 2.3 / ComfyUI graph (or set of graphs) that handles face-locking, lip-sync, and stable voice generation in a single pass.
2. An example project that proves it works: one talking-head clip and one multi-scene sequence, each rendered without visual or audio drift.
3. Clear documentation: installation notes, recommended GPU/driver settings, and step-by-step usage so I can hand it to any editor on my team.
4. Source presets or checkpoints for maintaining character appearance across episodes or campaigns.
Acceptance criteria
• Same character model matches reference stills at ±1 pixel tolerance across every shot.
• Voice fingerprint deviation ≤ 3 % on standard speaker-verification metrics for clips up to 10 minutes.
• Rendered clips play without desync on a vanilla playback test (VLC, Safari, Android).
If you already ship bulletproof talking-head systems and can demo an end-to-end run tonight, message me (908-204-6313 is fastest). Let’s lock this in quickly and start producing.
• UGC-style ads that feel authentic yet never drift
• Short narrative pieces
• Full-length films
My priority is absolute character and audio consistency—no jitter, morphing, or timbre shifts even in multi-minute takes. Platform-wise, I want the solution to stay flexible enough to deploy on web, mobile, and desktop front-ends as the projects demand.
Key expectations
1. A repeatable LTX 2.3 / ComfyUI graph (or set of graphs) that handles face-locking, lip-sync, and stable voice generation in a single pass.
2. An example project that proves it works: one talking-head clip and one multi-scene sequence, each rendered without visual or audio drift.
3. Clear documentation: installation notes, recommended GPU/driver settings, and step-by-step usage so I can hand it to any editor on my team.
4. Source presets or checkpoints for maintaining character appearance across episodes or campaigns.
Acceptance criteria
• Same character model matches reference stills at ±1 pixel tolerance across every shot.
• Voice fingerprint deviation ≤ 3 % on standard speaker-verification metrics for clips up to 10 minutes.
• Rendered clips play without desync on a vanilla playback test (VLC, Safari, Android).
If you already ship bulletproof talking-head systems and can demo an end-to-end run tonight, message me (908-204-6313 is fastest). Let’s lock this in quickly and start producing.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.