ComfyUI workflow adds lip-sync dubs in six languages
This document describes a workflow for creating lip-synced video dubs and rephrases using a ComfyUI-based system, leveraging LTX-2.3-22b-IC-LoRA-LipDub models from Hugging Face. The workflow allows for translating speech or rephrasing dialogue in a video while generating new lip movements and audio to match the provided text, supporting multiple languages and emphasizing the importance of matching original dialogue length for natural-sounding output.
Key Takeaways
- The workflow uses LTX-2.3-22b-IC-LoRA-LipDub from Hugging Face inside ComfyUI.
- It supports both dubbing, which translates speech, and rephrasing, which changes dialogue without changing language.
- The model generates new lip movements and audio to match the provided text.
- The notes say the current LoRA supports only one speaker.
- The document recommends using native script for the target language and keeping the translated line close to the original length to avoid skipped words or unnatural pacing.
Why It Matters
This is a practical pipeline for localized video edits: the workflow does not just swap audio, it also regenerates lip motion to fit the new line. That matters for dubbing and rephrasing use cases where visible mouth movement is part of the quality bar. The document also makes the operating constraints clear: one speaker only, native script, and similar line length. For teams evaluating AI video tooling, those limits are as important as the generation itself. Watch for how the same workflow behaves across the listed language prompts and whether longer or shorter translated lines degrade output.
Additional Context
The release of the LTX-2.3 LipDub workflow aligns with a broader industry shift toward diffusion-based architectures for video localization. Per lipsync.com (February 2026), diffusion models are rapidly replacing Generative Adversarial Networks (GANs) due to their superior ability to preserve subject identity and render fine mouth detail. While GAN-based tools like Wav2Lip established the category, they often suffered from training instability and visual artifacts that modern diffusion pipelines successfully resolve. This shift is fueling a market expansion for AI dubbing tools, which Market.us (May 2026) projects will grow from $2.75 billion in 2025 to nearly $19 billion by 2035. In the commercial sector, specialized platforms are increasingly moving toward all-in-one multimodal solutions. According to NYBreakings (May 2026), tools such as Magic Hour and HeyGen are winning market share by combining video generation, voice cloning, and lip-syncing into unified workflows. These platforms cater to a growing enterprise demand—highlighted by Synthesia (April 2026)—where high-fidelity lip-sync is now a standard requirement for professional business communications. The democratization of these tools allows small creators and marketing teams to bypass traditional studio costs, which RWS (January 2026) estimates can reduce localization expenses by approximately 90%. Furthermore, the technical implementation of the LTX-2.3 workflow reflects a trend toward hybrid editing. Per TodaysMagazine (May 2026), production teams are increasingly blending traditional post-production tools with custom AI nodes rather than relying on monolithic software. This modular approach, supported by open-source repositories on Hugging Face and platforms like ComfyUI, enables granular control over specific elements such as vocal timbre and frame consistency. As more companies adopt these systems, the use of proprietary watermarking and verification tools is also rising to ensure transparency in AI-modified media.
Read full article at drive.google.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source