Meta Muse Video model enters closed beta with native audio sync
Meta has announced Muse Video, a generative AI model currently in closed beta that creates video clips with natively synchronized audio and music. The model aims to improve temporal consistency and motion rendering, positioning Meta to compete with other major AI video generation platforms.
Key Takeaways
- Muse Video generates audio and visuals simultaneously rather than adding soundtracks to silent clips post-production.
- The model shares a pretraining foundation with Muse Image, which is already integrated into Instagram and WhatsApp.
- Meta Superintelligence Labs designed the tool to compete directly with OpenAI’s Sora and Google’s Veo.
- A companion tool called Muse Spark is being developed to handle multi-step agentic functionalities for users.
Why It Matters
The introduction of native audio synchronization addresses a significant technical hurdle in generative video, where sound and motion often feel disconnected. By integrating this capability directly into the model, Meta is attempting to streamline the production workflow for creators who currently rely on multiple AI tools for a single clip. This move intensifies the competition with OpenAI and Google, shifting the focus from simple visual fidelity to functional, multi-modal output. As these tools move toward public release, the industry should monitor how Meta integrates Muse Video into its existing creator suites on Instagram and Facebook to drive platform engagement. Watch for the specific release date of the open beta to gauge the model's scalability.
Additional Context
Meta's Muse Video enters a competitive field where Google and OpenAI have already shipped public video-generation products. Google DeepMind announced Veo 2 in December 2024 as a successor to the original Veo model, claiming the ability to produce clips exceeding two minutes at resolutions up to 4K in theory, though the publicly available VideoFX tool initially capped output at 720p and eight seconds. DeepMind VP of Product Eli Collins acknowledged at launch that coherence over long durations and character consistency remained areas for growth, and the company embedded its SynthID watermarking technology into every generated frame to mitigate deepfake risk.
OpenAI launched Sora publicly on December 9, 2024 as part of its 12-day product release series, offering text-to-video generation at up to 1080p and 20 seconds for ChatGPT Pro subscribers at $200 per month. The product included features such as storyboards, remixing, and blending, and all outputs carried C2PA metadata and visible watermarks. However, OpenAI's own blog confirms that as of April 26, 2026, the Sora product is no longer available, meaning Meta's Muse Video now competes in a space where one of its most prominent rivals has already exited the standalone product market.
On the technical side, Google's Veo 2 documentation on Vertex AI specifies output limited to 720p at 24 FPS with clip lengths of 5 to 8 seconds and no sound generation support, a notable gap that Meta's native audio synchronization directly addresses. The model launched as generally available on May 27, 2025 with a fixed-quota consumption model and a stated retirement date of June 30, 2026, indicating Google is already planning a successor. Meta's emphasis on temporally consistent audio within Muse Video positions it to fill a capability gap that neither Veo 2's current API nor the now-discontinued Sora product offered at scale. To further combat the rise of synthetic media, Vloggi video verification layer has recently launched to provide provenance for generated content, while EU AI Act transparency rules now mandate labeling for synthetic video content. The industry is also seeing a production workflows as these models mature, with new solutions like continuing to drive this trend.
Read full article at cryptobriefing.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source