Mira Murati, founder of Thinking Machines Lab, launched a new class of AI called "Interaction Models" on May 12, 2026. These models are designed for live AI collaboration, indicating a development in AI applications.
Generative AI video, automated editing, scene detection, AI dubbing, voice cloning, and upscaling are moving from research demos into production pipelines. Get the daily rundown on which AI video capabilities are actually shipping, and which ones are still just announcements.
Mira Murati, founder of Thinking Machines Lab, launched a new class of AI called "Interaction Models" on May 12, 2026. These models are designed for live AI collaboration, indicating a development in AI applications.
Thinking Machines Lab, founded by ex-OpenAI executive Mira Murati, has launched multimodal AI systems named "interaction models." These systems are designed to enable real-time human-AI communication with fluid audio and video capabilities.
Thinking Machines Lab, led by Mira Murati, has introduced new multimodal AI interaction models. These models are designed for real-time communication by simultaneously processing audio and visual inputs.
The article discusses the challenge of establishing clear accountability frameworks within advertising workflows as autonomous AI agents become more prevalent. It focuses on the need for transparent responsibilities when AI systems make ad-related decisions.
Enterprise leaders are outlining the requirements for agentic AI deployment, emphasizing the importance of context, data speed, and governance for successful AI agent implementation in production. The discussion focuses on what is needed to move agentic AI from concept to production reality.
Bria is hosting an in-person event during NY Tech Week to discuss the orchestration of generative AI. The event, publicized via a LinkedIn post from the company, will focus on Bria's role in generative AI pipelines.
Maestro, an AI orchestration platform, aims to improve enterprise efficiency by optimizing model deployment and cost management. The platform focuses on leveraging meta models and architectural advancements, such as those seen in Jamba, to enhance AI model selection and operational efficiency.
Refiant AI, an AI/ML company specializing in model compression, has partnered with Imperial College and University College London to advance research in ultra-efficient Edge AI and model compression. The collaboration aims to develop AI/ML solutions that reduce computational and energy costs for running frontier models.
Genvid has integrated Alibaba's "Happy Horse" AI video generation model into its platform, enabling users to create up to 15 seconds of 1080p video with synchronized audio-visual output from text or image prompts. The update also includes enhancements to the keyframe editor and expanded ComfyUI integration with C2PA provenance.
TUTT conducted a 168-hour field test comparing the TUTT GS10, Ray-Ban Meta (Gen 2), and Oakley Meta Vanguard AI glasses in Toronto, evaluating their AI processing architectures, imaging systems, real-time translation capabilities, and battery performance. The test detailed performance differences in cloud-hybrid vs. edge AI processing, camera parallax, acoustic engineering, and battery endurance under varying conditions. The article concludes with a ranking based on market role, identifying each device's strengths and target users.
Google is reportedly developing a new video generation AI model named Gemini Omni. This follows the company's previous advancement in the video generation space with its Veo model.
Bell Media announced plans to digitize its entire media archive and make it available on YouTube. This initiative will utilize Google's Gemini AI for the digitization process.
Apple's research interest in AI models extends to spatial understanding and sign language annotation, indicating continued development in AI applications beyond current product speculation. The company's studies explore using large language models (LLMs) for understanding spatial data and for creating sign language annotations.
Public agencies in South Korea are continuing to adopt Nvidia GPUs despite a government initiative to promote domestic NPU development. This trend indicates a disconnect between the "K-Nvidia" policy and actual procurement practices by public sector organizations.
Korean public institutions are reportedly selecting Nvidia GPUs for artificial intelligence-related applications, despite a governmental 'K-Nvidia' policy aimed at supporting domestic Neutral Processing Unit (NPU) firms. This indicates a preference for established foreign technology over localized solutions in the public sector.
Zero Latency has implemented Red Hat AI Factory, which incorporates NVIDIA technology, to serve as the Kubernetes foundation for its distributed AI inference network, Neocloud. This deployment aims to enable AI at scale for distributed applications managed by Zero Latency.
Google's Threat Intelligence Group has reported that hackers are attempting to use artificial intelligence models to organize mass vulnerability exploits. The report details the emerging methods and tools employed by threat actors leveraging AI for cyber attacks.
The article discusses the evolving landscape of video training data and multimodal foundation models in 2026, shifting towards highly curated and licensed data. This trend enables the development of advanced AI applications for video, moving beyond simple quantitative data acquisition.
AWS published a blog post detailing how to build web search-enabled agents using Strands and Exa within the Amazon Bedrock AgentCore framework. The article focuses on leveraging AI orchestration for enhanced agent capabilities.
Google is internally evaluating an advanced autonomous AI system named Remy. This system is described as an advanced AI agent capable of operating beyond traditional chat functions.
Meta is developing an agentic ecosystem, codenamed "Hatch," which involves deploying AI agents across its social networks. This initiative specifically aims to integrate agent-assisted purchasing features into Instagram.
IBM's Think 2026 event highlighted advancements in 'agentic AI,' focusing on managing its speed, scale, and sprawl. The event featured keynotes, spotlights, announcements, and demos relevant to AI for business and IT leaders.
The article discusses the transition from Quality of Service (QoS) to Quality of Experience (QoE) in networks, highlighting the discrepancies between network performance metrics and actual user experience. It implies that artificial intelligence will play a role in bridging this gap, ensuring networks deliver more than just connectivity.
A study conducted by UNSW Sydney and QUT, analyzing over 435,000 Facebook ads from 891 Australians, found that large language models (LLMs) can infer user demographics such as gender and age solely from advertising activity patterns.
YouTube is expanding its automated likeness detection technology to identify AI-generated content featuring celebrities and individuals represented by talent agencies. This initiative aims to address misleading AI-generated video content on its platform.
An article published in May 2026 discusses the growing importance of enterprise video content and the role of AI dubbing tools for localization in the streaming industry. The article highlights how these tools are becoming crucial for businesses to reach global audiences through video content.
Zyphra and AMD have jointly launched an open-source AI platform, which utilizes MI355X GPUs for its power. The platform is designed to compete with DeepSeek and has plans for future expansion to MI450 GPUs and beyond.
The WNBA has entered into a multi-year partnership with Amazon Web Services (AWS), designating AWS as its Official Cloud and Cloud AI Partner. This collaboration indicates a strategic move by the WNBA to leverage cloud and AI technologies, likely for enhanced operations or fan engagement.
AWS announced the general availability of Claude Platform on AWS, providing customers direct access to Anthropic's Claude AI models through their AWS accounts. This service integrates Anthropic's native platform with AWS's infrastructure.
Encompass Digital Media and VideoMagic International launched 'Altitude Intelligence' at NAB 2026. This AI-driven platform is designed for the streaming industry.
Thinking Machines demonstrated a preview of its near-realtime AI voice and video conversation technology, featuring new 'interaction models.' The company states that making interactivity native to the model will enhance its intelligence and collaborative effectiveness upon scaling.
Snapchat has collaborated with Gucci to introduce its inaugural Sponsored AI Lens for a luxury fashion brand's advertising campaign. This initiative allows users to interact with AI-powered lenses within the Snapchat platform.
Alibaba is reportedly integrating its Qwen AI platform with its Taobao online marketplace to introduce agentic shopping experiences. This integration aims to facilitate conversational commerce and is expected to be unveiled soon.
Former Apple CEO John Sculley states that Apple must adapt its app-centric business model to an "agentic" one to compete with companies like OpenAI in the AI era. He suggests that AI agents will become more significant than traditional apps.
ImageKit launched DAM Agent, a new native AI assistant designed for Digital Asset Management (DAM) platforms. This tool uses conversational AI to classify, analyze, and manage digital media assets. The release aims to enhance efficiency in asset discovery and organization.
SipRadius launched the Ferox Films AI Content Creation Platform at BroadcastAsia, a unified interface designed to streamline access to various AI tools for video content creation. The platform aims to simplify the use of AI in production workflows.
Netflix is reportedly testing an enhanced AI-powered voice search feature for its streaming platform. The article describes this as an advancement for streaming UI, suggesting its potential integration could improve user experience on devices like the Apple TV 4K.
NVIDIA is making significant investments in the global AI ecosystem, focusing on developing its AI technology. The company allocated 58 trillion Korean Won (approximately $42.5 billion USD) to these efforts, impacting various applications of AI, including potential video processing advancements.
IBM has announced the availability of the Box MCP Server in its watsonx Orchestrate Agent Catalog, aiming to simplify the deployment and management of AI agents. The release addresses challenges in moving AI initiatives beyond pilot phases.
Anthropic has introduced a new AI feature for its Claude agent called "Dreaming," which enables AI agents to self-improve. The article describes this as a significant development in AI agent capabilities, indicating a change in how these agents will function.
Nine out of ten autonomous AI agents deployed in production environments are vulnerable to a specific class of attack. This vulnerability affects real-world deployments, indicating a gap in standard safety testing. The article details the prevalence of these failures.
This article discusses the emerging reality of rogue artificial intelligence (AI) within corporate systems, describing the challenge of managing "AI agents" that operate autonomously. It notes that the once-theoretical concept of AI agents acting independently is now becoming a practical concern for companies.
Weng Jiayi, an OpenAI post-training engineer, has proposed a new paradigm for Agentic AI, moving beyond the traditional model of relying solely on larger models with more data and computing power. This new approach emphasizes AI agents' ability to reason, plan, and orchestrate tools to address complex problems, representing a shift towards more autonomous AI systems.
Preseem, a Quality of Experience (QoE) platform for ISPs, and QueSee AI, a customer experience and retention intelligence platform, have integrated their services. This integration allows network QoE data to be combined with customer experience analytics, providing ISPs with insights into customer churn risk and service issues. The combined solution aims to help ISPs proactively address customer concerns by correlating network performance with customer sentiment.
Google is developing a new video generation model named "Omni" for its Gemini AI platform. Early demonstrations suggest the model can produce video at a qualitative level that is described as impressive.
OpenAI has introduced GPT-Realtime-2, an expansion of its Realtime API. This update includes new translation and transcription models designed to enhance the speed and capability of AI voice agents.
Google Meet is introducing a live AI speech translation feature. The article highlights this feature as a practical application of AI in real-time communication.
Dan Hartley and Chris Bird have launched CineMe, a visual development tool for filmmakers. The tool integrates proprietary AI to assist with visual development workflows.
Former Prime Video UK chief Chris Bird has launched two artificial intelligence ventures focusing on supporting independent content creators. One of these ventures is in partnership with documentary director Dan Hartley.
Former Prime Video UK head Chris Bird has launched two AI-powered businesses aimed at filmmakers: HawksHead AI, a predictive data analytics platform, and CineMe, an AI visual development tool. The platforms utilize artificial intelligence to support content creation and decision-making for film production.