Researchers at Mitsubishi Electric Research Laboratories have introduced a streaming multimodal Q-Former framework designed for real-time robot action generation from unsegmented audiovisual inputs. The method achieves low-latency performance with less than 10% accuracy degradation compared to offline processing, offering a scalable approach for interactive human-robot collaboration.
This development addresses the critical latency bottleneck in human-robot collaboration by shifting from batch-processed offline analysis to continuous streaming inference. By maintaining 90% of offline accuracy while processing unsegmented audiovisual data, the framework provides a viable path for deploying humanoid robots in dynamic environments like kitchens or factories. Within the broader AI ecosystem, this demonstrates that Q-Former architectures can be adapted for temporal streaming without retraining the underlying large language model. Industry observers should monitor whether Mitsubishi Electric Research Laboratories integrates streaming automatic speech recognition to further refine action descriptions in future iterations.
Researchers at Mitsubishi Electric Research Laboratories have developed a streaming multimodal Q-Former framework that allows robots to generate actions from unsegmented audiovisual inputs in real-time. By processing data in one-second chunks, the system achieves 90% of offline accuracy, significantly reducing latency and enabling more effective human-robot collaboration in dynamic environments.
It is a framework that enables robots to generate action sequences and confirmation messages from unsegmented audiovisual inputs in real-time using a frozen OPT-2.7B large language model.
The system uses loss-based alignment selection and processes video in one-second chunks, which reduced average latency by 2.92 seconds compared to traditional methods.
The framework maintains high performance, achieving less than 10% accuracy loss compared to offline processing while eliminating the need for manual clip segmentation.
The model was tested on the YouCook2 dataset, which demonstrated that interleaving multimodal features does not negatively impact baseline performance.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source