CrowdRMLE framework cuts video quality assessment time for large-scale datasets
Researchers at Politecnico di Torino have introduced CrowdRMLE, a computationally efficient framework for analyzing quality assessment datasets from crowdsourcing experiments. By utilizing an analytical solution to model subject-level scale quantization and inconsistency, the method significantly reduces processing time for large-scale datasets compared to previous approaches.
Key Takeaways
- Processed a dataset of 11,000+ stimuli and 2,200+ subjects in under two seconds, compared to six days for the legacy RMLE method.
- Introduced the Scale Quantization Index (SQI) to identify systematic rater behaviors such as central scoring or bimodal bias.
- Demonstrated higher robustness to noise in tests across ten datasets, including laboratory and crowdsourcing environments.
- Identified that scale quantization is significantly more prevalent in crowdsourced settings than in controlled laboratory experiments.
Why It Matters
Accurate subjective quality scores are the bedrock for training the AI models that drive modern adaptive bitrate (ABR) and encoding ladders. As platforms move toward user-generated content (UGC) at scale, the reliance on massive, often noisy crowdsourced datasets has become a bottleneck. CrowdRMLE provides a scalable way to clean these datasets without the heavy computational overhead of nonlinear optimization. For engineers and data scientists, this means faster iteration on quality-of-experience (QoE) metrics and more reliable ground-truth data for tuning encoders. Watch for the integration of this framework into automated quality-monitoring pipelines by major streaming service providers.
Additional Context
The release of CrowdRMLE aligns with a broader shift in the industry toward standardizing Quality of Experience (QoE) in uncontrolled environments. Per ITU-T updates in late 2023 and throughout 2024, Recommendation P.910 has increasingly emphasized flexibility in rating scales and environment to accommodate mobile and streaming-first viewing. Research published in late 2025 by Babak Naderi and colleagues on arXiv noted that while crowdsourcing is cost-effective, method-specific biases like compressed scale use in Absolute Category Rating (ACR) can shift saturation points in bitrate-ladder recommendations, emphasizing the need for the structured noise modeling CrowdRMLE provides. Technological progress in 2026 has further complicated quality assessment with the rise of generative AI (GenAI) and short-form video. According to recent IEEE ICIP 2026 reporting, new metrics like the Quality Consistency Score (QCS) are now emerging to complement average quality scores by highlighting transient drops that traditional Mean Opinion Score (MOS) averages miss. Meanwhile, University of Konstanz researchers have expanded datasets like KonViD to include KonViD-150k in 2025, providing a massive 150,000-video benchmark that traditional recovery methods. such as the legacy RMLE, simply cannot handle. Industry leaders are also pivoting toward "machine-oriented" quality assessment. As per the MoIQA challenge at ACM MM 2026, there is a growing trend of assessing video quality not just for human viewers, but for the computer vision models that handle automated content moderation and face detection. This evolution suggests that robust, efficient rater-cleaning frameworks like CrowdRMLE will be essential not only for streaming delivery but for the metadata-driven processing layers of the media supply chain.
Read full article at iris.polito.it
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source