ETSI launches 18 data quality metrics to standardize AI training
ETSI has published Technical Report 104 180, which establishes 18 standardized metrics for assessing data quality in AI ecosystems. The framework provides a common language and mathematical methods for evaluating datasets, supported by an open-source validation system developed by a multi-institutional research team.
Key Takeaways
- Technical Report 104 180 introduces 18 standardized metrics including accuracy, completeness, and representation bias.
- A multi-institutional team including Sejong University and CNIT developed an open-source Data Quality Validation System to score datasets.
- Proof-of-concept testing validated the framework using Industrial IoT sensor data and demographic datasets.
- Diego Lopez, Chair of ETSI TC DATA, stated the metrics provide a common language for establishing trustworthy AI.
Why It Matters
The introduction of these 18 metrics provides a technical baseline for streaming companies to audit the datasets used in recommendation engines and automated content moderation. By moving away from proprietary quality assessments toward a standardized mathematical framework, the industry can better address algorithmic bias and data integrity issues that plague large-scale AI deployments. This standardization is critical as the streaming ecosystem shifts toward more automated, data-driven decision-making where 'black box' models require transparent validation. Watch for whether major cloud providers and streaming platforms integrate this open-source validation system into their existing MLOps pipelines to demonstrate compliance with emerging AI transparency requirements.
Additional Context
ETSI's push to standardize data quality assessment arrives as telecom operators accelerate production deployments of AI-driven network automation that depend on high-quality training data. In June 2026, Ericsson launched its AI in RAN commercial software subscription claiming up to 20% higher downlink throughput and up to 10% better spectral efficiency across more than 15 live deployments, while Nokia and Indosat Ooredoo Hutchison announced a GPU-accelerated AI-RAN partnership in Indonesia on June 8, expanding the Nokia-NVIDIA architecture already adopted by T-Mobile US, SoftBank, and Vodafone. These deployments generate the operational datasets that ETSI TR 104 180's 18 metrics are designed to evaluate, making the framework directly relevant to the telecom sector's AI ambitions.
The business case for standardized data quality metrics is sharpened by the competitive dynamics among vendors building agentic AI platforms for operators. At DTW Ignite 2026 in Copenhagen, Nokia teamed up with Google Cloud to build six specialized AI agents using Gemini technology capable of slashing network problem-solving times by 50% to 80%, with plans to launch the agentic platform on Google Cloud Marketplace in September. Separately, Nokia announced work with AWS and Databricks to build the data, cloud, and control layers for autonomous networks under its Autonomous Network Fabric architecture, claiming operators are already achieving automation rates higher than 90% and service delivery times of four hours or less. Each of these platforms ingests and processes vast volumes of network telemetry, creating demand for the kind of rigorous data quality validation that ETSI's framework provides.
The technical architecture choices underlying these AI deployments highlight why data quality standardization matters at the infrastructure level. Ericsson and Nokia are diverging on AI-RAN implementation, with Nokia running all Layer 1 functions on Nvidia GPUs via CUDA while Ericsson confines only the FEC function to the GPU, meaning each vendor's training data pipelines will differ substantially in format, scale, and provenance. Verizon has publicly called for industry-wide interoperability standards for agentic systems, and the TM Forum's Autonomous Networks L4/5 roadmap will need to incorporate agentic AI interoperability as a core requirement. ETSI TR 104 180's open-source validation system, developed with contributions from Sejong University, EGM, TTA, Daejeon University, and CNIT, offers a vendor-neutral starting point for ensuring that the datasets feeding these divergent architectures meet consistent quality thresholds before models are trained and deployed.
Read full article at telecomtv.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source