StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoTechnical DevelopmentJuly 27, 2026

KOREATECH and ETRI optimize Qwen3-VL for 25W edge video monitoring

KOREATECH and ETRI optimize Qwen3-VL for 25W edge video monitoring
MDPI

Researchers from KOREATECH and ETRI have developed a real-time video anomaly detection framework that optimizes the Qwen3-VL-2B model for edge hardware. By utilizing 4-bit weight quantization and a novel prompt optimization technique, the system achieves sub-second latency on NVIDIA Jetson Orin NX hardware within a 25W power envelope.

Key Takeaways

  • Achieved a 7.4x speedup and 55% memory reduction using INT4 Activation-aware Weight Quantization (AWQ).
  • Developed a C++ runtime stack via TensorRT-LLM that enables sub-second latency on NVIDIA Jetson Orin NX (16GB).
  • Automated prompt optimization increased zero-shot accuracy to 76.39% AUC, surpassing manually tuned prompts and those generated by GPT-4 or Gemini-Pro.
  • Validated single-frame input as the optimal performance-to-latency ratio compared to five- or eight-frame windows.

Why It Matters

This development moves Vision-Language Models (VLMs) from centralized data centers to localized, power-constrained streaming endpoints. By achieving sub-second processing on a 25W budget, the framework enables complex visual reasoning in privacy-sensitive sectors like public security and industrial monitoring without high-bandwidth cloud uplinks. While its 76.39% accuracy trails server-side benchmarks, the move toward interpretable, on-device AI signals a shift for hardware vendors prioritizing the 'smart edge' over cloud-only pipelines. Watch for the integration of this 4-bit optimization into upcoming NVIDIA JetPack releases for wider industrial adoption.

Additional Context

The research coincides with a broader push by South Korea to establish itself as a dominant force in specialized AI hardware and software. In late 2025, the Ministry of Science and ICT (MSIT) unveiled a national 2026 business plan aimed at positioning the country as one of the world's top three AI powers. Per Korea.net (December 2025), this strategy includes a KRW 3 trillion fund for AI startups and the procurement of 37,000 GPUs to bolster domestic R&D in sectors such as defense, manufacturing, and smart city infrastructure. ETRI remains a central figure in this ecosystem, leading governmental initiatives to standardize open-source AI and security governance, according to Digital Today (July 2026). Concurrently, the infrastructure for edge LLMs has matured rapidly. Performance benchmarks from April 2026 via IOT Digital Twin PLM indicate that the NVIDIA Jetson Orin platform has transitioned from academic testing to practical deployment for offline, privacy-first applications. TensorRT-LLM 0.13 and newer kernels have provided 30-70% faster token throughput than standard portable runtimes like llama.cpp on the same hardware. This software optimization is critical as the industry awaits the release of Llama 4 and Qwen3.5 variants, which are expected to offer native vision-language support for 8GB and 16GB edge systems by late 2026. Technically, the use of Qwen3-VL-2B reflects the growing utility of compact, multi-billion-parameter models. Originally open-sourced in late 2025, the Qwen3-VL series features deep-stack vision encoders and text-timestamp alignment, enabling precise event localization within streaming video. According to LM Studio (November 2025), the 2B-parameter version was specifically architected for edge-to-cloud flexibility, providing a path for researchers to apply high-density spatial reasoning to localized hardware like the Jetson Orin NX used in the KOREATECH-ETRI study.


Read full article at mdpi.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

Arxiv: Framework cuts video bandwidth requirements by 99% using generative AI
ayushchat: Whisper runs locally on Apple Silicon with no network access
NVIDIA Technical Blog: NVIDIA TensorRT converts FP8 checkpoints to high-efficiency video inference engines
Qiang Zhang: DeltaToken cuts video tokens from 180K to under 1,000

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 3 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 3 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →