DoTCoM vision transformer achieves 81.6% accuracy on mobile hardware
Researchers from Dankook University and LIG Defense & Aerospace have developed DoTCoM, a lightweight hybrid vision transformer architecture designed for mobile hardware. The model utilizes a Quarter-Inverted Bottleneck and Co-optimization Bias to improve accuracy and latency for computer vision tasks like object detection and semantic segmentation on mobile CPUs.
Key Takeaways
- The flagship DoTCoM-L variant reached 81.6% Top-1 accuracy on ImageNet-1K using 11.1 million parameters.
- A compact 1.7-million parameter version achieved 74.0% accuracy, suitable for low-power mobile environments.
- Co-optimization Bias uses per-channel parameters to align feature distributions between hybrid network branches.
- Testing on Samsung mobile CPUs confirmed the model's efficiency for object detection and semantic segmentation tasks.
Why It Matters
The development of DoTCoM addresses the persistent trade-off between model intelligence and mobile battery life by optimizing the statistical language between hybrid neural layers. For the streaming and mobile ecosystems, this enables more sophisticated on-device video processing, such as real-time semantic segmentation and object tracking, without relying on cloud-based inference. By reducing the parameter count to as low as 1.7 million while maintaining high accuracy, the architecture allows manufacturers like Samsung to implement advanced computer vision features in mid-range hardware. Industry observers should watch for the integration of these lightweight transformers into next-generation augmented reality applications and computational photography suites where low-latency visual understanding is critical.
Additional Context
The push to run vision transformers directly on mobile silicon has intensified as chipmakers and research labs race to close the gap between cloud-grade accuracy and edge-device constraints. Qualcomm's Snapdragon 8 Elite, announced in late 2024, introduced on-device support for multimodal AI agents capable of processing visual inputs locally, signaling that flagship mobile processors are being designed with transformer workloads as a first-class requirement. Apple similarly expanded its Neural Engine throughput with the A18 Pro chip, which delivers 35 trillion operations per second for on-device machine learning tasks including real-time image segmentation. These hardware moves create the deployment target that architectures like DoTCoM are designed to exploit.
On the business side, Samsung has been integrating lightweight vision models into its Galaxy AI suite, which launched in January 2024 with features including generative photo editing and live translate powered by on-device inference. The company's Exynos processors have historically lagged Qualcomm in AI benchmark scores, making architectures that reduce parameter count while preserving accuracy particularly relevant for Samsung's mid-range lineup. LIG Defense & Aerospace, the co-developer of DoTCoM, is a South Korean defense contractor that supplies guidance and sensor systems to the Republic of Korea military, suggesting the architecture may also target embedded vision for defense applications where cloud connectivity is unavailable.
Technical benchmarks from competing lightweight transformer designs illustrate the competitive landscape DoTCoM enters. Meta's EfficientViT family, published in 2023, achieved 80.7% top-1 accuracy on ImageNet at 224x224 resolution with as few as 6.2 million parameters, establishing a baseline that DoTCoM's 81.6% at 1.7 million parameters surpasses on a parameter-efficiency basis. Meanwhile, Apple's MobileViT v2, which uses a separable self-attention mechanism to reduce computational overhead on mobile GPUs, demonstrated that attention-based architectures can match or exceed convolutional networks on segmentation tasks when properly optimized for mobile memory hierarchies. The DoTCoM contribution of resolving statistical mismatches between convolutional and attention pathways addresses a known failure mode that these earlier designs handled through separate normalization strategies rather than a unified co-optimization bias.
Read full article at bioengineer.org
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source