Re:Pixel's state-space model wins ECCV 2026 all-in-one image restoration challenge
The LoViF 2026 Challenge at ECCV 2026 evaluated 20 teams on unified image restoration tasks, including blur, low-light, haze, rain, and snow removal. The winning team, Re:Pixel, introduced a wavelet-driven state-space model that sets a new performance benchmark for practical, all-in-one image restoration.
Key Takeaways
- Winning team Re:Pixel achieved a composite score of 42.28 using a 1.6M-parameter model requiring 131.7G FLOPs.
- The ReMamba model replaces traditional task-specific architectures with a unified wavelet-driven state-space framework.
- Performance was evaluated on the FoundIR-LoViF benchmark, a dataset comprising 24,500 paired real-world images.
- Runner-up REDnoteMediaLab and third-place LucidWorld followed closely with scores of 41.81 and 41.40, respectively.
- The evaluation metrics combined PSNR, SSIM, and LPIPS to measure pixel-level fidelity and perceptual quality.
Why It Matters
The shift toward all-in-one restoration models marks a departure from fragmented, task-specific video processing pipelines. By using state-space models like Mamba, developers can achieve global receptive fields with linear complexity, offering a more efficient alternative to quadratic-cost Transformers for high-resolution video streams. This approach directly addresses the training-to-deployment domain gap by leveraging large-scale real-world datasets like FoundIR. As streaming platforms face increasing pressure to deliver high-quality HDR content under varying network and environmental conditions, these unified frameworks suggest a path toward lower inference costs without sacrificing perceptual fidelity. Watch for the integration of these state-space models into edge-side AI decoding chips by 2027.
Additional Context
The LoViF 2026 Challenge highlights the rapid adoption of State Space Models (SSMs) as a replacement for both Convolutional Neural Networks and Transformers in low-level vision tasks. Per Arxiv reporting in February 2024, the original Mamba architecture introduced selective scan mechanisms that allowed for long-range dependency modeling at linear complexity. This technological foundation was refined by researchers throughout 2025, with MambaIRv2 being accepted at CVPR 2025 (per ETH Zurich, June 2025) for its ability to outperform previous Transformer-based baselines like SwinIR while using significantly fewer parameters and lower computational overhead. The scalability of these models is supported by the emergence of million-scale real-world datasets. According to IEEE/CVF (September 2025), the FoundIR dataset was released to provide the diversity of real-world degradations — such as coupled non-uniform blur and complex illumination — that synthetic datasets typically fail to replicate. This move toward 'foundation models' for image restoration mirrored trends in large language models, where scaling laws suggest that increasing data volume and model parameters yields significant gains in generalization across diverse sub-tasks. Industry interest in these all-in-one frameworks is also driven by the needs of real-time applications. Per reports from the 2026 Embedded Vision Summit (May 2026), SSM-based models are particularly well-suited for edge deployments in autonomous systems and mobile video processing because they maintain constant inference speed. Recent developments in 2026, such as the introduction of FoundIR-v2, have focused on optimizing data mixture proportions to prevent 'catastrophic forgetting' when a single model is trained to simultaneously handle disparate tasks like dehazing and deraining.
Read full article at arxiv.org
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source