Tag
The AI City Challenge at ECCV 2026 showcased winners developing robust visual AI systems, with top solutions in video forecasting using NVIDIA Cosmos world foundation models for tasks like scene understanding and prediction.
ZipDepth is a lightweight zero-shot monocular depth estimation model that achieves the best accuracy-efficiency trade-off, running in real time on any device from mobile phones to server GPUs, and has been accepted at ECCV 2026.
PriorEye introduces geospatial visual priors for end-to-end autonomous driving, enhancing anticipatory behavior and robustness through a dual-memory architecture, as presented at ECCV 2026.
360 AI Research Institute's MoSA (Motion-Grounded Segment Anything), accepted at ECCV 2026, trains the Segment Anything Model to segment objects using motion cues from unlabeled video, generating millions of pseudo-labels without manual annotation.
Announcing RT-DocLayout, the world's first layout analysis model capable of pixel-level multi-point polygon boxes for real-world documents, accepted to ECCV 2026. The paper argues that layout analysis remains critical despite end-to-end models, and presents an efficient 33M parameter model running at 132.1 FPS.