vision-foundation-models

Tag

Cards List
#vision-foundation-models

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Hugging Face Daily Papers · 4d ago Cached

Presents Self-Geometry, a plug-and-play test-time adaptation pipeline that enforces explicit multi-view geometric constraints using 2D pixel correspondences to improve geometrically consistent 3D vision foundation models.

0 favorites 0 likes
#vision-foundation-models

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

Hugging Face Daily Papers · 2026-06-22 Cached

RaysUp is an ultra-lightweight, task-agnostic feature upsampling framework that uses geometry-aware ray domain techniques to reconstruct high-resolution features from low-resolution VFM outputs, achieving state-of-the-art performance with 84% fewer parameters than prior work and 7x faster inference.

0 favorites 0 likes
#vision-foundation-models

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

Hugging Face Daily Papers · 2026-06-09 Cached

IDEAL proposes an in-depth alignment framework for discrete representation autoencoding, jointly aligning quantized tokens with shallow and deep VFM features to achieve superior reconstruction and generation performance.

0 favorites 0 likes
#vision-foundation-models

Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models

arXiv cs.AI · 2026-06-08 Cached

Proposes an attention-guided encoder-decoder for longitudinal medical visual question answering, using a frozen DINO-based mask generator and auxiliary losses to improve consistency and interpretability, achieving strong results on the Medical-Diff-VQA benchmark.

0 favorites 0 likes
#vision-foundation-models

SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models

Hugging Face Daily Papers · 2026-05-29 Cached

SOCO benchmark evaluates structured object understanding in vision models through consistent part-level annotations and keypoint descriptions, revealing gaps between language-grounded localization and visual correspondence while demonstrating strong prediction of downstream task performance.

0 favorites 0 likes
#vision-foundation-models

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

Hugging Face Daily Papers · 2026-04-13 Cached

Re2Pix is a hierarchical video prediction framework that improves future video generation by first predicting semantic representations using frozen vision foundation models, then conditioning a latent diffusion model on these predictions to generate photorealistic frames. The approach addresses train-test mismatches through nested dropout and mixed supervision strategies, achieving improved temporal semantic consistency and perceptual quality on autonomous driving benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback