latent-visual-reasoning

Tag

Cards List
#latent-visual-reasoning

Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents

Hugging Face Daily Papers · 2026-06-04 Cached

The paper challenges the assumption that cosine alignment between supervised latents and visual targets improves accuracy in vision-language models, finding a strong negative correlation. It introduces PRISM diagnostics revealing that answers are decoded downstream from latents, not within them, and that the auxiliary loss reshapes the language model via shared parameters.

0 favorites 0 likes
#latent-visual-reasoning

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

Hugging Face Daily Papers · 2026-06-04 Cached

Introduces Future-L1, an interleaved latent visual reasoning framework that improves video event prediction by maintaining visual semantics in latent space. Achieves state-of-the-art results on FutureBench and TwiFF-Bench benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback