VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals
Summary
VCR is a self-supervised framework that learns robust representations from incomplete wearable signals using orthogonal tokenization and missing-aware mixture-of-experts, improving performance under modality missingness.
Similar Articles
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
InternVideo3 introduces Multimodal Contextual Reasoning (MCR) and efficient attention mechanisms to enhance long-horizon multimodal tasks, achieving strong results on video understanding benchmarks and demonstrating video agent capabilities.
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
This paper introduces a two-stage neuro-symbolic framework that uses weak supervision (as little as 1% labels) with a slot-based VAE to learn interpretable symbols for object-centric visual reasoning, outperforming foundation models in domain generalization.
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
ViSAGE is a multimodal agentic memory framework for long-form video understanding that builds self-correcting, entity-centric memories via cross-modal binding, bidirectional memory refinement, and multi-agent cross-verification, achieving 5.9% higher accuracy than baselines.
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
O-VAD introduces a training-free agentic framework for industrial video anomaly detection that tracks object state evolution over time and reasons over temporal trajectories to identify abnormal objects, outperforming existing VLM and VAD methods on three datasets.
MARCUS: Missing-Aware Region Representation with Contextual Urban Signals for Rent Prediction
The paper proposes MARCUS, a missing-aware region representation model that treats missing data as contextual urban signals for rent prediction, achieving state-of-the-art performance with significant error reduction on real-world datasets.