Tag
This paper presents Tactus, an open-vocabulary tactile recognition model that maps low-cost pressure-array data to text embeddings, matching or exceeding a supervised closed-set CNN baseline on the STAG benchmark with only 187 training recordings and no classifier head.
Introduces OVEarth-Bench, a benchmark for open-vocabulary Earth observation that broadens category coverage and query diversity, revealing that current methods remain limited and MLLM-based approaches perform best.
This paper introduces MER-R1, a reinforcement learning framework that synergizes fast and slow thinking for multimodal emotion recognition. It achieves state-of-the-art performance by jointly optimizing recall and precision through dual-objective disentanglement and slow-fast confidence calibration.
Proposes a hierarchical semantic-constrained heterogeneous graph model for open-vocabulary audio-visual event localization, addressing cross-modal consistency at multiple temporal scales and hierarchical semantic constraints between segment and video levels. Achieves state-of-the-art results on OV-AVEL benchmark.
VoLoAgent integrates vision-language models with robot capabilities for open-vocabulary long-horizon manipulation tasks, introducing a physical orchestrator that plans, monitors, and recovers using interruptible tools, and a benchmark called RoboVoLo for evaluation.
This paper introduces DiGSeg, a framework that repurposes pretrained diffusion models for state-of-the-art semantic and open-vocabulary segmentation by leveraging latent space conditioning and text-guided alignment.
Grounding DINO is an open-vocabulary object detection model that can detect arbitrary objects based on text descriptions, now available on Replicate.