open-vocabulary

Tag

Cards List
#open-vocabulary

Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces a preference-guided adaptation framework for open-vocabulary semantic segmentation that leverages prompt disagreement to generate binary preferences, eliminating the need for dense pixel-level annotations and improving performance in specialized domains.

0 favorites 0 likes
#open-vocabulary

Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper proposes an Affect-Prototype-Conditioned Fusion (APCF) framework for open-vocabulary multimodal emotion recognition with incomplete modalities, using an affect-prototype library to guide feature fusion and an LLM decoder for generating natural language emotion labels.

0 favorites 0 likes
#open-vocabulary

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs

Hugging Face Daily Papers ↗ · 2026-09-11 Cached

RelateAnything is a lightweight, real-time open-vocabulary relation prediction model that accepts arbitrary predicate vocabularies and region sources, trained on a large geometrically verified dataset and evaluated on new cross-dataset benchmarks, showing significant performance gains over comparable methods.

0 favorites 0 likes
#open-vocabulary

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper introduces PromptKWS, a novel prompt-guided open-vocabulary keyword spotting framework that uses prompt embeddings and cross-attention to improve accuracy, achieving over 10% improvement in wakeup rate and over 15% in accuracy compared to baseline systems.

0 favorites 0 likes
#open-vocabulary

Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper presents Tactus, an open-vocabulary tactile recognition model that maps low-cost pressure-array data to text embeddings, matching or exceeding a supervised closed-set CNN baseline on the STAG benchmark with only 187 training recordings and no classifier head.

0 favorites 0 likes
#open-vocabulary

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Hugging Face Daily Papers ↗ · 2026-07-29 Cached

Introduces OVEarth-Bench, a benchmark for open-vocabulary Earth observation that broadens category coverage and query diversity, revealing that current methods remain limited and MLLM-based approaches perform best.

0 favorites 0 likes
#open-vocabulary

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper introduces MER-R1, a reinforcement learning framework that synergizes fast and slow thinking for multimodal emotion recognition. It achieves state-of-the-art performance by jointly optimizing recall and precision through dual-objective disentanglement and slow-fast confidence calibration.

0 favorites 0 likes
#open-vocabulary

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

arXiv cs.AI ↗ · 2026-06-08 Cached

Proposes a hierarchical semantic-constrained heterogeneous graph model for open-vocabulary audio-visual event localization, addressing cross-modal consistency at multiple temporal scales and hierarchical semantic constraints between segment and video levels. Achieves state-of-the-art results on OV-AVEL benchmark.

0 favorites 0 likes
#open-vocabulary

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Hugging Face Daily Papers ↗ · 2026-06-05 Cached

VoLoAgent integrates vision-language models with robot capabilities for open-vocabulary long-horizon manipulation tasks, introducing a physical orchestrator that plans, monitors, and recovers using interruptible tools, and a benchmark called RoboVoLo for evaluation.

0 favorites 0 likes
#open-vocabulary

Diffusion Model as a Generalist Segmentation Learner

Hugging Face Daily Papers ↗ · 2026-04-27 Cached

This paper introduces DiGSeg, a framework that repurposes pretrained diffusion models for state-of-the-art semantic and open-vocabulary segmentation by leveraging latent space conditioning and text-guided alignment.

0 favorites 0 likes
#open-vocabulary

adirik/grounding-dino

Replicate Explore ↗ · 2026-05-08 Cached

Grounding DINO is an open-vocabulary object detection model that can detect arbitrary objects based on text descriptions, now available on Replicate.

0 favorites 0 likes
← Back to home

Submit Feedback