Tag
This paper introduces SAPO, a segment-level automatic prompt optimization method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on weak and strong examples. Evaluated across several benchmarks, SAPO outperforms strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO on GPT-3.5-Turbo and GPT-4o-mini.
This paper introduces ConCor-1, a grounding model that treats vision-language grounding as bidirectional concept correspondence, jointly recovering text spans, image segments, and cross-modal matches without prespecified phrases. It unifies phrase grounding, referring expression grounding, and open-vocabulary detection, achieving significant F1 improvements on long-caption and zero-shot LVIS benchmarks.
iFAN is a training framework that improves mask transformers for segmentation by aligning query ranking with mask quality and distilling intermediate predictions to the final layer, yielding consistent gains across benchmarks.
This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.
GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.
This paper addresses the challenge of maintaining boundary faithfulness in backbone models when processing subsequent video frames beyond the first.
This paper presents a systematic comparison of two geospatial foundation models, TerraMind and THOR, developed under ESA's φ-lab, analyzing how architectural choices like patch size and decoder type affect performance across ten use cases in Earth observation tasks.
The tweet praises SAM3.1 for improved segmentation with faint borders and better tail coverage, leading to more accurate length measurements.
This paper introduces VIP-SAM for instance-level garment segmentation and CtrlVTON, a controllable virtual try-on framework that treats try-on as an image editing problem, allowing precise control over garment layout, style, and placement. Both methods achieve state-of-the-art results on their respective tasks.
REBASE is a training-free framework that suppresses spurious contextual correspondences in in-context segmentation by projecting features onto the orthogonal complement of a low-rank background subspace, achieving state-of-the-art results among training-free methods on several datasets.
This paper introduces SAM-MT, an extension of SAM2 for real-time interactive multi-target video segmentation, achieving high FPS independent of target count.
This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.
This paper presents CALHippo, a framework for 3D mapping of neurons and glial cells in the human hippocampus using state-of-the-art segmentation and density estimation models.
This paper presents MAOAM, a unified vision-language model framework that enables precise object and material selection through text or click interactions for interactive image editing. It introduces a scalable data generation pipeline and shows emergent improvement when combining text and clicks at inference.
Group Prompting introduces a training-free framework for cell instance segmentation that requires only one click per cell type, using the Segment Anything Model's feature space to recursively expand prompts, achieving competitive performance without training.
InstructSAM presents a unified framework for multi-instance segmentation using instruction-driven queries that bridge vision-language models and SAM3, achieving strong results across complex benchmarks.
Introduces Semantic Generative Tuning (SGT), a paradigm that uses image segmentation as a generative proxy to align visual understanding and generation in unified multimodal models, improving both comprehension and fidelity.
AuralSAM2 integrates audio into SAM2 via an AuralFuser module that generates sparse and dense prompts from audio-visual features, enhancing cross-modal segmentation while maintaining interactive efficiency.
Introduces CAFE, a benchmark for evaluating whether promptable segmentation models truly understand concepts by using counterfactual attribute manipulation, revealing that accurate mask prediction does not guarantee faithful semantic grounding.
TwinTrack is a post-hoc calibration framework for pancreatic cancer segmentation that aligns ensemble model probabilities with the empirical mean human response across multiple annotators, improving interpretability and calibration metrics on multi-rater benchmarks.