segmentation

Tag

Cards List
#segmentation

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

arXiv cs.AI · 4d ago Cached

This paper introduces SAPO, a segment-level automatic prompt optimization method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on weak and strong examples. Evaluated across several benchmarks, SAPO outperforms strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO on GPT-3.5-Turbo and GPT-4o-mini.

0 favorites 0 likes
#segmentation

Vision-Language Grounding as Bidirectional Concept Correspondence

Hugging Face Daily Papers · 2026-08-08 Cached

This paper introduces ConCor-1, a grounding model that treats vision-language grounding as bidirectional concept correspondence, jointly recovering text spans, image segments, and cross-modal matches without prespecified phrases. It unifies phrase grounding, referring expression grounding, and open-vocabulary detection, achieving significant F1 improvements on long-caption and zero-shot LVIS benchmarks.

0 favorites 0 likes
#segmentation

iFAN: Inference-Aware Learning for Plain Mask Transformers

Hugging Face Daily Papers · 2026-08-07 Cached

iFAN is a training framework that improves mask transformers for segmentation by aligning query ranking with mask quality and distilling intermediate predictions to the final layer, yielding consistent gains across benchmarks.

0 favorites 0 likes
#segmentation

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

Hugging Face Daily Papers · 2026-08-03 Cached

This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.

0 favorites 0 likes
#segmentation

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

Hugging Face Daily Papers · 2026-07-27 Cached

GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.

0 favorites 0 likes
#segmentation

A boundary faithful backbone still has to survive frame two

Reddit r/ArtificialInteligence · 2026-07-26

This paper addresses the challenge of maintaining boundary faithfulness in backbone models when processing subsequent video frames beyond the first.

0 favorites 0 likes
#segmentation

Now We Know? A Systematic Comparison of TerraMind and THOR

arXiv cs.LG · 2026-07-22 Cached

This paper presents a systematic comparison of two geospatial foundation models, TerraMind and THOR, developed under ESA's φ-lab, analyzing how architectural choices like patch size and decoder type affect performance across ten use cases in Earth observation tasks.

0 favorites 0 likes
#segmentation

@lillyguisnet: Not quite done with the comparison, but I already love the small extra segmentation improvements from SAM3.1! It seems …

X AI KOLs Following · 2026-07-15 Cached

The tweet praises SAM3.1 for improved segmentation with faint borders and better tail coverage, leading to more accurate length measurements.

0 favorites 0 likes
#segmentation

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

Hugging Face Daily Papers · 2026-07-10 Cached

This paper introduces VIP-SAM for instance-level garment segmentation and CtrlVTON, a controllable virtual try-on framework that treats try-on as an image editing problem, allowing precise control over garment layout, style, and placement. Both methods achieve state-of-the-art results on their respective tasks.

0 favorites 0 likes
#segmentation

REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

Hugging Face Daily Papers · 2026-07-10 Cached

REBASE is a training-free framework that suppresses spurious contextual correspondences in in-context segmentation by projecting features onto the orthogonal complement of a low-rank background subspace, achieving state-of-the-art results among training-free methods on several datasets.

0 favorites 0 likes
#segmentation

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

Hugging Face Daily Papers · 2026-07-09 Cached

This paper introduces SAM-MT, an extension of SAM2 for real-time interactive multi-target video segmentation, achieving high FPS independent of target count.

0 favorites 0 likes
#segmentation

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

arXiv cs.CL · 2026-06-26 Cached

This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.

0 favorites 0 likes
#segmentation

CALHippo - Mapping neurons and glial cells in the human brain hippocampus in 3D using SOTA segmentation and density estimation models [R]

Reddit r/MachineLearning · 2026-06-25

This paper presents CALHippo, a framework for 3D mapping of neurons and glial cells in the human hippocampus using state-of-the-art segmentation and density estimation models.

0 favorites 0 likes
#segmentation

MAOAM: Unified Object and Material Selection with Vision-Language Models

Hugging Face Daily Papers · 2026-06-02 Cached

This paper presents MAOAM, a unified vision-language model framework that enables precise object and material selection through text or click interactions for interactive image editing. It introduces a scalable data generation pipeline and shows emergent improvement when combining text and clicks at inference.

0 favorites 0 likes
#segmentation

One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation

Hugging Face Daily Papers · 2026-05-28 Cached

Group Prompting introduces a training-free framework for cell instance segmentation that requires only one click per cell type, using the Segment Anything Model's feature space to recursively expand prompts, achieving competitive performance without training.

0 favorites 0 likes
#segmentation

InstructSAM: Segment Any Instance with Any Instructions

Hugging Face Daily Papers · 2026-05-25 Cached

InstructSAM presents a unified framework for multi-instance segmentation using instruction-driven queries that bridge vision-language models and SAM3, achieving strong results across complex benchmarks.

0 favorites 0 likes
#segmentation

Semantic Generative Tuning for Unified Multimodal Models

Hugging Face Daily Papers · 2026-05-18 Cached

Introduces Semantic Generative Tuning (SGT), a paradigm that uses image segmentation as a generative proxy to align visual understanding and generation in unified multimodal models, improving both comprehension and fidelity.

0 favorites 0 likes
#segmentation

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting

Hugging Face Daily Papers · 2026-05-14 Cached

AuralSAM2 integrates audio into SAM2 via an AuralFuser module that generates sparse and dense prompts from audio-visual features, enhancing cross-modal segmentation while maintaining interactive efficiency.

0 favorites 0 likes
#segmentation

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Hugging Face Daily Papers · 2026-05-10 Cached

Introduces CAFE, a benchmark for evaluating whether promptable segmentation models truly understand concepts by using counterfactual attribute manipulation, revealing that accurate mask prediction does not guarantee faithful semantic grounding.

0 favorites 0 likes
#segmentation

TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation

Hugging Face Daily Papers · 2026-04-17 Cached

TwinTrack is a post-hoc calibration framework for pancreatic cancer segmentation that aligns ensemble model probabilities with the empirical mean human response across multiple annotators, improving interpretability and calibration metrics on multi-rater benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback