Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
Summary
This paper introduces a preference-guided adaptation framework for open-vocabulary semantic segmentation that leverages prompt disagreement to generate binary preferences, eliminating the need for dense pixel-level annotations and improving performance in specialized domains.
View Cached Full Text
Cached at: 09/30/26, 04:14 AM
Paper page - Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
Source: https://huggingface.co/papers/2609.34528
Abstract
Open-vocabularysemanticsegmentation(OVSS)enablespixel-levelpredictionoverarbitrarytext-specifiedvocabulariesandhasshownstronggeneralizationoncommonbenchmarks.However,OVSSperformanceoftendegradesinspecializeddomainssuchasmedicalimaging,remotesensing,andindustrialinspection,wheredensepixel-levelmasksforadaptationarecostlytoobtainandrequiredomain-specificexpertise.Weproposeapreference-guidedadaptationframeworkthatreplacesdensemasksupervisionwithbinarypreferences.Weobservethatdifferentprompttemplatesproducesystematicallydifferentsegmentationsforthesameimage,aphenomenonwecallpromptdisagreement,andwerepurposeitasabuilt-insourceofpreferencesupervision.Buildingonthis,weminelocalizedpreferencequeriesfromregionsofhighcross-templateuncertainty,andadapttheOVSSmodelwithRegion-LocalizedPreferenceOptimization(RLPO)togetherwithconsistencyregularizationthatstabilizesupdatesoutsidethequeriedregion.AcrossextensiveexperimentsontheMESSbenchmark,theproposedmethodachievesconsistentgainsacrossdiverseOVSSbackboneswithoutanypixel-levelannotation,andremainseffectiveundernoisypreferences.Ourcodeisavailableathttps://github.com/blue-531/pref-ovss.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2609\.34528
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34528 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34528 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34528 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation
PixCon proposes a clean-positive pixel-contrastive framework for semi-supervised semantic segmentation that guarantees contamination-free positive sets via per-class memory banks, improving accuracy over existing methods on benchmarks like Pascal VOC, Cityscapes, and ADE20K.
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
This paper introduces SemanticSeg, a large-scale dataset for semantic segmentation of long texts, and block distillation, a training framework that enables block attention models to approach full-attention performance, improving KV cache reuse in RAG and long-context scenarios.
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
This paper proposes DASO, a tree-aware post-training method for generative recommendation that addresses difficulty mismatch in GRPO by profiling rollout groups and reallocating based on prefix-match depth, improving performance on public benchmarks.
You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models
This paper investigates prompt ensembling and domain shift in agrifood vision-language models, introducing Prompt-based Inconsistency Detection (PID) to enhance reliability by using prompt disagreement as an uncertainty proxy.
Diffusion Model as a Generalist Segmentation Learner
This paper introduces DiGSeg, a framework that repurposes pretrained diffusion models for state-of-the-art semantic and open-vocabulary segmentation by leveraging latent space conditioning and text-guided alignment.