Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Hugging Face Daily Papers Papers

Summary

This paper introduces a preference-guided adaptation framework for open-vocabulary semantic segmentation that leverages prompt disagreement to generate binary preferences, eliminating the need for dense pixel-level annotations and improving performance in specialized domains.

Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences. Our code is available at https://github.com/blue-531/pref-ovss.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:14 AM

Paper page - Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Source: https://huggingface.co/papers/2609.34528

Abstract

Open-vocabularysemanticsegmentation(OVSS)enablespixel-levelpredictionoverarbitrarytext-specifiedvocabulariesandhasshownstronggeneralizationoncommonbenchmarks.However,OVSSperformanceoftendegradesinspecializeddomainssuchasmedicalimaging,remotesensing,andindustrialinspection,wheredensepixel-levelmasksforadaptationarecostlytoobtainandrequiredomain-specificexpertise.Weproposeapreference-guidedadaptationframeworkthatreplacesdensemasksupervisionwithbinarypreferences.Weobservethatdifferentprompttemplatesproducesystematicallydifferentsegmentationsforthesameimage,aphenomenonwecallpromptdisagreement,andwerepurposeitasabuilt-insourceofpreferencesupervision.Buildingonthis,weminelocalizedpreferencequeriesfromregionsofhighcross-templateuncertainty,andadapttheOVSSmodelwithRegion-LocalizedPreferenceOptimization(RLPO)togetherwithconsistencyregularizationthatstabilizesupdatesoutsidethequeriedregion.AcrossextensiveexperimentsontheMESSbenchmark,theproposedmethodachievesconsistentgainsacrossdiverseOVSSbackboneswithoutanypixel-levelannotation,andremainseffectiveundernoisypreferences.Ourcodeisavailableathttps://github.com/blue-531/pref-ovss.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2609\.34528

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34528 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34528 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34528 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Diffusion Model as a Generalist Segmentation Learner

Hugging Face Daily Papers

This paper introduces DiGSeg, a framework that repurposes pretrained diffusion models for state-of-the-art semantic and open-vocabulary segmentation by leveraging latent space conditioning and text-guided alignment.