Group Preference Collapse in Personalized Multimodal Large Language Models

arXiv cs.AI Papers

Summary

This paper identifies group preference collapse in personalized multimodal large language models and proposes PrefMoE, a preference-centric framework that separates profile information from preferences to improve personalization and reduce collapse.

arXiv:2607.22603v1 Announce Type: new Abstract: Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:26 AM

# Group Preference Collapse in Personalized Multimodal Large Language Models
Source: [https://arxiv.org/html/2607.22603](https://arxiv.org/html/2607.22603)
Fan Lyu, Wenqi Zhang, Joost van de Weijer Computer Vision Center, Universitat Autònoma de Barcelona, Barcelona, Spain, 08193 fanlyu@cvc\.uab\.cat, wenqi\.zhang@autonoma\.cat, joost@cvc\.uab\.es

Project Page:[https://prefmoe\.github\.io/](https://prefmoe.github.io/)

###### Abstract

Personalized multimodal large language models \(MLLMs\) aim to generate user\-specific responses, but existing methods mainly rely on profile\-level information and overlook diverse user preferences\. We identify group preference collapse, where multi\-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population\-level choices due to suppressed preference signals and unreliable preference use during generation\. We propose PrefMoE, a preference\-centric framework that separates stable profile information from preference\-related representations\. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance\-aware learning, counterfactual pseudo\-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths\. Experiments across multiple MLLM backbones show that PrefMoE improves preference\-sensitive personalization while substantially reducing preference collapse\.

## 1Introduction

Multimodal Large Language Models \(MLLMs\)Alayracet al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib31)\); Liet al\.\([2023](https://arxiv.org/html/2607.22603#bib.bib32)\); Daiet al\.\([2023](https://arxiv.org/html/2607.22603#bib.bib33)\); Zhuet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib34)\)have shown strong performance across a wide range of multimodal tasks, including visual question answeringAntolet al\.\([2015](https://arxiv.org/html/2607.22603#bib.bib46)\); Fanget al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib49)\); Wanget al\.\([2025b](https://arxiv.org/html/2607.22603#bib.bib50)\), visual groundingTanget al\.\([2026](https://arxiv.org/html/2607.22603#bib.bib3)\), and multimodal reasoningHuanget al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib51)\)\. Building on this progress, personalized MLLMs have recently attracted increasing attentionAlalufet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib1)\); Cohenet al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib38)\); Yehet al\.\([2023](https://arxiv.org/html/2607.22603#bib.bib39)\); Nguyenet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib40)\), aiming to generate responses tailored to a target user by conditioning on user\-related informationSeifiet al\.\([2026](https://arxiv.org/html/2607.22603#bib.bib41)\)\. Existing efforts mainly incorporate user\-related profile signalsAlalufet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib1)\); Piet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib6)\); Phamet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib7)\); Seifiet al\.\([2026](https://arxiv.org/html/2607.22603#bib.bib41)\), such as portraits, demographic cues, or textual descriptions, to ground responses in user identity and attributesPiet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib6)\)\.However, these methods primarily model who the user is, while paying much less attention to what the user prefers\.

Preference personalization differs from profile\-based personalization because preferences describe context\-dependent user choices rather than relatively stable identity cuesRendleet al\.\([2009](https://arxiv.org/html/2607.22603#bib.bib42)\)\. Users with similar profiles may still prefer different clothing styles, travel plans, food options, or lifestyle decisions under the same visual query\. Therefore, personalization should not be profile modeling alone, but should also preserve fine\-grained preference diversity and use the relevant preference cues during generation\. Despite its practical importance, preference\-level personalization remains largely underexplored\.

![Refer to caption](https://arxiv.org/html/2607.22603v1/x1.png)Figure 1:Illustration of*group preference collapse*\. Although users have diverse preferences, existing personalized MLLMs tend to predict the dominant preference across users \(“Yoga”\) rather than reflecting the target user’s specific preference\. Bars indicate the percentage of users from each ground\-truth preference group whose Top\-1 preference is the dominant preference \(“Yoga”\)\.When preference diversity is learned across many users, personalized MLLMs may suffer from a critical failure mode, which we term*group preference collapse*\. We find that existing MLLMs and common fine\-tuning strategies struggle with highly imbalanced user preferences\. When answering user\-specific questions, models often ignore individual preference signals and instead follow dominant population\-level preferences\. As illustrated in Fig\.[1](https://arxiv.org/html/2607.22603#S1.F1), even when user preferences are explicitly injected as textual conditions and the models are fine\-tuned, generated responses still collapse toward similar outputs across users and fail to reflect the target user’s preference\.

We attribute group preference collapse to two key failures\. First, individualized preference signals can be suppressed during user representation learning, as sparse and low\-frequency preferences are easily absorbed by dominant population\-level patterns\. Second, preference use during generation can be unreliable\. Even when preference information is provided, the model may fail to select the preference cue relevant to the current image\-question pair and instead follow generic or population\-dominant reasoning patterns\. Together, these failures make responses drift toward broadly plausible group\-level preferences rather than fine\-grained user\-specific choices\.

To address these failures, we proposePrefMoE, a preference\-centric personalization framework for MLLMs\. PrefMoE first learns a structured user state that separates stable profile cues from preference\-related information\. It decomposes preference representations into shared population\-level prototypes and personalized residuals, allowing dominant group tendencies and user\-specific deviations to be modeled separately\. To prevent minority preferences from being absorbed by shared patterns, we preserve individualized residuals through imbalance\-aware contrastive learning, counterfactual preference augmentation, and decorrelation\. In addition, PrefMoE introduces a factorized hierarchical MoE router to ensure that learned preferences are used during generation: preference residuals activate preference\-specific LoRA experts, profile cues activate a profile\-aware adaptation branch, and a query\-dependent router adaptively fuses these updates for the current visual question\. In this way, PrefMoE not only preserves fine\-grained preference diversity in user representations, but also turns structured user states into factorized computation paths for preference\-sensitive response generation\. Our work makes three main contributions:

1. \(1\)We identify group preference collapse as a critical failure mode in multi\-user personalized MLLMs, where dominant population\-level preferences suppress user\-specific preference signals even when explicit user information is available\.
2. \(2\)We propose PrefMoE, a preference\-centric framework that combines residual\-based preference factor learning with factorized hierarchical routing to preserve individualized preference offsets and activate profile\- and preference\-aware adaptation paths during generation\.
3. \(3\)Extensive experiments across multiple MLLM backbones show that PrefMoE consistently improves preference\-sensitive personalization and reduces preference collapse\. On LLaVA\-1\.5\-7B, PrefMoE improves preference accuracy from 44\.13% to 67\.33% and reduces collapse from 34\.25% to 12\.33% over full fine\-tuning, while reaching 78\.93% preference accuracy and 10\.96% collapse\.

## 2Related Work

MLLM Personalization\. Early personalization studies in language models mainly adapted generated text to personas, speaker roles, or dialogue histories, focusing on stylistic and contextual consistency in text\-only interactionsLiet al\.\([2016](https://arxiv.org/html/2607.22603#bib.bib10)\); Zhanget al\.\([2018](https://arxiv.org/html/2607.22603#bib.bib11)\); Chiet al\.\([2017](https://arxiv.org/html/2607.22603#bib.bib12)\)\. With the rise of MLLMsLi and Lyu \([2025](https://arxiv.org/html/2607.22603#bib.bib52)\); Tanet al\.\([2026](https://arxiv.org/html/2607.22603#bib.bib54)\); Wanget al\.\([2025a](https://arxiv.org/html/2607.22603#bib.bib53)\), personalization has been extended to multimodal inputs, where most methods focus on associating user\-specific or instance\-specific concepts with visual entities\. For example, existing approaches bind personalized concepts to visual representations through learnable embeddings, special concept tokens, or soft promptsNguyenet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib5)\); Anet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib8)\); Seifiet al\.\([2026](https://arxiv.org/html/2607.22603#bib.bib41)\)\. Recent methods further improve scalability and flexibility through multimodal in\-context conditioning, reference\-based identification, and online concept learningPiet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib6)\); Phamet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib7)\); Baiet al\.\([2025a](https://arxiv.org/html/2607.22603#bib.bib9)\)\. Despite these advances, current MLLM personalization remains largely centered on descriptive or profile\-like information, such as who the user is or which personalized entity is being referred to\. Preference Learning\. Preference learning aims to infer preferred outputs from pairwise comparisonsBradley and Terry \([1952](https://arxiv.org/html/2607.22603#bib.bib13)\), rankings, or choice\-based feedback and has been widely used for recommendationsRendleet al\.\([2009](https://arxiv.org/html/2607.22603#bib.bib42)\)and decision makingChristianoet al\.\([2017](https://arxiv.org/html/2607.22603#bib.bib14)\)\. Beyond classical ranking settings, preference feedback has also been used to train agentsChristianoet al\.\([2017](https://arxiv.org/html/2607.22603#bib.bib14)\)and generative models from human judgmentsZiegleret al\.\([2019](https://arxiv.org/html/2607.22603#bib.bib15)\)\. It has become a central tool for aligning large language models, from RLHFOuyanget al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib16)\); Baiet al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib17)\)to direct or RL\-free optimization methods such as DPORafailovet al\.\([2023](https://arxiv.org/html/2607.22603#bib.bib18)\), contrastive preference learningHejnaet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib19)\), and KTOEthayarajhet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib20)\)\. Most existing studies focus on exploiting preference signals to optimize model outputs toward shared or aggregated human preferencesOuyanget al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib16)\); Baiet al\.\([2022](https://arxiv.org/html/2607.22603#bib.bib17)\); Rafailovet al\.\([2023](https://arxiv.org/html/2607.22603#bib.bib18)\); Geet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib21)\)\. However, this objective differs from personalized preference modeling, where multiple users with heterogeneous preferences must be modeled jointly\. In such settings, aggregating preference signals can suppress minority or user\-specific preferences and make different users receive increasingly similar responses\. Our work studies this failure mode in personalized MLLMs and explicitly models preference diversity to reduce group preference collapse\.

## 3Preference Diversity and Group Preference Collapse

We study personalized MLLMs, where users’ information is available during training, and the model is expected to generate user\-specific responses for these users at inference time\. For each useruiu\_\{i\}, we denote the available personalization information as𝒳i=\{Ii,Di,𝒫i\}\\mathcal\{X\}\_\{i\}=\\\{I\_\{i\},D\_\{i\},\\mathcal\{P\}\_\{i\}\\\}, whereIiI\_\{i\}andDiD\_\{i\}denote user\-related image information and textual profile description, and𝒫i\\mathcal\{P\}\_\{i\}denotes the user preference set over different preference facets\. Here, preference facets refer to semantic dimensions of user preferences, such as fashion, lifestyle, travel, shopping, and entertainment\. The profile informationIiI\_\{i\}andDiD\_\{i\}captures relatively stable identity\-related cues, while𝒫i\\mathcal\{P\}\_\{i\}captures user\-specific tendencies across these facets and answers what the user prefers\.

We take visual question answering as an example downstream task\. Given an input imageVnV\_\{n\}, a questionQnQ\_\{n\}, for the target useruiu\_\{i\}, the goal is to learn a personalized MLLM that is able to predict the user\-specific answerAn,iA\_\{n,i\}:

ℒn,ivqa=CE​\(MLLM​\(Vn,Qn;𝒳i\),An,i\),\\mathcal\{L\}^\{\\mathrm\{vqa\}\}\_\{n,i\}=\\mathrm\{CE\}\\big\(\\mathrm\{MLLM\}\(V\_\{n\},Q\_\{n\};\\mathcal\{X\}\_\{i\}\),A\_\{n,i\}\\big\),\(1\)whereCE\\mathrm\{CE\}denotes the answer cross\-entropy prediction loss\. This formulation defines the desired user\-conditioned prediction, but does not prescribe how the user condition should be implemented\.

A straightforward implementation is to encode the raw user information together with the visual input and the question, and concatenate them into a single multimodal token sequence\[ϕv​\(Vn\);ϕq​\(Qn\);ϕu​\(𝒳i\)\]\[\\phi\_\{v\}\(V\_\{n\}\);\\phi\_\{q\}\(Q\_\{n\}\);\\phi\_\{u\}\(\\mathcal\{X\}\_\{i\}\)\], whereϕv​\(⋅\)\\phi\_\{v\}\(\\cdot\),ϕq​\(⋅\)\\phi\_\{q\}\(\\cdot\), andϕu​\(⋅\)\\phi\_\{u\}\(\\cdot\)denote the visual encoder, question tokenizer, and user\-information encoder, respectively\. Although simple, this raw\-conditioning paradigm often fails under imbalanced multi\-user preferences\. First, during representation learning, frequent preferences provide more stable optimization signals and can dominate the shared user condition, causing sparse user\-specific preferences to be averaged out\. Second, during generation, the model may follow visual priors or common answer patterns instead of routing attention to the preference cue relevant to the current image\-question pair\. As a result, responses become insensitive to individualized preferences and drift toward dominant population\-level choices\. We term this failure mode group preference collapse, and next present our framework for mitigating it\.

## 4Method

### 4\.1Factorized User Representation and Profile Learning

We first introduce the factorized user representation used by our framework\. Given the training\-time personalization information𝒳i=\{Ii,Di,𝒫i\}\\mathcal\{X\}\_\{i\}=\\\{I\_\{i\},D\_\{i\},\\mathcal\{P\}\_\{i\}\\\}of useruiu\_\{i\}, we represent each user with a factorized state that contains both profile\-related and preference\-related factors\. The profile factors are learned from relatively stable user information, such as the user\-related image and textual description\. The preference factors describe the user’s preferences overFFdifferent semantic facets \(e\.g\., lifestyle, fashion\) and will be further used for preference\-sensitive personalization\. For useruiu\_\{i\}, the factorized user representation consists of\{𝐳iimg,𝐳ides,\{𝐳i,fpref\}f=1F\}\\\{\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\},\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\},\\\{\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}\\\}\_\{f=1\}^\{F\}\\\}, where𝐳iimg\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\}and𝐳ides\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\}denote image\-based and description\-based profile factors, and𝐳i,fpref\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}denotes the representation of theff\-th preference facet\. This factorized form allows profile information and preference information to be modeled separately, rather than being compressed into a single user representation\.

We then learn the profile factors from two complementary profile views\. Given the user\-related imageIiI\_\{i\}and textual profile descriptionDiD\_\{i\}, we obtain

𝐳iimg=gimg​\(Ii\),𝐳ides=gdes​\(Di\),\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\}=g\_\{\\mathrm\{img\}\}\(I\_\{i\}\),\\quad\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\}=g\_\{\\mathrm\{des\}\}\(D\_\{i\}\),\(2\)wheregimg​\(⋅\)g\_\{\\mathrm\{img\}\}\(\\cdot\)andgdes​\(⋅\)g\_\{\\mathrm\{des\}\}\(\\cdot\)are the visual and textual profile encoders\. SinceIiI\_\{i\}andDiD\_\{i\}describe the same user from different modalities, their embeddings should be consistent for the same user and distinguishable across different users\. We therefore use a profile learning objective:

ℒprof=ℒalignprof\+ℒcontraprof\.\\mathcal\{L\}\_\{\\mathrm\{prof\}\}=\\mathcal\{L\}\_\{\\mathrm\{align\}\}^\{\\mathrm\{prof\}\}\+\\mathcal\{L\}\_\{\\mathrm\{contra\}\}^\{\\mathrm\{prof\}\}\.\(3\)Here,ℒalignprof\\mathcal\{L\}\_\{\\mathrm\{align\}\}^\{\\mathrm\{prof\}\}aligns the image\-based and description\-based profile embeddings of the same user, whileℒcontraprof\\mathcal\{L\}\_\{\\mathrm\{contra\}\}^\{\\mathrm\{prof\}\}separates profile embeddings from different users\. The learned profile factors provide stable user profile information, while preference\-related variations are modeled by the preference factors\.

### 4\.2Residual\-based Preference Factor Learning

Unlike profile information, user preferences are often sparse and unevenly distributed across users\. Frequent preferences can dominate the learned representation, causing less common user\-specific preferences to be absorbed into group\-level patterns\. To separate common preference tendencies from individual deviations, we decompose each facet\-specific preference factor into a shared prototype and a personalized residual:

𝐳i,fpref=𝐳¯f\+Δi,f,f=1,…,F,\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}=\\bar\{\\mathbf\{z\}\}\_\{f\}\+\\Delta\_\{i,f\},\\qquad f=1,\\dots,F,\(4\)where𝐳¯f\\bar\{\\mathbf\{z\}\}\_\{f\}denotes the shared prototype of facetff, andΔi,f\\Delta\_\{i,f\}denotes the residual preference offset of useruiu\_\{i\}\. The prototype captures the population\-level tendency of a preference facet, while the residual represents how an individual user deviates from this shared tendency\. However, this decomposition alone does not ensure that the residuals preserve meaningful user\-specific preferences\. Under imbalanced preference distributions, residuals may still be pulled toward frequent preference groups, depend on spurious group\-level correlations, or become redundant across different facets\. We therefore regularize the residual space with three complementary designs: imbalance\-aware residual contrast, counterfactual pseudo\-user augmentation, and preference residual decorrelation\.

![Refer to caption](https://arxiv.org/html/2607.22603v1/x2.png)Figure 2:Overview of PrefMoE\. PrefMoE separates profile and preference representations, models preferences with shared prototypes and personalized residuals, and regularizes the residuals with contrastive learning and decorrelation\. A hierarchical MoE router then activates profile\- and preference\-aware LoRA experts for query\-dependent personalized reasoning\.Imbalance\-aware Residual Contrast\.Under imbalanced preferences, residuals of minority preference groups receive weaker supervision and can be pulled toward frequent preference patterns\. We therefore apply an imbalance\-aware contrastive objective to the query\-conditioned activations of personalized residuals\. For each image\-question pair\(Vn,Qn\)\(V\_\{n\},Q\_\{n\}\), let𝐡n=Backbone​\(Vn,Qn\)\\mathbf\{h\}\_\{n\}=\\mathrm\{Backbone\}\(V\_\{n\},Q\_\{n\}\)denote the hidden representation produced by the MLLM backbone\. For useruiu\_\{i\}and facetff, we compute the residual\-conditioned activation as𝐚i,fn=Ef​\(𝐡n;Δi,f\)\\mathbf\{a\}\_\{i,f\}^\{n\}=E\_\{f\}\(\\mathbf\{h\}\_\{n\};\\Delta\_\{i,f\}\), whereEfE\_\{f\}is a facet\-specific activation module that maps the current visual\-question context and the personalized residual into a preference\-aware activation\.

We use facet\-level preference annotations to construct contrastive supervision\. For useruiu\_\{i\}and facetff, letℬ\\mathcal\{B\}denote the candidate users in the mini\-batch, and let𝒩i,f⊆ℬ\\mathcal\{N\}\_\{i,f\}\\subseteq\\mathcal\{B\}denote users that share the same annotation withuiu\_\{i\}on facetff\. We define

pi,fn=∑uj∈𝒩i,fexp⁡\(cos​\(𝐚i,fn,𝐚j,fn\)/τ\)∑uj∈ℬexp⁡\(cos​\(𝐚i,fn,𝐚j,fn\)/τ\),p\_\{i,f\}^\{n\}=\\frac\{\\sum\_\{u\_\{j\}\\in\\mathcal\{N\}\_\{i,f\}\}\\exp\\left\(\\mathrm\{cos\}\(\\mathbf\{a\}\_\{i,f\}^\{n\},\\mathbf\{a\}\_\{j,f\}^\{n\}\)/\\tau\\right\)\}\{\\sum\_\{u\_\{j\}\\in\\mathcal\{B\}\}\\exp\\left\(\\mathrm\{cos\}\(\\mathbf\{a\}\_\{i,f\}^\{n\},\\mathbf\{a\}\_\{j,f\}^\{n\}\)/\\tau\\right\)\},\(5\)wherecos​\(⋅,⋅\)\\mathrm\{cos\}\(\\cdot,\\cdot\)denotes cosine similarity andτ\\tauis the temperature\. This contrastive term pulls together users with the same facet\-level preference under the current visual\-question context and contrasts them with other users\. Inspired by focal lossLinet al\.\([2017](https://arxiv.org/html/2607.22603#bib.bib43)\)and reweighting strategies for class\-imbalanced learningCuiet al\.\([2019](https://arxiv.org/html/2607.22603#bib.bib44)\), we define the imbalance\-aware residual contrast loss for the image\-question pairnn:

ℒres=−1B​F​∑i=1B∑f=1F\(1−ρi,f\)η​\(1−pi,fn\)​log⁡pi,fn,\\mathcal\{L\}\_\{\\mathrm\{res\}\}=\-\\frac\{1\}\{BF\}\\sum\\nolimits\_\{i=1\}^\{B\}\\sum\\nolimits\_\{f=1\}^\{F\}\(1\-\\rho\_\{i,f\}\)^\{\\eta\}\(1\-p\_\{i,f\}^\{n\}\)\\log p\_\{i,f\}^\{n\},\(6\)whereB=\|ℬ\|B=\|\\mathcal\{B\}\|,ρi,f\\rho\_\{i,f\}is the normalized frequency of the preference group of useruiu\_\{i\}under facetff, andη≥0\\eta\\geq 0controls the reweighting strength\. The frequency weight\(1−ρi,f\)η\(1\-\\rho\_\{i,f\}\)^\{\\eta\}assigns larger weights to low\-frequency preference groups, while the focal term\(1−pi,fn\)\(1\-p\_\{i,f\}^\{n\}\)emphasizes hard residual activations\. In practice, the loss is averaged over samples in the mini\-batch\.

Counterfactual Pseudo\-user Augmentation\.Real user data may contain spurious correlations between user groups and population\-dominant preferences, allowing the model to rely on group\-level shortcuts\. To reduce this dependency, we construct counterfactual pseudo users by recombining facet\-level preference entries from different real users\. Specifically, for each pseudo user, we form a counterfactual preference set𝒫~\\tilde\{\\mathcal\{P\}\}by independently sampling preference entries across facets from real users\. The recombined preference set is processed by the same preference factor learning module, producing pseudo preference factors and residuals\. We then add these pseudo users to the candidate setℬ\\mathcal\{B\}in the residual contrastive objective\. This augmentation increases preference diversity and encourages residual activations to be aligned by facet\-level preference semantics rather than by original user\-group correlations\. Since pseudo users are used only for residual contrastive learning, no additional answer\-level labels are required\.

Preference Residual Decorrelation\.The above contrastive objectives encourage preference\-specific activations, but they do not explicitly prevent different facets from encoding similar user\-specific offsets\. If multiple facets carry redundant residual information, the learned preference factors remain entangled despite the facet\-wise decomposition\. We therefore impose two soft decorrelation constraints on the residual space\. The first separates residuals across different preference facets, while the second separates each residual from its corresponding shared prototype, encouraging the residual to capture personalized deviations rather than population\-level tendencies\. LetΔ~i,f=Δi,f/‖Δi,f‖2\\tilde\{\\Delta\}\_\{i,f\}=\\Delta\_\{i,f\}/\\\|\\Delta\_\{i,f\}\\\|\_\{2\}denote the normalized residual, and let𝐳~f=𝐳¯f/‖𝐳¯f‖2\\tilde\{\\mathbf\{z\}\}\_\{f\}=\\bar\{\\mathbf\{z\}\}\_\{f\}/\\\|\\bar\{\\mathbf\{z\}\}\_\{f\}\\\|\_\{2\}denote the normalized prototype\. We define

ℒdec​\(ui\)=1\(F2\)​∑1≤f<g≤F\(Δ~i,f⊤​Δ~i,g\)2\+1F​∑f=1F\(Δ~i,f⊤​𝐳~f\)2\.\\mathcal\{L\}\_\{\\mathrm\{dec\}\}\(u\_\{i\}\)=\\frac\{1\}\{\{F\\choose 2\}\}\\sum\\nolimits\_\{1\\leq f<g\\leq F\}\\left\(\\tilde\{\\Delta\}\_\{i,f\}^\{\\top\}\\tilde\{\\Delta\}\_\{i,g\}\\right\)^\{2\}\+\\frac\{1\}\{F\}\\sum\\nolimits\_\{f=1\}^\{F\}\\left\(\\tilde\{\\Delta\}\_\{i,f\}^\{\\top\}\\tilde\{\\mathbf\{z\}\}\_\{f\}\\right\)^\{2\}\.\(7\)The first term reduces redundancy across preference facets, while the second encourages residuals to capture personalized deviations beyond the shared prototype\.

Averagingℒdec​\(ui\)\\mathcal\{L\}\_\{\\mathrm\{dec\}\}\(u\_\{i\}\)over users in the mini\-batch, we obtain the residual decorrelation termℒdec\\mathcal\{L\}\_\{\\mathrm\{dec\}\}\. The counterfactual pseudo users are included in the candidate set ofℒres\\mathcal\{L\}\_\{\\mathrm\{res\}\}, so they contribute to preference learning through the residual contrast objective\. Combining query\-conditioned residual contrast with residual\-space decorrelation, the overall preference factor learning objective is

ℒpref=ℒres\+ℒdec\.\\mathcal\{L\}\_\{\\mathrm\{pref\}\}=\\mathcal\{L\}\_\{\\mathrm\{res\}\}\+\\mathcal\{L\}\_\{\\mathrm\{dec\}\}\.\(8\)

### 4\.3Factorized User\-aware Reasoning with Hierarchical MoE

Although the previous modules separate profile and preference factors in representation space, these factors may still be underused if all user information is injected through a single reasoning path\. We therefore introduce a hierarchical MoE routing mechanism based on LoRA experts, which turns factorized user representations into factorized reasoning paths\. Specifically, image\-based profile factors, description\-based profile factors, and facet\-wise preference factors are routed to separate LoRA experts, while a query\-dependent router determines how to combine their updates for the current visual question\. Given an image\-question pair\(V,Q\)\(\{V\},Q\), we first obtain a query representation𝐪=\[ϕv​\(V\);ϕq​\(Q\)\]\\mathbf\{q\}=\[\\phi\_\{v\}\(V\);\\phi\_\{q\}\(Q\)\]\. Let𝐡=Backbone​\(V,Q\)\\mathbf\{h\}=\\mathrm\{Backbone\}\(V,Q\)denote the hidden representation produced by the MLLM backbone\. The query𝐪\\mathbf\{q\}is used for routing, while𝐡\\mathbf\{h\}is adapted by LoRA experts\. For a target useruiu\_\{i\}, we use the profile factors𝐳iimg\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\}and𝐳ides\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\}, and the facet\-wise preference factors\{𝐳i,fpref\}f=1F\\\{\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}\\\}\_\{f=1\}^\{F\}learned in the previous sections\.

We maintain separate LoRA experts for profile and preference reasoning\. The preference branch containsFFfacet\-specific expertsℰpref=\{Efpref\}f=1F\\mathcal\{E\}^\{\\mathrm\{pref\}\}=\\\{E\_\{f\}^\{\\mathrm\{pref\}\}\\\}\_\{f=1\}^\{F\}\. The profile branch contains two profile expertsℰprof=\{Eimgprof,Edesprof\}\\mathcal\{E\}^\{\\mathrm\{prof\}\}=\\\{E\_\{\\mathrm\{img\}\}^\{\\mathrm\{prof\}\},E\_\{\\mathrm\{des\}\}^\{\\mathrm\{prof\}\}\\\}, corresponding to the image\-based and description\-based profile factors\. Given the query𝐪\\mathbf\{q\}, the routers compute factor\-specific mixture weights:

𝜶pref=σ​\(\{𝐰pref⊤​\[𝐪;𝐳i,fpref\]\}f=1F\),𝜶prof=σ​\(\{𝐰prof⊤​\[𝐪;𝐳i,kprof\]\}k∈\{img,des\}\),\\boldsymbol\{\\alpha\}^\{\\mathrm\{pref\}\}=\\sigma\\left\(\\left\\\{\\mathbf\{w\}\_\{\\mathrm\{pref\}\}^\{\\top\}\[\\mathbf\{q\};\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}\]\\right\\\}\_\{f=1\}^\{F\}\\right\),\\quad\\boldsymbol\{\\alpha\}^\{\\mathrm\{prof\}\}=\\sigma\\left\(\\left\\\{\\mathbf\{w\}\_\{\\mathrm\{prof\}\}^\{\\top\}\[\\mathbf\{q\};\\mathbf\{z\}\_\{i,k\}^\{\\mathrm\{prof\}\}\]\\right\\\}\_\{k\\in\\\{\\mathrm\{img\},\\mathrm\{des\}\\\}\}\\right\),\(9\)whereσ​\(⋅\)\\sigma\(\\cdot\)denotes the Softmax function,𝐳i,imgprof=𝐳iimg\\mathbf\{z\}\_\{i,\\mathrm\{img\}\}^\{\\mathrm\{prof\}\}=\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\}, and𝐳i,desprof=𝐳ides\\mathbf\{z\}\_\{i,\\mathrm\{des\}\}^\{\\mathrm\{prof\}\}=\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\}\. The preference router selects the preference facets relevant to the query, while the profile router selects the profile source that is more useful for the current reasoning context\. The factor\-aware MoE updates are computed as

𝐡pref=∑f=1Fαfpref​Efpref​\(𝐡;𝐳i,fpref\),𝐡prof=∑k∈\{img,des\}αkprof​Ekprof​\(𝐡;𝐳i,kprof\)\.\\mathbf\{h\}^\{\\mathrm\{pref\}\}=\\sum\\nolimits\_\{f=1\}^\{F\}\\alpha\_\{f\}^\{\\mathrm\{pref\}\}E\_\{f\}^\{\\mathrm\{pref\}\}\(\\mathbf\{h\};\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}\),\\quad\\mathbf\{h\}^\{\\mathrm\{prof\}\}=\\sum\\nolimits\_\{k\\in\\\{\\mathrm\{img\},\\mathrm\{des\}\\\}\}\\alpha\_\{k\}^\{\\mathrm\{prof\}\}E\_\{k\}^\{\\mathrm\{prof\}\}\(\\mathbf\{h\};\\mathbf\{z\}\_\{i,k\}^\{\\mathrm\{prof\}\}\)\.\(10\)The preference branch performs facet\-aware preference reasoning, while the profile branch performs profile\-aware reasoning from image\-based and description\-based user information\. By separating these reasoning paths, the model avoids relying on a single mixed user condition and can select the user factors relevant to the current visual question\. A branch\-level router further determines how much the model should rely on profile\-aware and preference\-aware updates:

𝜸=σ​\(𝐖br​\[𝐪;𝐡prof;𝐡pref\]\)\.\\boldsymbol\{\\gamma\}=\\sigma\\left\(\\mathbf\{W\}\_\{\\mathrm\{br\}\}\[\\mathbf\{q\};\\mathbf\{h\}^\{\\mathrm\{prof\}\};\\mathbf\{h\}^\{\\mathrm\{pref\}\}\]\\right\)\.\(11\)The final adapted representation is

𝐡~=𝐡\+γprof​𝐡prof\+γpref​𝐡pref\.\\tilde\{\\mathbf\{h\}\}=\\mathbf\{h\}\+\\gamma\_\{\\mathrm\{prof\}\}\\mathbf\{h\}^\{\\mathrm\{prof\}\}\+\\gamma\_\{\\mathrm\{pref\}\}\\mathbf\{h\}^\{\\mathrm\{pref\}\}\.\(12\)
Training objective and inference\.The overall training objective combines answer prediction, profile learning, and preference factor learning:

ℒ=ℒvqa\+ℒprof\+ℒpref\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{vqa\}\}\+\\mathcal\{L\}\_\{\\mathrm\{prof\}\}\+\\mathcal\{L\}\_\{\\mathrm\{pref\}\}\.\(13\)At inference time, for a target useruiu\_\{i\}, the model retrieves the learned profile factors𝐳iimg,𝐳ides\\mathbf\{z\}^\{\\mathrm\{img\}\}\_\{i\},\\mathbf\{z\}^\{\\mathrm\{des\}\}\_\{i\}and preference factors\{𝐳i,fpref\}f=1F\\\{\\mathbf\{z\}\_\{i,f\}^\{\\mathrm\{pref\}\}\\\}\_\{f=1\}^\{F\}, and applies the hierarchical router to the current image\-question pair\.

Table 1:Major comparisons with SOTAs under 0\-turn and 10\-turn settings\.MethodType0\-turn10\-turnOverall↑\\uparrowPreference↑\\uparrowProfile↑\\uparrowCollapse↓\\downarrowOverall↑\\uparrowPreference↑\\uparrowProfile↑\\uparrowCollapse↓\\downarrowLLaVA\-1\.5\-7BNT0\.35640\.33870\.37160\.62280\.31670\.30670\.33370\.6164LLaVA\-1\.5\-13BNT0\.40230\.36530\.42220\.60270\.37160\.33330\.39280\.6119LLaVA\-OV\-72BNT0\.51000\.48000\.52620\.52970\.46200\.43330\.47740\.5525DeepSeek\-VL2\-TinyNT0\.36830\.36130\.37200\.71690\.34970\.34400\.35270\.6895DeepSeek\-VL2\-SmallNT0\.38660\.38530\.38710\.68950\.36420\.36670\.36690\.6712DeepSeek\-VL2NT0\.46200\.45330\.46670\.58900\.41030\.40670\.41220\.5982Qwen2\.5\-VL\-7BNT0\.42760\.35870\.56530\.64840\.38850\.34330\.43150\.6530Qwen2\.5\-VL\-32BNT0\.49140\.43330\.52260\.57080\.44100\.40270\.46160\.5982Qwen2\.5\-VL\-72BNT0\.53010\.46000\.56770\.50230\.50160\.44530\.53190\.5525LLaVA\-1\.5\-7BFFT0\.50720\.44130\.54270\.34250\.50260\.44130\.53550\.3470LLaVA\-1\.5\-13BFFT0\.53010\.46000\.56770\.31050\.51000\.44000\.54770\.3196LLaVA\-OV\-72BFFT0\.62000\.59470\.63370\.31510\.56130\.53200\.57710\.3379DeepSeek\-VL2\-TinyFFT0\.47110\.42130\.49750\.58450\.46810\.41870\.49460\.5936DeepSeek\-VL2\-SmallFFT0\.48620\.38400\.54120\.65300\.48340\.37470\.54190\.6530DeepSeek\-VL2FFT0\.58140\.50930\.62010\.45210\.57480\.49870\.61580\.4566Qwen2\.5\-VL\-7BFFT0\.44520\.30130\.52260\.65300\.48070\.37070\.53980\.7854Qwen2\.5\-VL\-32BFFT0\.57180\.54800\.58490\.47950\.49140\.43330\.52260\.5114Qwen2\.5\-VL\-72BFFT0\.60140\.59470\.60570\.40180\.52170\.50530\.53050\.3927Yo’LLaVAPEFT0\.50400\.48800\.51250\.20750\.48400\.46800\.49250\.2146LLaVA\-NeXT\-34BPEFT0\.55990\.62000\.52760\.24660\.52990\.59000\.49760\.2054LOVA3PEFT0\.53290\.56800\.51400\.42920\.49790\.53300\.47900\.4247TG\-LLaVAPEFT0\.54130\.58000\.52040\.23290\.50130\.54000\.48040\.1506PrefMoE \(LLaVA\-1\.5\-7B\)PEFT0\.67510\.67330\.67600\.12330\.59860\.58400\.60650\.1553PrefMoE \(LLaVA\-1\.5\-13B\)PEFT0\.70120\.68000\.71250\.14160\.66010\.62670\.67810\.1370PrefMoE \(LLaVA\-OV\-72B\)PEFT0\.78930\.76130\.80430\.11420\.69140\.66270\.70680\.1187PrefMoE \(DeepSeek\-VL2\-Tiny\)PEFT0\.66010\.64670\.66740\.25110\.60330\.59330\.60860\.2740PrefMoE \(DeepSeek\-VL2\-Small\)PEFT0\.73050\.58130\.81080\.27400\.63820\.52000\.70180\.2694PrefMoE \(DeepSeek\-VL2\)PEFT0\.79910\.75200\.82440\.13240\.70120\.68000\.71250\.1826PrefMoE \(Qwen2\.5\-VL\-7B\)PEFT0\.76130\.68800\.80070\.17810\.65030\.62530\.66380\.1826PrefMoE \(Qwen2\.5\-VL\-32B\)PEFT0\.79020\.76400\.80430\.14160\.69000\.66530\.70320\.2100PrefMoE \(Qwen2\.5\-VL\-72B\)PEFT0\.81120\.78930\.82370\.10960\.73010\.58130\.81000\.1279

## 5Experiment

### 5\.1Experimental Setup

Dataset\.MMPBKimet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib30)\)is a multimodal personalization benchmark containing multimodal queries over 50 human users and 61 non\-human personalized concepts, including animals, objects, and characters\. The benchmark includes both profile\-recognition queries and preference\-aware queries\. The preference annotations cover five semantic facets: entertainment, travel, lifestyle, shopping, and fashion\. Since MMPB does not provide an official train/test split, we construct a cleaned preference\-sensitive evaluation split, termed MMPB\-Clean, from the original annotations\. During construction, we preserve the original user\-concept annotations and retain the original labels for recognition and multiple\-choice questions, while reconstructing preference yes/no labels when necessary to reduce question\-format shortcuts\. All methods are evaluated on the same split\. MMPB\-Clean is designed to test whether models capture target\-specific preferences and personalized concepts rather than exploit dataset priors alone\. We report 0\-turn results without dialogue history and 10\-turn results with a ten\-turn interaction history prepended to the target query\. More details are provided in Appendix[B](https://arxiv.org/html/2607.22603#A2)\. Implementation details\.We implement our method in PyTorch and train it on four NVIDIA A100 GPUs\. The default batch size is 4\. We use the AdamW optimizerLoshchilov and Hutter \([2019](https://arxiv.org/html/2607.22603#bib.bib45)\)with a learning rate of2×10−42\\times 10^\{\-4\}and a cosine learning rate scheduler\. Training is performed with bfloat16 precision and TF32 enabled, together with gradient checkpointing to reduce GPU memory consumption\. Metrics\.We report exact\-match answer accuracy on profile and preference questions, together with overall accuracy\. We also report preference collapse, which measures how often a model predicts an item as preferred by a target user when the item does not match that user’s ground\-truth preference\. A lower collapse rate indicates better preservation of user\-specific preference boundaries\. More metric details are provided in the Appendix[C](https://arxiv.org/html/2607.22603#A3)\. Compared methods\.We compare with representative general\-purpose MLLMs across different model scales, including LLaVA\-1\.5Liuet al\.\([2024a](https://arxiv.org/html/2607.22603#bib.bib22)\), LLaVA\-OVLiet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib24)\), DeepSeek\-VL2Wuet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib25)\), and Qwen2\.5\-VLBaiet al\.\([2025b](https://arxiv.org/html/2607.22603#bib.bib27)\)\. We also include recent personalized or multimodal baselines, including Yo’LLaVANguyenet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib5)\), LLaVA\-NeXTLiuet al\.\([2024b](https://arxiv.org/html/2607.22603#bib.bib23)\), LOVA3Zhaoet al\.\([2024](https://arxiv.org/html/2607.22603#bib.bib28)\), and TG\-LLaVAYanet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib29)\)\. For general\-purpose MLLMs, we report no\-training \(NT\) and full fine\-tuning \(FFT\) results\. For personalized multimodal methods and our method, we report PEFT\-based personalization results\. At inference time, NT and FFT baselines receive explicit user profile and preference descriptions as contextual prompts, while PrefMoE only receives the user identifier and retrieves the learned structured user state from memory\.

### 5\.2Major Comparisons and Ablation Studies

Table[1](https://arxiv.org/html/2607.22603#S4.T1)reports the main results under 0\-turn and 10\-turn settings\. Raw MLLMs exhibit high preference collapse, showing that general multimodal capability alone is insufficient for user\-specific preference modeling\. Full fine\-tuning improves accuracy but still fails to reliably preserve preference boundaries\. By contrast, our method consistently improves overall and preference accuracy while reducing collapse across backbones and model scales\. On LLaVA\-1\.5\-7B, it improves 0\-turn overall accuracy from 0\.5072 to 0\.6751 and reduces collapse from 0\.3425 to 0\.1233 over FFT\. The gains further extend to stronger backbones, with different PrefMoE variants achieving the best 0\-turn overall accuracy, the best 10\-turn preference accuracy, and the lowest collapse rates\. These results suggest that preference\-sensitive personalization benefits from factorized adaptation rather than relying solely on scaling or standard fine\-tuning\.

Table 2:Component\-wise ablation study of the proposed method\. E, P, I, C, D, and M denote the basic user embedding module, profile factor learning, imbalance\-aware residual preservation, counterfactual user augmentation, preference decorrelation, and hierarchical MoE router, respectively\.0\-turn10\-turnEPICDMOverall↑\\uparrowPreference↑\\uparrowProfile↑\\uparrowCollapse↓\\downarrowOverall↑\\uparrowPreference↑\\uparrowProfile↑\\uparrowCollapse↓\\downarrow✓0\.55620\.41200\.63370\.60270\.49280\.35840\.56510\.6347✓✓0\.60000\.57000\.61590\.33330\.52660\.49020\.54610\.3653✓✓✓0\.62790\.61330\.63580\.32880\.55060\.52750\.56310\.3607✓✓✓✓0\.63030\.62530\.63300\.22830\.55710\.53930\.56670\.2603✓✓✓✓✓0\.63820\.64670\.63370\.14570\.56880\.56220\.57230\.1781✓✓✓✓✓✓0\.67510\.67330\.67600\.12330\.59860\.58400\.60650\.1553

Table[2](https://arxiv.org/html/2607.22603#S5.T2)shows that each component contributes to our method\. The basic user embedding still suffers from severe collapse, indicating that storing user information alone is insufficient\. Profile factor learning brings the largest early improvement, while imbalance\-aware residual preservation further improves preference accuracy by protecting low\-density preference offsets\. Counterfactual augmentation and preference decorrelation mainly reduce collapse, suggesting that they help break group\-level shortcuts and separate preference facets\. The hierarchical MoE router achieves the best results, confirming the need to couple factorized user states with factorized computation paths\.

### 5\.3More Analysis

Number of pseudo users\. Fig\.[5](https://arxiv.org/html/2607.22603#S5.F5)shows that performance improves as the number of pseudo users increases from 10 to 50, while both Preference and Collapse become saturated afterward\. This suggests that 50 pseudo users are sufficient to capture representative personalized patterns in our setting, and we therefore use this value by default\. w/o vs\. w/ facet labels\. MMPB annotations organize preferences into five facets: entertainment, travel, lifestyle, shopping, and fashion\. To test whether our method relies on explicit facet\-name shortcuts, we remove these facet labels and merge all preference descriptions into a single unstructured text paragraph\. In this setting, facet structure is not directly provided and must be inferred from the preference content, mainly through the decorrelation\-based factor learning\. As shown in Fig\.[5](https://arxiv.org/html/2607.22603#S5.F5), our method remains robust without facet labels and still outperforms FFT baselines across different backbones\. This suggests that the gains do not mainly come from surface\-level facet names, but from learning structured target\-specific preference relations from the underlying descriptions\.

![Refer to caption](https://arxiv.org/html/2607.22603v1/x3.png)Figure 3:No\. of pseudo users\.
![Refer to caption](https://arxiv.org/html/2607.22603v1/x4.png)Figure 4:W/o facet labels\.
![Refer to caption](https://arxiv.org/html/2607.22603v1/x5.png)Figure 5:Routing analysis\.

Adaptive routing vs\. fixed routing\.Fig\.[5](https://arxiv.org/html/2607.22603#S5.F5)analyzes the routing behavior\. We first fix the preference–profile branch weights while keeping our preference facet router unchanged\. The results show a clear trade\-off: profile\-only routing achieves strong profile accuracy but weak preference accuracy and severe collapse, whereas preference\-only routing improves preference accuracy and reduces collapse but sacrifices profile accuracy and overall performance\. Although balanced fixed routing mitigates this trade\-off, it still underperforms our adaptive branch router\. We also replace the adaptive preference facet router with random or uniform facet selection while keeping the adaptive branch router\. Both variants degrade performance, showing that query\-dependent facet selection is necessary\. Overall, PrefMoE achieves the best overall accuracy and the lowest collapse in the 0\-turn setting, suggesting that effective personalized reasoning requires both adaptive profile–preference fusion and adaptive preference facet routing\. t\-SNE visualizations\.We use t\-SNE to qualitatively examine whether the learned preference representations preserve structured preference boundaries at both facet and population levels\. Fig\.[7](https://arxiv.org/html/2607.22603#S5.F7)visualizes representations from the five semantic facets, including entertainment, travel, lifestyle, shopping, and fashion\. Compared with standard fine\-tuning, our method produces more separated facet\-level distributions, suggesting that preference decorrelation helps different facets capture complementary user\-specific offsets rather than collapsing into redundant directions\. Fig\.[7](https://arxiv.org/html/2607.22603#S5.F7)further examines the top\-1 to top\-4 preference populations within each facet\. Standard fine\-tuning tends to mix these populations into shared regions, indicating that frequent and semantically related preferences are not well distinguished\. In contrast, our method preserves clearer local population structures within the same broad facet\. These results suggest that our method mitigates preference collapse not only by separating different preference facets, but also by maintaining fine\-grained boundaries among preference populations inside each facet\. Qualitative Case Study\. Fig\.[8](https://arxiv.org/html/2607.22603#S5.F8)shows representative profile and preference cases under 0\-turn and 10\-turn settings\. For profile\-centric questions, baseline MLLMs often produce plausible but target\-agnostic descriptions by focusing on salient visual content, while PrefMoE better conditions on the queried user and outputs profile\-consistent answers\. For preference\-grounded questions, baselines tend to follow generic visual cues or frequent preference patterns, whereas PrefMoE more often selects answers aligned with the target user’s preferences, such as bohemian fashion, solo travel, kombucha, sports TV, cooking shows, and K\-pop concerts\. These cases show that PrefMoE can use both profile and preference factors for personalized reasoning\. The failure cases reveal remaining limitations\. In Case 6, PrefMoE retrieves a relevant preference cue but does not fully ground the answer in the exact visual activity\. In Case 8, it captures the broad entertainment preference but confuses fine\-grained categories such as K\-pop concerts, live concerts, symphonies, and musicals\. This suggests that reliable preference reasoning still depends on explicit preference descriptions, clear question intent, and sufficient visual evidence\.

![Refer to caption](https://arxiv.org/html/2607.22603v1/figs/different_factors_tsne_500_grid.png)Figure 6:t\-SNE of 5 different preference facets\.
![Refer to caption](https://arxiv.org/html/2607.22603v1/figs/inner_factor_all_five_tsne_horizontal.png)Figure 7:t\-SNE visualization of top\-1 to top\-4 preference groups within five preference facets\.

![Refer to caption](https://arxiv.org/html/2607.22603v1/x6.png)Figure 8:Qualitative examples\.<SKS\>denotes the target personalized identity or concept\. Green boxes indicate correct predictions, while red marks indicate incorrect predictions\.

## 6Conclusion

In this paper, we identified group preference collapse as a critical failure mode in multi\-user personalized MLLMs, where responses drift toward dominant population\-level preferences rather than reflecting individualized user preferences\. We analyzed this problem from both representation and reasoning perspectives, showing that sparse preference signals can be suppressed in user representations and unreliably used during generation\. To address this issue, we proposed PrefMoE, a preference\-centric personalization framework that separates stable profile information from preference\-related representations, decomposes preferences into shared prototypes and personalized residuals, and preserves individualized residuals through imbalance\-aware learning, counterfactual pseudo\-user augmentation, and residual decorrelation\. We further introduced a hierarchical routing mechanism that routes profile and preference factors into separate LoRA adaptation paths for query\-dependent personalized reasoning\. Experiments across multiple MLLM backbones and evaluation settings show that PrefMoE improves preference\-sensitive personalization and reduces preference collapse\. Despite these improvements, our method still relies on explicit and well\-specified preference descriptions\. When user preferences are underspecified, weakly grounded in the visual scene, or when the question does not clearly indicate the relevant preference, personalized reasoning may still fail\.

## References

- \[1\]\(2024\)MyVLM: personalizing VLMs for user\-specific queries\.InECCV,pp\. 73–91\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[2\]J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds, R\. Ring, E\. Rutherford, S\. Cabi, T\. Han, Z\. Gong, S\. Samangooei, M\. Monteiro, J\. Menick, S\. Borgeaud, A\. Brock, A\. Nematzadeh, S\. Sharifzadeh, M\. Binkowski, R\. Barreira, O\. Vinyals, A\. Zisserman, and K\. Simonyan\(2022\)Flamingo: a visual language model for few\-shot learning\.InNeurIPS,Vol\.35,pp\. 23716–23736\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[3\]R\. An, S\. Yang, M\. Lu, K\. Zeng, Y\. Luo, Y\. Chen, J\. Cao, H\. Liang, Q\. She, S\. Zhang, and W\. Zhang\(2024\)MC\-LLaVA: multi\-concept personalized vision\-language model\.arXiv preprint arXiv:2411\.11706\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[4\]S\. Antol, A\. Agrawal, J\. Lu, M\. Mitchell, D\. Batra, C\. L\. Zitnick, and D\. Parikh\(2015\-12\)VQA: visual question answering\.InICCV,Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[5\]H\. Bai, R\. Wang, Z\. Du, Y\. Zhao, F\. Zhang, H\. Chen, X\. Zhu, B\. Zheng, and X\. Zhao\(2025\)Online\-PVLM: advancing personalized VLMs with online concept learning\.arXiv preprint arXiv:2511\.20056\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[6\]S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. Lin\(2025\)Qwen2\.5\-VL technical report\.arXiv preprint arXiv:2502\.13923\.Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[7\]Y\. Bai, A\. Jones, K\. Ndousse, A\. Askell, A\. Chen, N\. DasSarma, D\. Drain, S\. Fort, D\. Ganguli, T\. Henighan, N\. Joseph, S\. Kadavath, J\. Kernion, T\. Conerly, S\. El\-Showk, N\. Elhage, Z\. Hatfield\-Dodds, D\. Hernandez, T\. Hume, S\. Johnston, S\. Kravec, L\. Lovitt, N\. Nanda, C\. Olsson, D\. Amodei, T\. Brown, J\. Clark, S\. McCandlish, C\. Olah, B\. Mann, and J\. Kaplan\(2022\)Training a helpful and harmless assistant with reinforcement learning from human feedback\.arXiv preprint arXiv:2204\.05862\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[8\]R\. A\. Bradley and M\. E\. Terry\(1952\)Rank analysis of incomplete block designs: i\. the method of paired comparisons\.Biometrika39\(3–4\),pp\. 324–345\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[9\]T\. Chi, P\. Chen, S\. Su, and Y\. Chen\(2017\)Speaker role contextual modeling for language understanding and dialogue policy learning\.InIJCNLP,pp\. 163–168\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[10\]P\. F\. Christiano, J\. Leike, T\. B\. Brown, M\. Martic, S\. Legg, and D\. Amodei\(2017\)Deep reinforcement learning from human preferences\.InNeurIPS,Vol\.30\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[11\]N\. Cohen, R\. Gal, E\. A\. Meirom, G\. Chechik, and Y\. Atzmon\(2022\)“This is my unicorn, fluffy”: personalizing frozen vision\-language representations\.InECCV,pp\. 558–577\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[12\]Y\. Cui, M\. Jia, T\. Lin, Y\. Song, and S\. Belongie\(2019\)Class\-balanced loss based on effective number of samples\.InCVPR,pp\. 9268–9277\.Cited by:[§4\.2](https://arxiv.org/html/2607.22603#S4.SS2.p3.9)\.
- \[13\]W\. Dai, J\. Li, D\. Li, A\. M\. H\. Tiong, J\. Zhao, W\. Wang, B\. Li, P\. Fung, and S\. C\. H\. Hoi\(2023\)InstructBLIP: towards general\-purpose vision\-language models with instruction tuning\.InNeurIPS,Vol\.36\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[14\]K\. Ethayarajh, W\. Xu, N\. Muennighoff, D\. Jurafsky, and D\. Kiela\(2024\)KTO: model alignment as prospect theoretic optimization\.InICML,Vol\.235,pp\. 12634–12651\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[15\]W\. Fang, Q\. Wu, J\. Chen, and Y\. Xue\(2025\)Guided mllm reasoning: enhancing mllm with knowledge and visual notes for visual question answering\.InCVPR,pp\. 19597–19607\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[16\]L\. Ge, D\. Halpern, E\. Micha, A\. D\. Procaccia, I\. Shapira, Y\. Vorobeychik, and J\. Wu\(2024\)Axioms for AI alignment from human feedback\.InNeurIPS,Vol\.37\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[17\]J\. Hejna, R\. Rafailov, H\. Sikchi, C\. Finn, S\. Niekum, W\. B\. Knox, and D\. Sadigh\(2024\)Contrastive preference learning: learning from human feedback without reinforcement learning\.InICLR,Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[18\]Q\. Huang, W\. Dai, J\. Liu, W\. He, H\. Jiang, M\. Song, J\. Chen, C\. Yao, and J\. Song\(2025\)Boosting mllm reasoning with text\-debiased hint\-grpo\.InICCV,pp\. 4848–4857\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[19\]J\. Kim, W\. Kim, W\. Park, and J\. Do\(2025\)MMPB: it’s time for multi\-modal personalization\.InNeurIPS Datasets and Benchmarks Track,Cited by:[§B\.1](https://arxiv.org/html/2607.22603#A2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[20\]B\. Li, Y\. Zhang, D\. Guo, R\. Zhang, F\. Li, H\. Zhang, K\. Zhang, P\. Zhang, Y\. Li, Z\. Liu, and C\. Li\(2025\)LLaVA\-OneVision: easy visual task transfer\.TMLR\.Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[21\]J\. Li, M\. Galley, C\. Brockett, G\. P\. Spithourakis, J\. Gao, and B\. Dolan\(2016\)A persona\-based neural conversation model\.InACL,pp\. 994–1003\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[22\]J\. Li, D\. Li, S\. Savarese, and S\. C\. H\. Hoi\(2023\)BLIP\-2: bootstrapping language\-image pre\-training with frozen image encoders and large language models\.InICML,pp\. 19730–19742\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[23\]X\. Li and F\. Lyu\(2025\)MM\-prompt: cross\-modal prompt tuning for continual visual question answering\.arXiv preprint arXiv:2505\.19455\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[24\]T\. Lin, P\. Goyal, R\. Girshick, K\. He, and P\. Dollár\(2017\)Focal loss for dense object detection\.InICCV,pp\. 2980–2988\.Cited by:[§4\.2](https://arxiv.org/html/2607.22603#S4.SS2.p3.9)\.
- \[25\]H\. Liu, C\. Li, Y\. Li, and Y\. J\. Lee\(2024\)Improved baselines with visual instruction tuning\.InCVPR,Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[26\]H\. Liu, C\. Li, Y\. Li, B\. Li, Y\. Zhang, S\. Shen, and Y\. J\. Lee\(2024\)LLaVA\-NeXT: improved reasoning, OCR, and world knowledge\.Note:Technical blogCited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[27\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InICLR,Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[28\]T\. Nguyen, H\. Liu, Y\. Li, M\. Cai, U\. Ojha, and Y\. J\. Lee\(2024\)Yo’LLaVA: your personalized language and vision assistant\.InNeurIPS,Vol\.37\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1),[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[29\]T\. Nguyen, K\. K\. Singh, J\. Shi, T\. Bui, Y\. J\. Lee, and Y\. Li\(2025\)Yo’Chameleon: personalized vision and language generation\.InCVPR,pp\. 14438–14448\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[30\]L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. Lowe\(2022\)Training language models to follow instructions with human feedback\.InNeurIPS,Vol\.35,pp\. 27730–27744\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[31\]C\. Pham, H\. Phan, D\. Doermann, and Y\. Tian\(2025\)PLVM: a tuning\-free approach for personalized large vision\-language model\.InCVPR Workshops,pp\. 3671–3680\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1),[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[32\]R\. Pi, J\. Zhang, T\. Han, J\. Zhang, R\. Pan, and T\. Zhang\(2025\)Personalized visual instruction tuning\.InICLR,Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1),[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[33\]R\. Rafailov, A\. Sharma, E\. Mitchell, S\. Ermon, C\. D\. Manning, and C\. Finn\(2023\)Direct preference optimization: your language model is secretly a reward model\.InNeurIPS,Vol\.36\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[34\]S\. Rendle, C\. Freudenthaler, Z\. Gantner, and L\. Schmidt\-Thieme\(2009\)BPR: bayesian personalized ranking from implicit feedback\.InUAI,pp\. 452–461\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p2.1),[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[35\]S\. Seifi, V\. Dorovatas, M\. Cassinelli, F\. Despinoy, D\. O\. Reino, and R\. Aljundi\(2026\)Personalization toolkit: training free personalization of large vision language models\.TMLR\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1),[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[36\]J\. Tan, F\. Lyu, T\. Liu, F\. Hu, and W\. Feng\(2026\)Towards dynamic modality alignment in multimodal continual learning\.InCVPR,pp\. 39911–39921\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[37\]W\. Tang, Y\. Sun, Q\. Gu, and Z\. Li\(2026\)Visual position prompt for MLLM based visual grounding\.IEEE TMM\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[38\]G\. Wang, F\. Lyu, and C\. Ding\(2025\)Partition\-then\-adapt: combating prediction bias for reliable multi\-modal test\-time adaptation\.InNeurIPS,Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[39\]Z\. Wang, T\. Guan, P\. Fu, C\. Duan, Q\. Jiang, Z\. Guo, S\. Guo, J\. Luo, W\. Shen, and X\. Yang\(2025\)Marten: visual question answering with mask generation for multi\-modal document understanding\.InCVPR,pp\. 14460–14471\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[40\]Z\. Wu, X\. Chen, Z\. Pan, X\. Liu, W\. Liu, D\. Dai, H\. Gao, Y\. Ma, C\. Wu, B\. Wang, Z\. Xie, Y\. Wu, K\. Hu, J\. Wang, Y\. Sun, Y\. Li, Y\. Piao, K\. Guan, A\. Liu, X\. Xie, Y\. You, K\. Dong, X\. Yu, H\. Zhang, L\. Zhao, Y\. Wang, and C\. Ruan\(2024\)DeepSeek\-VL2: mixture\-of\-experts vision\-language models for advanced multimodal understanding\.arXiv preprint arXiv:2412\.10302\.Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[41\]D\. Yan, P\. Li, Y\. Li, H\. Chen, Q\. Chen, W\. Luo, W\. Dong, Q\. Yan, H\. Zhang, and C\. Shen\(2025\)TG\-LLaVA: text guided LLaVA via learnable latent embeddings\.Proceedings of the AAAI Conference on Artificial Intelligence39\(9\),pp\. 9076–9084\.Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[42\]C\. Yeh, B\. Russell, J\. Sivic, F\. C\. Heilbron, and S\. Jenni\(2023\)Meta\-personalizing vision\-language models to find named instances in video\.InCVPR,pp\. 19123–19132\.Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[43\]S\. Zhang, E\. Dinan, J\. Urbanek, A\. Szlam, D\. Kiela, and J\. Weston\(2018\)Personalizing dialogue agents: i have a dog, do you have pets too?\.InACL,pp\. 2204–2213\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.
- \[44\]H\. H\. Zhao, P\. Zhou, D\. Gao, Z\. Bai, and M\. Z\. Shou\(2024\)LOVA3: learning to visual question answering, asking and assessment\.InNeurIPS,Vol\.37\.Cited by:[§5\.1](https://arxiv.org/html/2607.22603#S5.SS1.p1.1)\.
- \[45\]D\. Zhu, J\. Chen, X\. Shen, X\. Li, and M\. Elhoseiny\(2024\)MiniGPT\-4: enhancing vision\-language understanding with advanced large language models\.InICLR,Cited by:[§1](https://arxiv.org/html/2607.22603#S1.p1.1)\.
- \[46\]D\. M\. Ziegler, N\. Stiennon, J\. Wu, T\. B\. Brown, A\. Radford, D\. Amodei, P\. Christiano, and G\. Irving\(2019\)Fine\-tuning language models from human preferences\.arXiv preprint arXiv:1909\.08593\.Cited by:[§2](https://arxiv.org/html/2607.22603#S2.p1.1)\.

Appendix

## Appendix AImplementation Details

We instantiate our framework on multiple MLLM backbones, including LLaVA\-1\.5, LLaVA\-OneVision, DeepSeek\-VL2, and Qwen2\.5\-VL\. Unless otherwise specified, we use LLaVA\-1\.5\-7B as the default backbone for ablation studies and detailed analysis\. For all backbone variants, we keep the same training objective and personalized modules, while adapting the LoRA insertion modules and batch configuration to the corresponding model architecture and memory constraints\. We apply rank\-64 LoRA to seven projection modules\. We insert these adapters into the top 8 decoder layers, with two insertion sites per selected layer: one after the self\-attention output and one after the MLP output\. This results in 16 personalized residual\-adapter insertion sites\. For other backbones, we use the same top\-layer insertion strategy and adapt the exact number of insertion sites according to the depth of the language decoder\. All experiments are implemented in PyTorch\. We use AdamW as the optimizer with a learning rate of2×10−42\\times 10^\{\-4\}and a cosine learning\-rate scheduler\. The default batch size is 4 for LLaVA\-1\.5\-7B\. Training is performed with bfloat16 precision and TF32 enabled, and gradient checkpointing is used to reduce memory consumption\. Unless otherwise noted, experiments are conducted on NVIDIA A100\-40G GPUs\.

## Appendix BDataset Construction

### B\.1Why MMPB\-Clean is Needed

We adopt MMPBKimet al\.\([2025](https://arxiv.org/html/2607.22603#bib.bib30)\)as the base benchmark because it provides a relatively systematic evaluation setting for multimodal personalization\. MMPB contains 10,017 image\-query pairs and 111 personalizable concepts spanning humans, animals, objects, and characters\. Its human category further includes preference\-grounded questions, making it suitable for studying both profile\-centric recognition and preference\-sensitive reasoning\. However, the original benchmark is not specifically designed to evaluate user\-conditioned personalization under shortcut\-controlled settings\.

To make the evaluation better reflect whether a model retrieves and applies the personalized information associated with the queried target, we derive a cleaned evaluation protocol from MMPB, termedMMPB\-Clean\. MMPB\-Clean does not introduce external images, users, concepts, or preference annotations\. Instead, it reuses the original MMPB resources and reorganizes them into a shortcut\-controlled protocol\. Specifically, we preserve the original personalized concepts, user profiles, preference annotations, images, and natural\-language questions whenever possible, while adjusting the train/evaluation split and reconstructing preference yes/no labels only when necessary\. The resulting number of QA pairs may differ from the original MMPB because additional preference negative pairs are constructed from existing MMPB users and preference annotations, rather than from newly collected data\.

MMPB\-Clean reduces two potential shortcut sources\. For recognition questions, we reduce image\-level shortcuts by grouping samples derived from the same underlying image and constructing the evaluation set mainly from name\-sensitive groups, where changing the queried user or concept name can change the answer\. Nearly uniform groups, where most queried targets lead to the same answer, are primarily kept for training support and are prevented from dominating evaluation\. For preference yes/no questions, we reduce question\-format shortcuts by reformulating the task as image\-semantics to user\-preference matching\. Instead of relying on the surface polarity of the question, such as explicit negation, the reconstructed labels indicate whether the semantic meaning of the image group matches the queried user’s corresponding preference field\. All methods are trained and evaluated using the same reconstructed split, labels, answer normalization, and evaluation protocol\. Therefore, MMPB\-Clean provides a fair and reproducible protocol for evaluating user\-conditioned personalized memory, with greater emphasis on name\-sensitive recognition and preference\-boundary reasoning\.

### B\.2Construction Protocol

MMPB\-Clean is constructed from the original MMPB annotations\. The reconstruction consists of three parts, corresponding to profile\-centric recognition questions, preference yes/no questions, and multiple\-choice questions\.

For profile\-centric recognition questions, we preserve the original task semantics and answer labels\. We first group samples by their underlying image\. For each image group, we examine whether changing the queried user or concept name can change the answer\. Image groups satisfying this condition are treated as name\-sensitive groups\. The recognition evaluation set is mainly constructed from these name\-sensitive groups, while nearly uniform groups are primarily kept for training support\. We remove internally inconsistent image groups and ensure that the same underlying image does not appear in both recognition training and recognition evaluation\.

For preference yes/no questions, we reconstruct the evaluation labels as image\-semantics to user\-preference matching\. Each user profile contains preference fields over several facets, including entertainment, fashion, lifestyle, shopping, and travel\. For each originally positive image group, we identify the stable preference elements shared by the associated users and treat them as the semantic meaning of the image group\. A query is labeled positive if this semantic meaning appears in the queried user’s corresponding preference field, and negative otherwise\. We keep the original positive pairs and construct additional negative pairs by selecting users whose corresponding preference fields do not contain the image\-related elements\. Originally negative groups are kept unchanged\.

For multiple\-choice questions, we keep the original questions and labels unchanged\. We only adjust the split by user or concept identity, so that each target contributes a limited number of multiple\-choice questions to evaluation while the remaining samples are used for training\. This preserves the original multiple\-choice task format while reducing repeated target\-specific patterns across training and evaluation\.

### B\.3Fairness and Reproducibility

MMPB\-Clean is used only to define the data split and evaluation labels\. All methods are trained and evaluated on exactly the same training and evaluation sets\. We do not apply method\-specific filtering, relabeling, or sample selection\. The reconstructed preference labels are independent of our model design and do not rely on factorized user representations, residual preference learning, or MoE routing\. Thus, MMPB\-Clean provides a shared evaluation protocol rather than an architecture\-specific benchmark\. To ensure fair comparison, all methods use the same visual inputs, natural\-language questions, target user or concept identifiers, candidate answer space, answer normalization procedure, and evaluation metrics\. For questions inherited from the original benchmark, we keep the original wording and labels\. For reconstructed preference yes/no questions, the image\-semantics to user\-preference matching rule is applied before model training and is shared by all methods\. To support reproducibility, we will release the reconstructed split files, image\-group identifiers, preference\-element mappings, answer\-label files, and evaluation scripts\. These files are sufficient to reproduce MMPB\-Clean from the original MMPB annotations and to evaluate future methods under the same protocol\.

### B\.4Dataset Statistics

Table[3](https://arxiv.org/html/2607.22603#A2.T3)summarizes the statistics of MMPB\-Clean\. The dataset contains 12,516 QA pairs in total, with 10,371 examples for training and 2,145 examples for evaluation\. The evaluation split preserves all 50 human preference users and all 61 non\-human recognition concepts, ensuring that both personalized preference reasoning and visual concept recognition are assessed under the same user/concept coverage as training\.

Compared with the original MMPB, MMPB\-Clean increases the number of QA pairs from 10,017 to 12,516\. The increase comes from preference QA, which is expanded from 5,001 to 7,500 pairs by constructing additional negative user–image pairs using existing MMPB users and preference annotations\. Recognition QA keeps the original 5,016 pairs, with only the train/evaluation split reorganized\. Thus, the changed statistics come from protocol reconstruction rather than external data collection\.

For preference\-oriented QA, MMPB\-Clean contains 7,500 examples evenly distributed across five preference facets: entertainment, fashion, lifestyle, shopping, and travel\. Each facet contains 1,500 QA pairs in total, while the train/evaluation split remains approximately balanced across facets\. For recognition\-oriented QA, the dataset contains 5,016 examples covering animal, character, human identity, and object recognition\. In this table, question subtypes are aggregated to provide a compact view of the overall data composition\.

Table 3:Detailed statistics of MMPB\-Clean\. The same evaluation split is used for both 0\-turn and 10\-turn evaluation settings\.StatisticTrainEvaluationTotalDataset scaleQA pairs10371214512516Images8272213110016Human users505050Non\-human concepts616161Preference QA by facetAll preference QA67507507500Entertainment13411591500Fashion13491511500Lifestyle13571431500Shopping13561441500Travel13471531500Recognition QA by concept typeAll recognition QA362113955016Animal600100700Character318482800Human identity22502502500Object4535631016

## Appendix CEvaluation Metrics

We evaluate model performance using two metrics: answer accuracy and preference\-boundary preservation\.

#### Exact\-match Answer Accuracy\.

We evaluate answer correctness with strict exact\-match accuracy after normalizing both predictions and ground\-truth answers into a canonical label space\. For multiple\-choice questions, we normalize answers toa,b,c,d; for binary questions, we normalize answers toyes,no\. The evaluator first extracts an explicit option or yes/no answer from the generated response\. If no explicit option is found for a multiple\-choice question, it further matches the response against the candidate option texts\. Predictions that cannot be mapped to a valid canonical label are counted as incorrect\. Given an evaluation set𝒮\\mathcal\{S\}, letgig\_\{i\}andg^i\\hat\{g\}\_\{i\}denote the normalized ground\-truth answer and prediction for sampleii, respectively\.

We define the per\-sample correctness indicator as

hiti=𝕀​\[g^i=gi∧g^i≠∅\]\.\\mathrm\{hit\}\_\{i\}=\\mathbb\{I\}\\left\[\\hat\{g\}\_\{i\}=g\_\{i\}\\land\\hat\{g\}\_\{i\}\\neq\\emptyset\\right\]\.The exact\-match accuracy on𝒮\\mathcal\{S\}is then computed by instance\-level micro\-averaging:

Acc​\(𝒮\)=1\|𝒮\|​∑i∈𝒮hiti\.\\mathrm\{Acc\}\(\\mathcal\{S\}\)=\\frac\{1\}\{\|\\mathcal\{S\}\|\}\\sum\_\{i\\in\\mathcal\{S\}\}\\mathrm\{hit\}\_\{i\}\.We report accuracy on three evaluation sets:

Overall=Acc​\(𝒮all\),Preference=Acc​\(𝒮pref\),Profile=Acc​\(𝒮prof\)\.\\mathrm\{Overall\}=\\mathrm\{Acc\}\(\\mathcal\{S\}\_\{\\mathrm\{all\}\}\),\\qquad\\mathrm\{Preference\}=\\mathrm\{Acc\}\(\\mathcal\{S\}\_\{\\mathrm\{pref\}\}\),\\qquad\\mathrm\{Profile\}=\\mathrm\{Acc\}\(\\mathcal\{S\}\_\{\\mathrm\{prof\}\}\)\.Here,𝒮all\\mathcal\{S\}\_\{\\mathrm\{all\}\}is the full evaluation set,𝒮pref\\mathcal\{S\}\_\{\\mathrm\{pref\}\}contains preference\-related questions, and𝒮prof\\mathcal\{S\}\_\{\\mathrm\{prof\}\}contains profile\-recognition questions that require user\-specific profile information\. Since Overall is micro\-averaged over all samples, it is weighted by the subset sizes and is not necessarily the arithmetic mean of Preference and Profile\.

#### Preference Collapse\.

Accuracy alone does not reveal whether a model preserves user\-specific preference boundaries\. A model may answer many samples correctly while still over\-expanding a user’s preference range by judging non\-preferred items as preferred\. We therefore reportCollapse, defined as the false positive rate on samples outside the target user’s true preference boundary\.

For each preference sampleii, letuiu\_\{i\}denote the target user andSiS\_\{i\}denote the true preference user set of the queried item, i\.e\., the users for whom the item should be considered preference\-matching\. Since preference questions may use different surface formulations, we first normalize yes/no answers into a unified preference\-semantic space\. Letyi∗∈\{0,1\}y\_\{i\}^\{\*\}\\in\\\{0,1\\\}be the normalized ground\-truth label andy^i∈\{0,1\}\\hat\{y\}\_\{i\}\\in\\\{0,1\\\}be the normalized model prediction, where11means preferred and0means not preferred\.

We define the boundary\-external set as

𝒪=\{i∣ui∉Si,yi∗=0\}\.\\mathcal\{O\}=\\\{\\,i\\mid u\_\{i\}\\notin S\_\{i\},\\;y\_\{i\}^\{\*\}=0\\,\\\}\.This set contains samples where the queried item should remain outside the target user’s preference boundary\. The Collapse score is computed as

Collapse=1\|𝒪\|∑i∈𝒪y^i=ℙ\(y^=1∣u∉S,y∗=0\)\.\\mathrm\{Collapse\}=\\frac\{1\}\{\|\\mathcal\{O\}\|\}\\sum\_\{i\\in\\mathcal\{O\}\}\\hat\{y\}\_\{i\}=\\mathbb\{P\}\\big\(\\hat\{y\}=1\\mid u\\notin S,\\;y^\{\*\}=0\\big\)\.A higher Collapse score means that the model more often absorbs non\-preferred items into the target user’s preferred set, while a lower value indicates better preservation of user\-specific preference boundaries\. Unlike Overall, Preference, and Profile, Collapse is not an accuracy score; lower values are better\.

## Appendix DMore Experimental Results

### D\.1Collapse Breakdown by Preference Popularity

Table[4](https://arxiv.org/html/2607.22603#A4.T4)reports preference collapse across different preference popularity groups\. The bucket size\|Si\|\|S\_\{i\}\|denotes the number of users who truly prefer the queried item\. Small buckets \(\|Si\|≤4\|S\_\{i\}\|\\leq 4\) correspond to more user\-specific preferences, while large buckets \(\|Si\|≥9\|S\_\{i\}\|\\geq 9\) correspond to broadly shared preferences\. The Overall collapse score is computed as the sample\-size\-weighted average of the three groups rather than their arithmetic mean\. Specifically, the three groups containn≤4=50n\_\{\\leq 4\}=50,n5​\-​8=121n\_\{5\\text\{\-\}8\}=121, andn≥9=48n\_\{\\geq 9\}=48boundary\-external samples, respectively\.

Across all settings, the large\-preference group has the highest collapse rate\. This indicates that models are more likely to over\-generalize broadly shared preference signals: when many users prefer an item, the model tends to incorrectly predict that the target user also prefers it, even when the target user is outside the item’s true preference set\. Scaling and full fine\-tuning reduce collapse compared with the no\-training setting, but they do not remove this popularity\-driven failure\. PrefMoE further reduces collapse across all preference popularity groups, showing that it better preserves individual preference boundaries and mitigates drift toward population\-level preferences\.

Table 4:Collapse breakdown for LLaVA\-family models across user frequency groups\.MethodSetting≤4\\leq 45–8≥9\\geq 9OverallLLaVA\-1\.5\-7BNT0\.60200\.61030\.67600\.6228LLaVA\-1\.5\-13BNT0\.57500\.59580\.64900\.6027LLaVA\-OV\-72BNT0\.51400\.53330\.53700\.5297LLaVA\-1\.5\-7BFFT0\.31800\.33810\.37900\.3425LLaVA\-1\.5\-13BFFT0\.29200\.30560\.34200\.3105LLaVA\-OV\-72BFFT0\.29700\.30560\.35800\.3151PrefMoE \(LLaVA\-1\.5\-7B\)PEFT0\.10500\.11750\.15700\.1233PrefMoE \(LLaVA\-1\.5\-13B\)PEFT0\.13900\.13300\.16600\.1416PrefMoE \(LLaVA\-OV\-72B\)PEFT0\.09700\.10790\.14800\.1142

### D\.2Detailed Comparisons for Different Preference Facets

Tables[5](https://arxiv.org/html/2607.22603#A4.T5)and[6](https://arxiv.org/html/2607.22603#A4.T6)provide a facet\-wise analysis of 0\-turn preference questions across entertainment, fashion, lifestyle, shopping, and travel\. PrefMoE consistently improves preference accuracy and reduces collapse across LLaVA\-family backbones, showing that its gains are not only reflected in aggregate metrics but also hold across different preference types\. For example, on LLaVA\-OV\-72B, PrefMoE improves the overall preference accuracy from 0\.5947 to 0\.7613 and reduces the overall collapse rate from 0\.3151 to 0\.1142 compared with FFT\.

The results also reveal clear facet\-level difficulty differences\. Shopping achieves particularly high preference accuracy under PrefMoE, suggesting that its cues are often more visually grounded and easier to associate with the image\. In contrast, lifestyle remains more challenging, with lower accuracy and relatively higher collapse across backbones\. This suggests that lifestyle preferences require more abstract and context\-dependent personalization rather than direct visual matching\. Overall, scaling and full fine\-tuning improve general performance, but they do not sufficiently preserve fine\-grained preference boundaries; PrefMoE provides more consistent preference\-sensitive reasoning across facets\.

Table 5:Facet\-wise preference collapse on the LLaVA family\.MethodSettingEntertainmentFashionLifestyleShoppingTravelOverallLLaVA\-1\.5\-7BNT0\.62500\.64580\.62220\.60000\.63640\.6228LLaVA\-1\.5\-13BNT0\.58330\.62500\.60000\.57780\.60610\.6027LLaVA\-OV\-72BNT0\.52080\.56250\.48890\.53330\.54550\.5297LLaVA\-1\.5\-7BFFT0\.35420\.37500\.33330\.28890\.36360\.3425LLaVA\-1\.5\-13BFFT0\.31250\.29170\.28890\.26670\.39390\.3105LLaVA\-OV\-72BFFT0\.33330\.35420\.31110\.24440\.30300\.3151PrefMoE \(LLaVA\-1\.5\-7B\)PEFT0\.14580\.04170\.17780\.08890\.12120\.1233PrefMoE \(LLaVA\-1\.5\-13B\)PEFT0\.16670\.10420\.20000\.11110\.09090\.1416PrefMoE \(LLaVA\-OV\-72B\)PEFT0\.12500\.06250\.15560\.06670\.03030\.1142

Table 6:Facet\-wise 0\-turn preference accuracy on the LLaVA family\.MethodTypeEntertainmentTravelLifestyleShoppingFashionOvr\.PreferenceLLaVA\-1\.5\-7BNT0\.33100\.34400\.33700\.34200\.33990\.3387LLaVA\-1\.5\-13BNT0\.35900\.36700\.36100\.37100\.36880\.3653LLaVA\-OV\-72BNT0\.47400\.48600\.47700\.48200\.48120\.4800LLaVA\-1\.5\-7BFFT0\.43700\.45600\.38700\.46200\.46260\.4413LLaVA\-1\.5\-13BFFT0\.45200\.47100\.40500\.48600\.48460\.4600LLaVA\-OV\-72BFFT0\.58500\.61300\.53100\.63100\.61210\.5947PrefMoE \(LLaVA\-1\.5\-7B\)PEFT0\.61200\.65220\.58600\.84200\.68100\.6733PrefMoE \(LLaVA\-1\.5\-13B\)PEFT0\.62300\.65570\.59200\.83100\.70400\.6800PrefMoE \(LLaVA\-OV\-72B\)PEFT0\.78000\.85940\.60400\.85800\.69900\.7613

### D\.3Computational Cost

Table 7:Computational cost comparison\. GPU\-hours are computed as the number of GPUs multiplied by wall\-clock time\.MethodTrain CostEval CostPeak MemoryLLaVA\-1\.5\-7B–0\.5–1 GPU\-hours15–20 GB/GPUYo’LLaVA\-7B5–10 GPU\-hours1–2 GPU\-hours28–34 GB/GPULOVA3\-7B6–12 GPU\-hours1–2 GPU\-hours28–34 GB/GPUTG\-LLaVA\-7B6–12 GPU\-hours1–2 GPU\-hours28–34 GB/GPUPrefMoE \(LLaVA\-1\.5\-7B\)8–16 GPU\-hours1–2 GPU\-hours28–34 GB/GPUPrefMoE \(LLaVA\-1\.5\-13B\)24–32 GPU\-hours2–3 GPU\-hours34–39 GB/GPUPrefMoE \(LLaVA\-OV\-72B\)450–600 GPU\-hours40–80 GPU\-hours38–40 GB/GPU

Table[7](https://arxiv.org/html/2607.22603#A4.T7)reports the computational cost of different methods\. The no\-training LLaVA\-1\.5\-7B baseline only requires inference, resulting in the lowest cost and memory usage\. Existing personalized VQA methods introduce additional training overhead, but their costs remain in a similar range under 7B backbones\. Compared with these methods, PrefMoE has a slightly higher training cost on LLaVA\-1\.5\-7B due to its preference\-aware modules and auxiliary objectives, while maintaining comparable evaluation cost and peak memory\. When scaling to larger backbones, the cost increases mainly with model size: PrefMoE on LLaVA\-1\.5\-13B requires more training resources, and the LLaVA\-OV\-72B variant is substantially more expensive due to the larger backbone\. Overall, PrefMoE introduces moderate overhead at the 7B scale, while the major cost increase comes from backbone scaling rather than the personalization mechanism itself\.

Similar Articles

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences

arXiv cs.LG

This position paper argues that large language models should learn from personalized rather than aggregated human preferences, highlighting theoretical limitations from social choice theory and practical issues from demographic diversity. It proposes bounded personalization frameworks that respect individual autonomy while maintaining universal safety constraints.

PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs

arXiv cs.LG

This paper introduces PFAdapter, a communication-efficient framework for personalized federated fine-tuning of Multimodal Large Language Models (MLLMs). It uses hierarchical LoRA decomposition to separate adapter parameters into global-shared and local-private components, achieving near 50% reduction in communication costs while improving personalization through orthogonality regularization.