SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

arXiv cs.CL Papers

Summary

Introduces SocialPersona, a benchmark for evaluating multimodal large language models on their ability to recover revealed preferences from longitudinal social-media timelines and use them in personalized dialogue.

arXiv:2606.26654v1 Announce Type: new Abstract: Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability -- inferring what users care about from the multimodal traces they naturally leave behind. We introduce SocialPersona, a benchmark for evaluating whether multimodal large language models (MLLMs) can recover revealed preferences from longitudinal social-media timelines and use them in dialogue. Built from longitudinal timelines of 171 everyday, non-promotional social-media users, SocialPersona contains text, images, timestamps, and 2,597 human-verified preference tags across seven interest domains, separating stable interests from recent interests. It supports two tasks: constructing structured user profiles from multimodal context and generating responses aligned with inferred profiles. Experiments with proprietary and open-weight MLLMs show that models can identify broad interest domains, yet their performance drops on fine-grained and recent interests and degrades further when inferred profiles must be used to personalize dialogue. Together with evidence that text and images provide complementary preference signals, these results indicate that robust cross-modal, long-horizon user modeling remains a key challenge, and that SocialPersona can help measure and advance progress toward assistants that infer and act on revealed preferences.
Original Article
View Cached Full Text

Cached at: 06/26/26, 05:18 AM

# SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
Source: [https://arxiv.org/html/2606.26654](https://arxiv.org/html/2606.26654)
Qinkai Zhang1, Yanyan Zhao1, Xin Lu1, Yulin Hu1, Pengtao Han1, Bing Qin1 1Harbin Institute of Technology \{qkzhang, yyzhao\}@ir\.hit\.edu\.cn

###### Abstract

Personalized language\-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability—inferring what users care about from the multimodal traces they naturally leave behind\. We introduceSocialPersona, a benchmark for evaluating whether multimodal large language models \(MLLMs\) can recover revealed preferences from longitudinal social\-media timelines and use them in dialogue\. Built from longitudinal timelines of 171 everyday, non\-promotional social\-media users,SocialPersonacontains text, images, timestamps, and 2,597 human\-verified preference tags across seven interest domains, separating stable interests from recent interests\. It supports two tasks: constructing structured user profiles from multimodal context and generating responses aligned with inferred profiles\. Experiments with proprietary and open\-weight MLLMs show that models can identify broad interest domains, yet their performance drops on fine\-grained and recent interests and degrades further when inferred profiles must be used to personalize dialogue\. Together with evidence that text and images provide complementary preference signals, these results indicate that robust cross\-modal, long\-horizon user modeling remains a key challenge, and thatSocialPersonacan help measure and advance progress toward assistants that infer and act on revealed preferences\.

SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social\-Media Context

Qinkai Zhang1, Yanyan Zhao1††thanks:Corresponding author\., Xin Lu1, Yulin Hu1, Pengtao Han1, Bing Qin11Harbin Institute of Technology\{qkzhang, yyzhao\}@ir\.hit\.edu\.cn

## 1Introduction

![Refer to caption](https://arxiv.org/html/2606.26654v1/x1.png)Figure 1:A user’s social\-media timeline provides textual, visual, and temporal evidence for stable and recent interests, which can guide personalized responses to new queries\.Personalized assistants are increasingly expected to account for a user’s long\-term interests, recent activities, and implicit preferences\(Chen et al\.,[2024](https://arxiv.org/html/2606.26654#bib.bib6); Liu et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib25); Purificato et al\.,[2024](https://arxiv.org/html/2606.26654#bib.bib35)\)\. Existing benchmarks, however, mainly test whether models remember preferences explicitly stated in dialogue, emphasizing*memory*rather than*insight*\. In practice, preferences are often revealed indirectly through what users create, share, photograph, discuss, and repeatedly engage with\(He et al\.,[2023](https://arxiv.org/html/2606.26654#bib.bib15); Huang et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib16)\)\. Social\-media timelines offer a rich source of such signals, but recovering them requires aggregating multimodal evidence over time, distinguishing stable hobbies from recent fixations, and applying the inferred profile in personalized interaction\. Figure[1](https://arxiv.org/html/2606.26654#S1.F1)illustrates how timeline evidence can be transformed into stable and recent interests for dialogue\.

Prior benchmarks mostly represent user context as dialogue\-derived stated preferences\(Salemi et al\.,[2023](https://arxiv.org/html/2606.26654#bib.bib39); Zhao et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib46); Jiang et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib17); Zhao et al\.,[2025b](https://arxiv.org/html/2606.26654#bib.bib47)\)\. Although recent work incorporates longer behavioral histories\(Huang et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib16)\), it still relies on synthetic or structured textual logs\. These settings bypass a key challenge for MLLMs: inferring user interests from noisy, unstructured, longitudinal social\-media traces, where evidence is weak, distributed across posts, and often available only through images, timestamps, or cross\-post patterns\.

We introduceSocialPersona, a benchmark for evaluating*MLLM personalization from multimodal social\-media context*\. Built from real timelines of everyday, non\-promotional users,SocialPersonacontains chronologically organized text, images, and timestamps\. From these timelines, we construct human\-validated interest profiles across seven domains: sports and outdoor activities, entertainment, gaming, food and drink, travel and city exploration, photography and creation, and pets\. Each profile separates stable interests from recent interests and grounds them in supporting evidence\.

SocialPersonasupports two evaluation settings\. In*profile construction*, models infer active domains and fine\-grained interest tags from raw multimodal timelines\. In*personalized dialogue generation*, models receive social\-media context with a current request and are evaluated on whether their responses align with the user’s stable or recent interests\. Together, these tasks test whether MLLMs can both recover implicit preferences and use them in downstream interaction\.

Concretely,SocialPersonacontains timelines from 171 real users, with an average of 176\.81 posts and 130\.38 images per user\. A semi\-automated pipeline followed by human verification yields 2,597 preference tags grounded in textual, visual, and temporal evidence\. Experiments with proprietary and open\-weight MLLMs show that current models still struggle to infer fine\-grained and recent interests, and to consistently use inferred profiles in personalized responses\.

Our contributions are three\-fold:

1. 1\.We introduce a new task formulation that challenges MLLMs to infer user preferences from longitudinal, multimodal social\-media behavior—aggregating sparse textual, visual, and temporal signals across long horizons—and to apply the inferred preferences in personalized dialogue generation\.
2. 2\.We constructSocialPersona, a real\-user benchmark with long\-horizon timelines, multimodal evidence, timestamps, and human\-validated profiles across seven domains, and publicly release benchmark code with a de\-identified evaluation subset111Available at[https://anonymous\.4open\.science/r/socialpersona\-6E9B](https://anonymous.4open.science/r/socialpersona-6E9B)\. Original images are excluded to reduce re\-identification risk\. Qualified researchers may request controlled access to the full benchmark for replication and follow\-up studies; please contact the authors atqkzhang@ir\.hit\.edu\.cnfor details\.\.
3. 3\.We evaluate proprietary and open\-weight MLLMs on profile construction and personalized dialogue generation, revealing gaps in cross\-modal evidence aggregation and user\-aligned response generation\.

## 2Related Work

BenchmarkContext sourceMulti\-modalRealDataRevealedPref\.ProfileEvalDialogueEvalLaMP\(Salemi et al\.,[2023](https://arxiv.org/html/2606.26654#bib.bib39)\)user text history×\\times✓×\\times×\\times×\\timesPrefEval\(Zhao et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib46)\)dialogue history×\\times×\\times×\\times×\\times✓PERSONAMEM\(Jiang et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib17),[b](https://arxiv.org/html/2606.26654#bib.bib18)\)dialogue history✓×\\times×\\times✓✓Mem\-PAL\(Huang et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib16)\)behavioral logs \+ dialogue×\\times×\\times✓✓✓ALPBench\(Ren et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib38)\)e\-commerce behavior×\\times✓✓✓×\\timesGISTBench\(Fostiropoulos et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib9)\)short\-video engagement×\\times×\\times✓✓×\\timesSocialPersona\(ours\)user social timeline✓✓✓✓✓Table 1:Comparison of personalization benchmarks across context source, modality, data provenance, preference source, and evaluation target\. “Real Data” denotes organically accumulated user\-generated evidence; “Revealed Pref\.” denotes preference signals inferred from behavioral traces\.SocialPersonais the only benchmark covering multimodal real\-user social timelines, revealed preferences, and both profile and dialogue evaluation\.### 2\.1Personalization Benchmarks

Recent personalization benchmarks mainly construct user context from dialogue histories or structured behavior logs\. LaMP\(Salemi et al\.,[2023](https://arxiv.org/html/2606.26654#bib.bib39)\)evaluates personalized language tasks from user\-specific textual histories, while PrefEval\(Zhao et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib46)\), PersonaMem\(Jiang et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib17)\), and PersonaLens\(Zhao et al\.,[2025b](https://arxiv.org/html/2606.26654#bib.bib47)\)study preference recognition, user memory, and personalized response generation from conversational context\. More recent benchmarks move toward longer\-term behavioral modeling, including Mem\-PAL\(Huang et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib16)\)for behavioral\-log\-grounded dialogue and ALPBench\(Ren et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib38)\)/ GISTBench\(Fostiropoulos et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib9)\)for e\-commerce or short\-video interest inference\. Agent\-oriented benchmarks further extend personalization to search, web, and mobile environments\(Kim et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib21); Cai et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib4); Kim et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib22); Yang et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib42); Chen et al\.,[2026](https://arxiv.org/html/2606.26654#bib.bib7)\)\.

As summarized in Table[1](https://arxiv.org/html/2606.26654#S2.T1),SocialPersonadiffers from prior benchmarks by combining multimodal input, real\-user data, revealed\-preference signals, profile evaluation, and dialogue evaluation in one setting\.

### 2\.2Multimodal Social\-media Understanding

Prior multimodal social\-media datasets study content\-level tasks such as sentiment and affect analysis\(Niu et al\.,[2016](https://arxiv.org/html/2606.26654#bib.bib29); Yu and Jiang,[2019](https://arxiv.org/html/2606.26654#bib.bib44); Sharma et al\.,[2020](https://arxiv.org/html/2606.26654#bib.bib40)\), sarcasm and humor detection\(Cai et al\.,[2019](https://arxiv.org/html/2606.26654#bib.bib5)\), crisis response\(Alam et al\.,[2018](https://arxiv.org/html/2606.26654#bib.bib1)\), misinformation verification\(Shu et al\.,[2020](https://arxiv.org/html/2606.26654#bib.bib41); Nakamura et al\.,[2020](https://arxiv.org/html/2606.26654#bib.bib27); Nielsen and McConville,[2022](https://arxiv.org/html/2606.26654#bib.bib28); Mishra et al\.,[2022](https://arxiv.org/html/2606.26654#bib.bib26); Yao et al\.,[2023](https://arxiv.org/html/2606.26654#bib.bib43)\), harmful\-content recognition\(Kiela et al\.,[2021](https://arxiv.org/html/2606.26654#bib.bib20); Lin et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib24)\), and broad MLLM evaluation on social\-networking scenarios\(Zhang et al\.,[2024](https://arxiv.org/html/2606.26654#bib.bib45); Jin et al\.,[2024](https://arxiv.org/html/2606.26654#bib.bib19); Guo et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib14)\)\. However, these benchmarks primarily label individual posts or interactions for predefined tasks\. User\-related signals, when included, are usually treated as demographic attributes, engagement prediction, or recommendation targets\.SocialPersonainstead treats a user’s timeline as external personalization context: models must aggregate sparse textual, visual, and temporal evidence across many posts, distinguish stable from recent interests, and generate responses aligned with the inferred profile\.

## 3SocialPersonaConstruction

### 3\.1Problem Setting

We study whether MLLMs can infer and use preferences from social\-media timelines\. For each useruu, a temporally ordered timeline𝒮u=⟨p1,…,pn⟩\\mathcal\{S\}\_\{u\}=\\langle p\_\{1\},\\dots,p\_\{n\}\\rangleconsists of postspi=\(xi,vi,τi\)p\_\{i\}=\(x\_\{i\},v\_\{i\},\\tau\_\{i\}\)with textxix\_\{i\}, visualsviv\_\{i\}, and timestampτi\\tau\_\{i\}, spanning at most 200 posts over two years\.

We define profiles over seven interest domains adapted from prior preference taxonomies\(Zhao et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib46)\)and platform\-level interest categories:\{\\\{sports\_outdoor, entertainment, gaming, food\_drink, travel\_city\_exploration, photography\_creation, pets\}\\\}\. For each active domain, the gold profile contains stable interests \(recurring patterns across the timeline\), recent interests \(emerging or time\-local signals near the end of the observation window\), and supporting evidence links retained for auditability\. We exclude demographic, identity\-related, health, political, and other sensitive attributes\.

SocialPersonasupports two tasks\. Inprofile construction, a model predicts stable and recent interest tags from the user timeline\. Inpersonalized dialogue generation, a model receives the timeline together with a natural user request and generates a response aligned with the user’s stable or recent interests\. The overall construction and evaluation pipeline is shown in Figure[2](https://arxiv.org/html/2606.26654#S3.F2):SocialPersonafirst converts raw social\-media timelines into human\-verified stable and recent interest profiles, and then evaluates whether MLLMs can recover these profiles and use them in personalized dialogue\.

![Refer to caption](https://arxiv.org/html/2606.26654v1/x2.png)Figure 2:Overview of SOCIALPERSONA\.SocialPersonais constructed from real multimodal social\-media timelines through user filtering, post\-level interest extraction, cross\-post aggregation, temporal profiling, LLM calibration, and human verification, yielding gold profiles with stable and recent interests\. The benchmark evaluates MLLMs on two tasks: inferring user profiles from social media timelines, measured by domain activation and interest\-tag F1, and generating personalized dialogue responses for stable\-interest recommendation and recent\-interest exploration, judged by interest coverage, concreteness, and fluency\.
### 3\.2User and Timeline Collection

We constructSocialPersonafrom real social\-media timelines of*long\-tail organic users*, rather than celebrities, brand accounts, or highly curated public profiles\. This design choice is intended to capture relatively natural, self\-expressive preference traces instead of broadcast\-oriented content\. Starting from 8,000 candidate accounts, we apply automatic filters based on follower count, follower–followee ratio, and image trace density, retaining accounts with 5–5,000 followers, FFR in\[0\.5,2\]\[0\.5,2\], and ITDR≥0\.3\\geq 0\.3\. These filters remove extremely sparse, highly public, or insufficiently multimodal accounts while reducing the presence of broadcaster\-style users\(Oshimo et al\.,[2022](https://arxiv.org/html/2606.26654#bib.bib34); Leavitt et al\.,[2009](https://arxiv.org/html/2606.26654#bib.bib23)\)\. We then manually inspect the remaining accounts to exclude commercial, repost\-heavy, or otherwise low\-quality cases, resulting in 250 candidate users\. Detailed definitions of FFR, ITDR, and the manual filtering criteria are provided in Appendix[B](https://arxiv.org/html/2606.26654#A2)\.

For each selected user, we collect up to 200 posts from the most recent two\-year window\. Each post is stored as a structured multimodal record containing its timestamp, textual content, hashtags, URLs, and attached visual content, including images or video cover frames\. As the original posts come from users across multiple countries and languages, we standardize all textual content by translating it into English, thereby enabling consistent profile construction and evaluation\. After profile construction and verification, we further remove users with fewer than three active interest domains, as such profiles provide insufficient personalization signals for reliable evaluation\. This yields the final benchmark of 171 users\.

### 3\.3Gold Profile Construction

Given each user’s multimodal timeline, we construct gold profiles with an LLM\-assisted but human\-verified pipeline\. First, we useGemini\-3\-Flash\(Google DeepMind,[2025](https://arxiv.org/html/2606.26654#bib.bib12); Google AI for Developers,[2026a](https://arxiv.org/html/2606.26654#bib.bib10)\)to perform conservative post\-level extraction from the original text and visual content\. The extractor is instructed to identify only observable, evidence\-grounded interest signals, record modality attribution, and avoid demographic, identity\-related, or speculative claims\.

Second, we aggregate post\-level candidates across each timeline\. Near\-duplicate posts are down\-weighted, semantically equivalent tags are merged into canonical interests, and each user\-domain pair is represented as an evidence pack\. Third, we compute preliminary stable and recent assignments using duplicate\-adjusted support, temporal dispersion, recency, and confidence\. Stable interests require repeated support across the timeline, while recent interests emphasize evidence concentrated in the most recent 90 days\.

Fourth, we apply an LLM\-based calibration stage usingGemini\-3\.1\-Pro\(Google DeepMind,[2026](https://arxiv.org/html/2606.26654#bib.bib13); Google AI for Developers,[2026b](https://arxiv.org/html/2606.26654#bib.bib11)\)\. This stage checks the aggregated evidence packs, canonical candidates, scores, and preliminary temporal buckets for weak support, over\-generalization, speculative labels, and bucket errors\. Finally, trained annotators manually verify each calibrated profile against its supporting posts\. Accepted interests must be concrete, domain\-appropriate, sufficiently supported, and assigned to the correct temporal bucket\.

#### Quality assurance\.

Five trained annotators independently verified each user profile, with disagreements resolved by an additional adjudicator\. The process required approximately 350 annotator\-hours\. On a 40\-user overlap subset, inter\-annotator agreement reached Krippendorff’sα=0\.72\\alpha=0\.72for tag acceptance andα=0\.63\\alpha=0\.63for stable/recent bucket assignment\. Human verification modified about 12% of pipeline\-proposed tags: 8% removed for insufficient evidence, 3% refined to more concrete labels, and 1% added as missed but supported interests, confirming that human oversight is essential for benchmark quality\.

## 4Evaluation and Experiments

### 4\.1Evaluation Setup

All experiments use a fixed 100\-user subset to ensure comparable cost and conditions\. We evaluate profile construction and personalized dialogue generation using a timeline representation consisting of post text, image captions, and timestamps\.

#### Evaluated models\.

Our main model suite is designed to cover both proprietary and open\-weight MLLMs\. The initial suite contains six models: Gemini\-2\.5\-Flash\(Comanici et al\.,[2025](https://arxiv.org/html/2606.26654#bib.bib8)\), GPT\-4o\-mini\(OpenAI,[2024](https://arxiv.org/html/2606.26654#bib.bib30)\), GPT\-5\.4\(OpenAI,[2026a](https://arxiv.org/html/2606.26654#bib.bib32)\), Qwen2\.5\-VL\-7B\-Instruct\(Bai et al\.,[2025b](https://arxiv.org/html/2606.26654#bib.bib3)\), Qwen3\-VL\-8B\-Instruct\(Bai et al\.,[2025a](https://arxiv.org/html/2606.26654#bib.bib2)\), and Qwen3\.5\-35B\-A3B\(Qwen Team,[2026a](https://arxiv.org/html/2606.26654#bib.bib36)\)\.

#### Profile construction evaluation\.

Given𝒮~uM\\widetilde\{\\mathcal\{S\}\}\_\{u\}^\{M\}, each method predicts stable and recent interests over seven domains\. A domain is active if it has at least one gold interest, and predicted active if the method outputs any interest in it\. For interest\-tag recovery, we evaluate each user, domain, and bucketb∈\{stable,recent\}b\\in\\\{\\mathrm\{stable\},\\mathrm\{recent\}\\\}using normalized exact match, followed byo3\(OpenAI,[2025](https://arxiv.org/html/2606.26654#bib.bib31)\)\-based semantic matching for unmatched tags as in Appendix[A\.3](https://arxiv.org/html/2606.26654#A1.SS3)\. Matches are one\-to\-one and define true positives; unmatched predictions and gold tags are false positives and false negatives\.

#### Personalized dialogue evaluation\.

The dialogue task evaluates whether a model can generate personalized but not over\-personalized recommendations from social\-media\-derived user information\. We evaluate four dialogue input settings\. The first is a*timeline\-conditioned*setting, where the model receives the user’s post text, image captions, and timestamps directly\. The remaining three are two\-stage*profile\-conditioned*settings: the model first constructs a profile using one of the three profile\-construction settings described below, and then generates a dialogue response using only that generated profile as personalization context\. These settings test whether a model\-generated profile can serve as an effective intermediate memory representation, rather than only being evaluated as a structured prediction\.

For each dialogue setting, we consider two intents: stable\-interest recommendation, which targets the user’s long\-term preferences, and recent\-interest exploration, which introduces mildly novel yet personally suitable items\. For each setting, we prepare ten natural requests and randomly sample one for each user–setting instance; the full pools are given in Appendix[A\.5](https://arxiv.org/html/2606.26654#A1.SS5)\. We useGPT\-5\.5\(OpenAI,[2026b](https://arxiv.org/html/2606.26654#bib.bib33)\)andQwen3\.7\-Max\(Qwen Team,[2026b](https://arxiv.org/html/2606.26654#bib.bib37)\)as judges, following the prompt in Appendix[A\.6](https://arxiv.org/html/2606.26654#A1.SS6)\. Given the gold profile, the judge scores each response on*interest coverage*,*concreteness*, and*fluency*\.

### 4\.2Main Results: Profile Construction Across Settings and Models

We evaluate profile construction across three settings and six models \(Table[4\.2](https://arxiv.org/html/2606.26654#S4.SS2.SSS0.Px1)\), separating model capability from the strategy used to handle long timelines\.

#### Profile\-construction settings\.

*Direct*feeds the full timeline \(text, image captions, timestamps, up to 200 posts\) in one pass\.*Hierarchical*splits the timeline into 20\-post chunks, summarizes each independently, then aggregates chunk summaries into the final profile\.*Extractive–abstractive*proceeds in two stages as detailed in Appendix[A\.9](https://arxiv.org/html/2606.26654#A1.SS9)\. First \(*extraction*\), the LLM is prompted to select up toK=12K=12representative posts per domain, guided by relevance, specificity, recurrence, recency, and multimodal grounding\. Second \(*abstractive synthesis*\), the LLM generates the profile using only the selected posts as evidence, with no access to the original full timeline\.

ModelActive F1Int\. Prec\.Int\. Rec\.Int\. F1Stable F1Recent F1\\rowcolortabgrayDirectGemini\-2\.5\-Flash0\.86630\.53230\.23790\.32880\.35160\.0219GPT\-4o\-mini0\.74540\.46270\.19920\.27850\.18320\.0858GPT\-5\.40\.83830\.60380\.23070\.33380\.34050\.0687Qwen2\.5\-VL\-7B\-Instruct0\.06470\.46430\.00850\.01670\.02280\.0000Qwen3\-VL\-8B\-Instruct0\.59690\.36360\.16250\.22460\.21480\.0390Qwen3\.5\-35B\-A3B0\.78750\.54900\.17630\.26690\.29890\.0133\\rowcolortabgrayHierarchicalGemini\-2\.5\-Flash0\.73370\.55560\.04260\.07910\.06630\.0300GPT\-4o\-mini0\.82600\.42010\.26870\.32770\.32830\.0653GPT\-5\.40\.82950\.51210\.34800\.41440\.39320\.1202Qwen2\.5\-VL\-7B\-Instruct0\.73950\.38180\.11530\.17720\.15490\.0291Qwen3\-VL\-8B\-Instruct0\.80380\.39710\.27060\.32190\.28020\.1129Qwen3\.5\-35B\-A3B0\.80660\.45560\.23200\.30740\.34090\.0679\\rowcolortabgrayExtractive–abstractiveGemini\-2\.5\-Flash0\.29710\.59300\.03340\.06330\.05140\.0270GPT\-4o\-mini0\.82490\.42970\.14420\.21590\.01950\.0746GPT\-5\.40\.84280\.45000\.32110\.37480\.33880\.1102Qwen2\.5\-VL\-7B\-Instruct0\.37350\.27170\.01640\.03090\.00530\.0123Qwen3\-VL\-8B\-Instruct0\.52470\.32590\.04780\.08340\.06400\.0188Qwen3\.5\-35B\-A3B0\.85990\.38290\.20580\.26770\.23490\.0906
Table 2:Profile construction results across three construction settings and six models\. Best results within each setting arebolded\.ModelCov\.Conc\.Flu\.Avg\.GPT\-5\.5Qwen3\.7GPT\-5\.5Qwen3\.7GPT\-5\.5Qwen3\.7GPT\-5\.5Qwen3\.7\\rowcolortabgrayTimelineGemini\-2\.5\-Flash1\.88002\.42003\.17003\.12003\.84004\.20002\.96333\.2467GPT\-4o\-mini1\.97002\.25003\.51002\.87003\.92003\.89003\.13333\.0033GPT\-5\.42\.36002\.81004\.02004\.16004\.32004\.73003\.56673\.9000Qwen2\.5\-VL\-7B\-Instruct1\.63001\.78003\.04002\.40002\.97002\.70002\.54672\.2933Qwen3\-VL\-8B\-Instruct2\.22002\.55003\.73003\.52003\.96004\.23003\.30333\.4333Qwen3\.5\-35B\-A3B1\.63002\.03003\.42003\.12003\.84003\.99002\.96333\.0467\\rowcolortabgrayDirectGemini\-2\.5\-Flash1\.64002\.06002\.63002\.28003\.76003\.87002\.67672\.7367GPT\-4o\-mini1\.95002\.23003\.26002\.74003\.82003\.96003\.01002\.9767GPT\-5\.42\.54003\.04003\.92003\.94004\.19004\.65003\.55003\.8767Qwen2\.5\-VL\-7B\-Instruct0\.88001\.08002\.47001\.83003\.15002\.95002\.16671\.9533Qwen3\-VL\-8B\-Instruct2\.56002\.80003\.57003\.24004\.03004\.29003\.38673\.4433Qwen3\.5\-35B\-A3B1\.93002\.33003\.02002\.65003\.79004\.01002\.91332\.9967\\rowcolortabgrayHierarchicalGemini\-2\.5\-Flash1\.40001\.61002\.63002\.15003\.68003\.96002\.57002\.5733GPT\-4o\-mini2\.03002\.21003\.27002\.61003\.89003\.91003\.06332\.9100GPT\-5\.42\.50002\.89003\.81003\.72004\.14004\.50003\.48333\.7033Qwen2\.5\-VL\-7B\-Instruct1\.71001\.77002\.97002\.22003\.37003\.20002\.68332\.3967Qwen3\-VL\-8B\-Instruct2\.07002\.23003\.29002\.96003\.94003\.96003\.10003\.0500Qwen3\.5\-35B\-A3B1\.88002\.22003\.12002\.69003\.75003\.93002\.91672\.9467\\rowcolortabgrayExtractive–abstractiveGemini\-2\.5\-Flash0\.93001\.22002\.48002\.04003\.58003\.85002\.33002\.3700GPT\-4o\-mini1\.62002\.01003\.20002\.54003\.84003\.90002\.88672\.8167GPT\-5\.42\.54502\.97004\.03503\.89004\.26504\.62003\.61503\.8267Qwen2\.5\-VL\-7B\-Instruct1\.24001\.37002\.67001\.94003\.25003\.03002\.38672\.1133Qwen3\-VL\-8B\-Instruct1\.53001\.68003\.08002\.64003\.97004\.00002\.86002\.7733Qwen3\.5\-35B\-A3B1\.93002\.19003\.20002\.76003\.79003\.98002\.97332\.9767
Table 3:Personalized dialogue generation results\. Profile\-conditioned settings use generated profiles from the corresponding construction setting in Table[4\.2](https://arxiv.org/html/2606.26654#S4.SS2.SSS0.Px1)\.#### Over\-generalization trumps recall\.

The dominant failure mode across all models and settings is systematic over\-generalization: models reduce 4–7 gold interest tags to 1–2 broad categories per active domain\. Averaged over the direct setting, gold profiles contain 3\.4 tags per active domain, while models predict only 1\.5\. This gap is consistent across all evaluated models and represents a fundamental limitation of current MLLMs in fine\-grained preference elicitation from social\-media evidence\. A prompt analysis in Appendix[E](https://arxiv.org/html/2606.26654#A5)confirms that this over\-generalization is not an artifact of the evaluation prompt: removing the conservative instruction to “prefer fewer, broader tags” does not materially change Interest F1, indicating that the recall gap reflects genuine model limitations rather than benchmark design\.

#### Domain blindness follows evidence modality\.

Domain detection varies with evidence modality\. Text\-dispersed interests, such as music habits, city exploration, and meals, are often missed: in Appendix[G](https://arxiv.org/html/2606.26654#A7), most models incorrectly mark text\-heavy travel and food\-drink domains as inactive\. In contrast, visually concentrated domains, such as pets and gaming screenshots, obtain the highest F1 scores across settings \(Table[13](https://arxiv.org/html/2606.26654#A9.T13)\)\. This asymmetry indicates a reliance on visual salience for domain activation: models reliably activate domains when images directly show relevant content, but often default to inactive when evidence is diffuse and textual\. Even GPT\-5\.4, despite the best overall tag F1, misses 12\.5% of gold\-active domains under direct profiling\.

#### Recent\-interest blindness is a temporal reasoning gap\.

Stable F1 exceeds Recent F1 across every model, setting, and input mode\. This is not simply because recent interests are “harder”—it reflects a fundamental limitation in how models process timelines\. Distinguishing an interest that recurs across 18 months \(stable\) from one concentrated in the most recent 90 days \(recent\) requires tracking temporal dispersion across posts\. Current models receive posts as a flattened sequence; timestamps are present in the input but models show no evidence of computing temporal distribution patterns from them\. Gemini\-2\.5\-Flash and Qwen3\.5\-35B\-A3B assign nearly all predicted tags to the stable bucket regardless of the actual temporal evidence, while GPT\-4o\-mini—the only model with non\-trivial Recent F1 \(0\.086\)—distributes tags more evenly between buckets but without temporal alignment, suggesting its higher Recent F1 reflects a flatter prior over buckets rather than genuine temporal discrimination\.

#### Hierarchical profiling amplifies model\-specific behaviors\.

Hierarchical profiling improves Interest F1 for strong models \(GPT\-5\.4, GPT\-4o\-mini\) but causes conservative models to collapse, as sparse chunk\-level summaries leave the aggregation step with insufficient evidence\. The same pattern recurs in the extractive setting: capable models benefit from the two\-stage decomposition while weaker models degrade further\. Detailed per\-model breakdowns are provided in Appendix[G](https://arxiv.org/html/2606.26654#A7)\.

### 4\.3Main Results: Personalized Dialogue Generation

Table[4\.2](https://arxiv.org/html/2606.26654#S4.SS2.SSS0.Px1)reports dialogue results across four input settings and six models\. The*timeline\-conditioned*setting provides the raw timeline directly\. The other three are*profile\-conditioned*: the model first constructs a profile using one of the three settings from Table[4\.2](https://arxiv.org/html/2606.26654#S4.SS2.SSS0.Px1), then generates a response using only that profile\. We report three judge\-scored dimensions \(0–5 scale\) and their unweighted averageAvg=\(Coverage\+Concreteness\+Fluency\)/3\\mathrm\{Avg\}=\(\\mathrm\{Coverage\}\+\\mathrm\{Concreteness\}\+\\mathrm\{Fluency\}\)/3\.

#### Profile conditioning is a double\-edged filter\.

Profile conditioning filters timeline noise but also propagates profiling errors\. When generated profiles are sparse, personalization coverage and concreteness drop; when profiles are richer, the intermediate profile can improve dialogue quality\. For example, Gemini\-2\.5\-Flash loses coverage under profile conditioning, while GPT\-5\.4 slightly improves coverage, suggesting that profile quality determines whether the profile acts as a useful memory representation or an information bottleneck\. The profile–dialogue correlation analysis \(Section[4\.4](https://arxiv.org/html/2606.26654#S4.SS4)\) confirms this dependency quantitatively, with 16 of 18 Interest F1–Avg\. correlations positive and stronger under the tighter information bottleneck of hierarchical and extractive profiling\. Fluency remains relatively stable across settings and models \(3\.0–4\.3 for most\), indicating that the core challenge is not generating readable text but producing responses that engage the correct interests\.

#### Human–LLM agreement\.

A 30\-response human agreement study confirms that the three dialogue dimensions are reliably judgeable and that both LLM judges track human rankings\. Inter\-annotator agreement is substantial for fluency, concreteness, and coverage \(α=0\.79/0\.71/0\.65\\alpha=0\.79/0\.71/0\.65\), and GPT\-5\.5 correlates most strongly with human mean scores \(ρ=0\.76/0\.65/0\.58\\rho=0\.76/0\.65/0\.58, allp<0\.01p<0\.01\)\. Full sampling and annotation details are in Appendix[J](https://arxiv.org/html/2606.26654#A10)\.

### 4\.4Additional Analyses

#### Input\-modality analysis\.

Input modality analysis on Qwen3\.5\-35B\-A3B and GPT\-4o\-mini shows that text provides the main profiling signal, while image captions supply complementary tacit cues: adding captions raises GPT\-4o\-mini’s Active F1 from 0\.503 to 0\.737\. Timestamps further improve Active F1 but leave Tag F1 largely unchanged, suggesting that temporal structure helps domain detection more than fine\-grained tag recovery\.

#### Per\-domain difficulty\.

Per\-domain breakdown \(Table[13](https://arxiv.org/html/2606.26654#A9.T13)\) confirms that domains with visually distinctive evidence \(pets, gaming\) are consistently easier across settings, while domains requiring synthesis of dispersed evidence \(travel, entertainment\) are hardest\.

#### Profile–dialogue correlation\.

The two\-stage design enables a direct test of whether profile quality translates to dialogue quality\. Per\-user Spearman rank correlations show that profile Interest F1 positively correlates with dialogue quality: 16 of 18 Interest F1–Avg\. correlations are positive \(6 significant atp<0\.05p<0\.05\), with stronger correlations under Hierarchical and Extractive profiling, where the profile acts as a tighter information bottleneck\. Interest F1 correlates most strongly with coverage \(15 of 18 pairs significant\), confirming that better profiling primarily improves interest engagement\.

## 5Conclusion

We introducedSocialPersona, a benchmark for personalized user profiling and response generation from multimodal social\-media timelines, built from real social\-media user data with human\-validated interest profiles across seven domains\. Our experiments reveal three central findings\. First, current models systematically over\-generalize, collapsing specific interests into broad categories—a genuine capability gap, not an artifact of conservative evaluation instructions \(Appendix[E](https://arxiv.org/html/2606.26654#A5)\)\. Second, models exhibit pronounced modality asymmetry: visually salient domains are reliably detected while text\-distributed domains are frequently missed, and this bias propagates into dialogue\. Third, distinguishing stable from recent interests remains beyond current capabilities; models lack the cross\-post temporal reasoning needed to separate persistent patterns from emerging ones\. The profile–dialogue correlation validates the two\-stage design but reveals that sparse profiles compound these failures downstream\.

## Limitations

Our evaluation uses image captions rather than raw images because full user timelines can contain hundreds of images, making raw\-image evaluation difficult under heterogeneous API limits, visual\-token budgets, and upload interfaces\. Captions provide a unified cross\-model input format but may lose pixel\-level details\. The modality analysis shows that captions still add substantial signal alongside text for the interest categories studied here\.

## Ethical Considerations

SocialPersonais designed as an aggregate benchmark for studying personalized modeling from publicly accessible social\-media content under a restricted research protocol\. Annotation and evaluation are limited to non\-sensitive, evidence\-grounded interests, and explicitly avoid demographic, identity\-related, health, political, or other sensitive inferences\. We release only a de\-identified text\-plus\-caption subset, exclude original images, and provide the full benchmark only through controlled research access\. The benchmark is intended solely for aggregate model evaluation and research purposes, and must not be used for identification, surveillance, targeting, or consequential decision\-making\.

## References

- Alam et al\. \(2018\)Firoj Alam, Ferda Ofli, and Muhammad Imran\. 2018\.CrisisMMD: Multimodal Twitter datasets from natural disasters\.In*Proceedings of the 12th International AAAI Conference on Web and Social Media*\.
- Bai et al\. \(2025a\)Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, and 1 others\. 2025a\.Qwen3\-VL technical report\.*arXiv preprint arXiv:2511\.21631*\.
- Bai et al\. \(2025b\)Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others\. 2025b\.Qwen2\.5\-VL technical report\.*arXiv preprint arXiv:2502\.13923*\.
- Cai et al\. \(2025\)Hongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu, Xiaoyu Shen, Wenjie Li, and Tat\-Seng Chua\. 2025\.[Large language models empowered personalized web agents](https://doi.org/10.1145/3696410.3714842)\.In*Proceedings of the ACM Web Conference 2025*, pages 198–215\. ACM\.
- Cai et al\. \(2019\)Yitao Cai, Huiyu Cai, and Xiaojun Wan\. 2019\.[Multi\-modal sarcasm detection in Twitter with hierarchical fusion model](https://doi.org/10.18653/v1/P19-1239)\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pages 2506–2515, Florence, Italy\. Association for Computational Linguistics\.
- Chen et al\. \(2024\)Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, Kai Zheng, Defu Lian, and Enhong Chen\. 2024\.[When large language models meet personalization: Perspectives of challenges and opportunities](https://doi.org/10.1007/s11280-024-01276-1)\.*World Wide Web*, 27\(4\):42\.
- Chen et al\. \(2026\)Tongbo Chen, Zhengxi Lu, Zhan Xu, Guocheng Shao, Shaohan Zhao, Fei Tang, Yong Du, Kaitao Song, Yizhou Liu, Yuchen Yan, Wenqi Zhang, Xu Tan, Weiming Lu, Jun Xiao, Yueting Zhuang, and Yongliang Shen\. 2026\.[KnowU\-Bench: Towards interactive, proactive, and personalized mobile agent evaluation](https://doi.org/10.48550/arXiv.2604.08455)\.*Preprint*, arXiv:2604\.08455\.
- Comanici et al\. \(2025\)Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others\. 2025\.Gemini 2\.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.*arXiv preprint arXiv:2507\.06261*\.
- Fostiropoulos et al\. \(2026\)Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan, Boyu Fang, Yuchen Liu, Jiayi Liu, Hanchao Yu, Qi Guo, Jianyu Wang, Fei Liu, and Xiangjun Fan\. 2026\.[GISTBench: Evaluating LLM user understanding via evidence\-based interest verification](https://doi.org/10.48550/arXiv.2603.29112)\.*Preprint*, arXiv:2603\.29112\.
- Google AI for Developers \(2026a\)Google AI for Developers\. 2026a\.Gemini 3 flash preview\.[https://ai\.google\.dev/gemini\-api/docs/models/gemini\-3\-flash\-preview](https://ai.google.dev/gemini-api/docs/models/gemini-3-flash-preview)\.Accessed: 2026\-05\-25\.
- Google AI for Developers \(2026b\)Google AI for Developers\. 2026b\.Gemini 3\.1 pro preview\.[https://ai\.google\.dev/gemini\-api/docs/models/gemini\-3\.1\-pro\-preview](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview)\.Accessed: 2026\-05\-25\.
- Google DeepMind \(2025\)Google DeepMind\. 2025\.Model evaluation: Gemini 3 flash\.[https://storage\.googleapis\.com/deepmind\-media/gemini/gemini\_3\_flash\_model\_evaluation\.pdf](https://storage.googleapis.com/deepmind-media/gemini/gemini_3_flash_model_evaluation.pdf)\.Accessed: 2026\-05\-25\.
- Google DeepMind \(2026\)Google DeepMind\. 2026\.Model evaluation: Gemini 3\.1 pro\.[https://storage\.googleapis\.com/deepmind\-media/gemini/gemini\_3\-1\_pro\_model\_evaluation\.pdf](https://storage.googleapis.com/deepmind-media/gemini/gemini_3-1_pro_model_evaluation.pdf)\.Accessed: 2026\-05\-25\.
- Guo et al\. \(2025\)Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang, Weiting Liu, Anjie Le, Lei Li, and Zhoujun Li\. 2025\.[SNS\-Bench\-VL: Benchmarking multimodal large language models in social networking services](https://doi.org/10.48550/arXiv.2505.23065)\.*arXiv preprint arXiv:2505\.23065*\.Withdrawn\.
- He et al\. \(2023\)Zhicheng He, Weiwen Liu, Wei Guo, Jiarui Qin, Yingxue Zhang, Yaochen Hu, and Ruiming Tang\. 2023\.[A survey on user behavior modeling in recommender systems](https://doi.org/10.24963/ijcai.2023/746)\.In*Proceedings of the Thirty\-Second International Joint Conference on Artificial Intelligence*, pages 6656–6664\.
- Huang et al\. \(2026\)Zhaopei Huang, Qifeng Dai, Guozheng Wu, Xiaopeng Wu, Xubin Li, Tiezheng Ge, Wenxuan Wang, and Qin Jin\. 2026\.Mem\-pal: Towards memory\-based personalized dialogue assistants for long\-term user\-agent interaction\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 40, pages 31229–31237\.
- Jiang et al\. \(2025a\)Bowen Jiang, Zhuoqun Hao, Young\-Min Cho, Bryan Li, Yuan Yuan, Sihao Chen, Lyle Ungar, Camillo J\. Taylor, and Dan Roth\. 2025a\.Know me, respond to me: Benchmarking LLMs for dynamic user profiling and personalized responses at scale\.*arXiv preprint arXiv:2504\.14225*\.
- Jiang et al\. \(2025b\)Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, Radha Poovendran, Gregory Wornell, Lyle Ungar, Dan Roth, Sihao Chen, and Camillo Jose Taylor\. 2025b\.PersonaMem\-v2: Towards personalized intelligence via learning implicit user personas and agentic memory\.*arXiv preprint arXiv:2512\.06688*\.
- Jin et al\. \(2024\)Yiqiao Jin, Minje Choi, Gaurav Verma, Jindong Wang, and Srijan Kumar\. 2024\.MM\-SOC: Benchmarking multimodal large language models in social media platforms\.In*Findings of the Association for Computational Linguistics: ACL 2024*, pages 6192–6210\.
- Kiela et al\. \(2021\)Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Casey A\. Fitzpatrick, Peter Bull, Greg Lipstein, Tony Nelli, Ron Zhu, Niklas Muennighoff, Riza Velioglu, Jewgeni Rose, Phillip Lippe, Nithin Holla, Shantanu Chandra, Santhosh Rajamanickam, Georgios Antoniou, Ekaterina Shutova, and 6 others\. 2021\.[The hateful memes challenge: Competition report](https://proceedings.mlr.press/v133/kiela21a.html)\.In*Proceedings of the NeurIPS 2020 Competition and Demonstration Track*, volume 133 of*Proceedings of Machine Learning Research*, pages 344–360\. PMLR\.
- Kim et al\. \(2025\)Hyunseo Kim, Sangam Lee, Kwangwook Seo, and Dongha Lee\. 2025\.[BESPOKE: Benchmark for search\-augmented large language model personalization via diagnostic feedback](https://doi.org/10.48550/arXiv.2509.21106)\.*Preprint*, arXiv:2509\.21106\.
- Kim et al\. \(2026\)Serin Kim, Sangam Lee, and Dongha Lee\. 2026\.[Persona2Web: Benchmarking personalized web agents for contextual reasoning with user history](https://doi.org/10.48550/arXiv.2602.17003)\.*Preprint*, arXiv:2602\.17003\.
- Leavitt et al\. \(2009\)Alex Leavitt, Evan Burchard, David Fisher, and Sam Gilbert\. 2009\.[The influentials: New approaches for analyzing influence on twitter](https://www.webecologyproject.org/wp-content/uploads/2009/09/influence-report-final.pdf)\.Web Ecology Project\.Publication 04\.
- Lin et al\. \(2025\)Hongzhan Lin, Ziyang Luo, Bo Wang, Ruichao Yang, and Jing Ma\. 2025\.[GOAT\-Bench: Safety insights to large multimodal models through meme\-based social abuse](https://doi.org/10.1145/3729239)\.*ACM Transactions on Intelligent Systems and Technology*\.
- Liu et al\. \(2025\)Jiahong Liu, Zexuan Qiu, Zhongyang Li, Quanyu Dai, Jieming Zhu, Minda Hu, Menglin Yang, and Irwin King\. 2025\.A survey of personalized large language models: Progress and future directions\.*arXiv preprint arXiv:2502\.11528*\.
- Mishra et al\. \(2022\)Shreyash Mishra, S\. Suryavardan, Amrit Bhaskar, Parul Chopra, Aishwarya N\. Reganti, Parth Patwa, Amitava Das, Tanmoy Chakraborty, Amit Sheth, Asif Ekbal, and Chaitanya Ahuja\. 2022\.[FACTIFY: A multi\-modal fact verification dataset](https://ceur-ws.org/Vol-3199/paper18.pdf)\.In*Proceedings of the First Workshop on Multimodal Fact\-Checking and Hate Speech Detection*\.
- Nakamura et al\. \(2020\)Kai Nakamura, Sharon Levy, and William Yang Wang\. 2020\.[Fakeddit: A new multimodal benchmark dataset for fine\-grained fake news detection](https://aclanthology.org/2020.lrec-1.755/)\.In*Proceedings of the Twelfth Language Resources and Evaluation Conference*, pages 6149–6157, Marseille, France\. European Language Resources Association\.
- Nielsen and McConville \(2022\)Dan Saattrup Nielsen and Ryan McConville\. 2022\.[MuMiN: A large\-scale multilingual multimodal fact\-checked misinformation social network dataset](https://doi.org/10.1145/3477495.3531744)\.In*Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval*, pages 3141–3153\. Association for Computing Machinery\.
- Niu et al\. \(2016\)Teng Niu, Shiai Zhu, Lei Pang, and Abdulmotaleb El Saddik\. 2016\.[Sentiment analysis on multi\-view social data](https://doi.org/10.1007/978-3-319-27674-8_2)\.In*MultiMedia Modeling: 22nd International Conference, MMM 2016*, pages 15–27\. Springer\.
- OpenAI \(2024\)OpenAI\. 2024\.GPT\-4o mini: Advancing cost\-efficient intelligence\.[https://openai\.com/index/gpt\-4o\-mini\-advancing\-cost\-efficient\-intelligence/](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/)\.Accessed: 2026\-05\-25\.
- OpenAI \(2025\)OpenAI\. 2025\.OpenAI o3 and o4\-mini System Card\.[https://openai\.com/index/o3\-o4\-mini\-system\-card/](https://openai.com/index/o3-o4-mini-system-card/)\.Accessed: 2026\-05\-25\.
- OpenAI \(2026a\)OpenAI\. 2026a\.GPT\-5\.4 Thinking System Card\.[https://openai\.com/index/gpt\-5\-4\-thinking\-system\-card/](https://openai.com/index/gpt-5-4-thinking-system-card/)\.Accessed: 2026\-05\-25\.
- OpenAI \(2026b\)OpenAI\. 2026b\.GPT\-5\.5 System Card\.[https://openai\.com/index/gpt\-5\-5\-system\-card/](https://openai.com/index/gpt-5-5-system-card/)\.Accessed: 2026\-05\-25\.
- Oshimo et al\. \(2022\)Hiroki Oshimo, Shun Hironaka, Masanori Yoshida, and 1 others\. 2022\.Follower–followee ratio category and user vector for analyzing following behavior\.In*2022 9th International Conference on Advanced Informatics: Concepts, Theory and Applications \(ICAICTA\)*, pages 1–6\. IEEE\.
- Purificato et al\. \(2024\)Erasmo Purificato, Ludovico Boratto, and Ernesto William De Luca\. 2024\.User modeling and user profiling: A comprehensive survey\.*arXiv preprint arXiv:2402\.09660*\.
- Qwen Team \(2026a\)Qwen Team\. 2026a\.Qwen3\.5\-35B\-A3B\.[https://huggingface\.co/Qwen/Qwen3\.5\-35B\-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B)\.Model card\. Accessed: 2026\-05\-25\.
- Qwen Team \(2026b\)Qwen Team\. 2026b\.Qwen3\.7: The agent frontier\.[https://qwen\.ai/blog?id=qwen3\.7](https://qwen.ai/blog?id=qwen3.7)\.Accessed: 2026\-05\-25\.
- Ren et al\. \(2026\)Lu Ren, Junda She, Xinchen Luo, Tao Wang, Xin Ye, Xu Zhang, Muxuan Wang, Xiao Yang, Chenguang Wang, Fei Xie, Yiwei Zhou, Danjun Wu, Guodong Zhang, Yifei Hu, Guoying Zheng, Shujie Yang, Xingmei Wang, Shiyao Wang, Yukun Zhou, and 7 others\. 2026\.[ALPBench: A benchmark for attribution\-level long\-term personal behavior understanding](https://doi.org/10.48550/arXiv.2602.03056)\.*Preprint*, arXiv:2602\.03056\.
- Salemi et al\. \(2023\)Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani\. 2023\.LaMP: When large language models meet personalization\.*arXiv preprint arXiv:2304\.11406*\.
- Sharma et al\. \(2020\)Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas PYKL, Amitava Das, Tanmoy Chakraborty, Viswanath Pulabaigari, and Björn Gambäck\. 2020\.[SemEval\-2020 task 8: Memotion analysis—the visuo\-lingual metaphor\!](https://doi.org/10.18653/v1/2020.semeval-1.99)In*Proceedings of the Fourteenth Workshop on Semantic Evaluation*, pages 759–773, Barcelona, Spain\. International Committee for Computational Linguistics\.
- Shu et al\. \(2020\)Kai Shu, Deepak Mahudeswaran, Suhang Wang, Dongwon Lee, and Huan Liu\. 2020\.[FakeNewsNet: A data repository with news content, social context, and spatiotemporal information for studying fake news on social media](https://doi.org/10.1089/big.2020.0062)\.*Big Data*, 8\(3\):171–188\.
- Yang et al\. \(2026\)Qinglong Yang, Haoming Li, Haotian Zhao, Xiaokai Yan, Jingtao Ding, Fengli Xu, and Yong Li\. 2026\.[FingerTip 20k: A benchmark for proactive and personalized mobile LLM agents](https://openreview.net/forum?id=n3iFV0gLMc)\.In*The Fourteenth International Conference on Learning Representations*\.
- Yao et al\. \(2023\)Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin\-Hee Cho, and Lifu Huang\. 2023\.[End\-to\-end multimodal fact\-checking and explanation generation: A challenging dataset and models](https://doi.org/10.1145/3539618.3591879)\.In*Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval*, pages 2733–2743\. Association for Computing Machinery\.
- Yu and Jiang \(2019\)Jianfei Yu and Jing Jiang\. 2019\.[Adapting BERT for target\-oriented multimodal sentiment classification](https://doi.org/10.24963/ijcai.2019/751)\.In*Proceedings of the Twenty\-Eighth International Joint Conference on Artificial Intelligence, IJCAI\-19*, pages 5408–5414\. International Joint Conferences on Artificial Intelligence Organization\.
- Zhang et al\. \(2024\)Xinnong Zhang, Haoyu Kuang, Xinyi Mou, Hanjia Lyu, Kun Wu, Siming Chen, Jiebo Luo, Xuanjing Huang, and Zhongyu Wei\. 2024\.[SoMeLVLM: A large vision language model for social media processing](https://doi.org/10.18653/v1/2024.findings-acl.140)\.In*Findings of the Association for Computational Linguistics: ACL 2024*, pages 2366–2389, Bangkok, Thailand\. Association for Computational Linguistics\.
- Zhao et al\. \(2025a\)Siyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika, and Kaixiang Lin\. 2025a\.Do LLMs recognize your preferences? Evaluating personalized preference following in LLMs\.*arXiv preprint arXiv:2502\.09597*\.
- Zhao et al\. \(2025b\)Zheng Zhao, Clara Vania, Subhradeep Kayal, Naila Khan, Shay B\. Cohen, and Emine Yilmaz\. 2025b\.PersonaLens: A benchmark for personalization evaluation in conversational AI assistants\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 18023–18055\.

## Appendix APrompt Templates

### A\.1Single\-Post Analysis Prompt

`Single\-Post Analysis — System Prompt Single\-Post Analysis — User Prompt`

`A\.2 Profile Evaluation Model Prompt Profile Evaluation — System Prompt Profile Evaluation — User Prompt A\.3 Profile Evaluation Matching Judge Prompt Profile Matching Judge — System Prompt Profile Matching Judge — User Prompt A\.4 Dialogue Generation Prompt Dialogue Generation — System Prompts Dialogue Generation — User Prompt Templates A\.5 Dialogue User Request Pool For dialogue evaluation, we sample one user request from the corresponding setting\-specific pool\. The stable\-interest pool targets stable preference use, while the recent\-interest pool targets recent exploration that remains compatible with the user’s stable preferences\. Stable\-interest recommendation requests\. 1\. I want something that fits my usual taste\. Could you recommend one option for me? 2\. Could you suggest one thing I’d probably enjoy based on what I usually like? 3\. I’m looking for a recommendation that feels very me\. What’s one good option? 4\. Choose one option that fits what I’ve liked for a while\. 5\. I want a safe choice that matches my usual preferences\. What should I try? 6\. Recommend one activity or item that seems close to my regular taste\. 7\. Based on what I tend to enjoy, what is one practical suggestion? 8\. I’m not trying to branch out today; give me one recommendation that fits my normal style\. 9\. What’s one personalized option that would likely suit my everyday interests? 10\. Give me one recommendation grounded in what I’ve consistently liked before\. Recent\-interest exploration requests\. 1\. I want to try something a bit new, but still something that feels like me\. Any suggestion? 2\. Could you recommend one fresh option that connects to what I’ve been into lately? 3\. I’m open to exploring something new\. What’s one suggestion that still matches my taste? 4\. Choose one option that builds on what has caught my attention recently, without feeling random\. 5\. I’d like a small change from my usual choices\. What should I try? 6\. Recommend one new\-ish activity or item that fits what I seem to be into right now\. 7\. What’s one recommendation that reflects what I’ve been paying attention to lately? 8\. I want something slightly outside my routine, but not totally unfamiliar\. Any idea? 9\. Suggest one option that feels current for me while still matching my usual taste\. 10\. Give me one practical recommendation that feels timely for me, not just my old favorites\. A\.6 Dialogue Evaluation Judge Prompt Dialogue Evaluation Judge — System Prompt Dialogue Evaluation Judge — Rubrics Dialogue Evaluation Judge — User Prompt A\.7 Caption Prompt Image Captioning — System Prompt A\.8 Hierarchical Profile Construction Prompts Hierarchical — Chunk Summarization — System Prompt Hierarchical — Global Aggregation — System Prompt A\.9 Extractive Profile Construction Prompts The extractive–abstractive method proceeds in two stages\. First, the LLM receives the full user timeline and is prompted to select up to KK representative posts per domain\. The selection criteria include relevance, specificity \(concrete interest signals rather than vague topics\), recurrence \(preferring posts consistent with repeated behavior\), recency \(capturing potential recent interests\), and multimodal grounding \(using image captions as evidence\)\. Second, the LLM receives only the selected posts and synthesizes the final profile, without access to the original full timeline\. The prompts for both stages are shown below\. Extractive — Post Selection — System Prompt Extractive — Abstractive Synthesis — System Prompt A\.10 Calibration and Gold Rewrite Prompts Calibration — Evidence\-First Extraction — System Prompt Calibration — Domain LLM Calibration — System Prompt Calibration — Benchmark Gold Rewrite — System Prompt Appendix B User Filtering and Profile Verification Details This appendix provides implementation details that are summarized in Section 3\.2 and Section 3\.3\. Automatic account filtering\. We use three account\-level heuristics before manual inspection\. First, we retain accounts with 5–5,000 followers, excluding extremely inactive accounts and highly public accounts\. Second, we compute the follower–followee ratio \(FFR\) as FFR=NfollowerNfollowee\+ϵ,\\mathrm\{FFR\}=\\frac\{N\_\{\\mathrm\{follower\}\}\}\{N\_\{\\mathrm\{followee\}\}\+\\epsilon\}, \(1\) where ϵ\\epsilon avoids division by zero\. We retain accounts with FFR∈\[0\.5,2\]\\mathrm\{FFR\}\\in\[0\.5,2\], which favors relatively reciprocal social neighborhoods and filters highly asymmetric broadcaster\-style accounts\. Third, we compute the image trace density ratio \(ITDR\) as ITDR=Nposts​with​imagesNtotal​posts,\\mathrm\{ITDR\}=\\frac\{N\_\{\\mathrm\{posts\\ with\\ images\}\}\}\{N\_\{\\mathrm\{total\\ posts\}\}\}, \(2\) and retain users with ITDR≥0\.3\\mathrm\{ITDR\}\\geq 0\.3 to ensure sufficient visual evidence for multimodal profiling\. Manual account inspection\. After automatic filtering, annotators remove accounts that are commercial, celebrity\-like, organization\-operated, repost\-heavy, dominated by low\-information content, or lacking sufficient personal and preference\-relevant signals\. This stage yields 250 candidate users for timeline collection and profile construction\. Cross\-post aggregation\. Post\-level interest candidates are aggregated into canonical interests before temporal scoring\. Near\-duplicate posts are clustered and down\-weighted so that repeated captions, repost\-like content, or bursty discussions do not count as independent evidence\. Semantically equivalent tags are normalized and merged into canonical interests\. For each canonical interest, we retain supporting posts, modality attribution, duplicate\-adjusted support, temporal distribution, and extraction confidence\. Human profile verification\. Human verification is conducted over calibrated domain\-level profiles and their supporting evidence\. Annotators remove unsupported interests, revise overly broad labels into more concrete tags, and add missing tags when the evidence clearly supports them\. Accepted labels must be concrete, non\-sensitive, domain\-appropriate, and supported by the timeline\. After this verification stage, users with fewer than three active domains are removed to ensure enough positive personalization signals for evaluation\. Our annotators are trained to be conservative and evidence\-grounded, avoiding over\-interpretation or inference beyond what the timeline supports\. This process yields the final set of 100 users for the benchmark\. Appendix C Annotator Recruitment, Instructions, and Data Consent This section supplements the profile verification \(Section 3\.3\) and dialogue evaluation \(Appendix J\) with details on annotator recruitment, compensation, instructions, and data consent\. Recruitment and payment\. All annotators were recruited from the undergraduate and graduate student population at the authors’ institution\. Annotators were compensated at 50 CNY per hour\. This rate exceeds the typical student hourly wage at the authors’ university and is consistent with compensation for comparable annotation tasks in the region\. Instructions to participants\. All annotators received written annotation guidelines before beginning work, covering task definitions, quality rubrics, and annotated examples\. For profile verification, annotators were instructed to be conservative and evidence\-grounded, accepting interests only when directly supported by timeline evidence \(see the criteria in Appendix B\)\. For dialogue evaluation, annotators received the full 0–5 scoring rubrics shown in Appendix A\.6 and completed a calibration round before the agreement study\. The complete instruction documents are available upon request\. Data consent\. The social\-media posts used in this benchmark were collected from publicly accessible accounts\. We exclude private or restricted\-access content\. The released benchmark subset contains only de\-identified text and captions; original images are excluded to prevent re\-identification, and the full dataset is provided only through controlled research access \(see Ethical Considerations\)\. No annotator personal information was collected or retained\. Appendix D Profile Scoring Details This appendix provides the detailed scoring rules used in the algorithmic temporal scoring stage of profile construction\. These scores are used only to produce preliminary stable/recent assignments before LLM\-based calibration and human verification\. Canonical interest evidence\. After post\-level extraction and cross\-post aggregation, each canonical interest cc in domain dd is associated with an evidence set ℰu,d,c=\{e1,e2,…,em\},\\mathcal\{E\}\_\{u,d,c\}=\\\{e\_\{1\},e\_\{2\},\\dots,e\_\{m\}\\\}, \(3\) where each evidence item corresponds to a supporting post\. Each evidence item records the post timestamp, evidence modality, duplicate\-adjusted weight, and post\-level extraction confidence\. If a post belongs to a near\-duplicate cluster CC, its support weight is discounted by 1/\|C\|1/\|C\|, so that repeated or highly similar posts do not artificially inflate an interest\. Let wjw\_\{j\} denote the duplicate\-adjusted weight of evidence item eje\_\{j\}, and let cj∈\[0,1\]c\_\{j\}\\in\[0,1\] denote its extraction confidence\. The effective support count of a canonical interest is defined as Seff=∑ej∈ℰu,d,cwj\.S\_\{\\mathrm\{eff\}\}=\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}\}w\_\{j\}\. \(4\) The average confidence is computed as a weighted average: c¯=∑ej∈ℰu,d,cwj​cj∑ej∈ℰu,d,cwj\.\\bar\{c\}=\\frac\{\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}\}w\_\{j\}c\_\{j\}\}\{\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}\}w\_\{j\}\}\. \(5\) Temporal bins\. To estimate whether an interest is persistent over time, we divide each user’s timeline into monthly bins\. Let BdistinctB\_\{\\mathrm\{distinct\}\} be the number of distinct monthly bins that contain at least one supporting evidence item for the canonical interest\. Let BspanB\_\{\\mathrm\{span\}\} be the total number of monthly bins covered by the user’s collected timeline\. We set Breq=min⁡\(6,Bspan\),B\_\{\\mathrm\{req\}\}=\\min\(6,B\_\{\\mathrm\{span\}\}\), \(6\) so that long timelines require evidence spread across multiple periods, while shorter timelines are not penalized excessively\. Stable\-interest score\. The stable score estimates whether a canonical interest reflects a stable preference\. It combines three factors: effective support, temporal dispersion, and extraction confidence: scorestable\\displaystyle\\mathrm\{score\}\_\{\\mathrm\{stable\}\} =0\.60​min⁡\(1,Seff3\)\\displaystyle=60\\min\\\!\\left\(1,\\frac\{S\_\{\\mathrm\{eff\}\}\}\{3\}\\right\) \(7\) \+0\.20​c¯\+0\.20​min⁡\(1,Bdistinctm\),\\displaystyle\\quad\+20\\bar\{c\}\+20\\min\\\!\\left\(1,\\frac\{B\_\{\\mathrm\{distinct\}\}\}\{m\}\\right\), where m=max⁡\(2,Breq−1\)m=\\max\(2,B\_\{\\mathrm\{req\}\}\-1\)\. The first term rewards repeated support after duplicate discounting\. The second term incorporates the confidence of post\-level extraction\. The third term rewards evidence distributed across multiple time periods\. This score is designed to favor interests that appear repeatedly and persistently across the user’s timeline\. A canonical interest is marked as a preliminary stable interest if it satisfies scorestable≥θstable\\mathrm\{score\}\_\{\\mathrm\{stable\}\}\\geq\\theta\_\{\\mathrm\{stable\}\} \(8\) and has sufficient temporal support: Seff≥3,Bdistinct≥2\.S\_\{\\mathrm\{eff\}\}\\geq 3,\\qquad B\_\{\\mathrm\{distinct\}\}\\geq 2\. \(9\) In our implementation, we set θstable=0\.65\.\\theta\_\{\\mathrm\{stable\}\}=0\.65\. \(10\) Recent\-interest score\. The recent score estimates whether a canonical interest reflects a recent or emerging preference\. We focus on the most recent 90 days of the user’s timeline\. Let ℰu,d,c90\\mathcal\{E\}\_\{u,d,c\}^\{90\} denote the subset of evidence items whose timestamps fall within this window\. We define S90=∑ej∈ℰu,d,c90wj,S\_\{90\}=\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}^\{90\}\}w\_\{j\}, \(11\) and c¯90=∑ej∈ℰu,d,c90wj​cj∑ej∈ℰu,d,c90wj\.\\bar\{c\}\_\{90\}=\\frac\{\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}^\{90\}\}w\_\{j\}c\_\{j\}\}\{\\sum\_\{e\_\{j\}\\in\\mathcal\{E\}\_\{u,d,c\}^\{90\}\}w\_\{j\}\}\. \(12\) If no evidence appears in the most recent 90 days, we set S90=0S\_\{90\}=0 and c¯90=0\\bar\{c\}\_\{90\}=0\. The recent score is defined as scorerecent=0\.65​min⁡\(1,S902\)\+0\.35​c¯90\.\\mathrm\{score\}\_\{\\mathrm\{recent\}\}=0\.65\\min\\\!\\left\(1,\\frac\{S\_\{90\}\}\{2\}\\right\)\+0\.35\\bar\{c\}\_\{90\}\. \(13\) Compared with the stable score, the recent score places more emphasis on recent support and does not require broad temporal dispersion across the full timeline\. A canonical interest is marked as a preliminary recent interest if it does not satisfy the stable\-interest condition but satisfies scorerecent≥θrecent\\mathrm\{score\}\_\{\\mathrm\{recent\}\}\\geq\\theta\_\{\\mathrm\{recent\}\} \(14\) and has sufficient recent evidence: In our implementation, we set θrecent=0\.60\.\\theta\_\{\\mathrm\{recent\}\}=0\.60\. \(16\) Appendix E Prompt Analysis: Conservative vs\. Neutral Evaluation Instructions The profile evaluation prompt \(Appendix A\.2\) instructs models to “prefer fewer, broader tags” and limit most active domains to “1–2 reliable tags\.” This conservative design intentionally prioritizes precision over recall: without such guidance, models in pilot experiments produced noisy tag lists containing near\-duplicates and one\-off mentions that did not reflect genuine user interests\. We note a potential concern: if the prompt constrains output volume, the benchmark may measure prompt compliance rather than profiling capability, creating an artificial ceiling on recall\. To rule this out, we conduct a prompt analysis replacing the conservative instructions with a neutral variant that removes all constraints on tag count and breadth\. The full neutral system prompt is: Profile Evaluation — Neutral System Prompt The neutral user prompt mirrors the same changes: Profile Evaluation — Neutral User Prompt We evaluate Qwen3\.5\-35B\-A3B under the direct profile construction setting on the full 100\-user evaluation subset, using identical gold profiles, anchor\-matching, and evaluation protocol across both prompt variants\. Metric Conservative Neutral Δ\\Delta Interest Precision 0\.490 0\.417 −\-0\.073 Interest Recall 0\.157 0\.180 \+0\.023 Interest F1 0\.238 0\.252 \+0\.014 Active F1 0\.788 0\.812 \+0\.025 Table 4: Prompt analysis comparing a Conservative evaluation prompt \(prefers fewer, broader tags\) against a Neutral variant \(no tag\-count constraints\)\. Δ=Neutral−Conservative\\Delta=\\text\{Neutral\}\-\\text\{Conservative\}; positive values favor the neutral prompt\. Results on Qwen3\.5\-35B\-A3B under the direct setting with 100 users\. Removing the conservative constraints changes metrics only marginally\. Interest Recall improves slightly \(\+0\.023\), but Interest Precision drops by more \(–0\.073\) as models produce specific tags that misalign with gold labels\. The net effect on Interest F1 is negligible \(\+0\.014\)\. These results confirm that the conservative prompt is not the primary driver of low recall—the gap reflects genuine model limitations in recovering fine\-grained interests from behavioral traces, not an artifact of benchmark design\. Appendix F Input\-Modality Analysis To quantify the relative contribution of each input modality to profiling performance, we evaluate two models under progressively richer input configurations: text only, image captions only, text plus image captions, and the full setting adding timestamps\. Table F reports active\-domain detection and interest\-tag recovery across all four conditions\. Text provides the primary profiling signal, while image captions and timestamps contribute complementary gains, as discussed in Section 4\.4\. Input mode Active F1 Tag F1 Stable F1 Recent F1 \\rowcolortabgray Qwen3\.5\-35B\-A3B Text only 0\.6869 0\.2198 0\.2415 0\.0088 Image captions only 0\.6724 0\.2352 0\.2782 0\.0093 Text \+ image captions 0\.7692 0\.2682 0\.2846 0\.0171 Text \+ img\. caps\. \+ timestamp 0\.7875 0\.2669 0\.2989 0\.0133 \\rowcolortabgray GPT\-4o\-mini Text only 0\.5033 0\.1712 0\.0745 0\.0592 Image captions only 0\.4992 0\.1527 0\.1122 0\.0352 Text \+ image captions 0\.7367 0\.2796 0\.1705 0\.0812 Text \+ img\. caps\. \+ timestamp 0\.7454 0\.2785 0\.1832 0\.0858 Table 5: Input\-modality analysis on the fixed 100\-user evaluation subset\. Timestamps are removed in the first three ablation settings for each model\. The Text \+ img\. caps\. \+ timestamp row corresponds to the direct profile construction setting from Table 4\.2 and serves as the full\-input reference\. Stable F1 and Recent F1 measure interest\-tag recovery within the stable and recent temporal buckets respectively, computed via within\-bucket optimal bipartite matching with cached LLM anchor judgments\. Best values per column within each model group are bolded\. Appendix G Case Studies: Error Analysis This appendix presents detailed case studies supporting the error analysis in Sections 4\.2 and 4\.3\. We examine predictions from representative users across all seven domains, comparing gold profiles against model outputs under the direct, hierarchical, and extractive settings\. Tables 7–10 organize examples by failure pattern\. Notation\. In all case\-study tables that follow: S = stable interest tags; R = recent interest tags; D = direct, H = hierarchical, E = extractive–abstractive profiling\. \\rowcolortabgray User Posts Active Gold profile summary \(stable \|\| recent interests, with evidence modality\) Inactive User A 115 5/7 sports\_outdoor: outdoor recreation, hiking \[t\+v\] \|\| walking, park visit \[t\+v\] entertainment: music listening, book, creative writing \[t\+v\] \|\| poetry, reading, writing \[t\+v\] food\_drink: dining out, birthday cake, dessert, coffee \[t\+v\] \|\| restaurant dining, casual dining, confectionery \[t\+v\] travel\_city: sightsee, road trip, historic site, landmark \[t\+v\] \|\| Virginia Beach, roadside attraction, historical landmark \[t\+v\] photography: craft, photo sharing \[t\+v\] \|\| vintage photography \[t\+v\] gaming, pets User B 100 6/7 sports\_outdoor: football fandom \[t\+v\] \|\| soccer, Africa Cup, fitness \[t\+v\] entertainment: music listen, Arabic music \[t\+v\] \|\| Marwan Moussa, Spotify, movy \[t\+v\] gaming: eFootball \[t\+v\] \|\| video games, mobile gaming \[t\+v\] food\_drink: — \|\| iftar, breakfast, meal \[t\+v\] travel\_city: city exploration \[t\+v\] \|\| city walks, urban exploration \[t\+v\] pets: — \|\| cat \[v\] photography User C 187 5/7 sports\_outdoor: wrestl, indie wrestling \[t\+v\] entertainment: wrestling, AEW, live event \[t\+v\] \|\| wrestle kingdom \[t\+v\] food\_drink: cocktail, home cooking \[t\+v\] travel\_city: sightsee, city exploration \[t\+v\] photography: event photography, photo editing \[t\+v\] gaming, pets User D 185 6/7 sports\_outdoor: walk, snorkeling, birdwatching \[t\+v\] \|\| nature observation, outdoor recreation \[t\+v\] entertainment: classic rock, the grinch, music listening \[t\+v\] \|\| concert attendance, music \[t\+v\] food\_drink: pub, pub visit, coffee, breakfast \[t\+v\] \|\| banana bread \[t\+v\] travel\_city: sightsee, city exploration, cruise travel, city walks \[t\+v\] \|\| Volendam, Netherlands, Sinai desert \[t\+v\] photography: landscape photography, photography, nature photography, night photography \[t\+v\] \|\| bird photography, black and white photography, outdoor photography \[t\+v\] pets: dog walk, dog ownership, dog care, pet ownership \[t\+v\] \|\| pet friendly pub \[t\+v\] gaming Table 6: Overview of the four case\-study users\. For each user we report the number of posts in the observation window, the count of active domains out of seven, a compact gold profile summary with evidence modality annotations \(t = text, v = visual\), and the inactive domains\. A dash \(—\) in the stable or recent slot means no interests of that type were annotated for the domain\. Detailed per\-domain breakdowns appear in the individual case\-study tables below\. G\.1 Over\-Generalization and Cross\-Category Confusion Table 7 illustrates the most pervasive failure mode: models collapsing multiple specific gold tags into one or two broad categories, and confusing semantically adjacent but factually incorrect categories\. \\rowcolortabgray User Domain Gold tags Model predictions User A food\_drink S: dining out, birthday cake, dessert, coffee R: restaurant dining, casual dining, confectionery Gemini\-2\.5 \(D\): S=\{Dining out, Baking\} GPT\-5\.4 \(D\): S=\{sweets and desserts, restaurants and dining out\} GPT\-4o\-mini \(D\): INACTIVE Qwen3\-VL\-8B \(D\): INACTIVE User B entertainment S: music listen, Arabic music R: Marwan Moussa, Spotify, movy Gemini\-2\.5 \(D\): S=\{music, movies and TV\} GPT\-5\.4 \(D\): S=\{Arabic music\} Table 7: Over\-generalization and cross\-category confusion\. Models reduce 4–7 gold tags to 1–2 broad labels, hallucinate factually wrong categories \(e\.g\., “Baking” for a user who only dines out and buys desserts\), or predict INACTIVE for domains with abundant evidence\. \(D\) = direct setting\. Representative posts: User A food\_drink evidence \(dining out and store\-bought desserts, no home cooking\) 2026\-01\-01 New Year dining out Last night, for \#NewYear2026, we went out to eat\. As I was biting my delicious chicken strip, I thought about all who couldn’t be out, due to being sick, bedridden with cancer, frail unable to walk\. 2026\-01\-13 store\-bought birthday cake Hubs birthday soon\. My mom always bought Pepperidge Farms cakes for birthday celebrations\. Sure miss her\. 2026\-01\-05 restaurant dining I miss eating Paradiso with my mom and hubs together\. These three posts illustrate the user’s food\_drink profile: dining out on New Year’s Eve, purchasing store\-bought Pepperidge Farms cakes for a birthday, and eating at a restaurant \(Paradiso\)\. The user’s gold profile includes home cooking as a negative interest—this user does not cook at home\. Yet Gemini\-2\.5 predicts “Baking” with no evidence of any baking activity anywhere in the timeline\. GPT\-5\.4 over\-generalizes seven specific tags into two broad categories\. GPT\-4o\-mini and Qwen3\-VL\-8B predict INACTIVE despite seven gold tags supported by posts spanning the full four\-month window\. The User A food\_drink case exhibits three distinct but related failure modes in a single domain\. First, over\-generalization: GPT\-5\.4 collapses seven specific gold tags—dining out, birthday cake, dessert, coffee, restaurant dining, casual dining, and confectionery—into just two broad labels \(sweets and desserts, restaurants and dining out\)\. While these are not factually wrong, they discard the granularity needed for downstream personalization: recommending a confectionery shop is qualitatively different from recommending a birthday cake bakery, and a coffee shop recommendation differs from a casual\-dining suggestion\. Second, cross\-category confusion: Gemini\-2\.5 predicts Baking despite the user having zero baking activity anywhere in the timeline\. The gold profile explicitly includes home cooking as a negative interest—this user purchases prepared food and eats at restaurants, but does not cook\. The model appears to conflate “engages with food content” with “prepares food at home,” a category error analogous to confusing “attends concerts” with “plays an instrument\.” Third, false inactive: GPT\-4o\-mini and Qwen3\-VL\-8B classify the entire domain as INACTIVE, missing all seven gold tags\. This is a severe false\-inactive error: the domain is supported by posts spanning the full four\-month observation window across multiple modalities \(text, images, and mixed\), yet two models fail to activate it at all\. The User B entertainment case exhibits the same over\-generalization pattern\. Gemini\-2\.5 predicts music, movies and TV, adding a film/television interest absent from the gold profile, while GPT\-5\.4 reduces five tags to one \(Arabic music\)\. Across both cases, the pattern is consistent: models default to broad, safe hypernyms and resist committing to the specific subcategories that make personalized recommendations actionable\. G\.2 Domain Blindness: False Inactive Errors Table 8 shows cases where models incorrectly classify an active domain as inactive, missing all gold interest tags\. These errors concentrate in domains whose evidence is primarily textual and dispersed across many posts\. \\rowcolortabgray User Domain Gold tags Which models missed it User B travel\_city S: city exploration R: city walks, urban exploration All four models \(D\) User B food\_drink R: iftar, breakfast, meal GPT\-5\.4, Qwen3\.5, Qwen2\.5 \(D\) Table 8: False inactive errors\. Domains supported primarily by scattered textual mentions are frequently missed entirely, even by strong models\. \(D\) = direct setting\. Representative posts: User B city\-exploration evidence \(text\-dispersed domain\) 2026\-01\-07 text only, no image I wanna migrate illegally and the guys are like let’s go to Ifrane hhhhhhhhhhhhh 2026\-02\-19 city walk \(4 images\) Late night walk 2026\-02\-23 urban exploration \(1 image\) Bars These three posts collectively support city exploration \(stable\) and city walks / urban exploration \(recent\)\. However, the evidence is distributed across short, casual text mentions with no hashtags, no location tags in the text, and no visually distinctive landmarks\. All four models classified this domain as inactive when using the direct setting, because without a visually anchoring post, the weak textual signal fails to reach the activation threshold\. The User B travel case is illustrative: all four models predict inactive for a domain where gold lists city exploration, city walks, and urban exploration\. These interests are expressed through casual text mentions across posts \(e\.g\., “went for a walk downtown,” “exploring a new neighborhood”\) rather than through prominent images or hashtags\. Without a visually anchoring post, the evidence fails to reach the model’s activation threshold\. G\.3 Hallucinated Domains: False Active Errors The inverse failure—activating a domain the user does not actually engage with—occurs when models over\-interpret incidental posts as preference signals\. Table 9 presents two representative cases\. The first involves a user whose timeline is dominated by professional\-wrestling content \(187 posts, gold\-active domains: sports\_outdoor and entertainment with wrestling\-related interests\)\. GPT\-5\.4 under hierarchical profiling activates the gaming domain with tags that explicitly name wrestling—“video game references in wrestling\-related memes” and “Pokémon\-themed wrestling events\.” The model’s own labels concede the content is wrestling\-related, yet it places them in gaming\. Five posts out of 187 mention video games at all, and all five are wrestling\-context posts: three document a CMLL×\\timesPokémon crossover wrestling show, one jokes about a wrestler taking time off to play a game, and one is an incidental mention\. The model conflates wrestling content that references gaming with the user having a gaming interest\. \\rowcolortabgray User Domain Gold status Model predictions User C gaming inactive GPT\-5\.4 \(H\): R=\{ video game references in wrestling\-related memes, Pokémon\-themed wrestling events \} GPT\-5\.4 \(E\): R=\{ video game releases and references \} User B photography inactive GPT\-5\.4 \(D\): S=\{ selfie and portrait photography \} GPT\-5\.4 \(H\): R=\{ selfie portraits and self\-image/profile visual curation \} Qwen3\-VL\-8B \(H\): R=\{ image editing and AI\-generated visuals \} Table 9: False active errors: category confusion and systemic over\-interpretation of incidental post content\. D = direct, H = hierarchical, E = extractive–abstractive\. Representative posts: User C wrestling content that triggered gaming hallucination 2025\-09\-26 CMLL×\\timesPokémon crossover wrestling show Checking out that \#CMLL Pokémon show for a bit\. This already looks like so much fun and I can’t believe they got approval from Nintendo for it\. \#LeyendasPokémonZA 2025\-09\-26 same event The commitment from this fan to rock the Umbreon gimp mask for the show … \#CMLL \#LeyendasPokémonZA 2025\-03\-13 wrestling joke referencing Assassin’s Creed It’s okay, Will\. We know you need time off to play the new Assassin’s Creed game coming out next week\. Don’t need to use your wife as an excuse\. \#AEWDynamite 2025\-10\-19 AEW stage design compared to Borderlands The St\. Louis arch on the \#AEWWrestleDream stage is now making me think of the Borderlands vaults and all the wrestlers are just different Vault Hunters\. \#AEW 2025\-10\-24 Undertale meme, wrestling reaction “Hopes and Dreams” intensifies\. \#AEW \#Undertale These five posts are the only video\-game mentions across 187 posts\. All are contextual to professional wrestling: three document a wrestling show with a Pokémon promotional crossover, one jokes about a wrestler taking time off, one compares an AEW stage design to Borderlands, and one uses an Undertale meme to react to a match\. GPT\-5\.4’s own predicted tags concede the content is “wrestling\-related”—yet the model activates the gaming domain rather than recognizing this as wrestling\-fan content that belongs in the user’s already\-active sports\_outdoor and entertainment domains\. The User B photography case reveals a broader systemic pattern: across the full benchmark, posting selfies—a common behavior on social media—is frequently misinterpreted as photography enthusiasm\. This mirrors the over\-generalization pattern from Section G: models lack the pragmatic judgment to distinguish “taking a photo to document an experience” from “photography as a sustained interest\.” The User C case adds a further dimension: even when the model correctly identifies the topic of a post \(wrestling\), it can assign it to the wrong domain, revealing a category\-boundary problem in how models map post content to interest domains\. G\.4 Recent\-Interest Blindness and Temporal Bucket Confusion Table 10 illustrates a pervasive failure: models cannot distinguish long\-standing interests from recently emerged ones, assigning nearly all predictions to the stable bucket regardless of when evidence appears in the timeline\. User D provides a representative case: a frequent traveler with a clear pattern of general, recurring travel behavior \(cruises, city walks, sightseeing\) punctuated by specific recent destinations \(Volendam, the Netherlands, the Sinai Desert\)\. \\rowcolortabgray User Domain Gold temporal split Predictions \(stable \|\| recent\) User D travel\_city S: 4 tags: sightsee, city exploration, cruise travel, city walks R: 3 tags: volendam, netherland, sinai desert GPT\-5\.4 \(H\): S=\{Caribbean cruise, city exploration, Amsterdam\} R=\{Egypt, Sinai, Netherlands\} Qwen3\.5 \(H\): S=\{International Travel, Cruise Travel\}, R=0 Qwen3\-VL \(E\): S=\{cruise\-based Caribbean exploration\}, R=0 Table 10: Recent\-interest blindness and temporal bucket confusion\. Models fail to distinguish stable from recent interests, assigning predictions indiscriminately to the stable bucket\. H = hierarchical, E = extractive–abstractive\. Representative posts: User D travel evidence illustrating stable vs\. recent temporal structure 2025\-11\-28 stable: cruise travel Marella Discovery 2, our home for the next two weeks at Berth in Bridgetown Port, Barbados as we and a multitude of other new passengers wait to board\. 2025\-12\-19 stable: city walks Happy \#FingerpostFriday from \#Budapest… I came across these cyclist friendly fingerposts on a late evening walk by the River \#Danube\. 2026\-01\-16 recent: Netherlands/Volendam Happy \#FingerpostFriday… Here’s a set of Fingerposts from the lovely town of \#Volendam in the Netherlands that we visited on a day trip out from Amsterdam\. 2026\-02\-15 recent: Sinai Desert Sunset Sinai Desert style\. A fabulous excursion out into the desert this afternoon, to see a stark but stunning landscape that just seems utterly timeless\. The temporal structure of User D’s travel domain is clear from the timeline: cruises, city walks, and sightseeing recur across the full four\-month timeline \(stable\), while Volendam \(January\), the Netherlands \(January–February\), and the Sinai Desert \(February\) are specific destinations visited only in the final two months\. GPT\-5\.4 partially captures this—placing Egypt and Netherlands travel in the recent bucket—but contradicts itself by assigning “Amsterdam city exploration” to the stable bucket, even though Amsterdam and the Netherlands are the same trip\. Qwen3\.5 and Qwen3\-VL exhibit complete temporal blindness: they collapse all evidence into one or two broad stable tags and assign zero predictions to the recent bucket\. The failure pattern extends beyond User D\. Across the benchmark, models assign 70–90% of predictions to the stable bucket regardless of temporal evidence distribution\. When recent tags are predicted, they often correspond to interests that the gold profile classifies as stable, and vice versa\. GPT\-5\.4’s Amsterdam/Netherlands contradiction is especially revealing: the model recognizes that “the Netherlands” is a recent topic but places the capital city of that same country in the stable bucket, demonstrating that these models perform surface\-level topic labeling rather than genuine temporal reasoning about when and how frequently evidence appears\. G\.5 Dialogue Case Studies The following examples illustrate how profiling failures propagate into downstream dialogue \(Section 4\.3\)\. Profile under\-generation case \(food\-drink user\)\. In the direct\-profile\-conditioned setting, a user’s gold food\-drink profile contains seven specific interests spanning dining out, regional cuisines, and holiday meals\. Gemini\-2\.5\-Flash’s profile reduces this to two tags \(Dining out, Asian cuisine\); its dialogue response ignores food entirely, instead recommending astrophotography based on the user’s photography interests\. The judge assigns coverage == 1\.0/5, noting the model “ignores the required interests\.” GPT\-5\.4’s profile for the same user captures restaurant dining and sushi; its dialogue recommends a specific Japanese restaurant, earning coverage == 3\.0/5\. The difference illustrates a compounding failure: conservative profiling removes the very tags that would enable diverse, personalized responses\. Visual dominance case \(astrophotography fixation\)\. In the timeline\-conditioned setting, models can access all posts directly, yet they exhibit the same modality bias observed in profile construction\. For the user described above, Gemini\-2\.5\-Flash overlooks seven supported food interests and a hiking interest to recommend astrophotography across both dialogue turns, fixating on the domain with the most visually prominent evidence \(sky and landscape photography\)\. This pattern recurs across users: when a domain generates abundant images, it crowds out text\-supported interests in downstream dialogue, even when those text\-supported interests are equally or more relevant to the user’s request\. These two cases illustrate a compounding failure chain\. Conservative profiling strips away fine\-grained interest tags \(profile under\-generation\); the resulting sparse profile provides insufficient hooks for the dialogue model, which then defaults to visually dominant domains regardless of the user’s actual request\. The chain can be broken at either stage—better profiling yields richer profiles, and direct timeline access bypasses profile sparsity—but the modality bias persists in both paths, suggesting that balanced cross\-modal attention is a prerequisite for either approach to succeed\. Appendix H Profile–Dialogue Correlation Tables This appendix provides the full correlation tables referenced in Section 4\.4\. Model Interest F1 ↔\\leftrightarrow Avg\. Interest F1 ↔\\leftrightarrow Coverage Interest F1 ↔\\leftrightarrow Concreteness \\rowcolortabgray Direct Gemini\-2\.5\-Flash −0\.016\-0\.016 \+0\.013\+0\.013 \+0\.039\+0\.039 GPT\-4o\-mini \+0\.221∗\+0\.221^\{\*\} \+0\.245∗\+0\.245^\{\*\} \+0\.189\+0\.189 GPT\-5\.4 \+0\.128\+0\.128 \+0\.113\+0\.113 \+0\.087\+0\.087 Qwen2\.5\-VL\-7B\-Instruct \+0\.094\+0\.094 \+0\.299∗∗\+0\.299^\{\*\*\} \+0\.319∗∗\+0\.319^\{\*\*\} Qwen3\-VL\-8B\-Instruct \+0\.092\+0\.092 \+0\.240∗\+0\.240^\{\*\} \+0\.224∗\+0\.224^\{\*\} Qwen3\.5\-35B\-A3B −0\.042\-0\.042 \+0\.089\+0\.089 \+0\.054\+0\.054 \\rowcolortabgray Hierarchical Gemini\-2\.5\-Flash \+0\.330∗⁣∗∗\+0\.330^\{\*\*\*\} \+0\.472∗⁣∗∗\+0\.472^\{\*\*\*\} \+0\.177\+0\.177 GPT\-4o\-mini \+0\.148\+0\.148 \+0\.319∗∗\+0\.319^\{\*\*\} \+0\.028\+0\.028 GPT\-5\.4 \+0\.214∗\+0\.214^\{\*\} \+0\.244∗\+0\.244^\{\*\} \+0\.306∗∗\+0\.306^\{\*\*\} Qwen2\.5\-VL\-7B\-Instruct \+0\.102\+0\.102 \+0\.477∗⁣∗∗\+0\.477^\{\*\*\*\} \+0\.256∗\+0\.256^\{\*\} Qwen3\-VL\-8B\-Instruct \+0\.065\+0\.065 \+0\.251∗\+0\.251^\{\*\} \+0\.157\+0\.157 Qwen3\.5\-35B\-A3B \+0\.207∗\+0\.207^\{\*\} \+0\.311∗∗\+0\.311^\{\*\*\} \+0\.233∗\+0\.233^\{\*\} \\rowcolortabgray Extractive Gemini\-2\.5\-Flash \+0\.208∗\+0\.208^\{\*\} \+0\.484∗⁣∗∗\+0\.484^\{\*\*\*\} \+0\.153\+0\.153 GPT\-4o\-mini \+0\.198∗\+0\.198^\{\*\} \+0\.222∗\+0\.222^\{\*\} \+0\.055\+0\.055 GPT\-5\.4 \+0\.172\+0\.172 \+0\.204∗\+0\.204^\{\*\} \+0\.128\+0\.128 Qwen2\.5\-VL\-7B\-Instruct \+0\.138\+0\.138 \+0\.446∗⁣∗∗\+0\.446^\{\*\*\*\} \+0\.245∗\+0\.245^\{\*\} Qwen3\-VL\-8B\-Instruct \+0\.022\+0\.022 \+0\.413∗⁣∗∗\+0\.413^\{\*\*\*\} \+0\.722∗⁣∗∗\+0\.722^\{\*\*\*\} Qwen3\.5\-35B\-A3B \+0\.186\+0\.186 \+0\.318∗∗\+0\.318^\{\*\*\} \+0\.193\+0\.193 Table 11: Per\-user Spearman rank correlation between profile Interest Tag F1 and dialogue quality dimensions across three profile\-conditioned settings\. Each cell reports ρ\\rho over 89–93 users \(model\-dependent, after excluding generation or judge errors\)\. Significance: p∗<0\.05\{\}^\{\*\}p<0\.05, p∗∗<0\.01\{\}^\{\*\*\}p<0\.01, p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\. Model Stable F1 ↔\\leftrightarrow Stable Rec\. Recent F1 ↔\\leftrightarrow Recent Expl\. Stable F1 ↔\\leftrightarrow Recent Expl\. Recent F1 ↔\\leftrightarrow Stable Rec\. Gemini\-2\.5\-Flash \+0\.415∗⁣∗∗\+0\.415^\{\*\*\*\} \+0\.189\+0\.189 \+0\.154\+0\.154 \+0\.336∗⁣∗∗\+0\.336^\{\*\*\*\} GPT\-4o\-mini \+0\.172\+0\.172 \+0\.265∗\+0\.265^\{\*\} \+0\.254∗\+0\.254^\{\*\} \+0\.048\+0\.048 GPT\-5\.4 \+0\.038\+0\.038 \+0\.186\+0\.186 −0\.060\-0\.060 \+0\.317∗∗\+0\.317^\{\*\*\} Qwen2\.5\-VL\-7B\-Instruct \+0\.496∗⁣∗∗\+0\.496^\{\*\*\*\} \+0\.385∗∗\+0\.385^\{\*\*\} \+0\.364∗∗\+0\.364^\{\*\*\} \+0\.466∗⁣∗∗\+0\.466^\{\*\*\*\} Qwen3\-VL\-8B\-Instruct −0\.067\-0\.067 −0\.109\-0\.109 −0\.009\-0\.009 \+0\.217∗\+0\.217^\{\*\} Qwen3\.5\-35B\-A3B \+0\.214∗\+0\.214^\{\*\} \+0\.161\+0\.161 −0\.018\-0\.018 −0\.036\-0\.036 Table 12: Temporal breakdown of profile–dialogue correlation under the extractive–abstractive setting\. Same\-bucket pairs test temporal alignment; cross\-bucket pairs test whether general profile quality drives dialogue regardless of bucket\. n=68n=68–100 per correlation\. p∗<0\.05\{\}^\{\*\}p<0\.05, p∗∗<0\.01\{\}^\{\*\*\}p<0\.01, p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\. Appendix I Per\-Domain Difficulty Breakdown Domain Direct Hierarchical Extractive Pets 0\.522 0\.497 0\.575 Gaming 0\.385 0\.483 0\.399 Sports & Outdoor 0\.324 0\.488 0\.463 Food & Drink 0\.341 0\.358 0\.345 Photography & Creation 0\.314 0\.405 0\.387 Entertainment 0\.313 0\.378 0\.299 Travel & City Exploration 0\.272 0\.394 0\.339 Table 13: Per\-domain Interest Tag F1 for GPT\-5\.4 across the three profile\-construction settings\. Hierarchical profiling achieves the best Interest Tag F1 on six of seven domains; extractive–abstractive profiling is best on Pets\. Table 13 breaks down Interest Tag F1 by domain for GPT\-5\.4 across the three profile\-construction settings\. The seven domains form a clear difficulty spectrum\. Pets and Gaming sit at the top \(Direct F1: 0\.522 and 0\.385\), benefiting from concentrated, visually distinctive evidence—pet photos and game screenshots are unambiguous signals that models detect reliably\. Travel & City Exploration and Entertainment anchor the bottom \(Direct F1: 0\.272 and 0\.313\), reflecting the challenge of synthesizing evidence dispersed across many posts: a user’s travel interests may be scattered across dozens of posts mentioning different destinations, cuisines, and activities, with no single post fully defining the interest\. Hierarchical profiling improves Interest Tag F1 on six of seven domains, with the largest gains on difficult domains that require cross\-post aggregation\. Sports & Outdoor gains \+0\.164 \(0\.324 →\\rightarrow 0\.488\) and Travel gains \+0\.122 \(0\.272 →\\rightarrow 0\.394\), confirming that chunking helps aggregate weak but recurrent signals that the direct setting misses\. The exception is Pets, where Extractive profiling achieves the best result \(0\.575 vs\. Hierarchical 0\.497\)\. Pets evidence tends to be concentrated in a small number of high\-signal posts \(photos of the user’s own pets\), so selecting the top KK posts per domain preserves nearly all available evidence while filtering noise\. For text\-dispersed domains, however, Extractive profiling underperforms Hierarchical, as the selection stage must commit to a fixed set of posts before knowing which evidence will prove relevant\. The difficulty ranking is consistent across all three settings, suggesting that per\-domain hardness is an intrinsic property of how evidence is distributed across modalities and posts rather than an artifact of any particular profiling strategy\. Appendix J Human–LLM Dialogue Evaluation Agreement Study This appendix provides the detailed setup and full results of the human–LLM agreement study referenced in Section 4\.3\. J\.1 Sampling and Annotation Setup We sample 30 model responses from the full dialogue evaluation pool \(100 users ×\\times 4 settings ×\\times 6 models\) using stratified sampling across three dimensions to ensure coverage of the full quality and setting space: 1\. Dialogue setting: 7–8 responses from each of the four settings \(Timeline\-conditioned, Direct, Hierarchical, Extractive–abstractive\), ensuring representation of both raw\-timeline and profile\-conditioned generation paths\. 2\. User intent: 15 stable\-interest recommendation responses and 15 recent\-interest exploration responses\. 3\. Quality tier: 10 responses each from low \(GPT\-5\.5 Avg\. ≤2\.5\\leq 2\.5\), medium \(2\.5<Avg\.≤3\.52\.5<\\text\{Avg\.\}\\leq 3\.5\), and high \(Avg\. \>3\.5\>3\.5\) tiers, ensuring annotators see the full quality range\. Three annotators with prior experience on the SocialPersona annotation team \(see Section 3\.3\) independently score each response\. Each annotator receives a sheet containing, for every response: \(1\) the user’s gold profile summary \(stable and recent interests per domain\), \(2\) the user request text and intent type, \(3\) the model response text, and \(4\) the identical 0–5 scoring rubrics for coverage, concreteness, and fluency used by the LLM judges \(Appendix A\.6\)\. Annotators do not see LLM judge scores, model identities, or dialogue settings\. Total annotation time is approximately 5 annotator\-hours\. J\.2 Agreement Metrics We report two complementary agreement measures: Human–human agreement\. We compute Krippendorff’s α\\alpha for ordinal data across the three annotators on each dimension\. This validates whether the evaluation task itself is reliably human\-judgeable\. We interpret α\>0\.80\\alpha\>0\.80 as strong agreement, α∈\[0\.67,0\.80\]\\alpha\\in\[0\.67,0\.80\] as moderate, and α<0\.67\\alpha<0\.67 as tentative\. Human–LLM agreement\. For each response and dimension, we average the three annotator scores to form a human reference\. We then compute per\-dimension Spearman rank correlation ρ\\rho between this reference and each LLM judge \(GPT\-5\.5 and Qwen3\.7\-Max\)\. Spearman ρ\\rho captures monotonic ranking agreement without assuming linearity or interval\-scale properties of the 0–5 scores\. J\.3 Results Dimension Human–Human α\\alpha GPT\-5\.5 ρ\\rho Qwen3\.7\-Max ρ\\rho Fluency 0\.79 0\.76 0\.72 Concreteness 0\.71 0\.65 0\.58 Coverage 0\.65 0\.58 0\.52 Table 14: Human–human inter\-annotator agreement \(Krippendorff’s α\\alpha, 3 annotators\) and human–LLM judge Spearman correlations \(ρ\\rho\) on 30 sampled dialogue responses\. Human reference is the mean of three annotator scores\. All human–LLM correlations are significant at p<0\.01p<0\.01\. Table 14 reports the full results\. Three findings emerge: First, human–human agreement follows the expected difficulty ordering: fluency is most objective \(α=0\.79\\alpha=0\.79\), followed by concreteness \(α=0\.71\\alpha=0\.71\), with coverage being most subjective \(α=0\.65\\alpha=0\.65\)\. All three values fall within the moderate\-to\-substantial agreement range, confirming that the evaluation dimensions are reliably applicable by trained annotators\. Second, both LLM judges correlate positively and significantly with human judgments across all dimensions \(p<0\.01p<0\.01 for all ρ\\rho\), validating their use as automated proxies for dialogue quality assessment\. GPT\-5\.5 consistently outperforms Qwen3\.7\-Max in human alignment, consistent with its role as the primary judge in our main results \(Table 4\.2\)\. Third, the dimension\-level gap mirrors the human–human pattern: coverage shows the lowest agreement in both settings\. This is expected for a task requiring judges to assess whether a response meaningfully engages with specific user interests; borderline responses that mention a domain tangentially without clearly centering on a gold interest are inherently ambiguous\. The residual disagreement on coverage \(ρ=0\.52\\rho=0\.52–0\.580\.58\) sets a plausible upper bound on automated coverage evaluation precision, but the positive and significant correlations confirm that LLM judges capture the correct ranking signal for model comparison—sufficient for the benchmarking conclusions drawn in Section 4\.3\.`

Similar Articles

PersonaVLM: Long-Term Personalized Multimodal LLMs

Hugging Face Daily Papers

PersonaVLM introduces a personalized multimodal LLM framework that enables long-term user adaptation through memory retention, multi-turn reasoning, and response alignment, outperforming GPT-4o by 5.2% on the new Persona-MME benchmark.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv cs.AI

This paper introduces LUNAR, a benchmark for evaluating how large language models personalize responses from longitudinal app interaction histories across daily-life domains such as clothing, food, housing, and mobility. Experiments on 19 mainstream LLMs reveal that effective personalization depends on evidence selection and cross-domain integration, and that stronger personalization can come at the cost of privacy protection.