Creating an Atomic User Model for Personality-Aware Large Language Model Interaction

arXiv cs.CL Papers

Summary

This paper introduces the Atomic User Model (AUM), a structured representation for organizing user personality to enhance large language model interactions, showing improved personalization and context efficiency in simulations.

arXiv:2609.12086v1 Announce Type: cross Abstract: Assistants built on large language models are expected to write as their user would, and the dominant approach is single-channel: preferences summarised from conversation history and reinserted into context. This inverts the order of inference. Preferences are the task-dependent surface of a comparatively stable personality structure, so a system storing only preferences relearns the person whenever the task changes. First, we characterise personality seepage, where a prompt's linguistic surface carries a personality fingerprint the assistant mirrors without access to the personality behind it. Second, we propose the Atomic User Model (AUM), a human-readable representation organising a person as a stable identity nucleus with four interpretable shells (psychological, cognitive and experiential, behavioural, and social), plus cross-shell entries recording internal conflict and authenticity. Third, we treat AUM as a retrieval index over a person rather than a prompt prefix, with a pipeline where a task classifier, component-selection function and budgeted retriever return a small payload of fields at generation time. Fourth, we evaluate it with sixteen language-model-simulated participants, six style-sensitive tasks and three seeds, plus a synthetic scaling study of the retriever. Retrieving eight fields matched the style fidelity of the full user model on 23% of the context (211 tokens against 915), improved on flat preference notes by 0.24 points on a five-point scale (p < 0.001, dz = 0.50), and raised forced-choice identification of the participant's own voice from 14.9% to 42.7% (25% chance). Four pre-registered controls returned null, locating the effect in the representation rather than the search over it. The benefit is largest for participants the un-personalised assistant reproduces worst (rho = -0.61, p = 0.013): personalisation is worth most to those the default serves least.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:46 AM

# Creating an Atomic User Model for Personality-Aware Large Language Model Interaction
Source: [https://arxiv.org/html/2609.12086](https://arxiv.org/html/2609.12086)
B\. Sankar[0000\-0001\-5844\-6273](https://orcid.org/0000-0001-5844-6273)Email:[sankarb@iisc\.ac\.in](mailto:[email protected])Corresponding author:Corresponding author: B\. SankarCorresponding author:Corresponding author email: sankarb@iisc\.ac\.inAffiliation:Department of Mechanical Engineering, Indian Institute of Science \(IISc\), Bengaluru, 560012, Karnataka, IndiaDeepthika S[0009\-0009\-4386\-6510](https://orcid.org/0009-0009-4386-6510)Email:[deepthikas123@gmail\.com](mailto:[email protected])Affiliation:Department of Design and Manufacturing, Indian Institute of Science \(IISc\), Bengaluru, 560012, Karnataka, IndiaPawni YadavEmail:[pawniyadav3435@gmail\.com](mailto:[email protected])Affiliation:Department of Design and Manufacturing, Indian Institute of Science \(IISc\), Bengaluru, 560012, Karnataka, IndiaAmogh A SEmail:[amogh\.setty07@gmail\.com](mailto:[email protected])Affiliation:Department of Design and Manufacturing, Indian Institute of Science \(IISc\), Bengaluru, 560012, Karnataka, India

###### Abstract

Assistants built on large language models are increasingly expected to write as their user would write, and the dominant approach to that expectation operates on a single channel: past preferences are summarised out of conversation history and reinserted into the context window\. We argue that this inverts the natural order of inference\. Preferences are the task\-dependent surface of an underlying personality structure that is comparatively stable, so a system that stores only preferences must relearn the person whenever the task changes\. This paper makes four contributions\. First, we characterise a phenomenon we call*personality seepage*, in which the linguistic surface of a prompt carries a personality fingerprint that the assistant mirrors without any access to the personality that produced it\. Second, we propose the Atomic User Model \(AUM\), a structured, human\-readable representation that organises a person as a stable identity nucleus surrounded by four interpretable shells covering psychological, cognitive and experiential, behavioural, and social content, together with cross\-shell entries recording internal conflict and authenticity\. Third, we treat AUM as a retrieval index over a person rather than a prompt prefix, and specify a personality\-aware pipeline in which a task classifier, a component\-selection function and a budgeted retriever return a small payload of fields at generation time\. Fourth, we evaluate the pipeline in a simulation study with sixteen language\-model\-simulated participants, six style\-sensitive tasks and three seeds, and in a synthetic scaling study of the retriever itself\. Retrieving eight fields matched the style fidelity of injecting the entire user model while using 23% of the context \(211 tokens against 915\), improved on flat preference notes by 0\.24 points on a five\-point scale \(p<0\.001p<0\.001,dz=0\.50d\_\{z\}=0\.50\), and raised forced\-choice identification of the participant’s own voice from 14\.9% to 42\.7% against 25% chance\. Four pre\-registered controls returned null, including the swarm retriever against forward greedy, which locates the effect in the structured representation rather than in the search over it\. A supplementary analysis shows the benefit is largest for participants whose voice the un\-personalised assistant reproduces worst \(ρ=−0\.61\\rho=\-0\.61,p=0\.013p=0\.013\), which is a design result rather than a statistical one: personalisation is worth most to the people the default serves least\. We discuss what a budgeted payload buys for privacy\-by\-architecture, and what a simulation can and cannot establish about whether real people recognise themselves\.

###### Keywords:

user modelling , personalisation , personality , large language models , retrieval augmented generation , context budget , scrutability , privacy by architecture , simulation study , artificial fish swarm algorithm

## 1Introduction

Adaptive systems have modelled their users for as long as they have adapted to them\. The generic user modelling tradition in human\-computer interaction established that a user model is a reusable component with an explicit schema, maintained separately from the application that consumes it\([Kobsa, 2001](https://arxiv.org/html/2609.12086#bib.bib11);[Brusilovsky and Millán, 2007](https://arxiv.org/html/2609.12086#bib.bib12)\), and personalised search built profiles from observed behaviour: interests and activities harvested from a desktop index\([Teevan et al\., 2005](https://arxiv.org/html/2609.12086#bib.bib28)\), click and query histories evaluated at scale\([Dou et al\., 2007](https://arxiv.org/html/2609.12086#bib.bib29)\), and short\-term session signals combined with long\-term interest\([Bennett et al\., 2012](https://arxiv.org/html/2609.12086#bib.bib30)\)\. Assistants built on large language models \(LLMs\) have inherited this design in a compressed form\. Deployed memory features distil past conversations into short preference notes and reinsert them into later contexts, and the research literature has followed with memory architectures\([Packer et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib3);[Zhou et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib31);[Liao et al\., 2026](https://arxiv.org/html/2609.12086#bib.bib36)\), per\-user parameter\-efficient adaptation\([Tan et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib32)\), and surveys that map the resulting space\([Tan and Jiang, 2023](https://arxiv.org/html/2609.12086#bib.bib10);[Zhang et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib33)\)\.

What these approaches share is not an algorithm but a unit of storage\. They store what a user has done, or has said they want\. They do not store the person from whom those wants emerge\. The difference is easiest to see when a stored preference goes stale\.

> A user plans a six\-day conference trip to Paris and discusses cost of living and itineraries with an assistant\. Two weeks later the user privately decides not to travel, and does not say so\. A month after that, while drafting a project timeline, the assistant blocks out the original travel week as an absence\.

The system is not wrong about its source: it remembers the preference faithfully, and the retrieved item was exactly the right item\. It is wrong about the person\. It holds no representation of the user as a decision maker whose plans get revised, so it has no basis on which to flag a stale assumption\. Retrieval quality is not the failure here, and better retrieval over the same unit would not have prevented it\.

This paper argues that the missing layer is a different unit of storage\. The argument rests on a claim borrowed from personality psychology rather than from information retrieval\. Personality traits show substantial rank\-order stability across the lifespan, with test\-retest correlations rising from roughly 0\.3 in childhood to 0\.7 in later adulthood\([Roberts and DelVecchio, 2000](https://arxiv.org/html/2609.12086#bib.bib89)\), and mean\-level change, while real, is gradual and patterned\([Roberts et al\., 2006](https://arxiv.org/html/2609.12086#bib.bib90);[Specht et al\., 2011](https://arxiv.org/html/2609.12086#bib.bib91)\)\. Preferences, by contrast, are situational: they are what personality produces when it meets a particular decision context\. If the stable layer is the one that generalises across tasks, then a system that stores only the unstable layer is storing the wrong thing\.

We make four contributions\.

1. 1\.Personality seepage\.We characterise a phenomenon in which a user’s linguistic and behavioural patterns enter the assistant through the surface of their prompt, and the assistant mirrors that surface while having no access to the personality that produced it \(Section[3](https://arxiv.org/html/2609.12086#S3)\)\. Seepage is why generic assistants can appear briefly personalised and then fail on exactly the parts of a task the user did not spell out\.
2. 2\.The Atomic User Model\.We propose a structured, human\-readable representation that organises a person as a stable identity Nucleus surrounded by four shells covering psychological, cognitive and experiential, behavioural, and social content, plus cross\-shell entries recording internal conflict, growth trajectory and an authenticity index \(Section[4](https://arxiv.org/html/2609.12086#S4)\)\. The metaphor is not decorative: it encodes two design commitments, differential stability and differential observability, that a flat schema cannot express\.
3. 3\.Personality\-aware retrieval\.We treataumas an index over a person rather than a prompt prefix, and specify a pipeline in which a task classifier, a component\-selection functionσ⁡\(t\)\\sigma\(t\)and a budgeted retriever return a payload ofkkfields at generation time \(Section[5](https://arxiv.org/html/2609.12086#S5)\)\. The retrieval objective carries an explicit redundancy penalty, which is what makes payload selection a combinatorial problem rather than a sort\.
4. 4\.Two studies\.We evaluate the pipeline with sixteen LLM\-simulated participants across six style\-sensitive tasks and three seeds \(Sections[6](https://arxiv.org/html/2609.12086#S6)and[7](https://arxiv.org/html/2609.12086#S7)\), and we study the retrieval problem on its own in a synthetic sweep over instrument size, budget and redundancy weight \(Section[8](https://arxiv.org/html/2609.12086#S8)\)\.

The headline empirical result is a context result\. Eight retrieved fields matched the style fidelity of injecting the entire 32\-field user model while using 23% of the injected context, and improved on flat preference notes by 0\.24 points on a five\-point scale\. Forced\-choice identification of the participant’s own voice rose from 14\.9% under preference notes, which is not distinguishable from chance, to 42\.7% against a 25% baseline\.

Four controls returned null, and we report them in full because together they bound the claim\. Selecting the rightkkfields was not distinguishable from selectingkkat random from the same admissible pool; component selection contributed nothing beyond retrieval; the swarm retriever neither beat forward greedy on generated style nor reached a higher value of its own objective; and the redundancy term did not earn its place against a relevance\-only top\-kk\. The scaling study of Section[8](https://arxiv.org/html/2609.12086#S8)explains why, and the explanation is specific: at the operating point the simulation used, forward greedy is already within 0\.1% of the exhaustive optimum, so there was nothing for a population method to recover\. What the positive results measure is the structured representation, not the search over it\.

A supplementary analysis produced the finding we think matters most for a human\-computer interaction audience\. The benefit of personalisation was largest exactly for those participants whose voice the un\-personalised assistant reproduced worst \(Spearmanρ=−0\.61\\rho=\-0\.61,p=0\.013p=0\.013\)\. Personalisation is worth most to the people the default serves least, which is a statement about who a system is for rather than about how well it scores\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_inversion.png)Figure 1:The inversion this paper proposes\. Left: current personalisation accumulates preferences observed at the surface and reinserts them, so an unseen task has no covering note and a stale note cannot be re\-interpreted\. Right:aumstores the comparatively stable personality structure and derives the situational preference at generation time, retrieving only the fields the task needs\.Figure[1](https://arxiv.org/html/2609.12086#S1.F1)states the inversion in one picture\.

The remainder of the paper is organised as follows\. Section[2](https://arxiv.org/html/2609.12086#S2)situates the work\. Section[3](https://arxiv.org/html/2609.12086#S3)defines personality, preferences and persona and introduces seepage\. Section[4](https://arxiv.org/html/2609.12086#S4)specifiesaum\. Section[5](https://arxiv.org/html/2609.12086#S5)formalises personality\-aware retrieval\. Sections[6](https://arxiv.org/html/2609.12086#S6)and[7](https://arxiv.org/html/2609.12086#S7)present the simulation study and its results\. Section[8](https://arxiv.org/html/2609.12086#S8)reports the scaling study\. Section[9](https://arxiv.org/html/2609.12086#S9)discusses implications and limitations, and Section[11](https://arxiv.org/html/2609.12086#S11)concludes\. Appendices give the full specification, every prompt used, the persona set and the complete statistical tables\.

## 2Related Work

### 2\.1User models as components in adaptive systems

The idea that a user model should be a separable, inspectable component predates the systems this paper is about\.[Kobsa \(2001\)](https://arxiv.org/html/2609.12086#bib.bib11)surveyed generic user modelling systems and argued for a shared server holding assumptions about the user, queried by applications rather than duplicated inside them\.[Brusilovsky and Millán \(2007\)](https://arxiv.org/html/2609.12086#bib.bib12)developed the same position for adaptive hypermedia and educational systems, where an overlay model over a domain structure gives the adaptation something explicit to reason about\. Information retrieval developed a parallel but distinct tradition, in which the user model is a statistical summary of behaviour rather than a schema: term vectors and topical distributions built from a desktop index\([Teevan et al\., 2005](https://arxiv.org/html/2609.12086#bib.bib28)\), click and query histories whose value for ranking turns out to vary sharply across queries\([Dou et al\., 2007](https://arxiv.org/html/2609.12086#bib.bib29)\), and a decomposition of behavioural evidence into short\-term session signals and long\-term interest\([Bennett et al\., 2012](https://arxiv.org/html/2609.12086#bib.bib30)\)\. Diversity\-aware ranking developed alongside, first as maximal marginal relevance\([Carbonell and Goldstein, 1998](https://arxiv.org/html/2609.12086#bib.bib42)\), then as an explicit evaluation concern\([Clarke et al\., 2008](https://arxiv.org/html/2609.12086#bib.bib43)\)and an optimisation target\([Agrawal et al\., 2009](https://arxiv.org/html/2609.12086#bib.bib44)\)\.

aumbelongs to the first tradition and borrows from the second\. It is a schema, in Kobsa’s sense, whose slots are declared in advance rather than induced from data\. What it takes from information retrieval is the insight that a profile is only useful to the extent that a system can select from it, and that selecting well means trading relevance against redundancy\.

### 2\.2Memory and personalisation in LLM assistants

Recent work recasts personalisation of LLM assistants as memory management\. MemGPT treats the model as an operating system with paged memory\([Packer et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib3)\); generative agents combine memory, reflection and planning to produce believable behaviour\([Park et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib4)\); medical and lifelong assistants coordinate short\-term and long\-term stores\([Zhang et al\., 2024a](https://arxiv.org/html/2609.12086#bib.bib1);[Wang et al\., 2024b](https://arxiv.org/html/2609.12086#bib.bib2)\)\. Closer to retrieval, cognitive personalised search couples an LLM with an efficient memory mechanism over query logs\([Zhou et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib31)\), profile memory organises per\-user summaries for reuse\([Liao et al\., 2026](https://arxiv.org/html/2609.12086#bib.bib36)\), and agent\-driven retrieval personalises search results directly\([Chhetri et al\., 2026](https://arxiv.org/html/2609.12086#bib.bib35)\)\. A separate branch adapts model parameters per user rather than the context, through per\-user parameter\-efficient fine\-tuning\([Tan et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib32)\); a recent survey maps the whole space\([Zhang et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib33)\)\.

Evaluation has developed in step\. LaMP\([Salemi et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib13)\)and its long\-form extension\([Kumar et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib14)\)define personalisation as task\-level retrieval over user history\. PersonaBench shows that retrieval\-augmented systems answer questions about private user data poorly even when given direct document access\([Tan et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib15)\), and rating\-prediction studies find that LLMs capture user preferences but in opaque ways\([Kang et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib16);[Huang et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib34)\)\.

These systems answer how to store and fetch user\-relevant content\.aumanswers a prior question: what the stored items should be, and how they should be organised so that fetching a few of them is enough\. The two are complementary, and the pipeline of Section[5](https://arxiv.org/html/2609.12086#S5)is deliberately built from standard components so that the representation, rather than the machinery, is what is under test\.

### 2\.3Context budgets and compression

A structured user model of sixty fields cannot be injected on every turn\. The problem of fitting useful conditioning into a bounded context has been attacked from the compression side: LLMLingua and its long\-context successor learn to drop tokens from a prompt while preserving downstream performance\([Jiang et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib47);[Jiang et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib48)\), and selective\-context methods prune low\-information spans\([Li et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib49)\)\. Our approach is complementary and operates a level up\. Rather than compressing a fixed payload, we select which fields enter the payload at all, conditioned on the task\. The two could compose: compression could shorten the eight fields that selection returns\.

### 2\.4Personality: structure, stability, and its trace in language

The Five\-Factor Model provides the trait vocabulary most widely used in computational work\([McCrae and John, 1992](https://arxiv.org/html/2609.12086#bib.bib22)\), with modern instruments including the BFI\-2\([Soto and John, 2017](https://arxiv.org/html/2609.12086#bib.bib92)\), public\-domain item pools\([Goldberg et al\., 2006](https://arxiv.org/html/2609.12086#bib.bib93)\)and very brief measures for settings where length is prohibitive\([Gosling et al\., 2003](https://arxiv.org/html/2609.12086#bib.bib94)\)\. Cross\-cultural generalisation of the factor structure has been examined but remains a live concern\([McCrae, 2002](https://arxiv.org/html/2609.12086#bib.bib23)\)\. We include Myers\-Briggs type as a field inside Shell 1 because users often know their type, while noting that its psychometric properties are weak: the type dichotomies do not correspond to natural categories, and reliability across retests is poor\([Pittenger, 2005](https://arxiv.org/html/2609.12086#bib.bib96);[Stein and Swan, 2019](https://arxiv.org/html/2609.12086#bib.bib97);[McCrae and Costa, 1989](https://arxiv.org/html/2609.12086#bib.bib95)\)\. This is precisely the argument for a schema in which an inventory is one field among many rather than the representation itself\.

That personality leaves a measurable trace in language is well established\.[Pennebaker and King \(1999\)](https://arxiv.org/html/2609.12086#bib.bib98)showed that function\-word use is a stable individual difference;[Yarkoni \(2010\)](https://arxiv.org/html/2609.12086#bib.bib100)related Big Five scores to word use across 100,000 words of blog text;[Schwartz et al\. \(2013\)](https://arxiv.org/html/2609.12086#bib.bib101)moved from closed\-vocabulary counting to open\-vocabulary differential language analysis over social media, and the LIWC line of work supplies the standard closed\-vocabulary instrument\([Tausczik and Pennebaker, 2010](https://arxiv.org/html/2609.12086#bib.bib99);[Boyd and Pennebaker, 2017](https://arxiv.org/html/2609.12086#bib.bib102)\)\. Prediction of traits from digital footprints reaches useful accuracy\([Kosinski et al\., 2013](https://arxiv.org/html/2609.12086#bib.bib103);[Park et al\., 2015](https://arxiv.org/html/2609.12086#bib.bib105)\), sometimes exceeding human judges given enough behavioural evidence\([Youyou et al\., 2015](https://arxiv.org/html/2609.12086#bib.bib104)\), though meta\-analysis places the typical correlation more modestly\([Azucar et al\., 2018](https://arxiv.org/html/2609.12086#bib.bib106)\), and earlier attempts to model personality from social media profile data illustrate how far the field has moved\([Goyal and Tawde, 2022](https://arxiv.org/html/2609.12086#bib.bib27)\)\. Within LLM research, trait recognition from text is feasible\([Ji et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib5);[Rao et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib6)\)and implicit persona can be modelled as a latent variable in dialogue\([Cho et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib7)\), but the resulting representations are trait scores or latent vectors: neither modular nor inspectable\.

The reliance on self\-report in Section[6](https://arxiv.org/html/2609.12086#S6)inherits a known limitation\. Self and other observers know different things about a person, with the self better on internal states and observers better on externally visible traits\([Vazire, 2010](https://arxiv.org/html/2609.12086#bib.bib107);[Connelly and Ones, 2010](https://arxiv.org/html/2609.12086#bib.bib108)\), and self\-report has long been criticised as a substitute for observed behaviour\([Baumeister et al\., 2007](https://arxiv.org/html/2609.12086#bib.bib109)\)\. We return to this in Section[9\.5](https://arxiv.org/html/2609.12086#S9.SS5)\.

### 2\.5Structured user representations and human digital twins

The closest contemporary architecture is the General User Model, which accumulates confidence\-weighted natural\-language propositions about a user from screen observations\([Shaikh et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib8)\)\. Generative agents grounded in self\-report interviews simulate individual humans better than demographic conditioning alone\([Park et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib9)\), which is both the closest precedent for our representation and the precedent for our evaluation method\. Both populate a representation bottom up, from observation or interview\.aumis top down: it specifies the slots into which such observations should be organised, which is what makes a component\-selection function over shells definable at all\.

The human digital twin literature pursues a superficially similar goal from an engineering direction, and a recent survey maps its scope\([Lin et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib79)\)\. The personal informatics tradition in HCI supplies the complementary user\-centred account of what it is like to live with a model of oneself: staged models of collection and reflection\([Li et al\., 2010](https://arxiv.org/html/2609.12086#bib.bib76)\), the lived\-informatics revision that treats tracking as episodic rather than linear\([Epstein et al\., 2015](https://arxiv.org/html/2609.12086#bib.bib77)\), and evidence on how people without prior self\-tracking experience engage with their own data\([Rapp and Cena, 2016](https://arxiv.org/html/2609.12086#bib.bib78)\)\.aumdiffers from a digital twin in purpose\. It is built for style\-sensitive task completion, not identity substitution, and Section[9](https://arxiv.org/html/2609.12086#S9)argues that the distinction has consequences the twin framing obscures\.

### 2\.6Scrutability, control, and the privacy of a user model

If a system holds a model of a person, the person should be able to read it\.[Kay and Kummerfeld \(2012\)](https://arxiv.org/html/2609.12086#bib.bib70)set out the drivers and principles for personalised systems that users can scrutinise and control, building on earlier work on scrutable adaptive hypertext\([Czarkowski and Kay, 2002](https://arxiv.org/html/2609.12086#bib.bib71)\)and on interfaces for inspecting semantic user models\([Bakalov et al\., 2010](https://arxiv.org/html/2609.12086#bib.bib72)\)\.[Cramer et al\. \(2008\)](https://arxiv.org/html/2609.12086#bib.bib73)found that transparency affected acceptance of a content\-based recommender, and[Harper et al\. \(2015\)](https://arxiv.org/html/2609.12086#bib.bib74)showed that giving users direct control over their recommendations changed both behaviour and satisfaction\.[Jeromela \(2022\)](https://arxiv.org/html/2609.12086#bib.bib75)extends the question to intelligent personal assistants specifically\. Mental\-model work supplies the reason this matters: users arrive with expectations of conversational agents that the systems do not meet\([Luger and Sellen, 2016](https://arxiv.org/html/2609.12086#bib.bib80);[Mahmood et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib83)\), and their models of what an agent knows shape how they use it\([Gero et al\., 2020](https://arxiv.org/html/2609.12086#bib.bib81);[Ngo et al\., 2020](https://arxiv.org/html/2609.12086#bib.bib82)\)\.

The privacy dimension is sharper foraumthan for a preference list, because the content is more intimate\. Contextual integrity supplies the governing frame: what matters is not secrecy but whether an information flow matches the norms of the context it came from\([Nissenbaum, 2011](https://arxiv.org/html/2609.12086#bib.bib66)\)\. Empirically, users of LLM\-based conversational agents disclose a great deal and reason about the risks only partially\([Zhang et al\., 2024b](https://arxiv.org/html/2609.12086#bib.bib67)\), and language models can leak personal information present in training data\([Huang et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib68)\)\. On\-device and federated approaches offer one architectural answer\([Chen et al\., 2019](https://arxiv.org/html/2609.12086#bib.bib69)\)\. Our position, developed in Sections[5](https://arxiv.org/html/2609.12086#S5)and[10](https://arxiv.org/html/2609.12086#S10), is that a budgeted payload converts local residence from a policy promise into a measurable property: if eight fields suffice, the other twenty\-four never leave the device\.

### 2\.7Writing assistance and the user’s voice

The application setting for this work is writing assistance, where the question of whose voice appears in the output is not incidental\. CoAuthor documented human\-LLM collaborative writing at scale\([Lee et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib84)\); Wordcraft explored story writing with a model in the loop\([Yuan et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib85)\); a recent design space synthesises the field\([Lee et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib87)\)and a companion study aligns research directions with what writers actually ask for\([Reza et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib88)\)\. The finding that most directly motivatesaumis[Jakesch et al\. \(2023\)](https://arxiv.org/html/2609.12086#bib.bib86): co\-writing with an opinionated language model shifts the user’s own expressed views\. If a default assistant voice can move what a person says, then a system that instead conditions on a model of that person is not merely a convenience feature\.

### 2\.8Selecting a subset under a budget

Choosingkkfields to maximise a payload score with a redundancy penalty is a constrained subset\-selection problem\. When such an objective is monotone submodular, the forward greedy algorithm carries the classical\(1−1/e\)\(1\-1/e\)guarantee\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.12086#bib.bib41)\), which is why greedy is the right baseline rather than a straw man\. Our objective, defined in Section[5](https://arxiv.org/html/2609.12086#S5), uses a mean rather than a sum and subtracts a mean pairwise similarity, so it is not submodular in general and the guarantee does not transfer; the empirical question of how close greedy comes is therefore live, and Section[8](https://arxiv.org/html/2609.12086#S8)answers it directly\.

Population\-based metaheuristics are widely applied to subset selection, particularly feature selection\([Xue et al\., 2016](https://arxiv.org/html/2609.12086#bib.bib45);[Nguyen et al\., 2020](https://arxiv.org/html/2609.12086#bib.bib46)\)\. The artificial fish swarm algorithm\([Li et al\., 2002](https://arxiv.org/html/2609.12086#bib.bib38)\)models a population of agents executing prey, swarm and follow behaviours, and has accumulated a substantial literature of variants and applications\([Neshat et al\., 2014](https://arxiv.org/html/2609.12086#bib.bib39);[Pourpanah et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib40)\)\. We use it, and we report that at the scale our simulation ran it did not earn its cost\. We think reporting that is more useful than dropping the comparison, and Section[8](https://arxiv.org/html/2609.12086#S8)identifies the regime in which the answer changes\.

### 2\.9Simulated participants and model\-based judgement

Using language models in place of human participants is now a small field with a sharp internal debate\.[Argyle et al\. \(2023\)](https://arxiv.org/html/2609.12086#bib.bib55)showed that conditioning on demographic backstories reproduces aggregate patterns in survey data, and[Park et al\. \(2024\)](https://arxiv.org/html/2609.12086#bib.bib9)showed that grounding agents in structured self\-report improves individual\-level simulation\. The critical literature is equally developed:[Bisbee et al\. \(2024\)](https://arxiv.org/html/2609.12086#bib.bib56)document instability and misestimation in synthetic survey responses, and[Wang et al\. \(2025\)](https://arxiv.org/html/2609.12086#bib.bib57)show that replacing human participants can misportray and flatten identity groups\. We adopt simulation as a feasibility method with explicit limits, stated in Section[6\.1](https://arxiv.org/html/2609.12086#S6.SS1), and we do not present any result here as evidence about human users\.

Model\-based judgement carries its own hazards, all of which we control for by construction\. Judges exhibit position bias\([Wang et al\., 2024a](https://arxiv.org/html/2609.12086#bib.bib52)\), self\-preference for their own generations\([Panickssery et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib51)\), and imperfect but usable agreement with human raters\([Zheng et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib50);[Chiang and Lee, 2023](https://arxiv.org/html/2609.12086#bib.bib53);[Liu et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib54)\)\. Section[6\.7](https://arxiv.org/html/2609.12086#S6.SS7)describes the specific controls: a judge from a different model family than the generator, per\-trial option shuffling, and a judge that never sees the user model it is implicitly evaluating\.

For the automatic style measure we use authorship representations rather than semantic embeddings\. LUAR learns universal authorship representations by contrastive training over authors\([Rivera\-Soto et al\., 2021](https://arxiv.org/html/2609.12086#bib.bib37)\), extending earlier invariant representations of social\-media users\([Andrews and Bishop, 2019](https://arxiv.org/html/2609.12086#bib.bib58)\)\. The distinction matters here: a semantic embedding scores any two deadline\-extension emails as similar because they are about the same thing, which is exactly the wrong invariance\. Work on what authorship representations actually encode\([Wang et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib60);[Wegmann et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib59)\), on making style embeddings interpretable\([Patel et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib62)\), and on robustness\([Man and Nguyen, 2024](https://arxiv.org/html/2609.12086#bib.bib61)\)informs how we read the resulting numbers\. Style\-transfer evaluation supplies the broader methodological caution that automatic style metrics are easy to misreport\([Mir et al\., 2019](https://arxiv.org/html/2609.12086#bib.bib63);[Ostheimer et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib64)\), and fine\-grained linguistic control offers an alternative route to the same goal\([Alhafni et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib65)\)\.

### 2\.10Cognitive architectures

Finally, the design philosophy behind a fixed, modular schema comes from the cognitive architecture tradition\. ACT\-R\([Anderson et al\., 2004](https://arxiv.org/html/2609.12086#bib.bib18)\)and Soar\([Laird, 2012](https://arxiv.org/html/2609.12086#bib.bib19)\)are long\-standing arguments that explicit structure buys something that implicit competence does not, and[Sun \(2024\)](https://arxiv.org/html/2609.12086#bib.bib17)makes the case that such architectures and LLMs are complementary rather than competing\.aumapplies the philosophy not to cognition in general but to the narrower problem of representing the user inside an LLM\-mediated system\.

## 3Personality, Preferences, Persona, and Seepage

### 3\.1Three terms used in a specific way

Work on personalisation uses*personality*,*preference*and*persona*loosely and often interchangeably\. Because the argument of this paper is precisely about the relation between them, we fix the terms\.

Personalityis the structured set of stable, task\-independent traits spanning the cognitive, affective, behavioural and social\-contextual dimensions of a person\. Stability is meant in the rank\-order sense established empirically by[Roberts and DelVecchio \(2000\)](https://arxiv.org/html/2609.12086#bib.bib89): a person’s position relative to others changes slowly, even where absolute levels drift with age and life events\([Roberts et al\., 2006](https://arxiv.org/html/2609.12086#bib.bib90);[Specht et al\., 2011](https://arxiv.org/html/2609.12086#bib.bib91)\)\.

Preferencesare the task\-dependent outputs of personality applied to a particular decision context\. “Prefers bullet points in status updates” is a preference\. It is downstream of conscientiousness, of a professional register learned in a particular workplace, and of a belief about what respects a reader’s time; it is not itself any of those things\.

Personais the externally perceived projection of personality through one channel: a professional profile, a social media account, a conference biography\. Goffman’s account of self\-presentation\([Goffman, 1959](https://arxiv.org/html/2609.12086#bib.bib24)\)is the reference point, and the key property is that a persona is a selective and audience\-dependent rendering rather than a compressed copy\.

aummodels personality\. Preferences are derived from personality and context at generation time\. Personas are situational projections and are represented, when they are represented at all, as fields inside Shell 4\.

The ordering matters for a practical reason\. If preferences are the stored unit, then every new task requires either a stored preference that happens to cover it or a guess\. If personality is the stored unit, a preference for an unseen task is derivable, and a stale stored preference can be re\-interpreted against a structure that did not change\. The Paris example of Section[1](https://arxiv.org/html/2609.12086#S1)is exactly a case where the second operation was needed and unavailable\.

### 3\.2Personality seepage

Table 1:Personality seepage on a single task: an email requesting a deadline extension from a colleague\. Three users phrase the same underlying request differently and the assistant mirrors each surface\. The third column is the point: what the generic response misses is in each case a property of the user, not of the prompt\.Table[1](https://arxiv.org/html/2609.12086#S3.T1)illustrates a phenomenon we call*personality seepage*\. When a person writes a prompt, their linguistic and behavioural patterns go into the prompt with them, and the model amplifies those patterns in its output\. That the trace is there is not in doubt: function\-word use is a stable individual difference\([Pennebaker and King, 1999](https://arxiv.org/html/2609.12086#bib.bib98)\), and personality is recoverable from text at useful accuracy\([Yarkoni, 2010](https://arxiv.org/html/2609.12086#bib.bib100);[Schwartz et al\., 2013](https://arxiv.org/html/2609.12086#bib.bib101)\)\.

The trouble is what the assistant does with it\. It matches the surface without access to what produced the surface\. Three consequences follow, and each appears in Table[1](https://arxiv.org/html/2609.12086#S3.T1)\.

1. 1\.Gaps are not filled\.Whatever the user habitually omits from a prompt is also absent from the output, because the model has no independent source for it\. U3’s missing deadline is the example\.
2. 2\.Surface features are over\-extrapolated\.A hedging register in the prompt becomes a hedging register in the output, amplified, because the model treats the surface as the target rather than as evidence\. U2’s apology is the example\.
3. 3\.Stable habits are not reproduced\.Regularities the user applies across every instance of a genre, which never appear in any single prompt because the user assumes them, are invisible\. U1’s crediting of collaborators is the example\.

Seepage also explains a common subjective experience: a generic assistant feels briefly personalised, because it does echo something real about the user, and then fails on exactly the parts of a task the user did not spell out\. The echo is real; the model beneath it is absent\.

### 3\.3Why accumulating preferences treats the symptom

The remedy that deployed systems offer for seepage is to accumulate preference notes from past conversations and reinsert them\. This is a reasonable engineering response to the observation that the assistant forgets, and it does address forgetting\. It does not address seepage, because the notes are observed at the level of preferences while the cause of variation across users sits one layer below\.

Two properties of accumulated notes make this concrete\. First, they are*topically bound*: a note distilled from a conversation about choosing a laptop is about laptops, and generalises to an unrelated writing task only by accident\. Section[7](https://arxiv.org/html/2609.12086#S7)gives an empirical form of this observation, in that the preference\-notes condition was not distinguishable from chance on forced\-choice identification of the user’s own voice\. Second, they are*uninterpretable against each other*: a flat list has no structure that would let a system notice that two notes conflict, or that one has gone stale, or that a third is a surface expression of the same underlying disposition as the first two\. Structure is what makes those operations definable, which is the argument for a schema and the subject of the next section\.

## 4The Atomic User Model

### 4\.1The metaphor, and what it commits us to

aumrepresents a person as a Nucleus surrounded by four shells, on the metaphor of an atom with a stable centre and reactive outer layers \(Figure[3](https://arxiv.org/html/2609.12086#S4.F3)\)\. The metaphor is doing work rather than decorating, and it commits the representation to two properties that a flat schema cannot express\.

Differential stability\.Identity is stable while behaviour adapts\. The Nucleus changes on the timescale of years, the inner shells on months, the outer shells on weeks\. A representation that treats every field as equally durable will either refresh core values as often as it refreshes a communication habit, which is wasteful and invasive, or refresh a communication habit as rarely as it refreshes core values, which makes it wrong\.

Differential observability\.A stranger sees the outermost shell, a colleague sees more, a close friend or family member sees much more, and the Nucleus is visible only to the self\. This layering echoes the onion model of social penetration theory\([Altman and Taylor, 1973](https://arxiv.org/html/2609.12086#bib.bib20);[Carpenter and Greene, 2015](https://arxiv.org/html/2609.12086#bib.bib21)\), and it has a direct computational consequence: the fields most useful for deep personalisation are exactly the fields most costly to expose\. A representation that does not encode the gradient cannot reason about the trade\-off\.

Both properties bear on Section[5](https://arxiv.org/html/2609.12086#S5)\. The first motivates differential update rates as a field attribute\. The second is why a budgeted payload is a privacy mechanism and not only an efficiency one\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_update_rates.png)Figure 2:Differential stability across tiers\. The nominal update periodτ\\taufalls from years at the Nucleus to weeks at Shell 4, while observability by others rises in the opposite direction\. A schema that does not encode this gradient cannot express a refresh policy, and cannot reason about the cost of exposing a field\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/aum_model.png)Figure 3:The Atomic User Model\. A Nucleus of core identity is surrounded by four shells\. Example fields per shell are shown in italics\. Stability decreases and observability by others increases from the centre outwards\.
### 4\.2Formal definition

Theaumof a useruuis a set of typed fields

Au=\{f1,f2,…,fn\},A\_\{u\}=\\\{f\_\{1\},f\_\{2\},\\ldots,f\_\{n\}\\\},\(1\)where each fieldfif\_\{i\}is a tuple

fi=⟨shell⁡\(fi\),name⁡\(fi\),value⁡\(fi\),τ⁡\(fi\)⟩\.f\_\{i\}=\\langle\\,\\mathrm\{shell\}\(f\_\{i\}\),\\;\\mathrm\{name\}\(f\_\{i\}\),\\;\\mathrm\{value\}\(f\_\{i\}\),\\;\\tau\(f\_\{i\}\)\\,\\rangle\.\(2\)Hereshell⁡\(fi\)∈𝒮\\mathrm\{shell\}\(f\_\{i\}\)\\in\\mathcal\{S\}with

𝒮=\{Nucleus,S​h​e​l​l​1,S​h​e​l​l​2,S​h​e​l​l​3,S​h​e​l​l​4,CrossShell\},\\mathcal\{S\}=\\\{\\text\{Nucleus\},\\,Shell~1,\\,Shell~2,\\,Shell~3,\\,Shell~4,\\,\\text\{CrossShell\}\\\},\(3\)name⁡\(fi\)\\mathrm\{name\}\(f\_\{i\}\)is a fixed slot name drawn from the specification,value⁡\(fi\)\\mathrm\{value\}\(f\_\{i\}\)is a short natural\-language string, andτ⁡\(fi\)\\tau\(f\_\{i\}\)is a nominal update period\. Three properties of this definition are load\-bearing\.

1. 1\.Values are natural language, not vectors\.A field reads “I credit whoever reviewed a draft before I ask them for anything else”, not a coordinate\. This is what makes the model scrutable in the sense of[Kay and Kummerfeld \(2012\)](https://arxiv.org/html/2609.12086#bib.bib70): the user can read, edit or delete any field without an interpretation layer\.
2. 2\.Slots are declared in advance\.The set of\(shell,name\)\(\\mathrm\{shell\},\\mathrm\{name\}\)pairs is fixed by the specification, so two users’ models are structurally comparable and a selection function over shells is definable\. This is the property that distinguishesaumfrom bottom\-up accumulations of free\-form propositions\([Shaikh et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib8)\)\.
3. 3\.Update rate is a field attribute\.τ\\tauranges over years for Nucleus fields, months for Shell 1 and Shell 2, and weeks for Shell 3 and Shell 4\. We do not specify an update mechanism in this paper and return to the omission in Section[9\.5](https://arxiv.org/html/2609.12086#S9.SS5)\.

### 4\.3The five tiers

#### Nucleus: core identity\.

The Nucleus holds what changes most slowly and what, if changed, makes the person a different agent: core values, fundamental beliefs about self, others and life, identity anchors of the form “I am the kind of person who …”, a moral framework, existential orientation, and life purpose\. The framing follows Erikson’s account of identity\([Erikson, 1968](https://arxiv.org/html/2609.12086#bib.bib26)\)\. Operationally, a Nucleus field is one whose violation the user would describe as being untrue to themselves rather than as an inconsistency\.

#### Shell 1: psychological core\.

Shell 1 explains how the Nucleus is expressed\. It holds personality structure, incorporating but not reducing to inventories such as the Five\-Factor Model\([McCrae and John, 1992](https://arxiv.org/html/2609.12086#bib.bib22);[Soto and John, 2017](https://arxiv.org/html/2609.12086#bib.bib92)\); emotional architecture; the attachment system\([Bowlby, 1969](https://arxiv.org/html/2609.12086#bib.bib25)\); motivational drivers; a fear map; and stress response\. Trait inventories, including Myers\-Briggs type where a user knows it, are fields inside Shell 1, not substitutes for the model\. Given the psychometric objections to type\-based instruments\([Pittenger, 2005](https://arxiv.org/html/2609.12086#bib.bib96);[Stein and Swan, 2019](https://arxiv.org/html/2609.12086#bib.bib97)\), a schema that demotes an inventory to one field among thirty\-two is a feature rather than a compromise\.

#### Shell 2: cognitive and experiential\.

Shell 2 is life history as it bears on cognition: formative experiences, a record of difficult episodes, learning style, cognitive style, the personal narrative the user tells about their own life, and knowledge base\. It is the tier that explains why a given Shell 1 profile produces the particular Shell 3 behaviours it does in this particular person\.

#### Shell 3: behavioural patterns\.

Shell 3 holds daily habits, communication patterns, decision behaviour, coping mechanisms, relationship behaviour and work habits\. This is the tier most often mistaken for personality in deployed systems, because it is the tier that observation reaches\. It is the output of the inner shells, not their cause, which is why a system that models only Shell 3 can predict what a user did and not what they would do in a situation it has not seen\.

#### Shell 4: social and contextual\.

Shell 4 is the interface to the world: social roles, cultural conditioning, digital identity, reputation, the social mask a person presents\([Goffman, 1959](https://arxiv.org/html/2609.12086#bib.bib24)\), and register range\. Most LLM personalisation today operates on the outer half of Shell 4, because that is the part chat history exposes\.

#### Cross\-shell integration\.

Three entries sit across tiers rather than inside one\.*Internal conflicts*record cases where a Shell 3 behaviour contradicts a Nucleus value, which is the machinery that lets a system notice that a stored preference is out of character rather than merely old\.*Growth trajectory*records the direction of deliberate change\. The*authenticity index*measures alignment between what the inner shells hold and what the outer shells present, and is the field a system would consult before, for example, matching a user’s professional register in a private message\.

### 4\.4The full specification and the reduced instrument

The full specification holds approximately sixty fields across the five tiers plus cross\-shell entries; Appendix[A](https://arxiv.org/html/2609.12086#A1)lists it\. The studies in this paper use a reduced 32\-field instrument, six fields per tier plus two cross\-shell entries, shown in Table[2](https://arxiv.org/html/2609.12086#S4.T2)\. The reduction was made for a practical reason: a simulated participant completing sixty fields in character produces noticeably thinner answers per field than one completing thirty\-two, and we preferred fewer, denser fields\. Section[8](https://arxiv.org/html/2609.12086#S8)examines what happens to retrieval as the instrument grows towards the full size\.

Table 2:The reduced 32\-field instrument used in both studies\. Six fields per tier plus two cross\-shell entries\. Update periodτ\\tauis the nominal refresh interval for the tier\.
### 4\.5Two design commitments

#### Local residence\.

aumcontains inner\-shell content of a kind that has no analogue in a preference list\. It should reside on the user’s device or in a user\-controlled vault, never in a third\-party log\. We state this as an architectural commitment rather than a policy preference because the retrieval design of Section[5](https://arxiv.org/html/2609.12086#S5)makes it enforceable: only the selected payload need ever leave the device, and Section[7](https://arxiv.org/html/2609.12086#S7)measures how small that payload can be without loss\. Contextual integrity\([Nissenbaum, 2011](https://arxiv.org/html/2609.12086#bib.bib66)\)gives the frame for what such a boundary is for\.

#### User authorship\.

Every field is a human\-inspectable string that the user can read, edit or delete\. This placesaumin the scrutable\-personalisation tradition\([Kay and Kummerfeld, 2012](https://arxiv.org/html/2609.12086#bib.bib70);[Czarkowski and Kay, 2002](https://arxiv.org/html/2609.12086#bib.bib71);[Bakalov et al\., 2010](https://arxiv.org/html/2609.12086#bib.bib72)\)and answers the interpretability objection to trait\-vector user models\([Ji et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib5);[Kang et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib16)\)\. A budgeted payload strengthens the commitment: eight short strings are few enough to show the user before they are sent, where thirty\-two are not\.

## 5Personality\-Aware Retrieval

aumis a representation of a person; it requires a pipeline to be useful\. This section specifies one, in whichaumis treated as an index over a person and queried under a context budget rather than injected wholesale\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/pipeline.png)Figure 4:The personality\-aware retrieval pipeline\. A task classifier routes simple requests around the user model entirely\. Complex requests pass through component selection, which restricts the search to the shells the task type touches, and then through a budgeted retrieval that returnskkfields\. Theaumstore remains on the user’s device; only the payload leaves it\.### 5\.1The four stages

Given a promptqqfrom useruu, the pipeline of Figure[4](https://arxiv.org/html/2609.12086#S5.F4)proceeds as follows\.

#### \(i\) Task identification\.

A classifier assignsqqa task typettand a complexity label\. A request is*simple*if a correct answer is the same regardless of who asked, such as a unit conversion or a factual lookup, and*complex*if the form of an acceptable answer depends on who is asking\. Simple requests bypass the user model, which matters both for latency and for disclosure: a system that consults an intimate user model to convert kilometres into miles is leaking for no benefit\. Section[7\.7](https://arxiv.org/html/2609.12086#S7.SS7)reports how accurately this stage runs\.

#### \(ii\) Component selection\.

A selection functionσ:𝒯→2𝒮\\sigma:\\mathcal\{T\}\\rightarrow 2^\{\\mathcal\{S\}\}restricts the search space to the shells that a task type touches\. The mapping used throughout is given in Table[3](https://arxiv.org/html/2609.12086#S5.T3); CrossShell is always admissible\. The admissible pool is

pool⁡\(t\)=\{i:shell⁡\(fi\)∈σ⁡\(t\)∪\{CrossShell\}\}\.\\mathrm\{pool\}\(t\)=\\\{\\,i:\\mathrm\{shell\}\(f\_\{i\}\)\\in\\sigma\(t\)\\cup\\\{\\text\{CrossShell\}\\\}\\,\\\}\.\(4\)If\|pool⁡\(t\)\|<k\+4\|\\mathrm\{pool\}\(t\)\|<k\+4the pool is topped up with the most query\-relevant fields from outsideσ⁡\(t\)\\sigma\(t\)\. Without this top\-up a largekkwould silently collapse to “every fieldσ⁡\(t\)\\sigma\(t\)admits”, which would make thek=16k=16condition a different experiment from thek=4k=4andk=8k=8conditions rather than a third point on one curve\.

Table 3:The component\-selection functionσ⁡\(t\)\\sigma\(t\)\. CrossShell is admissible for every complex task type\. Factual requests are routed around the user model entirely, which is what the empty set denotes\.
#### \(iii\) Retrieval\.

A retriever selects a payloadP⁡\(q,u,k\)⊆AuP\(q,u,k\)\\subseteq A\_\{u\}with\|P\|=k≪n\|P\|=k\\ll nby maximising the objective of Section[5\.2](https://arxiv.org/html/2609.12086#S5.SS2)overpool⁡\(t\)\\mathrm\{pool\}\(t\)\.

#### \(iv\) Generation\.

The model is conditioned on\(q,P\)\(q,P\)\. The payload is inserted as a delimited block ahead of the user’s request; the exact template is in Appendix[B](https://arxiv.org/html/2609.12086#A2)\.

### 5\.2The payload objective

Let𝐟i\\mathbf\{f\}\_\{i\}denote an embedding of fieldiiand𝐪\\mathbf\{q\}an embedding of the query, bothℓ2\\ell\_\{2\}\-normalised\. The obvious objective is mean relevance,

Frel​\(S\)=1\|S\|​∑i∈Scos⁡\(𝐟i,𝐪\),F\_\{\\mathrm\{rel\}\}\(S\)=\\frac\{1\}\{\|S\|\}\\sum\_\{i\\in S\}\\cos\(\\mathbf\{f\}\_\{i\},\\mathbf\{q\}\),\(5\)but this is separable: it decomposes into per\-item scores, so its exact maximiser over a pool is obtained by sorting, and no search is required or useful\. Under a context budget, mean relevance is also the wrong objective\. Eight near\-duplicate fields waste seven slots, and a budgeted payload should therefore trade relevance against diversity in the manner of maximal marginal relevance\([Carbonell and Goldstein, 1998](https://arxiv.org/html/2609.12086#bib.bib42)\)and of result diversification more broadly\([Clarke et al\., 2008](https://arxiv.org/html/2609.12086#bib.bib43);[Agrawal et al\., 2009](https://arxiv.org/html/2609.12086#bib.bib44)\)\. We use

F⁡\(S\)=1\|S\|​∑i∈Scos⁡\(𝐟i,𝐪\)⏟relevance−λ​1\|S\|​\(\|S\|−1\)​∑i,j∈Si≠jcos⁡\(𝐟i,𝐟j\)⏟redundancy,F\(S\)\\;=\\;\\underbrace\{\\frac\{1\}\{\|S\|\}\\sum\_\{i\\in S\}\\cos\(\\mathbf\{f\}\_\{i\},\\mathbf\{q\}\)\}\_\{\\text\{relevance\}\}\\;\-\\;\\lambda\\,\\underbrace\{\\frac\{1\}\{\|S\|\\,\(\|S\|\-1\)\}\\sum\_\{\\begin\{subarray\}\{c\}i,j\\in S\\\\ i\\neq j\\end\{subarray\}\}\\cos\(\\mathbf\{f\}\_\{i\},\\mathbf\{f\}\_\{j\}\)\}\_\{\\text\{redundancy\}\},\(6\)withλ≥0\\lambda\\geq 0the redundancy weight, and the payload is

P⁡\(q,u,k\)=arg⁡maxS⊆pool⁡\(t\)\|S\|=k⁡F⁡\(S\)\.P\(q,u,k\)\\;=\\;\\arg\\max\_\{\\begin\{subarray\}\{c\}S\\subseteq\\mathrm\{pool\}\(t\)\\\\ \|S\|=k\\end\{subarray\}\}F\(S\)\.\(7\)
Atλ=0\\lambda=0, Equation[6](https://arxiv.org/html/2609.12086#S5.E6)reduces to Equation[5](https://arxiv.org/html/2609.12086#S5.E5)and the problem is a sort\. Forλ\>0\\lambda\>0the redundancy term couples the items and the problem becomes combinatorial, with\(\|pool⁡\(t\)\|k\)\\binom\{\|\\mathrm\{pool\}\(t\)\|\}\{k\}candidate payloads and no closed form\. We useλ=0\.15\\lambda=0\.15throughout, and Section[8](https://arxiv.org/html/2609.12086#S8)reports what changes asλ\\lambdavaries\.

Two remarks on the structure of the objective\. First, because it uses a mean rather than a sum and subtracts a mean pairwise similarity, it is not monotone submodular in general, so the classical greedy guarantee of[Nemhauser et al\. \(1978\)](https://arxiv.org/html/2609.12086#bib.bib41)does not transfer; how close greedy comes is an empirical question, answered in Section[8](https://arxiv.org/html/2609.12086#S8)\. Second, the redundancy term is defined over the payload only, not over the user’s history, which is what distinguishes it from novelty in a search result list\.

### 5\.3Retrievers

Four retrievers are implemented, plus exhaustive search wherever the search space is small enough to enumerate\.

#### Dense top\-kk\.

Take thekkhighest\-relevance fields, ignoring redundancy\. This is the retrieval\-augmented generation default and is the exact maximiser of Equation[5](https://arxiv.org/html/2609.12086#S5.E5)\. Cost: one objective evaluation\.

#### Forward greedy\.

Start fromS=∅S=\\emptysetand repeatedly add the field that most increasesFF\. Cost:O⁡\(k⋅\|pool\|\)O\(k\\cdot\|\\mathrm\{pool\}\|\)objective evaluations\. Strong, cheap and deterministic; it is the baseline that matters\.

#### Randomkk\.

Drawkkfields uniformly frompool⁡\(t\)\\mathrm\{pool\}\(t\)\. This is a control rather than a method\. Without it, an improvement for a retrieved payload is equally consistent with the hypothesis that anykkshort strings about the user help, which is not the claim under test\.

#### Artificial fish swarm\.

A population ofppfish, each akk\-subset of the pool, evolves forTTiterations under three behaviours drawn from[Li et al\. \(2002\)](https://arxiv.org/html/2609.12086#bib.bib38)\. Algorithm[1](https://arxiv.org/html/2609.12086#alg1)gives the procedure\. In*prey*, a fish attempts up to𝑡𝑟𝑦\\mathit\{try\}random local swaps and accepts the first improvement, falling back to a random walk\. In*swarm*, a fish moves toward the centre of its visual neighbourhood if that centre is better and the neighbourhood is not crowded\. In*follow*, it moves toward the best neighbour under the same conditions\. Distance between fish isd⁡\(a,b\)=1−\|a∩b\|/kd\(a,b\)=1\-\|a\\cap b\|/k, and a move toward a target replaces the lowest\-relevance member of the current subset with a member of the target\. The global best is retained, so the returned payload never degrades across iterations\. In the simulation study we usep=12p=12,T=8T=8, visual=6=6,𝑡𝑟𝑦=4\\mathit\{try\}=4and crowd factor0\.750\.75; Section[8](https://arxiv.org/html/2609.12086#S8)shows that this fixed budget is the source of the algorithm’s weakness rather than the behaviours are\.

Algorithm 1Artificial fish swarm overkk\-subsets1:pool

𝒫\\mathcal\{P\}, budget

kk, objective

FF, population

pp, iterations

TT, visual

vv, tries

𝑡𝑟𝑦\\mathit\{try\}, crowd

δ\\delta
2:payload

S⋆S^\{\\star\}with

\|S⋆\|=k\|S^\{\\star\}\|=k
3:

Pop←\{S1,…,Sp\}\\mathrm\{Pop\}\\leftarrow\\\{\\,S\_\{1\},\\ldots,S\_\{p\}\\,\\\}, each a uniform random

kk\-subset of

𝒫\\mathcal\{P\}
4:

S⋆←arg⁡maxS∈Pop⁡F⁡\(S\)S^\{\\star\}\\leftarrow\\arg\\max\_\{S\\in\\mathrm\{Pop\}\}F\(S\)
5:for

ι=1\\iota=1to

TTdo

6:for all

S∈PopS\\in\\mathrm\{Pop\}do

7:

C←\{Prey​\(S\)\}C\\leftarrow\\\{\\,\\textsc\{Prey\}\(S\)\\,\\\}
8:

N←\{S′∈Pop:S′≠S,1−\|S∩S′\|/k≤v/k\}N\\leftarrow\\\{\\,S^\{\\prime\}\\in\\mathrm\{Pop\}:S^\{\\prime\}\\neq S,\\;1\-\|S\\cap S^\{\\prime\}\|/k\\leq v/k\\,\\\}
9:if

N≠∅N\\neq\\emptysetand

\|N\|/p≤δ\|N\|/p\\leq\\deltathen

10:

Sc←S\_\{c\}\\leftarrowthe

kkmost frequent members across

NN⊳\\trianglerightswarm centre

11:if

F⁡\(Sc\)\>F⁡\(S\)F\(S\_\{c\}\)\>F\(S\)then

C←C∪\{MoveToward​\(S,Sc\)\}C\\leftarrow C\\cup\\\{\\,\\textsc\{MoveToward\}\(S,S\_\{c\}\)\\,\\\}
12:endif

13:

Sb←arg⁡maxS′∈N⁡F⁡\(S′\)S\_\{b\}\\leftarrow\\arg\\max\_\{S^\{\\prime\}\\in N\}F\(S^\{\\prime\}\)
14:if

F⁡\(Sb\)\>F⁡\(S\)F\(S\_\{b\}\)\>F\(S\)then

C←C∪\{MoveToward​\(S,Sb\)\}C\\leftarrow C\\cup\\\{\\,\\textsc\{MoveToward\}\(S,S\_\{b\}\)\\,\\\}
15:endif

16:endif

17:

S←arg⁡maxS′∈C∪\{S\}⁡F⁡\(S′\)S\\leftarrow\\arg\\max\_\{S^\{\\prime\}\\in C\\cup\\\{S\\\}\}F\(S^\{\\prime\}\)
18:endfor

19:

S⋆←arg⁡max⁡\{F⁡\(S⋆\),maxS∈Pop⁡F⁡\(S\)\}S^\{\\star\}\\leftarrow\\arg\\max\\\{\\,F\(S^\{\\star\}\),\\;\\max\_\{S\\in\\mathrm\{Pop\}\}F\(S\)\\,\\\}⊳\\trianglerightelitism

20:endfor

21:return

S⋆S^\{\\star\}
22:

23:functionPrey\(

SS\)

24:for

r=1r=1to

𝑡𝑟𝑦\\mathit\{try\}do

25:

S′←S^\{\\prime\}\\leftarrowSSwith one or two members replaced at random from

𝒫∖S\\mathcal\{P\}\\setminus S
26:if

F⁡\(S′\)\>F⁡\(S\)F\(S^\{\\prime\}\)\>F\(S\)thenreturn

S′S^\{\\prime\}
27:endif

28:endfor

29:return

SSwith one member replaced at random⊳\\trianglerightrandom walk

30:endfunction

31:

32:functionMoveToward\(

SS,

StS\_\{t\}\)

33:replace

arg⁡mini∈S⁡cos⁡\(𝐟i,𝐪\)\\arg\\min\_\{i\\in S\}\\cos\(\\mathbf\{f\}\_\{i\},\\mathbf\{q\}\)with a uniformly chosen member of

St∖SS\_\{t\}\\setminus S
34:returnthe result

35:endfunction

### 5\.4Why local residence becomes measurable

The design objective is to maximise style fidelity subject to\|P\|≤k\|P\|\\leq k\. That constraint is what separatesaumfrom a prompt prefix, and it is where the commitments of Section[4](https://arxiv.org/html/2609.12086#S4)stop being rhetorical\. Ifkkfields suffice, the remainingn−kn\-knever leave the device, and the disclosure per query is bounded by construction rather than by policy\. The size of the achievablekkis therefore not only an efficiency number but a privacy number, and Section[7](https://arxiv.org/html/2609.12086#S7)measures it\.

## 6Simulation Study: Design

### 6\.1Why simulate, and what a simulation can establish

Participants in this study are language\-model\-simulated, not human\. We state this first, and repeat it in the abstract, because it governs how every number in Section[7](https://arxiv.org/html/2609.12086#S7)should be read\.

The method has a precedent and a critique, and both are relevant\.[Argyle et al\. \(2023\)](https://arxiv.org/html/2609.12086#bib.bib55)showed that conditioning a language model on demographic backstories reproduces aggregate patterns found in human survey data, and[Park et al\. \(2024\)](https://arxiv.org/html/2609.12086#bib.bib9)showed that grounding agents in structured self\-report improves individual\-level simulation over demographic conditioning alone, which is the closest precedent for what we do here\. Against this,[Bisbee et al\. \(2024\)](https://arxiv.org/html/2609.12086#bib.bib56)document instability and misestimation in synthetic survey responses, and[Wang et al\. \(2025\)](https://arxiv.org/html/2609.12086#bib.bib57)show that replacing human participants can misportray and flatten identity groups\.

We therefore fix what the study is for\. A simulation of this kind*can*establish that a pipeline runs end to end; that a budgeted retrieval over a structured user model recovers fields that a downstream generator uses; that the resulting text differs measurably from text produced without the model; and that this holds under controls which rule out the most obvious alternative explanations\. It*cannot*establish that a real person recognises their own voice, because the entity doing the recognising here is a judge model reading a persona description, which is a considerably easier task\. It also cannot speak to how a person would feel about a system holding such a model, which is a question for the human study described in Section[9\.6](https://arxiv.org/html/2609.12086#S9.SS6)\.

There is a second and less obvious reason to simulate first\. The human study this substitutes for would require participants to disclose inner\-shell content about themselves: fears, difficult episodes, moral commitments\. Running that study before knowing whether the architecture works at all would spend real disclosure on a question that a simulation can answer\. We regard establishing feasibility in simulation as the ethically prior step, not merely the cheaper one\.

### 6\.2Simulated participants

Sixteen personas are declared in code rather than sampled by a model, so that a run is reproducible from a seed and the population is auditable\. Each persona carries Big Fivezz\-scores in\[−2,2\]\[\-2,2\], a one\-line life context, a formative\-experience seed on which Shell 2 can draw, and concrete idiolect markers, which are what make style measurable rather than merely asserted\. Thezz\-scores were spread deliberately: no two personas share a sign pattern across all five traits, and every trait has both tails represented, so that style differences are not confounded with a single dominant axis\. Table[4](https://arxiv.org/html/2609.12086#S6.T4)lists the set\. Appendix[C](https://arxiv.org/html/2609.12086#A3)gives the full declarations\.

Table 4:The sixteen simulated participants\. Big Fivezz\-scores are shown as O/C/E/A/N\. Idiolect markers are abbreviated; the full declarations are in Appendix[C](https://arxiv.org/html/2609.12086#A3)\.
### 6\.3Instrument population

Each persona completes the 32\-field instrument of Table[2](https://arxiv.org/html/2609.12086#S4.T2)in character, by guided self report: the agent is instructed to fill every field with one short first\-person sentence of roughly ten to twenty\-five words, specific and concrete rather than generic\. Responses are parsed against the declared schema and rejected if any shell is missing or any field is empty, with up to six retries and an explicit corrective instruction on each\. The exact prompt is in Appendix[B](https://arxiv.org/html/2609.12086#A2)\.

Populating a user model by self report inherits a known limitation: self and other observers know different things about a person, and the self is not the better source for every trait\([Vazire, 2010](https://arxiv.org/html/2609.12086#bib.bib107);[Connelly and Ones, 2010](https://arxiv.org/html/2609.12086#bib.bib108)\)\. For a simulated participant whose ground truth is a declared persona this is less severe than it would be for a human, since the agent has access to its own declaration, but the resulting instrument is still a self\-description rather than an observation\. Section[9\.5](https://arxiv.org/html/2609.12086#S9.SS5)returns to what this means for a human deployment\.

### 6\.4Task battery, prompt surfaces, and references

Six style\-sensitive tasks form the evaluation battery, chosen to span the task types of Table[3](https://arxiv.org/html/2609.12086#S5.T3): requesting a deadline extension \(communication\), declining an invitation from a colleague \(communication\), writing a short public post about one’s work \(self\-presentation\), giving written feedback on a junior collaborator’s draft \(relational\-evaluative\), planning a two\-day trip \(planning\-decision\), and explaining a technical result to a friend outside the field \(explanatory\)\. Four additional simple tasks, all factual lookups, exist only to exercise and score the classifier’s simple\-versus\-complex branch and are excluded from the main results\. Table[5](https://arxiv.org/html/2609.12086#S6.T5)lists the battery\.

Table 5:The task battery\. The six complex tasks form the evaluation set\. The four simple tasks are used only to score the classifier’s routing decision and never enter the main results\.For each complex task an agent produces two texts, in an order that is the single most important correctness property of the design\.

1. 1\.Theprompt surface: the message this person would type to an assistant, in their own voice, at their own level of detail and formality, capped at sixty words\. This is the personality\-seepage surface of Section[3](https://arxiv.org/html/2609.12086#S3), and it is the*only*user\-derived text the retriever ever sees\.
2. 2\.Thereference: what this person would write themselves, with no assistant involved, matching their usual length, structure and habits\. This is the ground truth against which every condition is scored, and it is generated without sight of the prompt surface\.

Passing the reference to the retriever would give everyaum\-rcondition an oracle that the Fullaumand preference\-notes baselines never receive, and would make the comparison meaningless\. The separation is enforced in code\.

### 6\.5The preference\-notes baseline

The baseline that matters is not the generic assistant but the flat preference notes that deployed memory features actually store, so it has to be built the way such features build it\. Each agent generates four short transcripts of past assistant conversations, four to six turns each, on topics drawn from a held\-out list that shares no item with the evaluation battery: choosing between two laptops, a recurring appliance problem, a week of simple meals, what to do with a free Saturday, picking a book, and replying to an awkward group message\. A separate summariser, with no access to the persona and no access to theaum, then reads those transcripts and writes at most twelve flat preference bullets\.

Two decisions here are deliberate\. The topics are held out so that the baseline cannot contain a rehearsal of the evaluation tasks\. The summariser is blind to the persona and the user model so that the baseline is a lossy distillation of observed conversation, which is what a deployed system has, rather than a partial copy of the treatment, which is what concatenating Shell 3 and Shell 4 fields would produce\.

### 6\.6Conditions

Ten conditions, all sharing a base model and a prompt, split into a main block and a control block \(Table[6](https://arxiv.org/html/2609.12086#S6.T6)\)\.

Table 6:Conditions\. The main block appears in Table[8](https://arxiv.org/html/2609.12086#S7.T8); the control block isolates one component at a time and appears in Table[10](https://arxiv.org/html/2609.12086#S7.T10)\. Only the four conditions markedidare shown to the identification judge\.ConditionBlockInjected beyond the promptGenericidmainNothingPreference notesidmainFlat bullets distilled from four held\-out prior conversationsFullaumidmainAll 32 fields as JSONaum\-r,k=4k=4mainσ⁡\(t\)\\sigma\(t\)\+ swarm retrieval, four fieldsaum\-r,k=8k=8idmainσ⁡\(t\)\\sigma\(t\)\+ swarm retrieval, eight fieldsaum\-r,k=16k=16mainσ⁡\(t\)\\sigma\(t\)\+ swarm retrieval, sixteen fieldsRandomk=8k=8controlEight fields drawn uniformly from the same admissible poolNoσ⁡\(t\)\\sigma\(t\),k=8k=8controlSwarm retrieval over all shells, no component selectionDense top\-kk,k=8k=8controlRelevance\-only top eight, redundancy term removedGreedy,k=8k=8controlForward greedy on the same objectiveEach control removes exactly one component, so that a null result localises\. Randomkkseparates “the right fields” from “anykkfields about the user”\. Noσ⁡\(t\)\\sigma\(t\)asks whether component selection contributes beyond retrieval\. Dense top\-kkasks whether the redundancy term earns its place\. Greedy asks whether the swarm earns its cost\.

### 6\.7Measures

#### Fidelity\.

A judge model rates, on a five\-point scale, the statement that a candidate text reads like something the person would have written\. The judge sees the persona and one text the person genuinely wrote, and never sees theaum\. Withholding the user model from the judge is not incidental: showing it would present the judge with the same fields theaum\-rpayload contained, allowing it to reward surface agreement with the payload rather than similarity to the person\.

#### Identification\.

A forced choice among the fouridconditions, presented in an order shuffled per trial, with chance at 25%\. The judge again sees the persona and the reference\. Position bias in LLM judges is documented\([Wang et al\., 2024a](https://arxiv.org/html/2609.12086#bib.bib52)\), so option order is randomised per trial and tested post hoc in Section[7\.7](https://arxiv.org/html/2609.12086#S7.SS7)\.

#### Style similarity\.

The cosine between LUAR authorship embeddings\([Rivera\-Soto et al\., 2021](https://arxiv.org/html/2609.12086#bib.bib37)\)of the candidate and the participant’s own reference, computed locally\. Authorship representations are used rather than semantic embeddings for a specific reason: a semantic embedding scores any two deadline\-extension emails as highly similar because they are about the same thing, and the property we need is sensitivity to*who*wrote a text and insensitivity to*what*it is about\([Wegmann et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib59);[Wang et al\., 2023](https://arxiv.org/html/2609.12086#bib.bib60)\)\. The pipeline raises rather than falling back if the authorship model cannot be loaded, because a silent substitution would change what the column means while leaving its heading intact\. A semantic cosine is computed in addition and reported separately, never as style similarity\.

#### Context\.

The number of tokens in exactly the string prepended to the prompt, counted with the generator’s own tokeniser\. This is a measurement, not a word count, so that the ratio reported in Section[7](https://arxiv.org/html/2609.12086#S7)is exact\.

#### Judge independence\.

The judge model differs from the generator by construction, and the pipeline refuses to run otherwise\. LLM evaluators favour their own generations\([Panickssery et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib51)\), so a run in which judge and generator coincide produces a fidelity column that mostly measures self\-preference\.

### 6\.8Models, seeds and cost

Table[7](https://arxiv.org/html/2609.12086#S6.T7)records the configuration\. Every API call is cached on disk under a hash of its arguments, so a re\-run after a crash or a code change costs nothing for work already completed, and the run is reproducible from the manifest plus a seed\.

Table 7:Reproducibility record for the simulation study\. Token and cost figures are read from the run ledger, not estimated\.
### 6\.9Analysis plan

The unit of analysis is the \(participant, task\) cell averaged over seeds, giving16×6=9616\\times 6=96paired observations per condition\. Seeds are repeated measures of the same cell rather than additional participants, so averaging before testing is what keeps the degrees of freedom honest; treating three seeds×\\times96 cells as 288 independent observations would inflate every test\.

For each planned contrast we report the mean paired difference, a 95% bootstrap confidence interval over 10,000 resamples, a pairedttstatistic, a Wilcoxon signed\-rank statistic, Cohen’sdzd\_\{z\}and Cliff’s delta\. The Wilcoxon test is primary for fidelity, because a five\-point rating scale is ordinal; the pairedttis primary for the continuous style measure\. The Holm procedure corrects within metric\. Identification uses an exact binomial test against 25% chance with a Wilson interval\.

Contrasts were fixed before the run and are those listed in Section[7](https://arxiv.org/html/2609.12086#S7)\. Supplementary analyses reported in Section[7\.7](https://arxiv.org/html/2609.12086#S7.SS7)were specified after inspecting the main results and are labelled as such; we do not present them as confirmatory\.

## 7Simulation Study: Results

### 7\.1Main comparison

Table[8](https://arxiv.org/html/2609.12086#S7.T8)gives the main block, and Table[10](https://arxiv.org/html/2609.12086#S7.T10)the planned contrasts\. Three findings carry the section\.

Table 8:Main results\. Fidelity and identification are judge ratings on behalf of simulated participants; style is LUAR authorship cosine against the participant’s own reference; semantic is a general\-purpose embedding cosine, reported for comparison only and never as style; context is tokens injected beyond the prompt\. Means over16×6=9616\\times 6=96cells and three seeds, with 95% bootstrap intervals on fidelity\.ConditionFidelity95% CIIdent\.StyleContext\(1 to 5\)\(%\)\(LUAR\)\(tokens\)Generic \(prompt only\)2\.65\[2\.45, 2\.84\]6\.90\.6740Preference notes2\.68\[2\.48, 2\.88\]14\.90\.684164Fullaum\(all 32 fields\)2\.90\[2\.71, 3\.08\]35\.40\.705915aum\-r,k=4k=42\.86\[2\.68, 3\.06\]n/a0\.701111aum\-r,k=8k=8\(ours\)2\.91\[2\.74, 3\.09\]42\.70\.698211aum\-r,k=16k=162\.98\[2\.78, 3\.17\]n/a0\.701406*Controls, all atk=8k=8*Randomkkfields2\.89\[2\.69, 3\.07\]n/a0\.697211Noσ⁡\(t\)\\sigma\(t\)2\.95\[2\.75, 3\.14\]n/a0\.703209Dense top\-kk2\.92\[2\.74, 3\.10\]n/a0\.699210Greedy2\.95\[2\.75, 3\.14\]n/a0\.700210#### Finding 1: a budgeted payload matches the whole model at under a quarter of the context\.

aum\-ratk=8k=8was not distinguishable from Fullaumon fidelity \(Δ=\+0\.02\\Delta=\+0\.02, 95% CI\[−0\.07,\+0\.11\]\[\-0\.07,\\,\+0\.11\],p=1p=1after Holm correction\) while injecting 211 tokens against 915, that is 23\.0%\. Figure[5](https://arxiv.org/html/2609.12086#S7.F5)plots the whole condition set against injected context on a logarithmic axis; the point of the figure is the vertical gap between preference notes andaum\-rat comparable cost, and the horizontal gap betweenaum\-rand Fullaumat comparable fidelity\.

#### Finding 2: structure beats flat notes, by a moderate effect\.

aum\-ratk=8k=8exceeded preference notes on fidelity by\+0\.24\+0\.24\(95% CI\[\+0\.14,\+0\.33\]\[\+0\.14,\\,\+0\.33\], WilcoxonW=376\.5W=376\.5,p=7\.2×10−5p=7\.2\\times 10^\{\-5\}after Holm correction,dz=\+0\.50d\_\{z\}=\+0\.50, Cliff’sδ=0\.35\\delta=0\.35,n=96n=96cells\), and exceeded the generic baseline by\+0\.27\+0\.27\(dz=\+0\.57d\_\{z\}=\+0\.57\)\. On LUAR style similarity the same ordering holds with a smaller effect:\+0\.013\+0\.013over preference notes \(p=0\.006p=0\.006\) and\+0\.024\+0\.024over generic \(p<10−5p<10^\{\-5\},dz=\+0\.53d\_\{z\}=\+0\.53\)\. Fullaumalso beat preference notes \(\+0\.22\+0\.22,p=4×10−4p=4\\times 10^\{\-4\}\), which is the expected ordering: the gain comes from having a structured model at all, and the retrieval stage recovers it cheaply rather than adding to it\.

#### Finding 3: identification separates the conditions sharply\.

Forced\-choice identification of the participant’s own voice reached 42\.7% foraum\-ratk=8k=8\(95% CI\[37\.1,48\.5\]\[37\.1,\\,48\.5\],p=4\.2×10−11p=4\.2\\times 10^\{\-11\}against 25% chance\) and 35\.4% for Fullaum\(\[30\.1,41\.1\]\[30\.1,\\,41\.1\],p=5\.3×10−5p=5\.3\\times 10^\{\-5\}\)\. Preference notes reached 14\.9% and generic 6\.9%, neither above chance \(Table[9](https://arxiv.org/html/2609.12086#S7.T9), Figure[7](https://arxiv.org/html/2609.12086#S7.F7)\)\. That both user\-model conditions sit above chance while both baselines sit below it is the cleanest separation in the study, and it is the result most directly relevant to the inversion argument of Section[3](https://arxiv.org/html/2609.12086#S3): notes distilled from unrelated past conversations carry topic, not person\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_pareto_context_fidelity.png)Figure 5:Fidelity against injected context on a logarithmic axis\. Eight retrieved fields reach the fidelity of the entire 32\-field user model at 23% of its context\. Error bars are 95% bootstrap intervals over the 96 \(participant, task\) cells\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_fidelity_vs_k.png)Figure 6:Fidelity against retrieval budgetkk, with the Fullaumand preference\-notes levels marked\. Fidelity is still rising at the largest budget tested, sok=8k=8is a trade rather than a saturation point\.Table 9:Forced\-choice identification\. Chance is 25%;n=288n=288presentations per condition \(96 cells×\\times3 seeds\)\. Intervals are Wilson;ppis an exact binomial test against chance in the upper tail\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_identification.png)Figure 7:Forced\-choice identification by the judge, with Wilson intervals and the 25% chance line\. Both user\-model conditions sit above chance; both baselines sit below it\.

### 7\.2The budget curve

Fidelity rose monotonically across the budgets tested: 2\.86 atk=4k=4, 2\.91 atk=8k=8, 2\.98 atk=16k=16\(Figure[6](https://arxiv.org/html/2609.12086#S7.F6)\)\. Neither adjacent step was individually significant after correction \(k=16k=16againstk=8k=8:Δ=\+0\.06\\Delta=\+0\.06,p=0\.74p=0\.74;k=8k=8againstk=4k=4:Δ=\+0\.05\\Delta=\+0\.05,p=1p=1\), so the curve is best read as a gradual improvement rather than a threshold\. Two things follow\. First,k=8k=8is not a saturation point, and a system with a looser budget should use more fields\. Second, and more usefully for the privacy argument, there is no evidence of degradation as further fields enter the payload: the concern that irrelevant fields would dilute a payload and hurt output quality is not borne out at these budgets, so the argument for a smallkkrests on cost and disclosure rather than on quality\.

### 7\.3Style similarity

Figure[8](https://arxiv.org/html/2609.12086#S7.F8)shows the distribution of LUAR authorship cosine by condition, and Figure[9](https://arxiv.org/html/2609.12086#S7.F9)the general\-purpose semantic cosine for the same trials\. The two behave differently in an instructive way\. The authorship metric separates the conditions in the same order as fidelity, with a compressed range \(0\.674 to 0\.705 across all ten conditions\)\. The semantic metric separates them far less \(0\.593 to 0\.620\), which is what one expects given that every candidate for a given trial is an attempt at the same task: semantic similarity is dominated by topic, and topic is held constant by design\. We report the semantic column only to make that contrast visible, and we would regard a paper that reported it as “style similarity” as having measured the wrong thing\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_dist_style_luar.png)Figure 8:Distribution of LUAR authorship\-embedding cosine between each generated text and the participant’s own reference, by condition\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_dist_style_semantic.png)Figure 9:The same trials under a general\-purpose semantic embedding\. The compression relative to Figure[8](https://arxiv.org/html/2609.12086#S7.F8)is the point: a semantic metric mostly measures that every candidate addresses the same task, which is the wrong invariance for a style claim\.
### 7\.4Four controls return null

Table 10:Planned contrasts, paired over the 96 \(participant, task\) cells\.Δ\\Deltais the mean paired difference; the interval is a 95% bootstrap interval over 10,000 resamples;ppis Holm\-corrected within metric, with Wilcoxon primary for fidelity and pairedttprimary for style\.MetricContrastΔ\\Delta95% CIdzd\_\{z\}ppfidelityaum\-rk=8k\{=\}8vs Preference notes\+0\.236\+0\.236\[\+0\.142\+0\.142,\+0\.330\+0\.330\]\+0\.50\+0\.500\.00010\.0001fidelityaum\-rk=8k\{=\}8vs Generic\+0\.267\+0\.267\[\+0\.177\+0\.177,\+0\.365\+0\.365\]\+0\.57\+0\.570\.00010\.0001fidelityFullaumvs Preference notes\+0\.219\+0\.219\[\+0\.115\+0\.115,\+0\.323\+0\.323\]\+0\.42\+0\.420\.00040\.0004fidelityaum\-rk=8k\{=\}8vs Fullaum\+0\.017\+0\.017\[−0\.069\-0\.069,\+0\.108\+0\.108\]\+0\.04\+0\.041\.00001\.0000fidelityaum\-rk=16k\{=\}16vsaum\-rk=8k\{=\}8\+0\.063\+0\.063\[−0\.031\-0\.031,\+0\.156\+0\.156\]\+0\.13\+0\.130\.74360\.7436fidelityaum\-rk=8k\{=\}8vsaum\-rk=4k\{=\}4\+0\.049\+0\.049\[−0\.031\-0\.031,\+0\.125\+0\.125\]\+0\.12\+0\.121\.00001\.0000fidelityaum\-rk=8k\{=\}8vs Randomk=8k\{=\}8\+0\.028\+0\.028\[−0\.052\-0\.052,\+0\.108\+0\.108\]\+0\.07\+0\.071\.00001\.0000fidelityaum\-rk=8k\{=\}8vs noσ⁡\(t\)\\sigma\(t\)−0\.035\-0\.035\[−0\.111\-0\.111,\+0\.045\+0\.045\]−0\.09\-0\.090\.98130\.9813fidelityaum\-rk=8k\{=\}8vs Dense top\-kk−0\.007\-0\.007\[−0\.083\-0\.083,\+0\.073\+0\.073\]−0\.02\-0\.021\.00001\.0000fidelityaum\-rk=8k\{=\}8vs Greedy−0\.035\-0\.035\[−0\.111\-0\.111,\+0\.042\+0\.042\]−0\.09\-0\.091\.00001\.0000styleaum\-rk=8k\{=\}8vs Preference notes\+0\.013\+0\.013\[\+0\.006\+0\.006,\+0\.021\+0\.021\]\+0\.36\+0\.360\.00570\.0057styleaum\-rk=8k\{=\}8vs Generic\+0\.024\+0\.024\[\+0\.015\+0\.015,\+0\.033\+0\.033\]\+0\.53\+0\.530\.00000\.0000styleFullaumvs Preference notes\+0\.020\+0\.020\[\+0\.012\+0\.012,\+0\.029\+0\.029\]\+0\.47\+0\.470\.00010\.0001styleaum\-rk=8k\{=\}8vs Fullaum−0\.007\-0\.007\[−0\.014\-0\.014,\+0\.001\+0\.001\]−0\.18\-0\.180\.58890\.5889styleaum\-rk=8k\{=\}8vs Randomk=8k\{=\}8\+0\.001\+0\.001\[−0\.005\-0\.005,\+0\.007\+0\.007\]\+0\.04\+0\.041\.00001\.0000styleaum\-rk=8k\{=\}8vs noσ⁡\(t\)\\sigma\(t\)−0\.006\-0\.006\[−0\.012\-0\.012,\+0\.001\+0\.001\]−0\.17\-0\.170\.60550\.6055styleaum\-rk=8k\{=\}8vs Dense top\-kk−0\.001\-0\.001\[−0\.007\-0\.007,\+0\.005\+0\.005\]−0\.03\-0\.031\.00001\.0000styleaum\-rk=8k\{=\}8vs Greedy−0\.002\-0\.002\[−0\.009\-0\.009,\+0\.005\+0\.005\]−0\.06\-0\.061\.00001\.0000All four controls returned null, on both metrics, and we take each in turn because each bounds the claim differently\.

#### The rightkkfields, or merelykkfields?

aum\-ratk=8k=8was not distinguishable from eight fields drawn uniformly at random from the same admissible pool \(Δ=\+0\.028\\Delta=\+0\.028,p=1p=1\)\. This is the most consequential null in the study\. It says that once component selection has narrowed the pool to roughly seventeen fields, which eight of them are injected did not measurably matter\. The gain in Findings 1 to 3 is therefore attributable to injecting structured, task\-plausible fields about the person, not to ranking within that set\.

#### Doesσ⁡\(t\)\\sigma\(t\)contribute beyond retrieval?

Removing component selection and retrieving over all thirty\-two fields did not hurt; it was very slightly better and not significantly so \(Δ=−0\.035\\Delta=\-0\.035,p=0\.98p=0\.98\)\. Given the preceding null this is coherent rather than surprising: if ranking within the pool does not matter, restricting the pool cannot matter either, so long as the restriction does not exclude something essential\. Section[7\.5](https://arxiv.org/html/2609.12086#S7.SS5)shows thatσ⁡\(t\)\\sigma\(t\)nonetheless selects sensibly, which is a weaker but still useful property\.

#### Does the redundancy term earn its place?

A relevance\-only top\-kkwas indistinguishable from the full objective \(Δ=−0\.007\\Delta=\-0\.007,p=1p=1\)\. Section[8](https://arxiv.org/html/2609.12086#S8)shows this is a property of the operating point rather than of the objective: atλ=0\.15\\lambda=0\.15and a pool of seventeen, the redundancy penalty rarely changes which eight fields win\.

#### Does the swarm earn its cost?

It did not\.aum\-rdid not beat forward greedy on generated style \(Δ=−0\.035\\Delta=\-0\.035,p=1p=1\), and it did not reach a higher value of its own objective\. Table[11](https://arxiv.org/html/2609.12086#S7.T11)gives the retrieval diagnostics: the swarm reachedF=0\.224F=0\.224using 1,736 objective evaluations, where greedy reachedF=0\.226F=0\.226using 108 and dense top\-kkreachedF=0\.225F=0\.225using one\. Random reachedF=0\.183F=0\.183, which confirms that the objective is meaningful even though optimising it harder did not help downstream\. Figure[10](https://arxiv.org/html/2609.12086#S7.F10)shows the quality and cost side by side\.

Table 11:Retrieval diagnostics atk=8k=8, averaged over all trials\.F⁡\(S\)F\(S\)is the value of Equation[6](https://arxiv.org/html/2609.12086#S5.E6)reached; evaluations counts calls toFF; the pool is the size ofpool⁡\(t\)\\mathrm\{pool\}\(t\)after component selection\. The swarm spent sixteen times greedy’s budget to reach a marginally lower objective value\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_retrieval_cost.png)Figure 10:Retriever comparison atk=8k=8: objective value reached and objective evaluations spent\. The swarm is both slightly worse and sixteen times more expensive than forward greedy at this scale\.
#### What the four nulls jointly establish\.

They localise the effect\. Something about injecting a structured, task\-plausible slice of a user model changes generated style substantially; nothing about which slice, or about how hard one searches for it, changes it further at this scale\. That is a narrower claim than the one we set out to test, and it is the claim the data supports\. Section[8](https://arxiv.org/html/2609.12086#S8)takes the optimisation question out of the language\-model setting entirely and asks where, if anywhere, the search does start to matter\.

### 7\.5Which shells each task type draws on

Although component selection did not change downstream quality, the payloads it produced are interpretable, and inspecting them is a check that the pipeline is doing what the specification says\. Figure[11](https://arxiv.org/html/2609.12086#S7.F11)shows the share of retrieved payload drawn from each shell, broken down by task type, and Table[23](https://arxiv.org/html/2609.12086#A4.T23)gives the same figures numerically\.

The pattern matches the mapping in Table[3](https://arxiv.org/html/2609.12086#S5.T3)without having been forced to\. Communication tasks drew 44\.4% of their payload from Shell 3 and 42\.3% from Shell 4\. Self\-presentation drew most heavily on the Nucleus \(37\.2%\) and Shell 4 \(41\.2%\), which is what one would expect of a task in which a person decides what to project\. Planning and decision tasks drew 49\.5% from Shell 3 and split the rest between Shell 1 and Shell 2\. Relational\-evaluative tasks, which involve judging another person’s work, drew 24\.7% from the Nucleus, the second\-highest Nucleus share in the study, alongside 45\.6% from Shell 3\. Explanatory tasks drew 30\.0% from Shell 2, which holds knowledge base and cognitive style, the highest Shell 2 share of any task type\. Cross\-shell fields appeared in every task type at between 7\.6% and 14\.8%\.

This is a weak result and we present it as one\. It shows that the representation is legible and that retrieval over it behaves sensibly, not that legibility improves output\. It matters mainly for the scrutability argument: a user inspecting why the assistant produced a particular draft can be shown a payload whose composition is defensible in ordinary language\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_shell_usage.png)Figure 11:Share of the retrieved payload drawn from each shell, by task type\. Rows sum to 100%\. The composition tracks Table[3](https://arxiv.org/html/2609.12086#S5.T3)without having been constrained to\.
### 7\.6Which tasks the effect holds for

Table 12:Effect ofaum\-ratk=8k=8over preference notes, by task\.Δ\\Deltais the mean paired difference in fidelity over the sixteen participants; “favouring” is the share of participants for whom the difference is positive\.ppis an uncorrected Wilcoxon signed\-rank test within task\.All six tasks favouredaum\-rin the mean \(Table[12](https://arxiv.org/html/2609.12086#S7.T12), Figure[12](https://arxiv.org/html/2609.12086#S7.F12)\), but the sizes differ in a way worth naming\. The largest effects are on tasks with the loosest genre conventions: explaining something to a friend outside the field \(\+0\.375\+0\.375\) and declining an invitation \(\+0\.312\+0\.312\)\. The smallest, and effectively zero, is the deadline\-extension email \(\+0\.021\+0\.021, favouring in only 18\.8% of participants\)\. That is the most conventionalised task in the battery: a deadline\-extension email has a widely shared expected form, and a generic assistant already produces it\. Where the genre supplies the structure, the user model has little left to contribute; where it does not, the model fills the gap\. This is the task\-level counterpart of the participant\-level pattern in Section[7\.8](https://arxiv.org/html/2609.12086#S7.SS8)\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_per_task.png)Figure 12:Fidelity by task and condition\. The deadline\-extension task, the most conventionalised in the battery, shows the smallest spread across conditions\.
### 7\.7Robustness

Four supplementary analyses address the objections a reader is most likely to raise\. All were specified after the main results and are reported as exploratory\.

#### The classifier stage works\.

Section[5](https://arxiv.org/html/2609.12086#S5)posits a routing stage, and an unevaluated first stage is a weak point in any pipeline\. Over the full battery the classifier assigned the correct complexity label in 98\.8% of trials \(83 of 84, Wilson interval\[93\.6,99\.8\]\[93\.6,\\,99\.8\]\), with all twelve simple trials routed to bypass and 71 of 72 complex trials routed to the user model\. Six\-way task\-type accuracy on complex trials was 90\.3% \(65 of 72,\[81\.3,95\.2\]\[81\.3,\\,95\.2\]\) against 16\.7% chance\. Table[13](https://arxiv.org/html/2609.12086#S7.T13)gives per\-type recall and Figure[13](https://arxiv.org/html/2609.12086#S7.F13)the confusion matrix\. The single systematic error is that explanatory requests are read as communication in a third of cases, which is intelligible: explaining a result to a friend*is*a communication act, and the twoσ⁡\(t\)\\sigma\(t\)sets share Shell 3\.

Table 13:Task classifier performance, pooled over three seeds\. Recall is within gold task type on complex trials only\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_classifier_confusion.png)Figure 13:Task\-type confusion for complex requests, as row percentages\. The one systematic error, explanatory read as communication, reflects a genuine overlap between the two categories\.
#### The effect is not a length artefact\.

A judge asked whether a text reads like a person’s own writing might be rewarded largely by length agreement\. It is: pooled across all trials, the absolute log ratio of candidate length to reference length correlates strongly and negatively with fidelity \(Spearmanρ=−0\.511\\rho=\-0\.511,p=3×10−191p=3\\times 10^\{\-191\},n=2880n=2880\)\. Matching the reference’s length is a large part of reading like the person, which is unsurprising and not itself a confound\.

The question is whether the treatment effect survives adjustment for it\. It does\. Regressing the paired fidelity difference \(aum\-rk=8k=8minus preference notes\) on the paired difference in absolute log length ratio leaves an intercept of\+0\.227\+0\.227\(95% CI\[\+0\.127,\+0\.327\]\[\+0\.127,\\,\+0\.327\],p=1\.9×10−5p=1\.9\\times 10^\{\-5\}\) against a raw difference of\+0\.236\+0\.236, and the covariate slope is not significant \(−0\.155\-0\.155,p=0\.519p=0\.519\)\. In other words the two conditions did not differ appreciably in length agreement, and the effect is not carried by it\. Table[14](https://arxiv.org/html/2609.12086#S7.T14)gives the per\-condition figures\. The generic condition is the outlier: it produced the shortest outputs \(96 words against a 142\-word reference mean\), the worst length agreement \(0\.007 on the absolute log ratio scale relative to the next condition\) and the highest rate of degenerate outputs \(7\.6% of generic trials produced fewer than fifteen words, typically a clarifying question rather than an attempt at the task\)\.

Table 14:Output length by condition, against a reference mean of 141\.6 words\. “Agreement” is the mean absolute log ratio of candidate to reference length, where lower is closer\. “Stub” is the share of trials producing fewer than fifteen words\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_length_by_condition.png)Figure 14:Output length by condition against the participants’ own reference length\. Every user\-model condition moves output length towards the reference; the generic condition is furthest from it\.
#### The effect is broad, not driven by outliers\.

Twelve of sixteen participants and six of six tasks favouredaum\-rin the mean\. Of the 96 cells, 51\.0% favouredaum\-r, 33\.3% were exact ties on the ordinal scale, and 15\.6% favoured preference notes\. Figure[15](https://arxiv.org/html/2609.12086#S7.F15)gives a per\-participant forest plot, and Table[15](https://arxiv.org/html/2609.12086#S7.T15)the numbers\. Two participants show small negative effects, the meticulous planner \(−0\.111\-0\.111\) and the reflective introvert \(−0\.056\-0\.056\), and neither interval excludes zero\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_by_persona.png)Figure 15:Effect ofaum\-ratk=8k=8over preference notes, by simulated participant, with 95% bootstrap intervals over that participant’s six tasks\. The dashed line is the pooled mean\.
#### The judge is usable but not sharp, and is not position\-biased\.

Across the three seeds the fidelity judge achieved ICC\(2,1\)=0\.57=0\.57over 960 targets, which by the conventional bands is moderate reliability\([Shrout and Fleiss, 1979](https://arxiv.org/html/2609.12086#bib.bib110);[Koo and Li, 2016](https://arxiv.org/html/2609.12086#bib.bib111)\)\. We report the intraclass correlation rather than a chance\-corrected agreement coefficient because the ratings are ordinal and the three seeds are exchangeable rather than fixed coders, which is the condition under which the coefficient families diverge\([Krippendorff, 2004](https://arxiv.org/html/2609.12086#bib.bib113);[Artstein and Poesio, 2008](https://arxiv.org/html/2609.12086#bib.bib112)\)\. Exact agreement between seed pairs ranged from 42\.0% to 43\.2% and agreement within one scale point from 85\.6% to 87\.4%, with Spearman correlations of 0\.55 to 0\.60\. The judge used the whole scale rather than collapsing to the middle: 13\.8% of ratings were 1, 25\.2% were 2, 25\.8% were 3, 30\.9% were 4 and 4\.3% were 5\. On identification, choices did not cluster significantly by option slot despite the per\-trial shuffling \(χ2=7\.37\\chi^\{2\}=7\.37, 3 d\.f\.,p=0\.061p=0\.061\), with slot shares of 22\.9%, 29\.2%, 27\.4% and 20\.5%\. Thatppis close enough to the conventional threshold that we would not describe position bias as absent, only as not detected at this sample size\. Figure[16](https://arxiv.org/html/2609.12086#S7.F16)shows both diagnostics\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_judge_diagnostics.png)Figure 16:Judge diagnostics\. Left: use of the five\-point scale across all trials\. Right: identification choices by option slot, against the uniform 25% expectation, with Wilson intervals\.

### 7\.8Personalisation buys most where the default voice is furthest away

The per\-participant breakdown suggested a pattern that we then tested directly\. Fidelity in the generic condition is a direct measure of how well the un\-personalised assistant already writes like a given participant: it is the score the default voice earns without any user model at all\. If personalisation is doing what it claims, its benefit should be largest where that baseline is lowest\.

It is\. Across the sixteen participants, generic\-condition fidelity correlates negatively with the treatment effect \(Spearmanρ=−0\.61\\rho=\-0\.61,p=0\.013p=0\.013; Pearsonr=−0\.60r=\-0\.60,p=0\.013p=0\.013\), and the relationship holds at cell level with more power and a smaller coefficient \(ρ=−0\.27\\rho=\-0\.27,p=0\.007p=0\.007,n=96n=96\)\. Splitting participants at the median generic fidelity, those the default serves worst gain\+0\.333\+0\.333\(95% CI\[\+0\.181,\+0\.486\]\[\+0\.181,\\,\+0\.486\]\) and those it serves best gain\+0\.139\+0\.139\(\[\+0\.035,\+0\.243\]\[\+0\.035,\\,\+0\.243\]\); the difference between the two groups is in the expected direction but does not itself reach significance \(Mann\-Whitneyp=0\.096p=0\.096\)\. Figure[17](https://arxiv.org/html/2609.12086#S7.F17)plots the relationship\.

The extremes are interpretable\. The blunt minimalist, whose lowercase fragments and absent sign\-offs are about as far from a default assistant register as the battery contains, scores 1\.78 under generic and gains\+0\.611\+0\.611, the largest effect in the study\. The meticulous planner, whose numbered points and explicit dates are close to what an assistant produces unprompted, scores 2\.94 under generic and gains nothing \(−0\.111\-0\.111\)\. The default voice already was that person\.

One caveat is important and we state it rather than burying it\. The same correlation computed on the LUAR style metric rather than on judge fidelity is in the same direction but does not reach significance \(ρ=−0\.12\\rho=\-0\.12,p=0\.242p=0\.242\)\. The pattern is therefore established in the judge’s ratings and not, on this sample, in the automatic measure\. We report it because its design implication is substantial and because suppressing a result that did not replicate across both metrics would be worse than reporting it with its caveat attached\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_prior_distance.png)Figure 17:The treatment effect against how well the un\-personalised assistant already matches each participant\. Points are the sixteen simulated participants; the line is an ordinary least squares fit\.Table 15:Effect ofaum\-ratk=8k=8over preference notes, by simulated participant, ordered by effect size\. “Generic” is that participant’s mean fidelity in the prompt\-only condition, which measures how well the default voice already matches them\.

## 8Scaling Study of the Retriever

### 8\.1Separating two questions

Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4)reported that the swarm retriever neither improved generated style over forward greedy nor reached a higher value of its own objective, and that it spent roughly sixteen times as many objective evaluations doing so\. That is a result about one operating point: a reduced 32\-field instrument, a pool of about seventeen admissible fields after component selection, andk=8k=8\. It is not a result about payload selection in general, and treating it as one would be as much of an error as the reverse\.

This section separates the two questions\. It removes the language model entirely and studies the optimisation problem on its own, sweeping the instrument sizenn, the context budgetkkand the redundancy weightλ\\lambda\. The question it answers is narrow and answerable: for which\(n,k,λ\)\(n,k,\\lambda\)does a population method reach a better value of Equation[6](https://arxiv.org/html/2609.12086#S5.E6)than forward greedy, and at what cost\. No API call is involved, so the whole sweep is free, deterministic from a seed, and can be re\-run by a reader\.

### 8\.2Instance model

Field embeddings from real text are not isotropic Gaussian vectors\. Fields belonging to the same shell talk about related things and therefore cluster, and that clustering is precisely what makes the redundancy term non\-trivial\. Each synthetic instance is generated with an explicit cluster structure\. Fornnfields distributed overmmshells indddimensions,

𝐜s\\displaystyle\\mathbf\{c\}\_\{s\}∼Uniform\(𝒮d−1\),s=1,…,m,\\displaystyle\\sim\\mathrm\{Uniform\}\(\\mathcal\{S\}^\{d\-1\}\),\\qquad s=1,\\ldots,m,\(8\)𝐟i\\displaystyle\\mathbf\{f\}\_\{i\}=normalise⁡\(ρ​𝐜s⁡\(i\)\+1−ρ2​𝜺i\),𝜺i∼𝒩⁡\(𝟎,σ2​I\),\\displaystyle=\\mathrm\{normalise\}\\\!\\left\(\\rho\\,\\mathbf\{c\}\_\{s\(i\)\}\+\\sqrt\{1\-\\rho^\{2\}\}\\;\\boldsymbol\{\\varepsilon\}\_\{i\}\\right\),\\qquad\\boldsymbol\{\\varepsilon\}\_\{i\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma^\{2\}I\),\(9\)𝐪\\displaystyle\\mathbf\{q\}=normalise⁡\(∑s∈𝒜ws​𝐜s\+σq​𝜺q\),\\displaystyle=\\mathrm\{normalise\}\\\!\\left\(\\sum\_\{s\\in\\mathcal\{A\}\}w\_\{s\}\\mathbf\{c\}\_\{s\}\+\\sigma\_\{q\}\\boldsymbol\{\\varepsilon\}\_\{q\}\\right\),\(10\)wheres⁡\(i\)s\(i\)is the shell of fieldii,𝒜\\mathcal\{A\}is a random subset of shells of size two, andws∼Uniform⁡\(0\.5,1\.5\)w\_\{s\}\\sim\\mathrm\{Uniform\}\(0\.5,1\.5\)\. The query construction mimics a task that draws on two shells, which is what Table[3](https://arxiv.org/html/2609.12086#S5.T3)specifies\. The parameterρ\\rhocontrols how tightly a shell clusters;ρ=0\\rho=0recovers the isotropic case and makes the redundancy term nearly inert\. We useρ=0\.65\\rho=0\.65,d=64d=64,m=max⁡\(2,⌊n/6⌋\)m=\\max\(2,\\lfloor n/6\\rfloor\)and 25 independent instances per cell\.

### 8\.3Methods and budgets

Six retrievers are compared, plus exhaustive search wherever\(nk\)≤60,000\\binom\{n\}\{k\}\\leq 60\{,\}000\.

Dense top\-kkthe relevance\-only maximiser of Equation[5](https://arxiv.org/html/2609.12086#S5.E5), one objective evaluation\.

Greedyforward greedy on the full objective\.

Randomkkfields drawn uniformly, the floor\.

AFSA \(study budget\)the swarm with exactly the population and iteration counts used in Section[6](https://arxiv.org/html/2609.12086#S6),p=12p=12andT=8T=8, held fixed asnngrows\. This is what the simulation actually ran\.

AFSA \(scaled budget\)population and iterations grown with the search space:p=clip⁡\(⌈2​log⁡\(nk\)⌉,12,60\)p=\\mathrm\{clip\}\(\\lceil 2\\log\\binom\{n\}\{k\}\\rceil,12,60\)andT=clip⁡\(⌈1\.2​log⁡\(nk\)⌉,8,40\)T=\\mathrm\{clip\}\(\\lceil 1\.2\\log\\binom\{n\}\{k\}\\rceil,8,40\)\. The constants are set so that\(n,k\)=\(32,8\)\(n,k\)=\(32,8\)reproduces roughly the study’s own budget and everything larger receives proportionally more\.

AFSA \(greedy warm start\)the scaled\-budget swarm with one member of the initial population replaced by the greedy solution\. This is the obvious hybrid and the fairest test of whether a population method can add anything on top of greedy\.

The sweep coversn∈\{16,32,64,128,256\}n\\in\\\{16,32,64,128,256\\\},k∈\{4,8,16,32\}k\\in\\\{4,8,16,32\\\}withk≤n/2k\\leq n/2, andλ∈\{0,0\.15,0\.30,0\.60\}\\lambda\\in\\\{0,0\.15,0\.30,0\.60\\\}, giving 68 cells and 1,700 instances\. Quality is reported as the relative gap to the best solution found for that instance by any method, which equals the exhaustive optimum wherever that is available\.

### 8\.4Results

Table 16:Relative gap to the best solution found, in percent, averaged overkkand 25 instances per cell, atλ=0\.15\\lambda=0\.15\. Lower is better\. Forward greedy is within 0\.1% of the best solution at every instrument size; the swarm at the study’s fixed budget degrades steeply\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_scaling_gap.png)Figure 18:Relative optimality gap against instrument size, one panel per budgetkk, atλ=0\.15\\lambda=0\.15\. Greedy and the warm\-started swarm are indistinguishable from the best solution found at every scale; the fixed\-budget swarm used in the simulation study degrades steeply\.#### Forward greedy is very hard to beat\.

Across the entire sweep, greedy stayed within 0\.10% of the best solution found \(Table[16](https://arxiv.org/html/2609.12086#S8.T16), Figure[18](https://arxiv.org/html/2609.12086#S8.F18)\)\. Where exhaustive search was tractable, it returned the exact optimum in 77\.7% of instances\. It is not a straw man, and the fact that Equation[6](https://arxiv.org/html/2609.12086#S5.E6)is not submodular in general\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.12086#bib.bib41)\)does not appear to cost it much in practice on instances with this cluster structure\.

#### The study’s swarm budget, not the swarm, is what failed\.

AFSA at the fixed budget of Section[6](https://arxiv.org/html/2609.12086#S6)degraded from a 0\.44% gap atn=16n=16to 43\.23% atn=256n=256, and found the exact optimum in only 46\.3% of the tractable instances against greedy’s 77\.7%\. It is a poor optimiser at anynnbeyond the smallest, and it was already worse than greedy at then=32n=32operating point the simulation used\. This is the direct explanation of the null in Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4): the swarm did not lose because population methods are unsuited to this problem, it lost because it was given a budget that does not grow with the search space\.

#### Growing the budget helps, up to a point\.

The scaled\-budget swarm stayed within 0\.6% up ton=64n=64and found the exact optimum in 84\.7% of tractable instances, exceeding greedy’s 77\.7%\. Beyond that it degrades again \(5\.46% atn=128n=128, 12\.59% atn=256n=256\), because even a budget growing withlog⁡\(nk\)\\log\\binom\{n\}\{k\}is a vanishing fraction of the search space\.

#### Warm\-starting from greedy dominates\.

The warm\-started swarm stayed within 0\.05% at every scale tested and found the exact optimum in 96\.3% of tractable instances\. Against greedy head to head it won or tied in 100% of instances \(16\.0% strict wins, 84\.0% ties, meanΔ=\+0\.001\\Delta=\+0\.001\), which is what one expects of a method that starts from greedy’s answer and keeps the better of the two\. It costs roughly seven times greedy’s evaluations \(Table[17](https://arxiv.org/html/2609.12086#S8.T17)\)\.

Table 17:Objective evaluations atk=8k=8,λ=0\.15\\lambda=0\.15, averaged over 25 instances\. Dense top\-kkand random require a single evaluation\. Exhaustive search is tractable only atn=16n=16\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_scaling_cost.png)Figure 19:Search cost in objective evaluations atk=8k=8,λ=0\.15\\lambda=0\.15\. The fixed\-budget swarm spends a roughly constant amount regardless of instrument size, which is exactly why its solution quality collapses asnngrows\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_scaling_winrate.png)Figure 20:Head\-to\-head strict win rate against forward greedy atλ=0\.15\\lambda=0\.15, averaged overkk\. Strict wins are rare because ties are common: greedy and the warm\-started swarm frequently return the same subset\.
#### The redundancy weight decides whether any of this matters\.

Figure[21](https://arxiv.org/html/2609.12086#S8.F21)and Table[18](https://arxiv.org/html/2609.12086#S8.T18)isolateλ\\lambdaatn=64n=64,k=8k=8\. Atλ=0\\lambda=0the objective is separable, dense top\-kkis exactly optimal by construction, and every search method is wasted effort\. Asλ\\lambdarises the relevance\-only solution degrades: a 0\.32% gap atλ=0\.15\\lambda=0\.15, 1\.51% atλ=0\.30\\lambda=0\.30and 7\.68% atλ=0\.60\\lambda=0\.60, at which point greedy also begins to lose ground \(1\.73%\)\. This locates the regime in which combinatorial search over a payload is worth performing at all: it is governed by how much a system penalises near\-duplicate fields, and at theλ=0\.15\\lambda=0\.15we used, the penalty is mild enough that dense top\-kkis nearly optimal\. That is the second half of the explanation for the null against dense top\-kkin Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4)\.

Table 18:Relative gap to the best solution found, in percent, against the redundancy weightλ\\lambdaatn=64n=64,k=8k=8\. Atλ=0\\lambda=0the objective is separable and sorting is optimal\.![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_scaling_lambda.png)Figure 21:Effect of the redundancy weight atn=64n=64,k=8k=8\. Atλ=0\\lambda=0the problem is a sort; asλ\\lambdarises, relevance\-only selection loses ground and combinatorial search starts to earn its cost\.

### 8\.5What this means for deployment

Three recommendations follow, and we state them as recommendations because they change what we would build\.

1. 1\.Use forward greedy\.At any instrument size we tested it is within 0\.1% of the best solution found, at a cost ofO⁡\(k​n\)O\(kn\)evaluations\. For an on\-device retriever over sixty fields this is a few hundred dot products and is not worth improving on\.
2. 2\.If a population method is used, warm\-start it from greedy\.A cold\-started swarm is strictly worse than greedy at every scale beyond the smallest\. A warm\-started one never loses, reaches the exhaustive optimum in 96% of tractable instances, and costs about seven times greedy\. Whether that is worth paying depends on whether the objective value is what one actually cares about, which Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4)suggests it may not be\.
3. 3\.Setλ\\lambdadeliberately, or drop the redundancy term\.Below roughlyλ=0\.3\\lambda=0\.3the redundancy penalty rarely changes which fields win, and a system that is not going to tuneλ\\lambdashould use dense top\-kkand save the search entirely\. Above it, the penalty matters and greedy is the right tool\.

The broader methodological point is that the null in Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4)was not uninformative\. Reading it required taking the optimisation question out of the language\-model setting, where a downstream metric with a moderate effect size cannot resolve a 0\.2% difference in objective value, and asking it where it could be answered\.

## 9Discussion

### 9\.1A context result, and why it is a privacy result

The finding we would put first is not the effect size but the ratio\. Eight retrieved fields, 211 tokens, matched the fidelity of injecting the entire 32\-field user model, 915 tokens\. Nothing in the study suggests that the remaining twenty\-four fields were doing work that the eight did not\.

For efficiency this is a modest saving\. For architecture it is the whole argument\. Section[4](https://arxiv.org/html/2609.12086#S4)committedaumto local residence: the store sits on the user’s device or in a user\-controlled vault, never in a third\-party log\. That commitment is easy to state and hard to keep if the whole model must be transmitted on every complex request, because a model that is transmitted in full on every turn has, for practical purposes, been uploaded\. A budgeted payload changes the character of the commitment\. If eight fields suffice, disclosure per query is bounded by construction, the bound is measurable, and the fields that never leave are disproportionately the inner\-shell ones, which are exactly the ones whose exposure carries the greatest cost\.

![Refer to caption](https://arxiv.org/html/2609.12086v1/Figures/fig_dataflow.png)Figure 22:What crosses the device boundary\. The full 32\-field model, 915 tokens in the study, never leaves the device\. A simple request leaves nothing\. A complex request sends a payload of eight fields, 211 tokens, which the user can be shown before it is sent\. The bound on disclosure per query is a measured quantity rather than a policy promise\.In the vocabulary of contextual integrity\([Nissenbaum, 2011](https://arxiv.org/html/2609.12086#bib.bib66)\), the design question is not whether the user model is secret but whether each flow out of it matches the norms of the context that produced it\. A retrieval stage that returns a small task\-conditioned slice is a mechanism for making flows contextual, and the size of the slice is the parameter that governs it\. This is what we mean by privacy by architecture rather than privacy by policy: the former is a property one can measure in tokens, and we measured it\.

There is a second consequence for scrutability\. Eight short natural\-language strings can be shown to a user before they are sent; thirty\-two cannot, and a JSON blob certainly cannot\. The scrutable\-personalisation literature has long argued that users should be able to inspect and correct the model a system holds of them\([Kay and Kummerfeld, 2012](https://arxiv.org/html/2609.12086#bib.bib70);[Czarkowski and Kay, 2002](https://arxiv.org/html/2609.12086#bib.bib71);[Bakalov et al\., 2010](https://arxiv.org/html/2609.12086#bib.bib72)\), and evidence that transparency and control change acceptance and behaviour is not new\([Cramer et al\., 2008](https://arxiv.org/html/2609.12086#bib.bib73);[Harper et al\., 2015](https://arxiv.org/html/2609.12086#bib.bib74)\)\. The obstacle has usually been that the model is too large or too statistical to show\. A budgeted payload of readable sentences removes that obstacle for the per\-request case: a user can be shown exactly what the assistant was told about them for this draft\.

### 9\.2What the null results bound, and what they do not

Four controls returned null, and Section[8](https://arxiv.org/html/2609.12086#S8)explained why\. We think the joint reading is more interesting than any of them individually\.

The positive results establish that injecting a structured, task\-plausible slice of a user model changes generated style substantially relative to both a generic baseline and the flat preference notes deployed systems store\. The nulls establish that, at this scale, nothing about*which*slice or about*how hard one searches for it*adds anything further\. Together they locate the effect in the representation rather than in the retrieval machinery\. That is a narrower claim than the one the architecture invites, and it has a practical implication we would act on: a first deployment should use forward greedy over a well\-specified schema, and should not spend engineering effort on the selection algorithm until the schema is larger and the redundancy penalty is doing real work\.

Two boundaries on the nulls deserve emphasis\. First, they are nulls at an operating point with a pool of about seventeen fields, and Section[8](https://arxiv.org/html/2609.12086#S8)shows the optimisation landscape is genuinely flat there: greedy is within 0\.1% of the exhaustive optimum, so there was nothing for a better search to recover\. They are not evidence that selection is irrelevant at the full sixty\-field specification, and we do not present them as such\. Second, a null against random selection within aσ⁡\(t\)\\sigma\(t\)\-filtered pool is not a null against random selection from the whole model\. The pool was already restricted to shells the task touches, so “random” here means random among plausible fields, which is a considerably stronger baseline than it sounds\.

### 9\.3Personalisation is worth most to the people the default serves least

The result with the clearest design consequence is the one in Section[7\.8](https://arxiv.org/html/2609.12086#S7.SS8): the benefit of the user model was largest exactly for those participants whose voice the un\-personalised assistant reproduced worst \(ρ=−0\.61\\rho=\-0\.61,p=0\.013p=0\.013at participant level\)\. The blunt minimalist gained\+0\.611\+0\.611; the meticulous planner, whose numbered and dated register a default assistant already produces, gained nothing\.

Read as a statistic this is a moderate correlation on sixteen points that did not replicate on the automatic style metric\. Read as a design observation it says something sharper\. A default assistant voice is not neutral\. It is a particular register, and being well served by an un\-personalised assistant means happening to write the way that register writes\. Personalisation, on this reading, is not a uniform improvement distributed evenly across users; it is a correction whose size depends on how far a person sits from the model’s prior\.

This connects to a finding that motivates the whole project\.[Jakesch et al\. \(2023\)](https://arxiv.org/html/2609.12086#bib.bib86)showed that co\-writing with an opinionated language model shifts users’ own expressed views\. If a default assistant voice can move what a person says, then users whose register is furthest from the default face the largest pull towards it, and are also the users for whom a model of their own voice does the most work\. Personalisation in that framing is not a convenience feature but a countermeasure, and the task\-level pattern in Table[12](https://arxiv.org/html/2609.12086#S7.T12)says the same thing at a different granularity: the effect was near zero on the deadline\-extension email, the most conventionalised genre in the battery, where the assistant’s default already is the correct answer\.

### 9\.4Implications for the design of personalised assistants

Four implications follow that we would apply to a system\.

1. 1\.Store the person, retrieve the preference\.The unit of storage should be the stable layer, and the situational preference should be derived at generation time\. This is what makes an unseen task addressable and a stale preference re\-interpretable rather than merely old\.
2. 2\.Budget the payload, and show it\.A small task\-conditioned payload costs nothing measurable in quality, bounds disclosure per request, and is small enough to display\. A system that shows the user the eight sentences it used has made scrutability concrete rather than aspirational\.
3. 3\.Route simple requests around the model\.The classifier stage routed factual lookups to bypass with 98\.8% accuracy\. This is cheap to implement and eliminates a class of disclosure that buys nothing, and we would regard omitting it as a design error rather than a missed optimisation\.
4. 4\.Expect uneven benefit, and design for the tail\.If the gain concentrates on users whose voice the default serves worst, then average\-case evaluation will systematically understate the value of personalisation to the users who need it most\. Reporting a mean effect across a user population is not sufficient; the distribution is the result\.

### 9\.5Limitations

#### Simulated participants\.

The most important limitation is the one stated first in Section[6\.1](https://arxiv.org/html/2609.12086#S6.SS1)\. Every participant here is language\-model\-simulated\. A judge model rating on behalf of a persona description has an easier task than a person recognising their own voice, and the absolute fidelity means sit near the middle of a five\-point scale for every condition, including the best\. The critical literature on synthetic participants\([Bisbee et al\., 2024](https://arxiv.org/html/2609.12086#bib.bib56);[Wang et al\., 2025](https://arxiv.org/html/2609.12086#bib.bib57)\)applies to this study in full, and no result here should be transferred to a claim about human users\.

#### The personas are not a sample\.

Sixteen personas written to be distinguishable are a friendlier population than any real one\. The idiolect markers were authored, so the style differences the study measures are differences an author put there\. In particular, conclusions about*which*users benefit should not be transferred to demographic groups, since the personas carry an author’s assumptions about how people with a given trait profile write\.

#### Judge reliability is moderate\.

ICC\(2,1\) of 0\.57 across seeds is moderate rather than good\([Koo and Li, 2016](https://arxiv.org/html/2609.12086#bib.bib111)\)\. Agreement within one scale point was 86%, so the judge is consistent about direction and imprecise about magnitude, which is adequate for paired contrasts and inadequate for absolute claims\. Position bias was not detected but the test was close to threshold \(p=0\.061p=0\.061\)\.

#### Acquisition is unspecified\.

aumspecifies the fields a user representation should contain and not how they are populated\. Self report is unreliable for inner\-shell entries\([Vazire, 2010](https://arxiv.org/html/2609.12086#bib.bib107);[Connelly and Ones, 2010](https://arxiv.org/html/2609.12086#bib.bib108);[Baumeister et al\., 2007](https://arxiv.org/html/2609.12086#bib.bib109)\); passive observation raises the privacy concerns the architecture exists to avoid; and clinical assessment does not scale\. We treat this as an open problem on whichaumas a target representation is agnostic, but it is the largest gap between this paper and a deployable system\.

#### Update is unspecified\.

Fields carry a nominal update periodτ\\tau, but no mechanism is given for detecting that a field has become stale or for revising it without re\-running the whole instrument\. Given that the motivating example in Section[1](https://arxiv.org/html/2609.12086#S1)is precisely a staleness failure, this is a conspicuous omission\.

#### Text only, and one language\.

The specification is text\-based\. Many personality cues are non\-verbal: tone of voice, facial expression, pacing, response latency\. The study also runs entirely in English, and the personality frameworksaumdraws on were developed predominantly in specific populations, so cross\-cultural validity of the inner tiers is untested\([McCrae, 2002](https://arxiv.org/html/2609.12086#bib.bib23)\)\.

#### Single generator and judge\.

One generator model and one judge model were used\. Whether the effect sizes transfer across model families is unknown, and a judge from a single family may carry systematic preferences that a panel would average out\.

#### Lifespan and capacity\.

aumas specified assumes an adult user with an articulable identity\. It does not represent young children, or persons with significant cognitive impairment, and we do not think it should be extended to them without a separate ethical analysis\.

### 9\.6Future work

The necessary next study is the human one this simulation substitutes for\. Its shape is determined by what the simulation cannot answer: participants populate their ownaum, write their own references for the same six tasks, and rate outputs from the same conditions blind, with the primary measure being self\-rated fidelity and forced\-choice identification of their own style\. Section[10](https://arxiv.org/html/2609.12086#S10)states the conditions under which we think that study can be run\.

Three further directions follow from the results rather than from the gaps\. First, acquisition: whether anaumcan be populated incrementally from ordinary interaction without the passive observation the architecture is meant to avoid, perhaps by asking the user a small number of well\-chosen questions at moments when the answer is cheap to give\. Second, update: detecting staleness by watching for conflict between a stored field and new evidence, which is what the cross\-shell internal\-conflicts entry exists to support\. Third, scale: whether the selection nulls of Section[7\.4](https://arxiv.org/html/2609.12086#S7.SS4)persist at the full sixty\-field specification, where Section[8](https://arxiv.org/html/2609.12086#S8)suggests the optimisation landscape is no longer flat\.

## 10Ethical Considerations

A structured user model that records inner\-shell content such as fears, difficult episodes and core values is among the most sensitive artefacts a person could hold\. Building one creates risks that a preference list does not, and we do not think those risks are adequately answered by noting that the model is opt\-in\.

#### Three binding constraints\.

We bindaumto three\.Local residence: the store sits on the user’s device or in a user\-controlled vault, never in a third\-party log, and only the retrieved payload leaves it\. Section[7](https://arxiv.org/html/2609.12086#S7)makes this quantitative rather than aspirational: a payload of eight fields reaches the fidelity of the whole model, so budgeted retrieval bounds what any single query can disclose at no measured cost in quality\.User authorship: every field is inspectable, editable and deletable, and eight short strings are few enough to display before they are sent\.Purpose restriction:aumis intended for task\-style personalisation under the user’s direction, not for profiling, advertising, surveillance, or persuasion against the user’s interest\.

#### The misuse the architecture cannot prevent\.

A high\-resolution model of a person is valuable to anyone who wants to influence that person, and an operator who ignores the local\-residence constraint gains a profiling asset of a kind that does not currently exist at scale\. Our own results sharpen this rather than softening it: Section[7\.8](https://arxiv.org/html/2609.12086#S7.SS8)found the effect largest for participants whose voice the default serves worst, and those are also the people for whom an accurate model would be most identifying\. Users of LLM\-based conversational agents already disclose a great deal and reason about the risks only partially\([Zhang et al\., 2024b](https://arxiv.org/html/2609.12086#bib.bib67)\), and language models can leak personal information present in training data\([Huang et al\., 2022](https://arxiv.org/html/2609.12086#bib.bib68)\)\.

We do not think this argues against specifying the representation\. Opaque, implicit user models already sit inside large\-scale systems with weaker user control, and an explicit schema at least makes the contents auditable and contestable\. But mitigation has to be architectural, in where the data lives, as well as procedural, in who may query it and for what, and a paper that proposes the representation without saying so would be incomplete\.

#### Human participants\.

No human participants were involved in this work and no human\-subjects data was collected\. Every participant in Sections[6](https://arxiv.org/html/2609.12086#S6)and[7](https://arxiv.org/html/2609.12086#S7)is language\-model\-simulated and every persona is fictional and declared in code \(Appendix[C](https://arxiv.org/html/2609.12086#A3)\)\. The human study described in Section[9\.6](https://arxiv.org/html/2609.12086#S9.SS6)would require informed consent, the right to withdraw with deletion of all collectedaumfields, and institutional ethics approval, since participants would disclose identity\-level content about themselves\. Inner\-shell fields should be optional by default and skippable without penalty, and participants should be able to review their populated instrument before any of it is used\.

#### A concern specific to simulation\.

Personas written by an author carry that author’s assumptions about how people of a given trait profile write\. Ours are deliberately varied but are not a sample of any population\. Conclusions about which users benefit from personalisation should not be transferred to demographic groups on the basis of this study, and we have tried to phrase Section[7\.8](https://arxiv.org/html/2609.12086#S7.SS8)so that it cannot be read that way\.

## 11Conclusion

Personalisation of language\-model assistants currently stores what a user has done\. We argued that it should store the person from whom those doings come, and that the two are not interchangeable because personality is comparatively stable while preferences are situational\. We named the phenomenon that makes the difference visible, personality seepage, in which a prompt carries a fingerprint the assistant mirrors without being able to read\. We specified the Atomic User Model, a five\-tier structured and human\-readable representation whose metaphor encodes differential stability and differential observability, and a personality\-aware retrieval pipeline that treats it as an index rather than a prompt prefix\.

In a simulation study with sixteen language\-model\-simulated participants, six style\-sensitive tasks and three seeds, retrieving eight fields matched the style fidelity of injecting the whole user model at 23% of the context, improved on the flat preference notes deployed systems store by 0\.24 points on a five\-point scale, and raised forced\-choice identification of the participant’s own voice from 14\.9% to 42\.7% against 25% chance\. Four controls returned null, and a synthetic scaling study of the retriever explained why: at the operating point used, forward greedy is already within 0\.1% of the exhaustive optimum, so there was nothing for a population method to recover\. The effect lives in the representation, not in the search over it\.

The result we would carry forward is neither the effect size nor the null\. It is that a small, readable, task\-conditioned payload is enough, because that is what makes a user model something a person can keep on their own device and read before it is sent\. And it is that the benefit was largest for the participants whose voice the default assistant served worst, which suggests that personalisation is least valuable to the users who resemble the model’s prior and most valuable to everyone else\.

## CRediT authorship contribution statement

B\. Sankar: Conceptualization, Methodology, Software, Formal analysis, Investigation, Writing, original draft, Writing, review and editing, Supervision, Project administration\.Deepthika S: Methodology, Software, Validation, Investigation, Data curation, Writing, review and editing\.Pawni Yadav: Methodology, Software, Validation, Investigation, Visualization, Writing, review and editing\.Amogh A S: Software, Validation, Data curation, Visualization, Writing, review and editing\.

## Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.

## Declaration of generative AI in the writing process

The studies reported in Sections[6](https://arxiv.org/html/2609.12086#S6)to[8](https://arxiv.org/html/2609.12086#S8)use large language models as the object of study and as simulated participants; this use is described in full in those sections\. In preparing this manuscript the authors additionally used a large language model to assist with drafting and editing of prose\. The authors reviewed and edited all content and take full responsibility for the content of the publication\.

## Data and code availability

The simulation code, the persona declarations, every prompt, and the run manifest and ledger, all per\-trial records, the analysis scripts, and the scaling benchmark are available at[https://github\.com/REPOSITORY\-URL\-HERE](https://github.com/REPOSITORY-URL-HERE)\. Each number reported in Sections[7](https://arxiv.org/html/2609.12086#S7)and[8](https://arxiv.org/html/2609.12086#S8)is reproducible from the released artefacts and a seed\.

## Funding declaration

This research did not receive any specific grant from funding agencies in the public, commercial, or not\-for\-profit sectors\.

## References

- Agrawalet al\.\(2009\)R\. Agrawal, S\. Gollapudi, A\. Halverson, and S\. IeongDiversifying search results\.InProceedings of the Second ACM International Conference on Web Search and Data Mining,pp\. 5–14\.External Links:[Document](https://dx.doi.org/10.1145/1498759.1498766)Cited by:[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.12086#S5.SS2.p1.2)\.
- Alhafniet al\.\(2024\)B\. Alhafni, V\. Kulkarni, D\. Kumar, and V\. RahejaPersonalized text generation with fine\-grained linguistic control\.InProceedings of the 1st Workshop on Personalization of Generative AI Systems \(PERSONALIZE 2024\),pp\. 88–101\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.personalize-1.8)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- Altman and Taylor \(1973\)I\. Altman and D\. A\. TaylorSocial penetration: the development of interpersonal relationships\.Holt, Rinehart and Winston,New York\.Cited by:[item Differential observability\.](https://arxiv.org/html/2609.12086#S4.I1.ix2.p1.1)\.
- Andersonet al\.\(2004\)J\. R\. Anderson, D\. Bothell, M\. D\. Byrne, S\. Douglass, C\. Lebiere, and Y\. QinAn integrated theory of the mind\.Psychological Review111\(4\),pp\. 1036–1060\.External Links:[Document](https://dx.doi.org/10.1037/0033-295X.111.4.1036)Cited by:[§2\.10](https://arxiv.org/html/2609.12086#S2.SS10.p1.1)\.
- Andrews and Bishop \(2019\)N\. Andrews and M\. BishopLearning invariant representations of social media users\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 1684–1695\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1178)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- Argyleet al\.\(2023\)L\. P\. Argyle, E\. C\. Busby, N\. Fulda, J\. R\. Gubler, C\. Rytting, and D\. WingateOut of one, many: using language models to simulate human samples\.Political Analysis31\(3\),pp\. 337–351\.External Links:[Document](https://dx.doi.org/10.1017/pan.2023.2)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p1.1),[§6\.1](https://arxiv.org/html/2609.12086#S6.SS1.p2.1)\.
- Artstein and Poesio \(2008\)R\. Artstein and M\. PoesioInter\-coder agreement for computational linguistics\.Computational Linguistics34\(4\),pp\. 555–596\.External Links:[Document](https://dx.doi.org/10.1162/coli.07-034-R2)Cited by:[§7\.7](https://arxiv.org/html/2609.12086#S7.SS7.SSS0.Px4.p1.1)\.
- Azucaret al\.\(2018\)D\. Azucar, D\. Marengo, and M\. SettanniPredicting the Big 5 personality traits from digital footprints on social media: a meta\-analysis\.Personality and Individual Differences124,pp\. 150–159\.External Links:[Document](https://dx.doi.org/10.1016/j.paid.2017.12.018)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Bakalovet al\.\(2010\)F\. Bakalov, B\. König\-Ries, A\. Nauerz, and M\. WelschIntrospectiveViews: an interface for scrutinizing semantic user models\.InUser Modeling, Adaptation, and Personalization,Lecture Notes in Computer Science,pp\. 219–230\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-13470-8%5F21)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px2.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p4.1)\.
- Baumeisteret al\.\(2007\)R\. F\. Baumeister, K\. D\. Vohs, and D\. C\. FunderPsychology as the science of self\-reports and finger movements: whatever happened to actual behavior?\.Perspectives on Psychological Science2\(4\),pp\. 396–403\.External Links:[Document](https://dx.doi.org/10.1111/j.1745-6916.2007.00051.x)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p3.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px4.p1.1)\.
- Bennettet al\.\(2012\)P\. N\. Bennett, R\. W\. White, W\. Chu, S\. T\. Dumais, P\. Bailey, F\. Borisyuk, and X\. CuiModeling the impact of short\- and long\-term behavior on search personalization\.InProceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’12,New York, NY, USA,pp\. 185–194\.External Links:[Document](https://dx.doi.org/10.1145/2348283.2348312)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1)\.
- Bisbeeet al\.\(2024\)J\. Bisbee, J\. D\. Clinton, C\. Dorff, B\. Kenkel, and J\. M\. LarsonSynthetic replacements for human survey data? the perils of large language models\.Political Analysis32\(4\),pp\. 401–416\.External Links:[Document](https://dx.doi.org/10.1017/pan.2024.5)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p1.1),[§6\.1](https://arxiv.org/html/2609.12086#S6.SS1.p2.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px1.p1.1)\.
- Bowlby \(1969\)J\. BowlbyAttachment\.Attachment and Loss, Vol\.1,Basic Books,New York\.Cited by:[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px2.p1.1)\.
- Boyd and Pennebaker \(2017\)R\. L\. Boyd and J\. W\. PennebakerLanguage\-based personality: a new approach to personality in a digital world\.Current Opinion in Behavioral Sciences18,pp\. 63–68\.External Links:[Document](https://dx.doi.org/10.1016/j.cobeha.2017.07.017)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Brusilovsky and Millán \(2007\)P\. Brusilovsky and E\. MillánUser models for adaptive hypermedia and adaptive educational systems\.InThe Adaptive Web: Methods and Strategies of Web Personalization,P\. Brusilovsky, A\. Kobsa, and W\. Nejdl \(Eds\.\),Lecture Notes in Computer Science, Vol\.4321,pp\. 3–53\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-72079-9%5F1)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1)\.
- Carbonell and Goldstein \(1998\)J\. Carbonell and J\. GoldsteinThe use of MMR, diversity\-based reranking for reordering documents and producing summaries\.InProceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 335–336\.External Links:[Document](https://dx.doi.org/10.1145/290941.291025)Cited by:[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.12086#S5.SS2.p1.2)\.
- Carpenter and Greene \(2015\)A\. Carpenter and K\. GreeneSocial penetration theory\.InThe International Encyclopedia of Interpersonal Communication,C\. R\. Berger and M\. E\. Roloff \(Eds\.\),pp\. 1–4\.External Links:[Document](https://dx.doi.org/10.1002/9781118540190.wbeic160)Cited by:[item Differential observability\.](https://arxiv.org/html/2609.12086#S4.I1.ix2.p1.1)\.
- Chenet al\.\(2019\)M\. Chen, A\. T\. Suresh, R\. Mathews, A\. Wong, C\. Allauzen, F\. Beaufays, and M\. RileyFederated learning of n\-gram language models\.InProceedings of the 23rd Conference on Computational Natural Language Learning \(CoNLL\),pp\. 121–130\.External Links:[Document](https://dx.doi.org/10.18653/v1/K19-1012)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p2.1)\.
- Chhetriet al\.\(2026\)G\. Chhetri, S\. Das, and T\. I\. ChowdhurySPARK: search personalization via agent\-driven retrieval and knowledge\-sharing\.InProceedings of the Nineteenth ACM International Conference on Web Search and Data Mining,WSDM ’26,New York, NY, USA,pp\. 84–92\.External Links:[Document](https://dx.doi.org/10.1145/3779211.3793173)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Chiang and Lee \(2023\)C\. Chiang and H\. LeeCan large language models be an alternative to human evaluations?\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15607–15631\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.870)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p2.1)\.
- Choet al\.\(2022\)I\. Cho, D\. Wang, R\. Takahashi, and H\. SaitoA personalized dialogue generator with implicit user persona detection\.InProceedings of the 29th International Conference on Computational Linguistics,Gyeongju, Republic of Korea,pp\. 367–377\.External Links:[Link](https://aclanthology.org/2022.coling-1.29/)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Clarkeet al\.\(2008\)C\. L\. A\. Clarke, M\. Kolla, G\. V\. Cormack, O\. Vechtomova, A\. Ashkan, S\. Büttcher, and I\. MacKinnonNovelty and diversity in information retrieval evaluation\.InProceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 659–666\.External Links:[Document](https://dx.doi.org/10.1145/1390334.1390446)Cited by:[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.12086#S5.SS2.p1.2)\.
- Connelly and Ones \(2010\)B\. S\. Connelly and D\. S\. OnesAn other perspective on personality: meta\-analytic integration of observers’ accuracy and predictive validity\.Psychological Bulletin136\(6\),pp\. 1092–1122\.External Links:[Document](https://dx.doi.org/10.1037/a0021212)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p3.1),[§6\.3](https://arxiv.org/html/2609.12086#S6.SS3.p2.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px4.p1.1)\.
- Crameret al\.\(2008\)H\. Cramer, V\. Evers, S\. Ramlal, M\. van Someren, L\. Rutledge, N\. Stash, L\. Aroyo, and B\. WielingaThe effects of transparency on trust in and acceptance of a content\-based art recommender\.User Modeling and User\-Adapted Interaction18\(5\),pp\. 455–496\.External Links:[Document](https://dx.doi.org/10.1007/s11257-008-9051-3)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p4.1)\.
- Czarkowski and Kay \(2002\)M\. Czarkowski and J\. KayA scrutable adaptive hypertext\.InAdaptive Hypermedia and Adaptive Web\-Based Systems,Lecture Notes in Computer Science,pp\. 384–387\.External Links:[Document](https://dx.doi.org/10.1007/3-540-47952-X%5F43)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px2.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p4.1)\.
- Douet al\.\(2007\)Z\. Dou, R\. Song, and J\. WenA large\-scale evaluation and analysis of personalized search strategies\.InProceedings of the 16th International Conference on World Wide Web,WWW ’07,New York, NY, USA,pp\. 581–590\.External Links:[Document](https://dx.doi.org/10.1145/1242572.1242651)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1)\.
- Epsteinet al\.\(2015\)D\. A\. Epstein, A\. Ping, J\. Fogarty, and S\. A\. MunsonA lived informatics model of personal informatics\.InProceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing,pp\. 731–742\.External Links:[Document](https://dx.doi.org/10.1145/2750858.2804250)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p2.1)\.
- Erikson \(1968\)E\. H\. EriksonIdentity: youth and crisis\.W\. W\. Norton,New York\.External Links:ISBN 9780393311440Cited by:[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px1.p1.1)\.
- Geroet al\.\(2020\)K\. I\. Gero, Z\. Ashktorab, C\. Dugan, Q\. Pan, J\. Johnson, W\. Geyer, M\. Ruiz, S\. Miller, D\. R\. Millen, M\. Campbell, S\. Kumaravel, and W\. ZhangMental models of AI agents in a cooperative game setting\.InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems,pp\. 1–12\.External Links:[Document](https://dx.doi.org/10.1145/3313831.3376316)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1)\.
- Goffman \(1959\)E\. GoffmanThe presentation of self in everyday life\.Doubleday,Garden City, New York\.Cited by:[item Persona](https://arxiv.org/html/2609.12086#S3.I1.ix3.p1.1),[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px5.p1.1)\.
- Goldberget al\.\(2006\)L\. R\. Goldberg, J\. A\. Johnson, H\. W\. Eber, R\. Hogan, M\. C\. Ashton, C\. R\. Cloninger, and H\. G\. GoughThe international personality item pool and the future of public\-domain personality measures\.Journal of Research in Personality40\(1\),pp\. 84–96\.External Links:[Document](https://dx.doi.org/10.1016/j.jrp.2005.08.007)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1)\.
- Goslinget al\.\(2003\)S\. D\. Gosling, P\. J\. Rentfrow, and W\. B\. SwannA very brief measure of the Big\-Five personality domains\.Journal of Research in Personality37\(6\),pp\. 504–528\.External Links:[Document](https://dx.doi.org/10.1016/S0092-6566%2803%2900046-1)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1)\.
- Goyal and Tawde \(2022\)M\. Goyal and P\. TawdeA research attempt to predict and model personalities through users’ social media details\.In2022 IEEE Bombay Section Signature Conference \(IBSSC\),pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/IBSSC56953.2022.10037272)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Harperet al\.\(2015\)F\. M\. Harper, F\. Xu, H\. Kaur, K\. Condiff, S\. Chang, and L\. TerveenPutting users in control of their recommendations\.InProceedings of the 9th ACM Conference on Recommender Systems,pp\. 3–10\.External Links:[Document](https://dx.doi.org/10.1145/2792838.2800179)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p4.1)\.
- Huanget al\.\(2025\)F\. Huang, Y\. Bei, Z\. Yang, J\. Jiang, H\. Chen, Q\. Shen, S\. Wang, F\. Karray, and P\. S\. YuLarge language model simulator for cold\-start recommendation\.InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining,WSDM ’25,New York, NY, USA,pp\. 261–270\.External Links:[Document](https://dx.doi.org/10.1145/3701551.3703546)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p2.1)\.
- Huanget al\.\(2022\)J\. Huang, H\. Shao, and K\. C\. ChangAre large pre\-trained language models leaking your personal information?\.InFindings of the Association for Computational Linguistics: EMNLP 2022,pp\. 2038–2047\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.148)Cited by:[§10](https://arxiv.org/html/2609.12086#S10.SS0.SSS0.Px2.p1.1),[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p2.1)\.
- Jakeschet al\.\(2023\)M\. Jakesch, A\. Bhat, D\. Buschek, L\. Zalmanson, and M\. NaamanCo\-writing with opinionated language models affects users’ views\.InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems,pp\. 1–15\.External Links:[Document](https://dx.doi.org/10.1145/3544548.3581196)Cited by:[§2\.7](https://arxiv.org/html/2609.12086#S2.SS7.p1.1),[§9\.3](https://arxiv.org/html/2609.12086#S9.SS3.p3.1)\.
- Jeromela \(2022\)J\. JeromelaScrutability of intelligent personal assistants\.InProceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization,pp\. 335–340\.External Links:[Document](https://dx.doi.org/10.1145/3503252.3534355)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1)\.
- Jiet al\.\(2023\)Y\. Ji, W\. Wu, H\. Zheng, Y\. Hu, X\. Chen, and L\. HeIs ChatGPT a good personality recognizer? A preliminary study\.External Links:2307\.03952,[Document](https://dx.doi.org/10.48550/arXiv.2307.03952),[Link](https://arxiv.org/abs/2307.03952)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px2.p1.1)\.
- Jianget al\.\(2023\)H\. Jiang, Q\. Wu, C\. Lin, Y\. Yang, and L\. QiuLLMLingua: compressing prompts for accelerated inference of large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 13358–13376\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.825)Cited by:[§2\.3](https://arxiv.org/html/2609.12086#S2.SS3.p1.1)\.
- Jianget al\.\(2024\)H\. Jiang, Q\. Wu, X\. Luo, D\. Li, C\. Lin, Y\. Yang, and L\. QiuLongLLMLingua: accelerating and enhancing LLMs in long context scenarios via prompt compression\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1658–1677\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.91)Cited by:[§2\.3](https://arxiv.org/html/2609.12086#S2.SS3.p1.1)\.
- Kanget al\.\(2023\)W\. Kang, J\. Ni, N\. Mehta, M\. Sathiamoorthy, L\. Hong, E\. Chi, and D\. Z\. ChengDo LLMs understand user preferences? evaluating LLMs on user rating prediction\.External Links:2305\.06474,[Document](https://dx.doi.org/10.48550/arXiv.2305.06474),[Link](https://arxiv.org/abs/2305.06474)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p2.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px2.p1.1)\.
- Kay and Kummerfeld \(2012\)J\. Kay and B\. KummerfeldCreating personalized systems that people can scrutinize and control: Drivers, principles and experience\.ACM Transactions on Interactive Intelligent Systems2\(4\),pp\. 1–42\.External Links:[Document](https://dx.doi.org/10.1145/2395123.2395129)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1),[item 1](https://arxiv.org/html/2609.12086#S4.I2.i1.p1.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px2.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p4.1)\.
- Kobsa \(2001\)A\. KobsaGeneric user modeling systems\.User Modeling and User\-Adapted Interaction11\(1–2\),pp\. 49–63\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1011187500863)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1)\.
- Koo and Li \(2016\)T\. K\. Koo and M\. Y\. LiA guideline of selecting and reporting intraclass correlation coefficients for reliability research\.Journal of Chiropractic Medicine15\(2\),pp\. 155–163\.External Links:[Document](https://dx.doi.org/10.1016/j.jcm.2016.02.012)Cited by:[§7\.7](https://arxiv.org/html/2609.12086#S7.SS7.SSS0.Px4.p1.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px3.p1.1)\.
- Kosinskiet al\.\(2013\)M\. Kosinski, D\. Stillwell, and T\. GraepelPrivate traits and attributes are predictable from digital records of human behavior\.Proceedings of the National Academy of Sciences110\(15\),pp\. 5802–5805\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1218772110)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Krippendorff \(2004\)K\. KrippendorffReliability in content analysis: some common misconceptions and recommendations\.Human Communication Research30\(3\),pp\. 411–433\.External Links:[Document](https://dx.doi.org/10.1111/j.1468-2958.2004.tb00738.x)Cited by:[§7\.7](https://arxiv.org/html/2609.12086#S7.SS7.SSS0.Px4.p1.1)\.
- Kumaret al\.\(2024\)I\. Kumar, S\. Viswanathan, S\. Yerra, A\. Salemi, R\. A\. Rossi, F\. Dernoncourt, H\. Deilamsalehy, X\. Chen, R\. Zhang, S\. Agarwal, N\. Lipka, C\. V\. Nguyen, T\. H\. Nguyen, and H\. ZamaniLongLaMP: a benchmark for personalized long\-form text generation\.External Links:2407\.11016,[Document](https://dx.doi.org/10.48550/arXiv.2407.11016),[Link](https://arxiv.org/abs/2407.11016)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p2.1)\.
- Laird \(2012\)J\. E\. LairdThe Soar cognitive architecture\.MIT Press,Cambridge, MA\.External Links:ISBN 9780262122962Cited by:[§2\.10](https://arxiv.org/html/2609.12086#S2.SS10.p1.1)\.
- Leeet al\.\(2024\)M\. Lee, K\. I\. Gero, J\. J\. Y\. Chung, S\. Buckingham Shum, V\. Raheja, H\. Shen, S\. Venugopalan, T\. Wambsganss, D\. Zhou, E\. A\. Alghamdi, T\. August, A\. Bhat, M\. Z\. Choksi, S\. Dutta, J\. L\. C\. Guo, M\. N\. Hoque, Y\. Kim, S\. Knight, S\. P\. Neshaei, A\. Shibani, D\. Shrivastava, L\. Shroff, A\. Sergeyuk, J\. Stark, S\. Sterman, S\. Wang, A\. Bosselut, D\. Buschek, J\. C\. Chang, S\. Chen, M\. Kreminski, J\. Park, R\. Pea, E\. H\. R\. Rho, Z\. Shen, and P\. SiangliulueA design space for intelligent and interactive writing assistants\.InProceedings of the CHI Conference on Human Factors in Computing Systems,pp\. 1–35\.External Links:[Document](https://dx.doi.org/10.1145/3613904.3642697)Cited by:[§2\.7](https://arxiv.org/html/2609.12086#S2.SS7.p1.1)\.
- Leeet al\.\(2022\)M\. Lee, P\. Liang, and Q\. YangCoAuthor: designing a human\-AI collaborative writing dataset for exploring language model capabilities\.InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems,pp\. 1–19\.External Links:[Document](https://dx.doi.org/10.1145/3491102.3502030)Cited by:[§2\.7](https://arxiv.org/html/2609.12086#S2.SS7.p1.1)\.
- Liet al\.\(2010\)I\. Li, A\. Dey, and J\. ForlizziA stage\-based model of personal informatics systems\.InProceedings of the SIGCHI Conference on Human Factors in Computing Systems,pp\. 557–566\.External Links:[Document](https://dx.doi.org/10.1145/1753326.1753409)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p2.1)\.
- Liet al\.\(2002\)X\. L\. Li, Z\. J\. Shao, and J\. X\. QianAn optimizing method based on autonomous animats: fish\-swarm algorithm\.Systems Engineering \- Theory & Practice22\(11\),pp\. 32–38\.Note:No DOI is registered for this record; verified on the publisher’s article pageCited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p2.1),[§5\.3](https://arxiv.org/html/2609.12086#S5.SS3.SSS0.Px4.p1.1)\.
- Liet al\.\(2023\)Y\. Li, B\. Dong, F\. Guerin, and C\. LinCompressing context to enhance inference efficiency of large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 6342–6353\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.391)Cited by:[§2\.3](https://arxiv.org/html/2609.12086#S2.SS3.p1.1)\.
- Liaoet al\.\(2026\)Y\. Liao, Y\. Deng, T\. Jiang, and J\. RenPersonalizing large language models with user profile memory\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2,KDD ’26,New York, NY, USA,pp\. 3051–3061\.External Links:[Document](https://dx.doi.org/10.1145/3770855.3817930)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Linet al\.\(2024\)Y\. Lin, L\. Chen, A\. Ali, C\. Nugent, I\. Cleland, R\. Li, J\. Ding, and H\. NingHuman digital twin: a survey\.Journal of Cloud Computing13\(1\),pp\. 131\.External Links:[Document](https://dx.doi.org/10.1186/s13677-024-00691-z)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p2.1)\.
- Liuet al\.\(2023\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-Eval: NLG evaluation using GPT\-4 with better human alignment\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 2511–2522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.153)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p2.1)\.
- Luger and Sellen \(2016\)E\. Luger and A\. Sellen“Like having a really bad PA”: the gulf between user expectation and experience of conversational agents\.InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems,pp\. 5286–5297\.External Links:[Document](https://dx.doi.org/10.1145/2858036.2858288)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1)\.
- Mahmoodet al\.\(2025\)A\. Mahmood, J\. Wang, B\. Yao, D\. Wang, and C\. HuangUser interaction patterns and breakdowns in conversing with LLM\-powered voice assistants\.International Journal of Human\-Computer Studies195,pp\. 103406\.External Links:[Document](https://dx.doi.org/10.1016/j.ijhcs.2024.103406)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1)\.
- Man and Nguyen \(2024\)H\. Man and T\. H\. NguyenCounterfactual augmentation for robust authorship representation learning\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 2347–2351\.External Links:[Document](https://dx.doi.org/10.1145/3626772.3657956)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- McCrae and Costa \(1989\)R\. R\. McCrae and P\. T\. CostaReinterpreting the Myers\-Briggs Type Indicator from the perspective of the five\-factor model of personality\.Journal of Personality57\(1\),pp\. 17–40\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-6494.1989.tb00759.x)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1)\.
- McCrae and John \(1992\)R\. R\. McCrae and O\. P\. JohnAn introduction to the five\-factor model and its applications\.Journal of Personality60\(2\),pp\. 175–215\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-6494.1992.tb00970.x)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1),[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px2.p1.1)\.
- McCrae \(2002\)R\. R\. McCraeCross\-cultural research on the five\-factor model of personality\.Online Readings in Psychology and Culture4\(4\)\.External Links:[Document](https://dx.doi.org/10.9707/2307-0919.1038)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px6.p1.1)\.
- Miret al\.\(2019\)R\. Mir, B\. Felbo, N\. Obradovich, and I\. RahwanEvaluating style transfer for text\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 495–504\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1049)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- Nemhauseret al\.\(1978\)G\. L\. Nemhauser, L\. A\. Wolsey, and M\. L\. FisherAn analysis of approximations for maximizing submodular set functions i\.Mathematical Programming14\(1\),pp\. 265–294\.External Links:[Document](https://dx.doi.org/10.1007/BF01588971)Cited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p1.1),[§5\.2](https://arxiv.org/html/2609.12086#S5.SS2.p3.1),[§8\.4](https://arxiv.org/html/2609.12086#S8.SS4.SSS0.Px1.p1.1)\.
- Neshatet al\.\(2014\)M\. Neshat, G\. Sepidnam, M\. Sargolzaei, and A\. N\. ToosiArtificial fish swarm algorithm: a survey of the state\-of\-the\-art, hybridization, combinatorial and indicative applications\.Artificial Intelligence Review42\(4\),pp\. 965–997\.Note:Published online 2012; print issue 2014External Links:[Document](https://dx.doi.org/10.1007/s10462-012-9342-2)Cited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p2.1)\.
- Ngoet al\.\(2020\)T\. Ngo, J\. Kunkel, and J\. ZieglerExploring mental models for transparent and controllable recommender systems: a qualitative study\.InProceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization,pp\. 183–191\.External Links:[Document](https://dx.doi.org/10.1145/3340631.3394841)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p1.1)\.
- Nguyenet al\.\(2020\)B\. H\. Nguyen, B\. Xue, and M\. ZhangA survey on swarm intelligence approaches to feature selection in data mining\.Swarm and Evolutionary Computation54,pp\. 100663\.External Links:[Document](https://dx.doi.org/10.1016/j.swevo.2020.100663)Cited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p2.1)\.
- Nissenbaum \(2011\)H\. NissenbaumA contextual approach to privacy online\.Daedalus140\(4\),pp\. 32–48\.External Links:[Document](https://dx.doi.org/10.1162/daed%5Fa%5F00113)Cited by:[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p2.1),[§4\.5](https://arxiv.org/html/2609.12086#S4.SS5.SSS0.Px1.p1.1),[§9\.1](https://arxiv.org/html/2609.12086#S9.SS1.p3.1)\.
- Ostheimeret al\.\(2023\)P\. Ostheimer, M\. K\. Nagda, M\. Kloft, and S\. FellenzA call for standardization and validation of text style transfer evaluation\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 10791–10815\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.687)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.External Links:2310\.08560,[Document](https://dx.doi.org/10.48550/arXiv.2310.08560),[Link](https://arxiv.org/abs/2310.08560)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Panicksseryet al\.\(2024\)A\. Panickssery, S\. Bowman, and S\. FengLLM evaluators recognize and favor their own generations\.InAdvances in Neural Information Processing Systems 37,Vol\.37,pp\. 68772–68802\.External Links:[Document](https://dx.doi.org/10.52202/079017-2197)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p2.1),[§6\.7](https://arxiv.org/html/2609.12086#S6.SS7.SSS0.Px5.p1.1)\.
- Parket al\.\(2015\)G\. Park, H\. A\. Schwartz, J\. C\. Eichstaedt, M\. L\. Kern, M\. Kosinski, D\. J\. Stillwell, L\. H\. Ungar, and M\. E\. P\. SeligmanAutomatic personality assessment through social media language\.Journal of Personality and Social Psychology108\(6\),pp\. 934–952\.External Links:[Document](https://dx.doi.org/10.1037/pspp0000020)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,UIST ’23,New York, NY, USA,pp\. 1–22\.External Links:[Document](https://dx.doi.org/10.1145/3586183.3606763)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Parket al\.\(2024\)J\. S\. Park, C\. Q\. Zou, J\. Kamphorst, N\. Egan, A\. Shaw, B\. M\. Hill, C\. Cai, M\. R\. Morris, P\. Liang, R\. Willer, and M\. S\. BernsteinLLM agents grounded in self\-reports enable general\-purpose simulation of individuals\.External Links:2411\.10109,[Document](https://dx.doi.org/10.48550/arXiv.2411.10109),[Link](https://arxiv.org/abs/2411.10109)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p1.1),[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p1.1),[§6\.1](https://arxiv.org/html/2609.12086#S6.SS1.p2.1)\.
- Patelet al\.\(2023\)A\. Patel, D\. Rao, A\. Kothary, K\. McKeown, and C\. Callison\-BurchLearning interpretable style embeddings via prompting LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 15270–15290\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.1020)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1)\.
- Pennebaker and King \(1999\)J\. W\. Pennebaker and L\. A\. KingLinguistic styles: language use as an individual difference\.Journal of Personality and Social Psychology77\(6\),pp\. 1296–1312\.External Links:[Document](https://dx.doi.org/10.1037/0022-3514.77.6.1296)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1),[§3\.2](https://arxiv.org/html/2609.12086#S3.SS2.p1.1)\.
- Pittenger \(2005\)D\. J\. PittengerCautionary comments regarding the Myers\-Briggs Type Indicator\.Consulting Psychology Journal: Practice and Research57\(3\),pp\. 210–221\.External Links:[Document](https://dx.doi.org/10.1037/1065-9293.57.3.210)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1),[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px2.p1.1)\.
- Pourpanahet al\.\(2023\)F\. Pourpanah, R\. Wang, C\. P\. Lim, X\. Wang, and D\. YazdaniA review of artificial fish swarm algorithms: recent advances and applications\.Artificial Intelligence Review56\(3\),pp\. 1867–1903\.Note:Published online 2022; print issue 2023External Links:[Document](https://dx.doi.org/10.1007/s10462-022-10214-4)Cited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p2.1)\.
- Raoet al\.\(2023\)H\. Rao, C\. Leung, and C\. MiaoCan ChatGPT assess human personalities? A general evaluation framework\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 1184–1194\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.84/)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Rapp and Cena \(2016\)A\. Rapp and F\. CenaPersonal informatics for everyday life: How users without prior self\-tracking experience engage with personal data\.International Journal of Human\-Computer Studies94,pp\. 1–17\.External Links:[Document](https://dx.doi.org/10.1016/j.ijhcs.2016.05.006)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p2.1)\.
- Rezaet al\.\(2025\)M\. Reza, J\. Thomas\-Mitchell, P\. Dushniku, N\. Laundry, J\. J\. Williams, and A\. KuzminykhCo\-writing with AI, on human terms: Aligning research with user demands across the writing process\.Proceedings of the ACM on Human\-Computer Interaction9\(7\),pp\. 1–37\.External Links:[Document](https://dx.doi.org/10.1145/3757566)Cited by:[§2\.7](https://arxiv.org/html/2609.12086#S2.SS7.p1.1)\.
- Rivera\-Sotoet al\.\(2021\)R\. A\. Rivera\-Soto, O\. E\. Miano, J\. Ordonez, B\. Y\. Chen, A\. Khan, M\. Bishop, and N\. AndrewsLearning universal authorship representations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,Online and Punta Cana, Dominican Republic,pp\. 913–919\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.70)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1),[§6\.7](https://arxiv.org/html/2609.12086#S6.SS7.SSS0.Px3.p1.1)\.
- Roberts and DelVecchio \(2000\)B\. W\. Roberts and W\. F\. DelVecchioThe rank\-order consistency of personality traits from childhood to old age: a quantitative review of longitudinal studies\.Psychological Bulletin126\(1\),pp\. 3–25\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.126.1.3)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p5.1),[item Personality](https://arxiv.org/html/2609.12086#S3.I1.ix1.p1.1)\.
- Robertset al\.\(2006\)B\. W\. Roberts, K\. E\. Walton, and W\. ViechtbauerPatterns of mean\-level change in personality traits across the life course: a meta\-analysis of longitudinal studies\.Psychological Bulletin132\(1\),pp\. 1–25\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.132.1.1)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p5.1),[item Personality](https://arxiv.org/html/2609.12086#S3.I1.ix1.p1.1)\.
- Salemiet al\.\(2024\)A\. Salemi, S\. Mysore, M\. Bendersky, and H\. ZamaniLaMP: when large language models meet personalization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7370–7392\.External Links:[Link](https://aclanthology.org/2024.acl-long.399/)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p2.1)\.
- Schwartzet al\.\(2013\)H\. A\. Schwartz, J\. C\. Eichstaedt, M\. L\. Kern, L\. Dziurzynski, S\. M\. Ramones, M\. Agrawal, A\. Shah, M\. Kosinski, D\. Stillwell, M\. E\. P\. Seligman, and L\. H\. UngarPersonality, gender, and age in the language of social media: the open\-vocabulary approach\.PLoS ONE8\(9\),pp\. e73791\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0073791)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1),[§3\.2](https://arxiv.org/html/2609.12086#S3.SS2.p1.1)\.
- Shaikhet al\.\(2025\)O\. Shaikh, S\. Sapkota, S\. Rizvi, E\. Horvitz, J\. S\. Park, D\. Yang, and M\. S\. BernsteinCreating general user models from computer use\.InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology,UIST ’25,New York, NY, USA\.External Links:[Document](https://dx.doi.org/10.1145/3746059.3747722)Cited by:[§2\.5](https://arxiv.org/html/2609.12086#S2.SS5.p1.1),[item 2](https://arxiv.org/html/2609.12086#S4.I2.i2.p1.1)\.
- Shrout and Fleiss \(1979\)P\. E\. Shrout and J\. L\. FleissIntraclass correlations: uses in assessing rater reliability\.Psychological Bulletin86\(2\),pp\. 420–428\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.86.2.420)Cited by:[§7\.7](https://arxiv.org/html/2609.12086#S7.SS7.SSS0.Px4.p1.1)\.
- Soto and John \(2017\)C\. J\. Soto and O\. P\. JohnThe next Big Five Inventory \(BFI\-2\): developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power\.Journal of Personality and Social Psychology113\(1\),pp\. 117–143\.External Links:[Document](https://dx.doi.org/10.1037/pspp0000096)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1),[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px2.p1.1)\.
- Spechtet al\.\(2011\)J\. Specht, B\. Egloff, and S\. C\. SchmukleStability and change of personality across the life course: the impact of age and major life events on mean\-level and rank\-order stability of the Big Five\.Journal of Personality and Social Psychology101\(4\),pp\. 862–882\.External Links:[Document](https://dx.doi.org/10.1037/a0024950)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p5.1),[item Personality](https://arxiv.org/html/2609.12086#S3.I1.ix1.p1.1)\.
- Stein and Swan \(2019\)R\. Stein and A\. B\. SwanEvaluating the validity of Myers\-Briggs Type Indicator theory: a teaching tool and window into intuitive psychology\.Social and Personality Psychology Compass13\(2\),pp\. e12434\.External Links:[Document](https://dx.doi.org/10.1111/spc3.12434)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p1.1),[§4\.3](https://arxiv.org/html/2609.12086#S4.SS3.SSS0.Px2.p1.1)\.
- Sun \(2024\)R\. SunCan a cognitive architecture fundamentally enhance LLMs? or vice versa?\.External Links:2401\.10444,[Document](https://dx.doi.org/10.48550/arXiv.2401.10444),[Link](https://arxiv.org/abs/2401.10444)Cited by:[§2\.10](https://arxiv.org/html/2609.12086#S2.SS10.p1.1)\.
- Tanet al\.\(2025\)J\. Tan, L\. Yang, Z\. Liu, Z\. Liu, R\. Murthy, T\. M\. Awalgaonkar, J\. Zhang, W\. Yao, M\. Zhu, S\. Kokane, S\. Savarese, H\. Wang, C\. Xiong, and S\. HeineckePersonaBench: evaluating AI models on understanding personal information through accessing \(synthetic\) private user data\.External Links:2502\.20616,[Document](https://dx.doi.org/10.48550/arXiv.2502.20616),[Link](https://arxiv.org/abs/2502.20616)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p2.1)\.
- Tan and Jiang \(2023\)Z\. Tan and M\. JiangUser modeling in the era of large language models: current research and future directions\.External Links:2312\.11518,[Document](https://dx.doi.org/10.48550/arXiv.2312.11518),[Link](https://arxiv.org/abs/2312.11518)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1)\.
- Tanet al\.\(2024\)Z\. Tan, Q\. Zeng, Y\. Tian, Z\. Liu, B\. Yin, and M\. JiangDemocratizing large language models via personalized parameter\-efficient fine\-tuning\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 6476–6491\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.372)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Tausczik and Pennebaker \(2010\)Y\. R\. Tausczik and J\. W\. PennebakerThe psychological meaning of words: LIWC and computerized text analysis methods\.Journal of Language and Social Psychology29\(1\),pp\. 24–54\.Note:Published online 2009; print issue 2010External Links:[Document](https://dx.doi.org/10.1177/0261927X09351676)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Teevanet al\.\(2005\)J\. Teevan, S\. T\. Dumais, and E\. HorvitzPersonalizing search via automated analysis of interests and activities\.InProceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’05,New York, NY, USA,pp\. 449–456\.External Links:[Document](https://dx.doi.org/10.1145/1076034.1076111)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12086#S2.SS1.p1.1)\.
- Vazire \(2010\)S\. VazireWho knows what about a person? The self\-other knowledge asymmetry \(SOKA\) model\.Journal of Personality and Social Psychology98\(2\),pp\. 281–300\.External Links:[Document](https://dx.doi.org/10.1037/a0017908)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p3.1),[§6\.3](https://arxiv.org/html/2609.12086#S6.SS3.p2.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px4.p1.1)\.
- Wanget al\.\(2023\)A\. Wang, C\. Aggazzotti, R\. Kotula, R\. Rivera Soto, M\. Bishop, and N\. AndrewsCan authorship representation learning capture stylistic features?\.Transactions of the Association for Computational Linguistics11,pp\. 1416–1431\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00610)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1),[§6\.7](https://arxiv.org/html/2609.12086#S6.SS7.SSS0.Px3.p1.1)\.
- Wanget al\.\(2025\)A\. Wang, J\. Morgenstern, and J\. P\. DickersonLarge language models that replace human participants can harmfully misportray and flatten identity groups\.Nature Machine Intelligence7\(3\),pp\. 400–411\.External Links:[Document](https://dx.doi.org/10.1038/s42256-025-00986-z)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p1.1),[§6\.1](https://arxiv.org/html/2609.12086#S6.SS1.p2.1),[§9\.5](https://arxiv.org/html/2609.12086#S9.SS5.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024a\)P\. Wang, L\. Li, L\. Chen, Z\. Cai, D\. Zhu, B\. Lin, Y\. Cao, L\. Kong, Q\. Liu, T\. Liu, and Z\. SuiLarge language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9440–9450\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.511)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p2.1),[§6\.7](https://arxiv.org/html/2609.12086#S6.SS7.SSS0.Px2.p1.1)\.
- Wanget al\.\(2024b\)T\. Wang, M\. Tao, R\. Fang, H\. Wang, S\. Wang, Y\. E\. Jiang, and W\. ZhouAI PERSONA: towards life\-long personalization of LLMs\.External Links:2412\.13103,[Document](https://dx.doi.org/10.48550/arXiv.2412.13103),[Link](https://arxiv.org/abs/2412.13103)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Wegmannet al\.\(2022\)A\. Wegmann, M\. Schraagen, and D\. NguyenSame author or just same topic? Towards content\-independent style representations\.InProceedings of the 7th Workshop on Representation Learning for NLP,pp\. 249–268\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.repl4nlp-1.26)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p3.1),[§6\.7](https://arxiv.org/html/2609.12086#S6.SS7.SSS0.Px3.p1.1)\.
- Xueet al\.\(2016\)B\. Xue, M\. Zhang, W\. N\. Browne, and X\. YaoA survey on evolutionary computation approaches to feature selection\.IEEE Transactions on Evolutionary Computation20\(4\),pp\. 606–626\.External Links:[Document](https://dx.doi.org/10.1109/TEVC.2015.2504420)Cited by:[§2\.8](https://arxiv.org/html/2609.12086#S2.SS8.p2.1)\.
- Yarkoni \(2010\)T\. YarkoniPersonality in 100,000 words: a large\-scale analysis of personality and word use among bloggers\.Journal of Research in Personality44\(3\),pp\. 363–373\.External Links:[Document](https://dx.doi.org/10.1016/j.jrp.2010.04.001)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1),[§3\.2](https://arxiv.org/html/2609.12086#S3.SS2.p1.1)\.
- Youyouet al\.\(2015\)W\. Youyou, M\. Kosinski, and D\. StillwellComputer\-based personality judgments are more accurate than those made by humans\.Proceedings of the National Academy of Sciences112\(4\),pp\. 1036–1040\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1418680112)Cited by:[§2\.4](https://arxiv.org/html/2609.12086#S2.SS4.p2.1)\.
- Yuanet al\.\(2022\)A\. Yuan, A\. Coenen, E\. Reif, and D\. IppolitoWordcraft: story writing with large language models\.In27th International Conference on Intelligent User Interfaces,pp\. 841–852\.External Links:[Document](https://dx.doi.org/10.1145/3490099.3511105)Cited by:[§2\.7](https://arxiv.org/html/2609.12086#S2.SS7.p1.1)\.
- Zhanget al\.\(2024a\)K\. Zhang, Y\. Kang, F\. Zhao, and X\. LiuLLM\-based medical assistant personalization with short\- and long\-term memory coordination\.External Links:2309\.11696,[Document](https://dx.doi.org/10.48550/arXiv.2309.11696),[Link](https://arxiv.org/abs/2309.11696)Cited by:[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Zhanget al\.\(2025\)Z\. Zhang, R\. A\. Rossi, B\. Kveton, Y\. Shao, D\. Yang, H\. Zamani, F\. Dernoncourt, J\. Barrow, T\. Yu, S\. Kim, R\. Zhang, J\. Gu, T\. Derr, H\. Chen, J\. Wu, X\. Chen, Z\. Wang, S\. Mitra, N\. Lipka, N\. Ahmed, and Y\. WangPersonalization of large language models: a survey\.Transactions on Machine Learning Research\.Note:Preprint available as arXiv:2411\.00027External Links:[Document](https://dx.doi.org/10.48550/arXiv.2411.00027)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.
- Zhanget al\.\(2024b\)Z\. Zhang, M\. Jia, H\. Lee, B\. Yao, S\. Das, A\. Lerner, D\. Wang, and T\. Li“It’s a fair game”, or is it? Examining how users navigate disclosure risks and benefits when using LLM\-based conversational agents\.InProceedings of the CHI Conference on Human Factors in Computing Systems,pp\. 1–26\.External Links:[Document](https://dx.doi.org/10.1145/3613904.3642385)Cited by:[§10](https://arxiv.org/html/2609.12086#S10.SS0.SSS0.Px2.p1.1),[§2\.6](https://arxiv.org/html/2609.12086#S2.SS6.p2.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-Bench and chatbot arena\.InAdvances in Neural Information Processing Systems 36,Vol\.36,pp\. 46595–46623\.External Links:[Document](https://dx.doi.org/10.52202/075280-2020)Cited by:[§2\.9](https://arxiv.org/html/2609.12086#S2.SS9.p2.1)\.
- Zhouet al\.\(2024\)Y\. Zhou, Q\. Zhu, J\. Jin, and Z\. DouCognitive personalized search integrating large language models with an efficient memory mechanism\.InProceedings of the ACM Web Conference 2024,WWW ’24,New York, NY, USA,pp\. 1464–1473\.External Links:[Document](https://dx.doi.org/10.1145/3589334.3645482)Cited by:[§1](https://arxiv.org/html/2609.12086#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.12086#S2.SS2.p1.1)\.

## Appendix AThe Full Atomic User Model Specification

Table[19](https://arxiv.org/html/2609.12086#A1.T19)gives the full specification\. Fields marked†\\daggerform the reduced 32\-field instrument used in both studies and listed in Table[2](https://arxiv.org/html/2609.12086#S4.T2); they are the only fields for which this paper reports evidence\. The remainder are part of the specification but were not exercised, and we flag that distinction rather than presenting the whole schema as validated\.

Table 19:The fullaumspecification, 60 fields across five tiers plus cross\-shell integration\.†\\daggermarks the 32 fields in the reduced instrument used in both studies\.τ\\tauis the nominal update period for the tier\.Four properties of the schema are worth restating in one place\.

1. 1\.Every field value is a short first\-person natural\-language sentence, not a score or a vector\.
2. 2\.The\(shell,name\)\(\\mathrm\{shell\},\\mathrm\{name\}\)pairs are fixed, so two users’ models are structurally comparable andσ⁡\(t\)\\sigma\(t\)is definable\.
3. 3\.Update periodτ\\tauis a tier property, so a refresh policy can be written without consulting field values\.
4. 4\.The CrossShell tier is not a sixth shell but a set of relations over the other five\.*Internal Conflicts*and*Contradiction Log*record where a Shell 3 behaviour contradicts a Nucleus value;*Authenticity Index*and*Shell Alignment Map*measure agreement between what the inner shells hold and what the outer shells present;*Confidence Annotation*and*Provenance Record*track how a field was populated and how much weight it should carry\.

## Appendix BPrompts

Every prompt used in the simulation study is reproduced here verbatim\. Placeholders in braces are substituted at run time\.

### B\.1Persona conditioning block

Prepended as the system message for every generation the simulated participant performs\.

> You are role\-playing a specific person\. Stay in character at all times\. Personality \(Big Five z\-scores, \-2 low to \+2 high\): \{traits\} Life context: \{context\} Formative background: \{narrative\} How you actually write: \{idiolect\} Write the way this person writes, including its rough edges\. Do not describe the persona; be it\.

### B\.2Instrument population

> You are completing a guided self\-report instrument called the Atomic User Model\. Fill in EVERY field below with one short first\-person sentence \(roughly 10 to 25 words\) describing yourself truthfully in character\. Be specific and concrete; avoid generic self\-description that could apply to anyone\. The JSON MUST contain all six shells \(Nucleus, Shell1\_Psychological, Shell2\_Cognitive, Shell3\_Behavioural, Shell4\_Social, CrossShell\) and all 32 fields as non\-empty strings\. Do not leave any shell empty\. Do not nest extra objects inside a field; each field value is a single sentence string\. Return ONLY a valid JSON object with exactly this structure and these keys: \{schema\} No commentary, no markdown fences\.

On a failed parse the following corrective is appended and the call retried, up to six attempts:

> Your previous JSON filled too few fields\. Fill ALL 32 fields in ALL six shells\. Every value must be a non\-empty first\-person sentence\. Empty shells are not acceptable\.

### B\.3Prompt surface

The only user\-derived text the retriever ever sees\. Generated with the persona block plus the participant’s completed instrument as system context\.

> You are typing a request to an AI assistant\. Write only the message you would type, in your own voice, with your own level of detail and formality\. Do not write the assistant’s answer\. What you want help with: \{task\} Keep it under 60 words\. Output the message only\.

### B\.4Reference text

The ground truth\. Generated without sight of the prompt surface\.

> Write your own response to the task below, exactly as you would actually write it yourself, with no assistant helping you\. Match your usual length, structure, formality and habits\. Task: \{task\} Output the text only, with no preamble\.

### B\.5Prior conversations for the preference\-notes baseline

Four transcripts per participant, on topics held out from the evaluation battery\.

> Write a short transcript of a past conversation you had with an AI assistant, four to six turns, labelled ’User:’ and ’Assistant:’\. The topic was: \{topic\}\. Write your own turns in your own voice\. Let your preferences and habits show through naturally rather than stating them outright\.

The six held\-out topics are: choosing between two laptops for everyday work; sorting out a recurring problem with a household appliance; putting together a week of simple meals; deciding what to do with a free Saturday; picking a book to read next; working out how to reply to a slightly awkward group message\.

### B\.6Preference\-note distillation

Run by a separate summariser with no access to the persona or theaum\.

> System:You are the memory component of an AI assistant\. You read past conversations and store durable user preferences as a flat list of short statements, the way a production memory feature does\. You have no other information about the user\. User:From the transcripts below, extract the durable preferences worth remembering about this user\. Write at most 12 short statements, one per line, each starting with ’\- ’\. Record preferences and habits only, not events\. No headings, no commentary\. \{transcripts\}

### B\.7Task classifier

> System:You route requests for a personalised assistant\. A request is ’simple’ if a correct answer is the same regardless of who asked \(a factual lookup, a unit conversion, an arithmetic result\)\. It is ’complex’ if the form of an acceptable answer depends on who is asking\. Answer with JSON only\. User:Classify this user request\. Request: \{prompt\_surface\} Return JSON: \{"complexity": "simple" or "complex", "task\_type": one of \[communication, self\_presentation, relational\_evaluative, planning\_decision, explanatory, creative, factual\]\}

### B\.8Generation

> System:You are a helpful AI assistant\. Complete the user’s task and return only the finished text they asked for, with no preamble, no explanation of your choices, and no options\. User:\{context\_block\}User request: \{prompt\_surface\}

The context block is empty for the generic condition, and otherwise a delimited section:\[Selected user model fields\]followed by one bulleted field per line and a closing tag, or\[Remembered user preferences\]for the preference\-notes condition, or\[User model\]followed by the full instrument as JSON for the Fullaumcondition\. Context tokens are counted on exactly this string\.

### B\.9Fidelity judge

> System:You are judging, on behalf of a specific person, whether a piece of text reads like something that person would have written\. You are given that person’s self\-description and one text they genuinely wrote\. Judge style, voice, length, structure and habits, not whether the text is good\. Answer with a single integer and nothing else\. User:The person: \- Life context: \{context\} \- How they write: \{idiolect\} A text this person genuinely wrote for the task "\{task\}": \-\-\- \{reference\} \-\-\- Now rate this other text for the same task on a 1 to 5 scale, where 5 means ’this reads exactly like something they would have written’ and 1 means ’this does not read like them at all’\. \-\-\- \{candidate\} \-\-\- Return ONLY a single integer from 1 to 5\.

The judge never receives theaum\. Only the persona’s life context and idiolect line are shown, plus the participant’s own reference\.

### B\.10Identification judge

> System:You are shown one text a specific person genuinely wrote, and four candidate texts for the same task\. Exactly one candidate was produced by a system that had access to a model of that person\. Choose the candidate whose style most closely matches the person’s own writing\. Answer with a single letter\. User:The person: \- Life context: \{context\} \- How they write: \{idiolect\} A text this person genuinely wrote for the task "\{task\}": \-\-\- \{reference\} \-\-\- Candidates: \{A, B, C, D in a per\-trial shuffled order\} Return ONLY the letter of the candidate that best matches this person’s own style\.

## Appendix CThe Simulated Participants

The sixteen personas are declared in code rather than sampled, so the population is fixed across seeds and auditable\. Each carries Big Fivezz\-scores, a life context, a formative\-experience seed and idiolect markers\. Table[4](https://arxiv.org/html/2609.12086#S6.T4)in the body gives the abridged form; Table[20](https://arxiv.org/html/2609.12086#A3.T20)gives the life context and formative background in full\.

Table 20:Life context and formative background for each simulated participant\.LabelLife contextFormative backgroundmeticulous plannerOperations lead at a mid\-size logistics firm; runs everything from checklists\.A shipment error early in their career cost the company a client; they have double\-checked everything since\.warm over\-explainerSecondary school teacher; the person colleagues go to when something needs smoothing over\.Grew up as the eldest of four and became the household mediator before they were twelve\.blunt minimalistBackend engineer; treats most meetings as an unpriced tax on the day\.Learned to code alone at fourteen from documentation and has never quite unlearned the register\.associative creativeFreelance illustrator; several projects always half\-open at once\.Dropped out of an economics degree to go to art school and treats that as the decision that defines them\.analytical flatQuantitative analyst; distrusts any claim without an interval attached\.A badly designed undergraduate experiment taught them how easy it is to fool yourself with data\.bright energeticCommunity manager for a hobbyist platform; genuinely likes the people they work with\.Found their footing running a fan forum as a teenager and never lost the habit of talking to a crowd\.dry scepticInvestigative reporter turned freelance editor; assumes the press release is lying\.Watched a story they had championed fall apart under fact\-checking and now leads with the caveat\.formal correctCivil servant in a planning department; twenty years of writing that must survive scrutiny\.Trained in an office where a misplaced comma in a notice could invalidate it\.scattered fastStartup founder wearing four hats; replies from three time zones in one week\.Left a stable job on two weeks’ notice and has been running on that momentum for three years\.nurturing relationalPaediatric nurse; the emotional temperature of a room is the first thing they read\.A long stretch caring for a grandparent taught them that people mostly want to be heard first\.competitive driverSales director; keeps a personal scoreboard nobody asked to see\.Came up through a commission\-only role and still measures the week in closed deals\.reflective introvertArchivist; the working day is mostly quiet and they prefer it that way\.Spent a year abroad where they spoke the language badly and learned to say less, better\.storytelling extrovertCorporate trainer; cannot make a point without an anecdote attached\.Grew up in a family where dinner was competitive storytelling and brevity lost\.pragmatic efficientSite foreman; the day is measured in what got finished\.Twenty years on sites where the person who over\-explained held everyone else up\.quirky metaphoricalResearch librarian with a sideline in cryptic crosswords\.Was the child who read the encyclopaedia for fun and has never found a better hobby\.stoic calmEmergency dispatcher; twelve\-hour shifts where panic is the one useless response\.Learned early that the calmest voice on the line is the one that gets things done\.
## Appendix DAdditional Statistical Detail

### D\.1Judge reliability across seeds

Table 21:Agreement between the fidelity judge’s ratings across seeds, over 960 \(condition, participant, task\) targets\. ICC\(2,1\) over all three seeds is 0\.573\.
### D\.2Use of the rating scale, and position of the chosen option

Table 22:Left: distribution of fidelity ratings over all 2,880 trials\. Right: share of identification trials in which the judge chose the option in each slot, against the uniform 25% expectation, with Wilson intervals\. A chi\-square test over slots givesχ2=7\.37\\chi^\{2\}=7\.37, 3 d\.f\.,p=0\.061p=0\.061\.
### D\.3Payload composition by task type

Table 23:Share of the retrieved payload drawn from each shell, by task type, over allaum\-rtrials\. Rows sum to 100% up to rounding\. This is the tabular form of Figure[11](https://arxiv.org/html/2609.12086#S7.F11)\.
### D\.4Optimality of each retriever where exhaustive search is tractable

Table 24:Share of instances in which each method returned the exhaustive optimum, over the twelve cells of the scaling sweep where\(nk\)≤60,000\\binom\{n\}\{k\}\\leq 60\{,\}000and exhaustive search was therefore computed\. 25 instances per cell\.

Similar Articles

How Well Do Large Language Models Capture Human Personality?

arXiv cs.AI

This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv cs.AI

This paper introduces LUNAR, a benchmark for evaluating how large language models personalize responses from longitudinal app interaction histories across daily-life domains such as clothing, food, housing, and mobility. Experiments on 19 mainstream LLMs reveal that effective personalization depends on evidence selection and cross-domain integration, and that stronger personalization can come at the cost of privacy protection.

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

arXiv cs.CL

This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.