CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation

Hugging Face Daily Papers Papers

Summary

CARD proposes a hierarchical framework for personalized text generation that clusters users and uses reward-guided decoding, demonstrating improved quality and efficiency on LaMP benchmarks.

Adapting large language models to individual users remains challenging due to the tension between fine-grained personalization and scalable deployment. We present CARD, a hierarchical framework that achieves effective personalization through progressive refinement. CARD first clusters users according to shared stylistic patterns and learns group-specific LoRA adapters, enabling robust generalization and strong low-resource performance. To capture individual differences within each cluster, we propose an implicit preference learning mechanism that contrasts user-authored text with cluster-level generations, allowing the model to infer user-specific style preferences without manual annotation. At inference time, CARD injects personalization exclusively at decoding via lightweight user preference vectors and low-rank logit corrections, while keeping the base model frozen. Experiments on the LaMP and LongLaMP benchmarks show that CARD achieves superior generation quality compared to baselines, while significantly improving efficiency and scalability for practical personalized text generation.
Original Article
View Cached Full Text

Cached at: 09/28/26, 08:03 AM

Paper page - CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation

Source: https://huggingface.co/papers/2601.06352 Published on Sep 20

·

Submitted byhttps://huggingface.co/hulehule

JWon Sep 28

Abstract

Adaptinglargelanguagemodelstoindividualusersremainschallengingduetothetensionbetweenfine-grainedpersonalizationandscalabledeployment.WepresentCARD,ahierarchicalframeworkthatachieveseffectivepersonalizationthroughprogressiverefinement.CARDfirstclustersusersaccordingtosharedstylisticpatternsandlearnsgroup-specificLoRAadapters,enablingrobustgeneralizationandstronglow-resourceperformance.Tocaptureindividualdifferenceswithineachcluster,weproposeanimplicitpreferencelearningmechanismthatcontrastsuser-authoredtextwithcluster-levelgenerations,allowingthemodeltoinferuser-specificstylepreferenceswithoutmanualannotation.Atinferencetime,CARDinjectspersonalizationexclusivelyatdecodingvialightweightuserpreferencevectorsandlow-ranklogitcorrections,whilekeepingthebasemodelfrozen.ExperimentsontheLaMPandLongLaMPbenchmarksshowthatCARDachievessuperiorgenerationqualitycomparedtobaselines,whilesignificantlyimprovingefficiencyandscalabilityforpracticalpersonalizedtextgeneration.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2601\.06352

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2601.06352 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2601.06352 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2601.06352 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Improving Text-to-Music Generation with Human Preference Rewards

Hugging Face Daily Papers

This paper presents a text-to-music generation system that leverages reward conditioning, expert iteration, and preference tuning to improve audio quality within a 120M-parameter model, submitted to the ATTM Grand Challenge at ICME 2026.

Hierarchical text-conditional image generation with CLIP latents

OpenAI Blog

OpenAI proposes a hierarchical two-stage model for text-conditional image generation using CLIP latents: a prior that generates CLIP image embeddings from text captions, and a diffusion-based decoder that generates images from embeddings. The approach improves image diversity and enables zero-shot language-guided image manipulations.