Tag
WebRider is a hierarchical framework that formalizes delegated web tasks as intent contracts, preserving persona-conditioned policies through every browsing step. It includes RiderBench, a benchmark of 4,096 live-web contracts, and an 8B action-policy model trained through its guarded interface.
Presents FOCUS, a fine-tuning method that decouples expert personas in LLMs via orthogonal decomposition and an expert gating module, improving domain-specific task accuracy across financial, legal, and medical benchmarks.
This paper introduces MemoryForge, a framework for synthesizing lifelong autobiographical memory from brief target personas to enable frozen LLMs to exhibit more human-like behaviors in role-play and user-simulation, outperforming descriptive conditioning baselines.
This paper investigates how the narrative framing of a task (e.g., disease investigation vs. murder mystery) acts as a stronger driver of LLM agent behavior than assigned personas, introducing the concept of 'narrative priors' that explain 5–31x more behavioral variance and are negatively associated with task success in two of three domains.
This paper shows that in chat models, refusal behavior is gated by a compliant model persona direction at late layers, rather than being an isolated mechanism. Steering persona suppresses refusal, and reintroducing refusal partially restores it only at late layers, revealing a coupling between persona and safety representations.
Proposes OpenAgent, a spec for defining AI agent identity (face, voice, writing style) in a single signed YAML file, enabling portability across different harnesses.
Anthropic is updating its privacy policy to require identity verification for certain capabilities starting July 8, 2026, using third-party vendor Persona, who previously had a data exposure incident with Discord.
Recommending the open-source repository awesome-human-distillation which organizes human distillation into reusable AI Skills, containing various persona skill packages such as Feng Ge, Zhang Xuefeng, and other typical figures.
This paper investigates whether role-playing in LLMs changes only outputs or also internal truth representations, using linear probes. It finds that roleplay shifts outputs more than internal beliefs, while emergent misalignment causes larger shifts in internal representations.
This paper investigates how instruction-tuned LLMs combine persona and task specifications in the residual stream, finding that near answer formation the combination is approximately additive, enabling substitution with minimal KL divergence, but this additive regime does not account for the full multi-token generation mechanism.
A study demonstrates that simply changing the formatting (prose vs bullet points) of a persona prompt dramatically flips an LLM's behavior in a Prisoner's Dilemma, from 96% cooperation to 20%, illustrating extreme sensitivity to format despite identical content (p < 0.001).