Spotify对话推荐代理引导:合成数据生成与自我改进循环

arXiv cs.CL 论文

摘要

本文提出了一种用于Spotify对话推荐代理的合成数据生成管道和自我改进循环,显著提升了用户参与度和性能。

arXiv:2609.30297v1 Announce Type: new Abstract: Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address this challenge. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn conversations, enabling systematic evaluation before launch. The self-improvement loop combines variance-based contrastive optimization with iterative refinement through a coding agent, automatically identifying and fixing planning and tool-use errors. Our approach improves quality by +8% on top of a highly optimized manual prompt. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience supporting only session refinement. This work provides a practical framework for accelerating the development of conversational recommendation agents in industry.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:35

# Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
Source: [https://arxiv.org/html/2609.30297](https://arxiv.org/html/2609.30297)
Conference:20th ACM Conference on Recommender Systems; September 27\-October 02, 2026; Minneapolis, MN, USA20th ACM Conference on Recommender Systems \(RecSys ’26\), September 27\-October 02, 2026, Minneapolis, MN, USADOI:[10\.1145/3773078\.3831910](https://doi.org/10.1145/3773078.3831910)ISBN:979\-8\-4007\-2284\-4/2026/09Enrico Palumbo1, Alexandre Tamborrino2, Victor Ode3, Ben Lacker3, Adrià Casas Escoda3, Jeremy Hopple3, Marcus Better3, James Leoni4, Hugo Galvão3, Hugues Bouchard5, Mounia Lalmas4, José Luis Redondo García5, Abenezer Abebe3, Ann Clifton3, Anton Blomberg6, Henrik Lindström6, Dani Doro6, Christine Doig Cardet5

© cc

###### Abstract\.

Conversational recommendation agents are emerging as a new paradigm for content discovery, enabling users to express complex intents through natural language \(e\.g\.,*“recommend Italian indie artists I haven’t heard before” or “explain why they fit my taste”*\)\. A central challenge in building such agents is optimizing agent planning, i\.e\., deciding how to select, sequence, and invoke tools\. This challenge is particularly acute in cold\-start settings, where real user interactions are not yet available\.

We introduce a pipeline for multi\-turn synthetic data generation and a self\-improvement loop to address the lack of interaction data and the difficulty of optimizing agent planning in cold\-start settings\. The synthetic data pipeline transforms single\-turn prompts into realistic multi\-turn user–agent conversations, enabling systematic evaluation of conversational capabilities before launch\. The self\-improvement loop then uses this data and evaluation feedback to combine variance\-based contrastive optimization with iterative refinement through a coding agent, allowing the system to automatically identify and fix planning and tool\-use errors\.

Our approach provides fine\-grained insights into conversational capabilities, uncovers issues before deployment, and improves quality by \+8% on top of a highly optimized manual prompt, automatically resolving several planning and tool\-use errors\. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify\.111[https://newsroom\.spotify\.com/2026\-04\-20/deeper\-music\-features\-about\-the\-song\-dj\-songdna/](https://newsroom.spotify.com/2026-04-20/deeper-music-features-about-the-song-dj-songdna/)\. Online A/B tests demonstrate its effectiveness, with \+14% user listening, \+5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience that only supports session refinement\.

Overall, this work provides a practical framework for accelerating the development of conversational recommendation agents in industry, addressing challenges that are becoming increasingly central to recommender systems and agentic applications as natural\-language interfaces reach widespread adoption\.

###### Keywords:

Conversational AI, Agents, Multi\-turn Dialog, Synthetic Data, Self\-Improvement, Prompt Engineering

††cc\-license:by\-nc\-nd## 1\.Introduction

![Refer to caption](https://arxiv.org/html/2609.30297v1/example_agent_paper.png)Figure 1\.Conversational Recommendation Agents enable users to actively ask for recommendations and continue the dialog using natural languageRecommender Systems are traditionally defined as functions that, given a user profile, retrieve items matching the user’s preferences\. This paradigm is often contrasted with search systems, which instead rely on explicit textual queries to retrieve relevant items\([Belkin and Croft, 1992](https://arxiv.org/html/2609.30297#bib.bib26)\)\. This distinction is increasingly blurring with the emergence of*conversational recommender systems*, which jointly model user preferences and natural language input, enabling more expressive and controllable personalization experiences driven by complex, multi\-faceted user intents\([Jannach et al\., 2021](https://arxiv.org/html/2609.30297#bib.bib24)\)\.

A key advantage of conversational recommenders is their multi\-turn nature, which allows users to iteratively refine, adjust, or shift their preferences through follow\-up interactions\. Moreover, when powered by large language models \(LLMs\) with tool\-use capabilities, conversational recommendation agents can go beyond item recommendation to support richer interactions within the same experience, such as answering questions about content \(*“who is this artist?”*\) or providing contextual explanations \(*“which artists influenced my favorite artist, and why?”*\)\.

Realizing this potential, however, requires overcoming key challenges\. While traditional recommender systems benefit from abundant user interaction data, large\-scale multi\-turn conversational data is typically unavailable, especially prior to product launch\. At Spotify, recent systems such as our agentic search architecture\([Palumbo et al\., 2025a](https://arxiv.org/html/2609.30297#bib.bib25)\)and the DJ222[https://newsroom\.spotify\.com/2025\-05\-13/dj\-voice\-requests/](https://newsroom.spotify.com/2025-05-13/dj-voice-requests/)experience support a range of user requests, but remain largely limited to single\-turn interactions and simple session refinement requests\.

In this paper, we address the problem of bootstrapping a conversational recommendation agent in a cold\-start setting, where both multi\-turn data and reliable evaluation signals are scarce\. In such settings, the lack of realistic conversational data makes it difficult to systematically evaluate multi\-turn capabilities, while the stochastic nature of agent planning makes failures difficult to diagnose and correct issues\. We therefore focus on two key challenges: \(1\) generating realistic multi\-turn user–agent conversations from existing single\-turn data, to enable pre\-launch evaluation, and \(2\) improving agent planning under stochastic and data\-scarce conditions to ensure reliable and consistent tool use\.

To address these challenges, we first define a taxonomy of multi\-turn conversational capabilities, grounded in prior work\([Deshpande et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib14);[Bai et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib23)\)and product requirements, together with a pipeline for generating synthetic user–agent conversations\. We then develop a self\-improvement loop based on a coding agent that uses evaluation variance to distinguish inconsistent planning behaviors from capability gaps\. The loop combines contrastive optimization, i\.e\., comparing successful and unsuccessful agent plans to identify reliable planning patterns, with iterative refinement for harder failure cases\.

Our approach produces high\-quality multi\-turn conversations, with human annotation scores above90%90\\%across all quality dimensions\. This data enables fine\-grained evaluation of conversational capabilities, revealing that performance varies across capabilities and degrades as conversations grow beyond four turns\. The self\-improvement loop further improves agent performance by \+8% on top of a highly optimized prompt, while uncovering and resolving subtle planning and tool\-use issues\. Together, these components significantly accelerate iteration on agent planning and enable reliable development in cold\-start settings\.

These methods have been instrumental in the launch of a conversational recommendation agent at Spotify\. Online A/B tests demonstrate significant improvements in user engagement, which led to its subsequent production deployment \(Figure[1](https://arxiv.org/html/2609.30297#S1.F1)\)\. More broadly, our results suggest that synthetic\-data\-driven evaluation combined with automated self\-improvement provides a practical and generalizable approach for developing conversational recommendation agents in data\-scarce settings\.

## 2\.Related Work

We review prior work in conversational recommendation, conversational data generation, and agent self\-improvement\.

##### Conversational Recommendation

Conversational recommendation is a form of conversational information seeking\([Zamani et al\., 2023](https://arxiv.org/html/2609.30297#bib.bib1)\)defined as*“a software system that supports users in achieving recommendation\-related goals through a multi\-turn dialogue”*\([Jannach et al\., 2021](https://arxiv.org/html/2609.30297#bib.bib24)\)\. The vision dates back to early work that framed recommendation as an interactive preference\-elicitation process driven by question\-answer exchanges\([Christakopoulou et al\., 2016](https://arxiv.org/html/2609.30297#bib.bib22)\)\. Beyond eliciting user needs, conversational recommenders are expected to recommend items, explain those recommendations, and answer follow\-up questions\([Jannach et al\., 2021](https://arxiv.org/html/2609.30297#bib.bib24)\)\.

Historically, conversational recommender systems have been built as pipelines of specialized components \(e\.g\., query understanding, item retrieval and ranking, dialogue management, and response generation\) coordinated through hand\-crafted rule\-based logic\([Jannach et al\., 2021](https://arxiv.org/html/2609.30297#bib.bib24)\)\. With the advent of LLMs and their general\-purpose reasoning abilities,*end\-to\-end*approaches have emerged along two complementary directions\.*Generative retrieval*approaches frame recommendation as the direct generation of item identifiers conditioned on conversational context\([Tay et al\., 2022](https://arxiv.org/html/2609.30297#bib.bib11);[Rajput et al\., 2023](https://arxiv.org/html/2609.30297#bib.bib12);[Palumbo et al\., 2025b](https://arxiv.org/html/2609.30297#bib.bib2);[He et al\., 2026](https://arxiv.org/html/2609.30297#bib.bib8)\)\.*Agentic*approaches instead allow LLMs to plan and call existing recommendation services as tools, leveraging production\-grade retrieval and ranking systems while keeping conversational reasoning within the LLM\([Palumbo et al\., 2025a](https://arxiv.org/html/2609.30297#bib.bib25);[Doh et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib3)\)\. Our work falls into the latter category and focuses on bootstrapping such agents in cold\-start settings\.

##### Conversational Data Generation

Evaluating recommendation agents only in the single\-turn setting is insufficient\. Recent studies show that even frontier LLMs degrade significantly as conversations grow, getting “lost” across turns and failing to maintain consistent grounding and instruction following\([Laban et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib27)\)\. Furthermore, key capabilities of conversational agents, such as refining a music session or asking follow\-up clarifications, only arise in a multi\-turn setting\. This motivates the need for realistic multi\-turn data that captures tool use, contextual reasoning, and intent shifts\.

For general\-purpose agents, recent benchmarks rely heavily on synthetic or semi\-synthetic multi\-turn data: MT\-Bench\-101\([Bai et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib23)\)and MultiChallenge\([Deshpande et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib14)\)probe fine\-grained conversational capabilities of frontier LLMs, whileτ\\tau\-bench\([Yao et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib16)\)simulates tool\-agent\-user interactions and evaluates how agents interact with simulated users\. In the more specific setting of conversational music and playlist recommendation, several datasets have been proposed\([Chaganty et al\., 2023](https://arxiv.org/html/2609.30297#bib.bib13);[Leszczynski et al\., 2023](https://arxiv.org/html/2609.30297#bib.bib4);[Choi et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib5)\), but are limited to recommendation as the sole intent\. In contrast, our pipeline seeds multi\-turn conversations from existing single\-turn intents and models a broader range of conversational behaviors, including refinement, follow\-up questions, and intent shifts\.

##### Agent Self\-Improvement

A growing body of work treats agent prompts and configurations as parameters to be optimized rather than hand\-tuned\. These approaches aim to improve agent behavior by automatically updating prompts or surrounding components based on feedback\. For example, methods such as*TextGrad*\([Yuksekgonul et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib19)\)propagate natural\-language “gradients” through LLM calls,*GEPA*\([Agrawal et al\., 2026b](https://arxiv.org/html/2609.30297#bib.bib17)\)evolves prompts through reflective optimization, and*ACE*\([Zhang et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib18)\)frames self\-improvement as context engineering, modifying the input context rather than model weights\. Beyond prompts, recent work also optimizes the broader agent*harness*, including tool routing, control flow, and decoding parameters\. For instance,*Meta\-Harness*\([Lee et al\., 2026](https://arxiv.org/html/2609.30297#bib.bib20)\)performs end\-to\-end optimization of agent pipelines, while*optimize\_anything*\([Agrawal et al\., 2026a](https://arxiv.org/html/2609.30297#bib.bib21)\)provides a general framework for optimizing textual components of LLM systems\.

Our self\-improvement loop is complementary to these approaches\. Rather than proposing a new optimization method, we focus on leveraging*data and evaluation signals*in a production setting, where feedback is often noisy due to infrastructure issues and inherent LLM variability\. Concretely, we use synthetic multi\-turn conversations to surface failure modes, analyze variance across sampled plans to identify their root causes, and use a coding agent to refine prompts and tool definitions when necessary\.

## 3\.Approach

We present our approach to multi\-turn conversation generation and agent self\-improvement, which enable evaluation and optimization of conversational recommendation agents in cold\-start scenarios\.

### 3\.1\.Multi\-turn Conversation Generation

Single\-turn evaluation captures whether the agent can resolve a user prompt, but it misses failure modes that emerge as conversations extend across multiple turns\. For instance, an agent may retrieve the right type of music for a given request, yet fail to refine it when the user updates their requirements\. In a cold\-start scenario, real multi\-turn traffic with sufficient scale and diversity is not available\. We therefore introduce a synthetic generation pipeline that bootstraps a multi\-turn dataset from existing single\-turn interactions \(Figure[2](https://arxiv.org/html/2609.30297#S3.F2)\)\.

Figure 2\.Multi\-turn Synthetic Data Generation Pipeline\. a\) We generalize single\-turn prompts into multi\-turn synthetic user\-agent dialogs\. b\) The dialogs are then evaluated through LLM\-as\-a\-judge with an instance\-level rubric\.As illustrated in Figure[2](https://arxiv.org/html/2609.30297#S3.F2)a, the key design choice behind the pipeline is to separate what the conversation should achieve from how the agent executes it\. We first define a*dialog plan*, i\.e\., an abstract sequence of conversational steps designed to exercise a target capability, and then*realize*this plan by interacting with the live agent through a UserLLM\.

Compared to directly generating user–agent conversations in a single step, this design offers several advantages\. First, it provides greater control over the dataset and facilitates versioning, since dialog plans can be audited, stored, and reused across evaluations\. Second, it naturally incorporates manually curated multi\-turn test cases, which can be represented as dialog plans and adapted to the agent’s responses\. Third, it extends naturally to real multi\-turn conversations once they become available, which can be logged and reused as dialog plans, with the UserLLM adapting them to the current agent version\.

The pipeline consists of five components: taxonomy, single\-turn data, dialog plan generation, conversation realization, LLM\-as\-a\-judge\.

#### 3\.1\.1\.Taxonomy

Multi\-turn conversations can fail in many different ways\. To make these failure modes measurable, we define a discrete set of conversational capabilities to evaluate, rather than relying on a vague notion of “multi\-turn quality\.” This allows both data generation and evaluation to target specific conversational capabilities explicitly\.

##### Approach\.

Drawing on prior work\([Bai et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib23);[Deshpande et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib14)\), as well as user testing and product priorities, we define a taxonomy𝒯=\{t1,…,t5\}\\mathcal\{T\}=\\\{t\_\{1\},\\ldots,t\_\{5\}\\\}of five multi\-turn capabilities\. Each capabilityti=\(namei,rubrici,ℰi\)t\_\{i\}=\(\\textsc\{name\}\_\{i\},\\textsc\{rubric\}\_\{i\},\\mathcal\{E\}\_\{i\}\)consists of a name, a natural\-language rubric used by the judge for instance\-level evaluation, and a small set of few\-shot examplesℰi\\mathcal\{E\}\_\{i\}used to guide dialog plan generation\. The five capabilities are described in Table[1](https://arxiv.org/html/2609.30297#S3.T1)\.

Task & RubricExample dialog planContent Refinement \(CR\)
Refine the response based on the user’s updated requirements\.U:Recommend some relaxing piano playlists\.A:Sure—here are a few calming playlists\.U:Narrow it down to songs in C major\.A:Sure, here are some C major relaxing piano songs\.Instruction Retention \(IR\)
Persist and follow user instructions and preferences across turns \(e\.g\., safety or style constraints\) without being reminded\.U:Please never include explicit lyrics\.A:Understood\. I’ll only suggest clean tracks\.U:Suggest songs for a family picnic\.A:Try clean versions of ‘Levitating’ \(Dua Lipa\)\.Anaphora Resolution \(AR\)
Identify pronoun referents throughout a multi\-turn dialogue\.U:Recommend three artists I should check out\.A:Billie Eilish, Justin Bieber, and The Weeknd\.U:Who is the third artist?A:The Weeknd is a Canadian singer\-songwriter\.Intent Shift \(IS\)
Recognize and focus on the new request when users unpredictably switch intent\.U:Recommend some upbeat workout songs\.A:Try Calvin Harris or David Guetta\.U:Actually, suggest a mindfulness podcast\.A:Try “The Daily Meditation Podcast\.”Separate Input \(SI\)
The first turn outlines the task requirements and the following turns specify the task input\.U:Make me a playlist\.A:Sure—what kind?U:Bon Iver hits\.A:Got it—building a Bon Iver top hits playlist\.Table 1\.The five multi\-turn tasks in our taxonomy\. Each task is defined by a natural\-language rubric used by the LLM\-as\-a\-judge and illustrated by a representative dialog plan\.U:andA:denote user and assistant turns\.The taxonomy serves two complementary purposes\. During data generation, the examplesℰi\\mathcal\{E\}\_\{i\}guide the DialogPlanLLM to produce conversations exercising a specific capabilitytit\_\{i\}\. During evaluation, the rubricrubrici\\textsc\{rubric\}\_\{i\}is incorporated into the judge prompt, ensuring that scoring is aligned with the intended capability rather than to a generic notion of helpfulness\.

#### 3\.1\.2\.Single\-Turn Data

Single\-turn prompts from real traffic are valuable for capturing the types of requests users are mostly interested in\. However, they are inherently limited by the capabilities of existing product surfaces\. For example, informational queries are relatively rare because they are not well supported in current products\. In this cold\-start scenario, we therefore complement real prompts with synthetic ones that reflect product priorities and anticipated conversational behaviors\.

##### Approach\.

The seed corpus𝒮\\mathcal\{S\}is defined as the union𝒮=𝒮real∪𝒮synth\\mathcal\{S\}=\\mathcal\{S\}\_\{\\text\{real\}\}\\cup\\mathcal\{S\}\_\{\\text\{synth\}\}, where𝒮real\\mathcal\{S\}\_\{\\text\{real\}\}is sampled from existing Spotify product traffic,333[https://newsroom\.spotify\.com/2025\-05\-13/dj\-voice\-requests/](https://newsroom.spotify.com/2025-05-13/dj-voice-requests/)and𝒮synth\\mathcal\{S\}\_\{\\text\{synth\}\}consists of synthetic prompts targeting underrepresented but important use cases\. These seed prompts are then used by the DialogPlanLLM to generate multi\-turn conversation plans\. The resulting training/validation/test split of𝒮\\mathcal\{S\}is propagated to the multi\-turn dataset to prevent leakage across splits at the seed level\.

#### 3\.1\.3\.Dialog Plan Generation

A straightforward way to obtain multi\-turn dialogs is to roll out a user simulator against the agent and rely on emergent behavior\. In practice, this produces conversations that vary across runs, are difficult to reuse, and depend on the specific agent version\.

Instead, we first define an abstract*dialog plan*, i\.e\., a turn\-by\-turn script designed to exercise a specific conversational capability, and only then realize it through interaction with the live agent\. This separation provides greater control over the dataset, enables reuse across evaluations, and ensures that each generated conversation is grounded in a well\-defined target capabilityti∈𝒯t\_\{i\}\\in\\mathcal\{T\}\. Moreover, dialog plans can be naturally extended with manually curated examples or real user conversations as they become available\.

##### Approach\.

Given a conversational capabilityti∈𝒯t\_\{i\}\\in\\mathcal\{T\}, a seed promptq∈𝒮q\\in\\mathcal\{S\}, and a target conversation lengthnturnsn\_\{\\text\{turns\}\}, the DialogPlanLLM generates a candidate conversation

d=\(\(u1,a1\),\(u2,a2\),…,\(unturns,anturns\)\),d=\\big\(\(u\_\{1\},a\_\{1\}\),\\,\(u\_\{2\},a\_\{2\}\),\\,\\ldots,\\,\(u\_\{n\_\{\\text\{turns\}\}\},a\_\{n\_\{\\text\{turns\}\}\}\)\\big\),whereuku\_\{k\}andaka\_\{k\}denotes the user and agent turns at stepkk\.

The DialogPlanLLM is conditioned on the task rubricrubrici\\textsc\{rubric\}\_\{i\}and the few\-shot examplesℰi\\mathcal\{E\}\_\{i\}, and is instructed to generate a conversation exercising the target capabilitytit\_\{i\}\. In particular, the firstnturns−1n\_\{\\text\{turns\}\}\-1provide the conversational context, while the final turn\(unturns,anturns\)\(u\_\{n\_\{\\text\{turns\}\}\},a\_\{n\_\{\\text\{turns\}\}\}\)defines the behavior under evaluation\. For example, given the Content Refinement capability and the seed*“rock vibes,”*the DialogPlanLLM may generate a two\-turn conversation in which the agent first returns a playlist and the user then asks for a more modern version \(Figure[2](https://arxiv.org/html/2609.30297#S3.F2)a\)\.

To cover a range of conversation lengths, we samplenturns∼𝒰⁡\{2,…,nmax\}n\_\{\\text\{turns\}\}\\sim\\mathcal\{U\}\\\{2,\\ldots,n\_\{\\max\}\\\}for each generated conversation\. The target lengthnturnsn\_\{\\text\{turns\}\}is provided to the DialogPlanLLM as a soft constraint, meaning that it guides generation rather than being enforced through truncation\. The train/validation/test assignment of each generated conversation inherits that of its seed promptqqin𝒮\\mathcal\{S\}, ensuring consistency across splits\.

#### 3\.1\.4\.Conversation Realization

In a dialog plan, the agent turnsaka\_\{k\}are synthetically generated by the DialogPlanLLM rather than by the live conversational recommendation agent, and therefore lack the structured outputs \(e\.g\., playlist URIs, album URIs, tool outputs\) present in real interactions\. Simply replaying the dialog plan would therefore evaluate the agent in an unrealistic and overly simplified setting\. To address this, we preserve the*structure*of the dialog plan , i\.e\., the sequence of interactions designed to exercise a target conversational capability, while grounding each turn in the agent’s actual behavior\.

##### Approach\.

We realize the dialog plan through a closed\-loop interaction between the UserLLM and the live agent\. The first user turnu1u\_\{1\}is sent verbatim to the agent, which produces a real responsea^1\\hat\{a\}\_\{1\}containing both text and structured outputs\. For each subsequent turnk\>1k\>1, the UserLLM generates the next user messageu^k\\hat\{u\}\_\{k\}conditioned on:

- •the original dialog plandd, which specifies the target conversational capability,
- •the observed conversation history\(\(u1,a^1\),…,\(u^k−1,a^k−1\)\)\\big\(\(u\_\{1\},\\hat\{a\}\_\{1\}\),\\ldots,\(\\hat\{u\}\_\{k\-1\},\\hat\{a\}\_\{k\-1\}\)\\big\),
- •and the planned user turnuku\_\{k\}, which serves as a reference but may be adapted\.

The UserLLM is instructed to follow the plan whenever it remains consistent with the agent’s previous responsea^k−1\\hat\{a\}\_\{k\-1\}, and to minimally adaptuku\_\{k\}otherwise, while preserving the intended conversational capability\. For example, in a Content Refinement scenario, if the planned user turn is*“make it more modern”*but the agent has already returned a modern playlist, the UserLLM may instead rewrite the request as*“make it more old school”*to preserve a meaningful refinement behavior \(Fig\.[2](https://arxiv.org/html/2609.30297#S3.F2)a\)\. The resulting realized conversationd^=\(\(u1,a^1\),…,\(u^nturns,a^nturns\)\)\\hat\{d\}=\\big\(\(u\_\{1\},\\hat\{a\}\_\{1\}\),\\ldots,\(\\hat\{u\}\_\{n\_\{\\text\{turns\}\}\},\\hat\{a\}\_\{n\_\{\\text\{turns\}\}\}\)\\big\)preserves the intended capability while remaining grounded in the agent’s actual behavior\.

#### 3\.1\.5\.LLM\-as\-a\-Judge

Once a conversationd^\\hat\{d\}has been generated against the live agent, we need a scalable way to assess whether the agent successfully performed the intended conversational capability and we use an LLM\-as\-a\-Judge approach\([Rahmani et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib6);[Zheng et al\., 2023](https://arxiv.org/html/2609.30297#bib.bib7)\)\. To have a more fine\-grained assessment of multi\-turn behavior, we use an instance\-level rubric tied to the specific capabilitytit\_\{i\}used to generate the dialog\. This evaluation process is illustrated in Figure[2](https://arxiv.org/html/2609.30297#S3.F2)b\.

##### Approach\.

For a realized conversationd^\\hat\{d\}generated under capabilitytit\_\{i\}, the LLM\-as\-a\-judge receives the firstnturns−1n\_\{\\text\{turns\}\}\-1turns as conversational context and evaluates only the final agent turna^nturns\\hat\{a\}\_\{n\_\{\\text\{turns\}\}\}\. The judge’s prompt is conditioned on the rubricrubrici\\textsc\{rubric\}\_\{i\}, ensuring that evaluation is aligned with the intended capability\.

The judge returns a labely∈\{green,red\}y\\in\\\{\\textsc\{green\},\\textsc\{red\}\\\}, wheregreenindicates that the final turn satisfies the rubric andredindicates a failure, along with a free\-form explanation\. While these explanations are not used in aggregate metrics, they are retained for error analysis and to support the self\-improvement loop described next\.

An alternative approach to evaluating only the final agent turn would be to score every turn in the conversation\. However, this would increase evaluation cost and latency linearly with the number of turns\. Instead, we focus on the final turn while varying conversation lengths across the dataset, allowing us to assess multi\-turn performance at different conversational horizons in aggregate\.

### 3\.2\.Self\-Improvement Loop

Figure 3\.Agent Self\-Improvement Loop\. We setup a process that leverages the variance in agent planning and the ability of coding agents to diagnose issues and propose improvements to self\-improve the agent prompt and tools\.Once synthetic data has been generated and validated, the remaining challenge is how to efficiently optimize agent planning using this data\. In practice, prompt engineering for conversational agents is a manual and iterative process: given a failing query, one inspects the execution trace, identifies the root cause, modifies the prompt or tool descriptions, and re\-runs evaluation to verify the fix\. This process is time\-consuming and does not scale\.

We therefore introduce a self\-improvement loop inspired by recent work on prompt and agent harness optimization \(Section[2](https://arxiv.org/html/2609.30297#S2)\)\. As illustrated in Figure[3](https://arxiv.org/html/2609.30297#S3.F3), the loop repeatedly samples agent plans, evaluates them using the LLM\-as\-a\-judge, attributes the source of failures through variance analysis, and applies targeted improvements through a coding agent\.

A key observation motivating the loop is that not all failures are of the same nature\. In practice, some failures correspond to inconsistent planning behaviors that occasionally succeed under sampling, while others reflect genuinely missing capabilities that cannot be recovered through exploration alone\.

Formally, letpass​@​k\\text\{pass\}@kdenote the probability that at least one ofkkindependent sampled plans is correct \(i\.e\., rated as GREEN by the judge\), and letpassk\\text\{pass\}^\{k\}denote the probability that*all*kksampled plans are correct\. We estimate these quantities using unbiased estimators following\([Yao et al\., 2024](https://arxiv.org/html/2609.30297#bib.bib16);[Chen et al\., 2021](https://arxiv.org/html/2609.30297#bib.bib15)\)\. Empirically, we observe a large gap between them \(e\.g\.,32%32\\%atk=5k=5\), which reveals two distinct failure modes:

- •Reliability failures\(pass​@​k≫passk\\text\{pass\}@k\\gg\\text\{pass\}^\{k\},pass​@​k\>0\\text\{pass\}@k\>0\): the correct behavior exists in the model’s output distribution but is not consistently selected\.
- •Capability gaps\(pass​@​k≈0\\text\{pass\}@k\\approx 0\): the correct behavior is absent from the model’s distribution and cannot be recovered through sampling alone\.

These modes require different treatment\. As shown in Figure[3](https://arxiv.org/html/2609.30297#S3.F3), reliability failures are handled through contrastive optimization over successful and unsuccessful plans, while capability gaps trigger iterative refinement through a coding agent\.

However, reliability failures are often noisy and may arise from multiple sources\. Variance may stem from genuinely different plans, variations in tool arguments, inconsistencies in tool outputs, or even disagreement from the judge itself\. We therefore structure the loop into gove stages: \(i\) sampling; \(ii\) root\-cause detection; \(iii\) contrastive optimization for reliability failures; and \(iv\) agentic refinement for capability gaps \(v\) batch editing and evaluation\.

The self\-improvement loop is implemented through a coding agent equipped with read–write access to the prompt artifacts \(e\.g\., system prompt, few\-shot examples, tool descriptions\), and read access to full execution and evaluation traces \(conversation, agent plan, tool calls, tool outputs, and judge feedback\)\. Auxiliary scripts support trace parsing and variance analysis throughout the process\.

#### 3\.2\.1\.Sampling

The first step in our self\-improvement loop is to distinguish reliability failures from capability gaps\. To do that, we need to sample multiple responses for the same query and run our evaluation\.

##### Approach\.

As illustrated in Figure[3](https://arxiv.org/html/2609.30297#S3.F3), we begin by sampling multiple agent plans for the same queryqqat elevated temperatureτ∈\[0\.7,1\.0\]\\tau\\in\[0\.7,1\.0\]:

πi∼pθ\(⋅∣q,𝒫\),i=1,…,k,\\pi\_\{i\}\\sim p\_\{\\theta\}\(\\cdot\\mid q,\\,\\mathcal\{P\}\),\\quad i=1,\\ldots,k,where𝒫\\mathcal\{P\}denotes the current prompt configuration \(system prompt, tool descriptions, and few\-shot examples\)\. Each sampled planπi\\pi\_\{i\}is executed and evaluated by the LLM\-as\-judge \(Section[3\.1\.5](https://arxiv.org/html/2609.30297#S3.SS1.SSS5), producing a labelyi∈\{green,red\}y\_\{i\}\\in\\\{\\textsc\{green\},\\textsc\{red\}\\\}\)\. This yields two sets of executions:

𝒟\+=\{πi:yi=green\},𝒟−=\{πk:yk=red\}\.\\mathcal\{D\}^\{\+\}=\\\{\\pi\_\{i\}:y\_\{i\}=\\textsc\{green\}\\\},\\qquad\\mathcal\{D\}^\{\-\}=\\\{\\pi\_\{k\}:y\_\{k\}=\\textsc\{red\}\\\}\.If all samples belong to𝒟\+\\mathcal\{D\}^\{\+\}, the query is consistently successful and no action is needed \(‘all pass’\)\. If some samples belong to𝒟\+\\mathcal\{D\}^\{\+\}and some belong to𝒟−\\mathcal\{D\}^\{\-\}\(‘mixed’\), we proceed with the Root Cause Detection step and Contrastive Optimization\. If all samples belong to𝒟−\\mathcal\{D\}^\{\-\}, we go directly to Agentic Iterative Refinement \(‘all fail’\)\.

#### 3\.2\.2\.Root Cause Detection

Variance in agent evaluation results can arise from multiple sources\. Some reflect meaningful signal for agent self\-improvement, while others are simply noise\. Treating all sources of variance uniformly would be both inefficient and potentially harmful, increasing the risk of incorrect updates\. A prerequisite for variance\-based optimization is therefore to attribute each instance of variance to its root cause\.

##### Approach\.

We compare successful and failing executions to identify the source of variance using a combination of trace parsing and LLM\-based reasoning\. The resulting root\-cause taxonomy includes:

- •planning\_tools— different tool selection \(e\.g\.SearchToolvsRecommendationTool\)\.
- •planning\_args— identical tools but different arguments \(e\.g\."mk gee top hits"vs"mk gee fresh songs"\)\.
- •action— different high\-level actions \(e\.g\., playback vs\. saving to library\)
- •tool\_failure— identical plan but downstream tool failure \(e\.g\., timeouts, connection issues\)\.
- •tool\_variance— identical plan and execution but non\-deterministic tool outputs
- •judge\_variance— identical execution but inconsistent judge outputs

Deterministic traces features \(e\.g\., tool names, argument hashes, exit codes\) are extracted programmatically, while the LLM is used only for cases requiring semantic interpretation\.

The root\-cause label determines how each execution is handled in the self\-improvement loop \(Figure[3](https://arxiv.org/html/2609.30297#S3.F3)\)\. Only failures attributed toplanning\_tools,planning\_args, oractionare passed to the contrastive optimization stage\. Other failure modes that are unrelated to agent planning are logged separately for analysis\.

#### 3\.2\.3\.Contrastive Optimization

Whenpass​@​k\>0\\text\{pass\}@k\>0, the model is already capable of producing the correct plan, but does so inconsistently\. The goal of contrastive optimization is therefore to increase the probability of selecting successful planning behaviors, converting stochastic success into more reliable execution\.

##### Approach\.

Given a queryqqwhose variance has been attributed to agent planning, the coding agent compares successful and unsuccessful plans from𝒟\+\\mathcal\{D\}^\{\+\}and𝒟−\\mathcal\{D\}^\{\-\}to identify the planning patterns associated with success\. These patterns typically involve differences in tool selection, tool ordering, or argument construction\. The coding agent then proposes a prompt improvemente∗e^\{\*\}that reinforces the successful planning behavior\. The updated prompt is𝒫′=𝒫∪\{e∗\}\\mathcal\{P\}^\{\\prime\}=\\mathcal\{P\}\\cup\\\{e^\{\*\}\\\}\.

#### 3\.2\.4\.Agentic Iterative Refinement

Whenpass​@​k=0\\text\{pass\}@k=0for allkk, the model never produces the correct plan, regardless of sampling\. In this mode, the required behavior lies outside the support of the current model distribution, and sampling\-based optimization is therefore insufficient\. Instead, we rely on a coding agent to analyze failures and propose targeted fixes based on the evaluation traces and prompt artifacts\.

##### Approach\.

Given a consistently failing execution trace for a queryqq, the coding agent analyzes why the current planπ^\\hat\{\\pi\}fails and what alternative planπ∗\\pi^\{\*\}could succeed\. The agent then proposes a minimal modification to the prompt configuration𝒫\\mathcal\{P\}, targeting the specific source of failure \(e\.g\., adding a tool\-use example, extending a tool description, or introducing a routing instruction\)\.

Unlike contrastive optimization, where successful plans can already be surfaced through sampling, iterative refinement requires the coding agent to reason about missing capabilities directly\. As a result, each iteration is more computationally expensive and requires more reasoning from the agent\. We therefore reserve this stage for queries where contrastive optimization has been attempted and exhausted \(i\.e\.,pass​@​k=0\\text\{pass\}@k=0\)\.

#### 3\.2\.5\.Batch editing and validation

To reduce the risk of overfitting, the agent operates on random*mini\-batches*of queries\. Given a batchQ=\{q1,…,qm\}Q=\\\{q\_\{1\},\\ldots,q\_\{m\}\\\}, the coding agent produces a single updateΔ​𝒫\\Delta\\mathcal\{P\}intended to resolve all the fixes identified in the entire batch:

Δ​𝒫∗=arg⁡minΔ​𝒫​1m​∑j=1mℒ⁡\(qj,𝒫\+Δ​𝒫\),\\Delta\\mathcal\{P\}^\{\*\}=\\arg\\min\_\{\\Delta\\mathcal\{P\}\}\\;\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\mathcal\{L\}\(q\_\{j\},\\,\\mathcal\{P\}\+\\Delta\\mathcal\{P\}\),whereℒ⁡\(q,𝒫\)\\mathcal\{L\}\(q,\\mathcal\{P\}\)denotes the “judge loss” for queryqqunder prompt configuration𝒫\\mathcal\{P\}\.

Batching encourages the coding agent to identify shared structural causes rather than overfitting to individual queries\. After each update, we re\-run the evaluation with temperatureτ=0\\tau=0to deterministically verify the effectiveness of the modification\. If successful, the update is committed and the process continues with the next batch\. At the end of the loop, a pull request is opened for human review and a final full evaluation is performed as a guardrail against potential regressions\.

## 4\.Results

In this section, we present results on the multi\-turn synthetic data generation pipeline, the self\-improvement loop, and the online performance of the conversational recommendation agent\.

The following experimental results are obtained using top\-tier frontier models with strong off\-the\-shelf capabilities for all the tasks described, i\.e\. synthetic data generation, LLM\-as\-a\-judge, coding agent, and for the conversational recommendation agent itself\.

### 4\.1\.Multi\-turn Evaluation Results

We first evaluate the quality of the synthetic multi\-turn data generated by our pipeline, and then use it to analyze the agent’s conversational capabilities\.

Figure 4\.Quality of synthetic multi\-turn conversations judged by human annotators##### What is the quality of the synthetic data?

Before using the synthetic multi\-turn data for evaluation and self\-improvement, we benchmark and validate its quality through human annotation\. We sample 100 synthetic user\-agent conversations stratified by conversational capability and ask five annotators to evaluate each conversation along four binary dimensions:

- •Q1\(task demonstration\):*does the conversation exercise the intended conversational capability, i\.e\. the multi\-turn task?*
- •Q2\(conversation realism\):*is the conversation natural, i\.e\., could it plausibly come from a real user?*
- •Q3\(adaptation quality\):*when the UserLLM adapts a turn, is the adaptation reasonable given the original dialog plan?*
- •Q4\(judge accuracy\):*is the judge verdict correct and aligned with human judgment?*

Results are shown in Figure[4](https://arxiv.org/html/2609.30297#S4.F4)\. All four dimensions score above 90%, indicating strong overall quality of the synthetic data generated by our pipeline\. Conversations almost always exercise the target capability \(Q1: 98%\)\. The only capability below 95% is Instruction Retention, where some conversations repeat the user’s constraint across turns, effectively reminding rather than testing retention\.

Conversations are also highly realistic \(Q2: 93%\)\. Most failures occur when the agent’s response diverges from the dialog plan, causing the UserLLM to generate an awkward adaptation linking Q2 failures directly to Q3\. Overall, UserLLM adaptation quality is strong with 94% of applicable cases rated as reasonable \(Q3\)\. Judge alignment with human annotation is also high, with 97% of samples rated as accurate \(Q4\)\. Notably, annotators never overruleREDverdicts, though a few disagreements arise onGREENcases\.

To further monitor judge alignment over time, we maintain a curated set of approximately150150*golden*examples on which we track near\-perfect agreement\. Note that we use different models for the judge and the agent itself to avoid same model biases\.

##### What are the most challenging conversational capabilities?

We analyze the agent’s multi\-turn capabilities by measuringp​a​s​s​@​1pass@1through LLM\-as\-a\-judge verdicts \(i\.e\., the GREEN rate\) across conversational capabilities\. Results are averaged overn=10n=10runs to estimate variability\.

We find that Intent Shift is the easiest capability, with the agent reliably recognizing changes in user intent during a conversation\. Anaphora Resolution performs next best, with the agent correctly identifying references to items mentioned in previous turns \(*e\.g\., “play the second one”*\)\. Separate Input follows, while Content Refinement and Instruction Retention are the most challenging capabilities\. The latter is particularly difficult, as it requires enforcing constraints introduced early in the conversation across turns\.

Beyond these aggregate results, multi\-turn evaluation also helped identify several failure modes before deployment\. For instance, in Anaphora Resolution, we observed cases where the agent returned incorrect answers when users referred to playlist items by position\. This behavior was traced to incorrect tool\-call arguments that altered track ordering\.

In Instruction Retention, we identified cases where multiple constraints were partially lost, with the agent*over\-summarizing*the prompt during music session creation\. Similarly, in Content Refinement, the agent sometimes failed to satisfy constraints due to missing or incomplete metadata\. These issues were subsequently addressed through manual fixes and automated improvements using the self\-improvement loop described in Section[3\.2](https://arxiv.org/html/2609.30297#S3.SS2)\.

Figure 5\.Multi\-Turn capabilities vary by task\. Relativep​a​s​s​@​1pass@1with respect to the best task Intent Shift\. See definition of each task in Tab\.[1](https://arxiv.org/html/2609.30297#S3.T1)\.
##### How does multi\-turn quality vary with conversation length?

We analyze howp​a​s​s​@​1pass@1changes as conversation length increases\. Performance remains stable from 2 to 4 turns, but degrades from 5 turns onward, consistent with prior findings\([Laban et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib27)\)\. Improving long\-context handling through techniques such as context compaction and history summarization\([Kang et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib10)\)remains future work\.

Figure 6\.Multi\-Turn capabilities vary by conversation length\. Relativep​a​s​s​@​1pass@1with respect to 2\-turns conversations\.

### 4\.2\.Self\-Improvement Results

We evaluate the effectiveness of the self\-improvement loop in automatically improving agent planning at scale\.

Given a dataset with approximately10001000examples, including both single\-turn and multi\-turn conversations, we run the self\-improvement loop to automatically identify and fix outstanding issues\. The data is processed in mini\-batches of size 8 to avoid overloading the coding agent with too many issues at once\. The pipeline typically runs for a few hours, iteratively fixing issues and committing updates to a code branch\.

In one run, the combined effect of these changes led to \+8% relative increase inp​a​s​s​@​1pass@1on top of a highly manually optimized prompt on the full test set \(statistically significant withp=0\.01p=0\.01using a pairwise sign test\)\. Examples of identified prompt improvements include:

- •intent clarification in multi\-turn intents:*when a user says ‘give me more artists’ after asking for a playlist, refine the existing playlist adding more variety rather than returning a list of artists*
- •genre disambiguation for music session:*When a genre name is ambiguous, disambiguate based on conversational context\. For instance, if paired with rock genres like ‘alt rock’, ‘hardcore’ means hardcore punk and not electronic hardcore\.*
- •time constraints:*‘Year in review’, ‘my year in music’ are ALL calendar\-year queries \- use listening history tools with start time end time for the full calendar year*
- •item type validation:*if the user is asking for ‘latest album’ but tools returned ONLY singles and no actual albums, you should plan again with the correct entity type filter*

All changes proposed by the self\-improvement loop are reviewed before being merged into the codebase\. In practice, the loop significantly accelerates agent planning optimization\. Compared to alternative approaches such as ACE\([Zhang et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib18)\), we find that contrastive optimization combined with variance\-based root\-cause analysis helps avoid spurious updates driven by noisy signals \(e\.g\., infrastructure failures, tool variability, or judge inconsistencies\)\.

Finally, using a coding agent rather than a standard LLM for reflection and refinement improves handling of long contexts and complex updates, enabling fixes that go beyond prompt edits, such as improvements to tool descriptions\. This is consistent with recent work on harness optimization\([Lee et al\., 2026](https://arxiv.org/html/2609.30297#bib.bib20)\)\.

### 4\.3\.Online Results

To validate the effectiveness of the conversational recommendation agent in a production setting, we conducted an A/B test comparing the new experience, which supports chat\-like, multi\-turn interactions across music session and informational intents, to a baseline limited to session refinement\. The experiment was run over two weeks across several markets, exposing approximately 15 million Spotify users across multiple device types\. Online results prove the effectiveness of general purpose conversational recommendation agents: we observed 14% additional user listening, 5% increase in weekly active users and a 5% reduction in skip\-rate when interacting with the agent compared to the prior experience\. Based on these results, the conversational recommendation agent was subsequently rolled out in production\.

## 5\.Conclusion and Discussion

In this paper, we presented an approach to bootstrapping a conversational recommendation agent at Spotify in a cold\-start setting\. We focused on two key contributions: a synthetic multi\-turn conversation generation pipeline and a self\-improvement loop that leverages evaluation feedback to iteratively improve agent planning\.

We show that the synthetic data pipeline produces high\-quality conversations, validated through human annotation with scores above 90% across all evaluation dimensions\. The pipeline enables fine\-grained analysis of conversational capabilities, revealing that tasks requiring the manipulation of lists under evolving constraints, such as Instruction Retention and Content Refinement, remain challenging even for frontier models\. We also observe that performance degrades as conversations extend beyond four turns\.

This result highlights that conversational recommendation agents are subject to the same long\-context limitations observed in agents and LLMs more broadly\([Laban et al\., 2025](https://arxiv.org/html/2609.30297#bib.bib27)\), and motivates techniques for context compaction that preserve key conversational details \(e\.g\., recommended items\) while compressing more generic verbal information\.

We also describe how the self\-improvement loop yields a \+8% quality improvement on top of a highly optimized prompt, while uncovering concrete planning and tool\-use issues\. Although outputs are always reviewed by humans before deployment, the loop significantly accelerates iteration on agent planning\.

More generally, these results align with the emerging trend of using coding agents to automate the improvement of models and agents through iterative “auto\-research” workflows\([Lee et al\., 2026](https://arxiv.org/html/2609.30297#bib.bib20);[Agrawal et al\., 2026a](https://arxiv.org/html/2609.30297#bib.bib21);[Karpathy, 2026](https://arxiv.org/html/2609.30297#bib.bib9)\), demonstrating their feasibility and effectiveness in a real industry setting\. In our experiments, contrastive optimization combined with variance\-based root\-cause analysis proved particularly important for improving the signal\-to\-noise ratio, reducing hallucinations, and increasing the efficiency of the self\-improvement loop\.

Finally, online experiments demonstrate the practical impact of our work: the conversational recommendation agent drives more than 14% additional user listening, increases weekly active users by 5%, and reduces skip rate by over 5% compared to the prior session\-refinement\-only experience\.

These results validate both the usefulness of the feature and the effectiveness of our approach, which builds on our previous work on scalable agentic query understanding at Spotify\([Palumbo et al\., 2025a](https://arxiv.org/html/2609.30297#bib.bib25)\)and extends it to multi\-turn conversational recommendation and information seeking\. More broadly, this work is part of an ongoing effort to enable natural and conversational exploration of the Spotify catalog through Agentic Search and Recommendations\.

Several directions remain open\. On the data side, we plan to extend the taxonomy to cover informational and meta\-conversational intents and to incorporate real production traffic as it becomes available\. On the optimization side, we aim to integrate harness optimization methods\([Lee et al\., 2026](https://arxiv.org/html/2609.30297#bib.bib20);[Agrawal et al\., 2026a](https://arxiv.org/html/2609.30297#bib.bib21)\)and extend self\-improvement beyond prompts and tool definitions to the broader agent pipeline\.

Overall, we believe the approach presented in this paper provides a scalable and practical framework for evaluating and improving conversational agents in the absence of real user interactions\. As natural\-language interfaces move from research prototypes to mainstream products, we expect the cold\-start challenges addressed here to become increasingly relevant across both recommender systems and agentic applications, and hope this work provides a useful starting point for teams facing similar constraints\.

## References

- Agrawalet al\.\(2026a\)L\. A\. Agrawal, D\. Lee, W\. Ma, K\. Elmaaroufi, S\. Tan, S\. A\. Seshia, K\. Sen, D\. Klein, I\. Stoica, J\. E\. Gonzalez, O\. Khattab, A\. G\. Dimakis, and M\. Zahariaoptimize\_anything: a universal API for optimizing any text parameter\.Note:[https://gepa\-ai\.github\.io/gepa/blog/2026/02/18/introducing\-optimize\-anything/](https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-optimize-anything/)Technical blog postCited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.30297#S5.p5.1),[§5](https://arxiv.org/html/2609.30297#S5.p8.1)\.
- Agrawalet al\.\(2026b\)L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang, C\. Potts, K\. Sen, A\. G\. Dimakis, I\. Stoica, D\. Klein, M\. Zaharia, and O\. KhattabGEPA: reflective prompt evolution can outperform reinforcement learning\.InInternational Conference on Learning Representations \(ICLR\),Note:Oral; arXiv:2507\.19457Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px3.p1.1)\.
- Baiet al\.\(2024\)G\. Bai, J\. Liu, X\. Bu, Y\. He, J\. Liu, Z\. Zhou, Z\. Lin, W\. Su, T\. Ge, B\. Zheng,et al\.Mt\-bench\-101: a fine\-grained benchmark for evaluating large language models in multi\-turn dialogues\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7421–7454\.Cited by:[§1](https://arxiv.org/html/2609.30297#S1.p5.1),[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1),[§3\.1\.1](https://arxiv.org/html/2609.30297#S3.SS1.SSS1.Px1.p1.1)\.
- Belkin and Croft \(1992\)N\. J\. Belkin and W\. B\. CroftInformation filtering and information retrieval: two sides of the same coin?\.Communications of the ACM35\(12\),pp\. 29–38\.Cited by:[§1](https://arxiv.org/html/2609.30297#S1.p1.1)\.
- Chagantyet al\.\(2023\)A\. T\. Chaganty, M\. Leszczynski, S\. Zhang, R\. Ganti, K\. Balog, and F\. RadlinskiBeyond single items: exploring user preferences in item sets with the conversational playlist curation dataset\.InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 2754–2764\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§3\.2](https://arxiv.org/html/2609.30297#S3.SS2.p4.1)\.
- Choiet al\.\(2025\)K\. Choi, S\. Doh, and J\. NamTalkplaydata 2: an agentic synthetic data pipeline for multimodal conversational music recommendation\.arXiv preprint arXiv:2509\.09685\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1)\.
- Christakopoulouet al\.\(2016\)K\. Christakopoulou, F\. Radlinski, and K\. HofmannTowards conversational recommender systems\.InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 815–824\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p1.1)\.
- Deshpandeet al\.\(2025\)K\. Deshpande, V\. Sirdeshmukh, J\. B\. Mols, L\. Jin, E\. Hernandez\-Cardona, D\. Lee, J\. Kritz, W\. E\. Primack, S\. Yue, and C\. XingMultichallenge: a realistic multi\-turn conversation evaluation benchmark challenging to frontier llms\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18632–18702\.Cited by:[§1](https://arxiv.org/html/2609.30297#S1.p5.1),[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1),[§3\.1\.1](https://arxiv.org/html/2609.30297#S3.SS1.SSS1.Px1.p1.1)\.
- Dohet al\.\(2025\)S\. Doh, K\. Choi, and J\. NamTalkplay\-tools: conversational music recommendation with llm tool calling\.arXiv preprint arXiv:2510\.01698\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Heet al\.\(2026\)R\. He, L\. Heldt, L\. Hong, R\. Keshavan, S\. Mao, N\. Mehta, Z\. Su, A\. Tsai, Y\. Wang, S\. Wang,et al\.Plum: adapting pre\-trained language models for industrial\-scale generative recommendations\.InProceedings of the ACM Web Conference 2026,pp\. 8093–8104\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Jannachet al\.\(2021\)D\. Jannach, A\. Manzoor, W\. Cai, and L\. ChenA survey on conversational recommender systems\.ACM Computing Surveys \(CSUR\)54\(5\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2609.30297#S1.p1.1),[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Kanget al\.\(2025\)M\. Kang, W\. Chen, D\. Han, H\. A\. Inan, L\. Wutschitz, Y\. Chen, R\. Sim, and S\. RajmohanAcon: optimizing context compression for long\-horizon llm agents\.arXiv preprint arXiv:2510\.00615\.Cited by:[§4\.1](https://arxiv.org/html/2609.30297#S4.SS1.SSS0.Px3.p1.1)\.
- Karpathy \(2026\)A\. KarpathyAutoresearch: autonomous ML experiment runner\.GitHub\.Note:[https://github\.com/karpathy/autoresearch](https://github.com/karpathy/autoresearch)GitHub repositoryCited by:[§5](https://arxiv.org/html/2609.30297#S5.p5.1)\.
- Labanet al\.\(2025\)P\. Laban, H\. Hayashi, Y\. Zhou, and J\. NevilleLLMs get lost in multi\-turn conversation\.arXiv preprint arXiv:2505\.06120\.Note:Microsoft ResearchCited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.30297#S4.SS1.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.30297#S5.p3.1)\.
- Leeet al\.\(2026\)Y\. Lee, R\. Nair, Q\. Zhang, K\. Lee, O\. Khattab, and C\. FinnMeta\-harness: end\-to\-end optimization of model harnesses\.arXiv preprint arXiv:2603\.28052\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.30297#S4.SS2.p5.1),[§5](https://arxiv.org/html/2609.30297#S5.p5.1),[§5](https://arxiv.org/html/2609.30297#S5.p8.1)\.
- Leszczynskiet al\.\(2023\)M\. Leszczynski, S\. Zhang, R\. Ganti, K\. Balog, F\. Radlinski, F\. Pereira, and A\. T\. ChagantyTalk the walk: synthetic data generation for conversational music recommendation\.arXiv preprint arXiv:2301\.11489\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1)\.
- Palumboet al\.\(2025a\)E\. Palumbo, M\. Isaksson, A\. Tamborrino, M\. Movin, C\. Dincu, A\. Vardasbi, L\. Nikeshkin, O\. Gorobets, A\. Nyman, P\. Newdick,et al\.You say search, i say recs: a scalable agentic approach to query understanding and exploratory search at spotify\.InProceedings of the Nineteenth ACM Conference on Recommender Systems,pp\. 1117–1121\.Cited by:[§1](https://arxiv.org/html/2609.30297#S1.p3.1),[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2609.30297#S5.p7.1)\.
- Palumboet al\.\(2025b\)E\. Palumbo, G\. Penha, A\. Damianou, J\. L\. R\. García, T\. C\. Heath, A\. Wang, H\. Bouchard, and M\. LalmasText2tracks: prompt\-based music recommendation via generative retrieval\.arXiv preprint arXiv:2503\.24193\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Rahmaniet al\.\(2024\)H\. A\. Rahmani, E\. Yilmaz, N\. Craswell, B\. Mitra, P\. Thomas, C\. L\. Clarke, M\. Aliannejadi, C\. Siro, and G\. FaggioliLLMJudge: llms for relevance judgments\.arXiv preprint arXiv:2408\.08896\.Cited by:[§3\.1\.5](https://arxiv.org/html/2609.30297#S3.SS1.SSS5.p1.1)\.
- Rajputet al\.\(2023\)S\. Rajput, N\. Mehta, A\. Singh, R\. Hulikal Keshavan, T\. Vu, L\. Heldt, L\. Hong, Y\. Tay, V\. Tran, J\. Samost,et al\.Recommender systems with generative retrieval\.Advances in Neural Information Processing Systems36,pp\. 10299–10315\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Tayet al\.\(2022\)Y\. Tay, V\. Tran, M\. Dehghani, J\. Ni, D\. Bahri, H\. Mehta, Z\. Qin, K\. Hui, Z\. Zhao, J\. Gupta,et al\.Transformer memory as a differentiable search index\.Advances in neural information processing systems35,pp\. 21831–21843\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p2.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px2.p2.1),[§3\.2](https://arxiv.org/html/2609.30297#S3.SS2.p4.1)\.
- Yuksekgonulet al\.\(2024\)M\. Yuksekgonul, F\. Bianchi, J\. Boen, S\. Liu, Z\. Huang, C\. Guestrin, and J\. ZouTextGrad: automatic “differentiation” via text\.arXiv preprint arXiv:2406\.07496\.Note:Published in Nature, 2025Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px3.p1.1)\.
- Zamaniet al\.\(2023\)H\. Zamani, J\. R\. Trippas, J\. Dalton, and F\. RadlinskiConversational information seeking\.Note:http://dx\.doi\.org/10\.1561/1500000081External Links:[Link](https://arxiv.org/abs/2201.08808)Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025\)Q\. Zhang, C\. Hu, S\. Upasani, B\. Ma, F\. Hong, V\. Kamanuru, J\. Rainton, C\. Wu, M\. Ji, H\. Li, U\. Thakker, J\. Zou, and K\. OlukotunAgentic context engineering: evolving contexts for self\-improving language models\.arXiv preprint arXiv:2510\.04618\.Cited by:[§2](https://arxiv.org/html/2609.30297#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.30297#S4.SS2.p4.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in Neural Information Processing Systems36,pp\. 46595–46623\.Cited by:[§3\.1\.5](https://arxiv.org/html/2609.30297#S3.SS1.SSS5.p1.1)\.

相似文章

将本地代理转变为自我优化代理

Reddit r/LocalLLaMA

一个自我优化的智能体管线,在TerminalBench上将基准性能从约30%提升至约90%,并且可以通过记录交互、使用本地模型进行反思、以及将经验注入未来的系统提示中,扩展应用到日常对话场景。

@neural_avb: https://x.com/neural_avb/status/2072294078805684613

X AI KOLs Timeline

本论文介绍了Autodata,这是一种利用智能“数据科学家”AI的方法,通过迭代生成、验证和优化来自动创建高质量合成数据集,该方法特别针对强化学习(GRPO)进行了优化,以提升语言模型的推理能力。

Shape Your Feed:基于LLM的对话式推荐智能体系统

arXiv cs.AI

Meta 推出了 Shape Your Feed (SYF),这是一个基于 LLM 的智能体框架,用于实时对话式推荐,通过多模态输入、智能体重排序和基于 DPO 的自我进化共同策展内容,在离线和在线评估中均取得了优异结果。