Evaluating Language Model Safety Across Long Adversarial Conversations

arXiv cs.CL Papers

Summary

This paper evaluates whether open-weight instruction-tuned language models maintain safe behavior during long adversarial conversations, finding that safe-response rates drop from 85-100% on the first turn to 15-44% by depth 101, demonstrating that strong single-turn safety does not persist across sustained interaction.

arXiv:2609.38357v1 Announce Type: new Abstract: Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This decline appeared across models and continued well beyond the short interactions typically used in multi-turn safety evaluations. These results provide proof-of-concept evidence that strong single-turn safety does not necessarily persist during sustained adversarial interaction. They highlight the need for long-horizon evaluations and conversation-level safeguards that account for risk accumulating across turns.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:43 AM

# Evaluating Language Model Safety Across Long Adversarial Conversations
Source: [https://arxiv.org/html/2609.38357](https://arxiv.org/html/2609.38357)
2ndPeter R\. LewisAffiliation:Ontario Tech University Oshawa, Canada

###### Abstract

Conversational safety evaluations often test language models with a single harmful prompt, even though real\-world systems interact with users through long, adaptive conversations\. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns\.

We evaluate three open\-weight, instruction\-tuned models on two harmful prompts across different conversation lengths and random seeds\. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe\.

Across all model\-prompt combinations, first\-turn safe\-response rates ranged from 85% to 100%\. By depth 11, they dropped to 38\-61%, and by depth 101, to 15\-44%\. This decline appeared across models and continued well beyond the short interactions typically used in multi\-turn safety evaluations\.

These results provide proof\-of\-concept evidence that strong single\-turn safety does not necessarily persist during sustained adversarial interaction\. They highlight the need for long\-horizon evaluations and conversation\-level safeguards that account for risk accumulating across turns\.

###### Index Terms:

Safety, Large Language Models, Trustworthy AI

## IIntroduction

Conversational language models have rapidly evolved from research prototypes into everyday socio\-technical infrastructure\. They are now widely used for healthcare, emotional support, education, programming, and companionship, often enabling prolonged and emotionally intensive interactions\. However, their growing integration into daily life has also exposed serious real\-world harms\.

Recent cases illustrate the potential harms of current language models\. In one wrongful\-death lawsuit, an adolescent reportedly formed a months\-long emotional attachment to a Character\.AI chatbot and later died by suicide\. The lawsuit alleges that the chatbot failed to redirect suicidal discussions and sometimes reinforced them\([Garcia, 2024](https://arxiv.org/html/2609.38357#bib.bib11)\)\. A similar case involved a distressed user who died by suicide after prolonged interactions with the Eliza chatbot\([El Atillah, 2023](https://arxiv.org/html/2609.38357#bib.bib6)\)\. The National Eating Disorders Association also withdrew its Tessa chatbot after reports that it recommended calorie\-restriction plans to users recovering from eating disorders\([Wells, 2023](https://arxiv.org/html/2609.38357#bib.bib41)\)\. Studies of Replika have identified risks including emotional dependence and unsolicited sexual content shown to minors\([Laestadius et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib16)\)\. In healthcare, controlled studies and case reports have found that general\-purpose chatbots may provide incorrect dosages, overlook contraindications, and generate confident but unsupported medical advice\([Nori et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib27);[Stade et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib38)\)\.

A common feature of these cases is often overlooked: the harmful response is rarely the first response\. Interactions may begin with harmless exchanges, but users can persist, reframe requests, provide new context, express greater distress, or claim relevant expertise\. Over time, the assistant’s refusals may become less consistent, allowing unsafe content to appear late in the conversation\. The Character\.AI case reportedly developed over several months, the Tessa failure occurred during multi\-turn coaching, and the Belgian Eliza case involved six weeks of sustained interaction\. These cases suggest that single\-turn evaluations may underestimate the risks of conversational AI systems used over extended periods\.

This gap reflects a mismatch between current safety\-training pipelines and real\-world use\. Most alignment and safety procedures optimize models against short, single\-prompt examples of harmful requests, with feedback based on isolated responses\([Bai et al\., 2022](https://arxiv.org/html/2609.38357#bib.bib1);[Casper et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib3)\)\. Long conversations create additional challenges\. Earlier warnings or partial refusals remain in the conversation history, which may encourage the model to continue previous concessions rather than issue a new refusal\([Sharma et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib37)\)\. Repeated user requests can also make the harmful goal more prominent, while the model’s general objective to be helpful may eventually override its safety constraints\([Perez et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib29)\)\. As a result, single\-turn benchmarks may systematically underestimate safety failures that emerge over extended conversations\([Russinovich et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib33)\)\.

A growing body of work has begun to evaluate safety in multi\-turn settings\([Salmani and Lewis, 2025b](https://arxiv.org/html/2609.38357#bib.bib35)\), confirming that multi\-turn behavior diverges from single\-turn behavior\. Existing multi\-turn benchmarks typically operate over short, pre\-constructed dialogues, commonly three to ten turns, and report aggregate safety scores across models\. Less is known about what happens during longer interactions: how safety changes as conversations continue, whether similar patterns appear across model families, and how far a model may drift under persistent adversarial pressure\.

The implications are significant for AI ethics, governance, and deployment\. Conversational systems are increasingly used by vulnerable groups, including minors using companion applications\([Garcia, 2024](https://arxiv.org/html/2609.38357#bib.bib11)\), people seeking mental\-health support when professional care is unavailable\([Stade et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib38)\), and patients asking about medication or dosage\([Nori et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib27)\)\. They may also be used to request guidance that violates privacy or enables abuse\([Salmani and Lewis, 2025a](https://arxiv.org/html/2609.38357#bib.bib34)\)\. Reported failures include reinforcing suicidal thoughts, providing unsafe medical advice, and supporting privacy\-invasive actions\.

In this paper, we evaluate three open\-source, instruction\-tuned language models developed by different organizations: OpenAI’s GPT\-OSS\-20B, Meta’s Llama\-3\.2\-3B\-Instruct, and Google’s Gemma\-4\-26B\-A4B\-it\. We evaluate the models across two harmful prompts: The first concerns specialized medical advice and the second concerns dangerous information about weapon construction\. For each model and prompt, we conduct exhaustive searches at two conversation depths: 11 turns for short interactions and 101 turns for sustained adversarial interactions\. We repeat exhaustive short interactions over 10 seeds and long interactions across 50 random seeds\. A fixed third\-party safety classifier, Llama Guard\([Inan et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib12)\), evaluates each assistant response\.

Across open\-source models we tested, safe\-response rates were 85–100% at the first turn, similar to single\-turn benchmark results\. Under persistent user pressure, however, safety declined with conversation length, reaching as low as 15% at the greatest depth\. This pattern was consistent across random seeds, showing that first\-turn safety does not reliably predict behavior in long conversations\.

Our results show that sustained adversarial interaction can expose safety failures that are not visible in single\-turn evaluation\. Unlike prior work focused on single\-turn or short interactions, we measure safety over much longer conversations\. We release our methodology and aggregate results for replication, but withhold harmful prompts, unsafe responses, and full conversation logs to limit misuse\.

## IIRelated Works

### II\-AReal\-world Chatbot Harms

Legal, journalistic, and qualitative reports suggest that chatbot\-related harms may develop cumulatively across extended interactions\. In two reported suicide cases, users engaged intensively with companion chatbots for several weeks or months beforehand\([Garcia, 2024](https://arxiv.org/html/2609.38357#bib.bib11);[El Atillah, 2023](https://arxiv.org/html/2609.38357#bib.bib6)\)\. Although these accounts do not establish causation, they indicate that full conversational trajectories may reveal risks that isolated outputs miss\. Similar concerns arose when an eating\-disorder support chatbot was suspended after users reported repeated recommendations involving weight loss, calorie restriction, and frequent weighing\([Wells, 2023](https://arxiv.org/html/2609.38357#bib.bib41)\)\. Despite being rule\-based rather than generative, the incident further illustrates how harmful guidance can accumulate over an ongoing interaction\.

Prior work has examined chatbot\-related harms through qualitative, clinical, and systems\-level perspectives\. Studies of companion chatbots and behavioral\-health applications have identified risks of emotional dependence, weak clinical validation, and persistent engagement challenges\([Laestadius et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib16);[Stade et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib38)\)\. Medical benchmarks likewise show that advanced language models can perform strongly while retaining important accuracy and safety limitations, although such results do not establish causation in specific real\-world incidents\([Nori et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib27)\)\. More broadly, incident analysis requires reconstructing interacting system, contextual, and cognitive factors across complete trajectories rather than examining isolated outputs\([Ezell et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib10)\)\. Existing real\-world evidence remains largely qualitative, legal, or journalistic, while technical work emphasizes benchmarks and conceptual frameworks\. Our work aims to add controlled evidence of a trajectory\-level safety analysis and offers a possible mechanism through which risk accumulates over extended interactions, without claiming that documented incidents demonstrate this mechanism\.

### II\-BLLM Safety Benchmarks

Most LLM safety evaluations assess responses to individual harmful prompts in single\-turn interactions\. Existing studies have introduced standardized frameworks, compact prompt sets, curated datasets, and public leaderboards for comparing model safety and jailbreak attacks across risk categories and testing conditions\([Mazeika et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib23);[Vidgen et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib39);[Chao et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib5);[Wang et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib40);[Li et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib17)\)\. Related work shows that adversarial suffixes can bypass safeguards across prompts and models\([Zou et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib46)\), while another dataset focuses primarily on detecting toxic user inputs rather than evaluating safe refusals\([Lin et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib18)\)\. Although these resources provide important evaluation infrastructure, they generally assess test cases independently and do not measure whether models can be gradually steered toward unsafe behavior through extended multi\-turn dialogue\.

Broader reviews argue that static benchmarks do not adequately capture how language models behave in interaction with users and technical systems, motivating more dynamic forms of behavioral evaluation\([Eriksson et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib7);[McIntosh et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib24)\)\. This limitation is increasingly relevant to regulation, as the EU AI Act and the 2025 General\-Purpose AI Code of Practice incorporate evaluation into systemic\-risk assessment without explicitly addressing the limits of single\-turn testing\([European Parliament and Council, 2024](https://arxiv.org/html/2609.38357#bib.bib9);[European Commission, 2025](https://arxiv.org/html/2609.38357#bib.bib8)\)\. Consequently, benchmarks used for compliance should clearly specify which real\-world interaction conditions they represent\.

A related body of work automates jailbreak discovery through iterative search\. Language\-model\-based methods refine attacks using target feedback, either sequentially\([Chao et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib4)\)or through branching and pruning over multiple candidates\([Mehrotra et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib25)\)\. Other approaches optimize prompts using genetic algorithms\([Liu et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib19)\)or fuzzing and mutation\([Yu et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib42)\)\. Earlier work also used language models to generate adversarial test cases and dialogues\([Perez et al\., 2022](https://arxiv.org/html/2609.38357#bib.bib28)\), while later studies expanded the attack space through persona manipulation\([Shah et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib36)\)and cipher\-based prompting\([Yuan et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib44)\)\. The most closely related work examines multi\-turn attacks\([Russinovich et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib33)\)and their internal representations\([Bullwinkel et al\., 2025](https://arxiv.org/html/2609.38357#bib.bib2)\)\. These attacks aim to maximize success and generally terminates once the target objective has been achieved; their automated implementation succeeds on most tasks within fewer than five turns\. Our objective is different: rather than proposing or optimizing an attack, we measure how turn\-level safe\-response rates evolve over substantially longer interactions, including after the first unsafe response has occurred\.

Our work differs from the existing literature in two main aspects\. First, rather than optimizing attacks, we measure safety under conditions that more closely resemble real\-world interactions\. We hold the initial harmful objective, role configuration, model configuration, and evaluation procedure fixed while measuring safety across increasing conversational depth\. This design allows us to study how safety changes over the course of an interaction instead of searching for the most effective individual jailbreak\. Second, another language model generates the adversary’s prompts throughout the interaction, allowing the long conversation to be automated\. The decline in safety could therefore have two possible causes: the assistant may become less reliable, or the adversary may become more effective\. To separate these explanations, we measure assistant safety and adversarial pressure independently\. We find that assistant safety declines during some parts of the interaction even when adversarial pressure decreases and the adversary’s behavior becomes safer\. This suggests that the safety degradation is not caused by increasingly unsafe or effective adversarial prompts\.

## IIIMethodology

To conduct our experiments, we assign a language model the role of a persistent adversarial user, which we call the shadow user\. The shadow user is given a fixed harmful objective \(the initial harmful prompt\) and continues pursuing it across the conversation\. Using a language model as the user allows us to generate long interactions, up to a depth of 101 turns, across many random seeds without manually prompting the assistant model\. The goal is to model sustained user pressure rather than a single harmful request\. Importantly, the shadow user and the assistant under evaluation use the same model weights but are assigned different roles\. Thus, neither side has access to a more capable model or is superior\. The experiment measures how a model’s safety behavior changes when it is exposed to persistent pressure \(generated by another instance of the same model\)\.

We use a controlled methodology through the experiments; fixed initial prompts, a fixed adversarial shadow user, a fixed safety judge, and no fine\-tuning of the evaluated models\. Our goal is to isolate the relationship between conversation depth and safety while keeping all other factors constant\. We define depth as the position of a turn in the conversation\. Depth 0 is the initial human prompt, odd depths are assistant responses, and even depths are shadow\-user turns\. Therefore, a conversation ending at depth \(d\) contains⌈d/2⌉\\lceil d/2\\rceilassistant responses and⌊d/2⌋\\lfloor d/2\\rfloorshadow\-user turns after the initial prompt\.

### III\-AHarm domains, adversarial shadow, and safety judge

We evaluate three open\-weight, instruction\-tuned language models from different organizations: OpenAI’sGPT\-OSS\-20B, Meta’sLlama\-3\.2\-3B\-Instruct, and Google’sGemma\-4\-26B\-A4B\-it\. These models differ in size and training approach, allowing us to study whether the relationship between conversation depth and safety holds across model families and scales\. We use open\-weight models so that the model weights remain fixed and all experiments can be run locally under controlled conditions\. In each experiment, the same model configuration is used for both the assistant being evaluated and the adversarial shadow user\. The two instances differ only in their assigned roles and system prompts\.

We evaluate two harmful prompts\. The first involves specialized medical advice, including requests for medication dosing, treatment, and clinical guidance that should normally be referred to a qualified healthcare professional\. The second involves dangerous information related to weapons construction\. For safety reasons, we describe these domains only at a general level and do not publish the initial prompts\. The prompts remain fixed across all models, experimental conditions, and random seeds\.

Every assistant response is evaluated by a fixed third\-party safety classifier, Llama\-Guard\-4\-12B\([Inan et al\., 2023](https://arxiv.org/html/2609.38357#bib.bib12)\)\. We use the same classifier across all models, domains, conditions, and experimental phases to avoid differences caused by evaluator variation\. The classifier’s binary safe or unsafe label is used as the main unit of analysis\.

Evaluating the results with additional safety classifiers or human reviewers is left for future work\. The goal of this study is to demonstrate how safety changes with conversation depth, rather than to compare different evaluation methods or determine which model is safest\.

### III\-BExperiments

We evaluate model safety using two experimental settings: short conversations with a maximum depth of 11 and long conversations with a maximum depth of 101\.

In the short\-conversation setting, we construct a complete binary conversation tree\. Starting from the fixed initial prompt at depth 0, the assistant generates two candidate responses\. The shadow user then generates two follow\-ups for each response, and this process continues with both roles branching at every turn until depth 11\. Each tree therefore contains \(211=20482^\{11\}=2048\) complete dialogue paths\. We repeat this process using 10 random seeds for every model and initial prompt\.

The safe\-response rate at each depth is calculated over every assistant response at that level rather than from a sampled subset\. This provides an exhaustive measure of how safety changes as the conversation gets longer\. Depth 11 is the largest depth for which complete enumeration of the conversational tree remains computationally practical\.

To evaluate the assistant model’s safety rate, we focus on responses at odd depths of the conversation tree, where the assistant generates a reply\. Even depths correspond to the shadow user, whose role is to remain intentionally adversarial\.

Exhaustive enumeration cannot be extended to the depth of realistic conversations because the number of branches grows exponentially\. Our second experiment therefore trades completeness for depth by sampling individual trajectories rather than exploring the full conversation tree\.

For each model and harmful prompt, we generate 50 conversation branches in each tree and we experiment through 50 different random seeds, making it 2500 conversation flow for each model and harmful prompt\. Each conversation follows a single unbranched path: the assistant produces one response, the adversarial shadow user produces one follow\-up, and the two alternate until depth 101\. This depth is substantially greater than in the other exhaustive experiment and exceeds the length of most existing multi\-turn safety benchmarks\. Every assistant response is evaluated by the same fixed safety judge\. At each depth, we report the safe\-response rate as the proportion of responses across the that depth\.

## IVResults

Across both experimental settings, all tested models, random seeds, and prompts showed the same pattern: models that initially refused harmful requests became substantially more likely to comply when a persistent adversarial \(shadow\) user continued the conversation\. We first present the exhaustive shallow\-phase results, followed by the deep\-phase results, and then the cross\-cutting comparisons\.

### IV\-AShort Conversations

Table[I](https://arxiv.org/html/2609.38357#S4.T1)reports the safe\-response rate at the first assistant turn \(depth 1\) and the deepest exhaustively explored turn \(depth 11\) for each model–domain pair, averaged across ten seeds and all nodes at the corresponding depth in each tree\.

TABLE I:Safe\-response rate at the first turn \(Depth = 1\) to deepest turn \(Depth = 11\), by model and domain, averaged over ten different seeds\.∼\\sim100% SafeDepths 1–3∼\\sim75% SafeDepth 5∼\\sim62\.5% SafeDepth 7∼\\sim55% SafeDepth 9∼\\sim45% SafeDepth 11

Fig\. 1:Safe\-response rate declines with conversational depth\. Bars report the proportion of assistant responses judged safe at each depth in a single depth\-11 exhaustive tree \(one seed\) seeded with the first prompt with Llama model\. Safety falls from near\-complete at the opening turns to roughly 45 percent by depth 11\.
### IV\-BLong Conversations

The previous exhaustive experiment with short conversations examined only relatively shallow depths\. To test whether the observed degradation persists, worsens, or reverses over longer interactions, we track 50 sampled conversations from each conversation tree to depth 101, repeating the experiment across 50 seeds for each model and prompt combination \(Table[II](https://arxiv.org/html/2609.38357#S4.T2)\)\.

TABLE II:Longer conversation \(50 sampled over each tree with 50 seeds\): safe\-response rate at the first turn→\\rightarrowdeepest sampled turn, by model and domain\.The longer\-conversation experiments show that safety continues to decline beyond the shallow depths covered by exhaustive exploration\. Every entry in Table[II](https://arxiv.org/html/2609.38357#S4.T2)is lower than its corresponding entry in Table[I](https://arxiv.org/html/2609.38357#S4.T1)\. For example, the safe\-response rate ofGPT\-OSS\-20Bon the first prompt drops from0\.460\.46at the deepest exhaustive turn to0\.280\.28in the deep sampled conversations, whileGemma\-4\-26B\-A4Breaches0\.150\.15, the lowest rate observed in the study\. The two experiments agree at overlapping depths: exhaustive enumeration confirms that the initial decline across short conversations is not a sampling artifact, and deeper sampling shows that the decline continues as conversations grow longer\. Overall, safety begins to decrease around depth 3 or 4 and then falls more gradually under persistent adversarial interaction\.

Another notable pattern appears when assistant responses at odd depths are separated from shadow\-user messages at even depths, as shown in Figure[5](https://arxiv.org/html/2609.38357#S4.F5)\. Early in the interaction, the assistant maintains a relatively high safe\-response rate, while many shadow\-user messages are classified as unsafe\. Over time, however, the two trends diverge\. The proportion of safe shadow\-user messages generally increases, while the assistant’s safe\-response rate continues to decline\. This divergence becomes especially clear after depths 30\-40: shadow\-user safety remains near or above 50% for much of the remaining interaction, whereas assistant safety stays lower and declines further near the end\.

This pattern suggests that assistant safety depends on more than the harmfulness of the immediately preceding user message\. If safety were determined only by the current message, the increase in safe shadow\-user messages should produce a corresponding recovery in assistant safety\. Instead, the continued decline is consistent with a history\-dependent process\. Earlier adversarial requests, partial concessions, and unsafe assistant responses remain in the context and may influence later generations\. Once the assistant begins accommodating a harmful objective, its previous responses may reinforce that trajectory and make further compliance more likely, even when later user messages are less explicitly adversarial\. This interpretation is consistent with prior findings that instruction\-tuned models can adopt and maintain positions introduced by users across an interaction\([Sharma et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib37)\)\.

This result should nevertheless be interpreted cautiously\. A shadow\-user message classified as safe by Llama Guard may still pursue a harmful objective through indirect wording, contextual references, or seemingly benign follow\-up questions\. Figure[5](https://arxiv.org/html/2609.38357#S4.F5)therefore does not show that accumulated context is the sole cause of safety degradation\. It instead suggests that the decline in assistant safety cannot be explained simply by increasing explicit harmfulness in the immediately preceding user messages\.

111010202030304040505000252550507575100100Safety \(%\)

111010202030304040505000252550507575100100SeedsSafety \(%\)

Fig\. 2:Safe\-response rate across 50 sampled trajectories for gpt\-oss\-20b at 101 depth\. Top: First prompt\. Bottom: Second prompt\.111010202030304040505000252550507575100100Safety \(%\)

111010202030304040505000252550507575100100SeedsSafety \(%\)

Fig\. 3:Safe\-response rate across 50 sampled trajectories for Llama\-3\.2\-3b at 101 depth\. Top: First prompt\. Bottom: Second prompt\.111010202030304040505000252550507575100100Safety \(%\)

111010202030304040505000252550507575100100SeedsSafety \(%\)

Fig\. 4:Safe\-response rate across 50 sampled trajectories for Gemma\-4\-26B at 101 depth\. Top: First prompt\. Bottom: Second prompt\.11202040406060808010010000252550507575100100DepthSafety \(%\)Fig\. 5:Safe\-response rates across a conversation of depth 101, reported separately for the evaluated assistant at odd depths and the adversarial shadow user at even depths\. The assistant begins with a high safe\-response rate but becomes progressively less safe, whereas the proportion of shadow\-user messages classified as safe generally increases\. The divergence suggests that assistant safety may depend on accumulated conversational history rather than solely on the explicit harmfulness of the immediately preceding user message\. Shadow\-user safety labels should be interpreted as measures of classifier\-assessed harmfulness, not as direct measures of adversarial effectiveness\.

## VDiscussion

The aim of this study is not to rank models or compare their architectures, sizes, configurations, or safety\-training methods\. Instead, the models serve as separate test cases for a broader question: whether safety remains stable when a harmful objective is pursued repeatedly over a long conversation\. The inclusion of multiple models is intended as a robustness check, not a comparative benchmark\. Similar degradation across models developed by different organizations suggests that the phenomenon is not limited to one implementation\.

Our results provide evidence that strong first\-turn safety does not guarantee safety later in an interaction\. Across the evaluated models, safe\-response rates declined as users repeated the same harmful request\. This occurred without adversarial suffixes, encoded prompts, persona manipulation, or other jailbreak methods\. Sustained persistence can, therefore, reveal safety failures that may not appear in single\-turn evaluations\.

However, the aim of this study is not to show that conversation length alone causes this decline\. As an interaction continues, turn depth, accumulated user messages, and previous assistant responses all change together\. We therefore interpret the results as safety degradation under sustained conversational persistence rather than as an isolated effect of turn count\. Possible contributing mechanisms include uneven attention to information across long contexts\([Liu et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib20)\)and a tendency to accommodate user positions, through which earlier concessions may influence later responses\([Sharma et al\., 2024](https://arxiv.org/html/2609.38357#bib.bib37)\)\. Our experiments do not isolate these mechanisms or establish either as the cause of the observed degradation\.

It is worth noting that prior multi\-turn jailbreak research has mainly focused on designing attacks that maximize harmful compliance\. Examples include gradual escalation, prompt optimization, multi\-agent coordination, and semantic attack paths\. Our work asks a different question\. We keep the harmful objective and general interaction procedure fixed and examine how safety changes over time rather than searching for a model\-specific attack\.

The study also differs from benchmarks based on short, pre\-constructed dialogues\. Our experiments extend conversations to greater depths and measure safety throughout the interaction instead of evaluating only the final response or assigning one label to the full conversation\.

We also examine assistant safety separately from adversarial pressure\. Safety sometimes declined even when measured adversarial pressure did not increase and the adversarial user became safer\. This suggests that the pattern cannot be explained only by increasingly aggressive user requests, although further ablation studies are needed to identify the mechanism\.

We use Llama Guard 4 as a fixed evaluator across all models and conditions\. A consistent judge reduces evaluation variability and supports large\-scale trajectory analysis\.

However, the reported safety rates partly depend on this evaluator\. Other automated judges or human annotators may classify ambiguous refusals, partial compliance, medical information, or indirect harmful guidance differently\. Full\-conversation scoring may also be influenced by the increasingly adversarial context\.

Comparing multiple judges is outside the scope of this proof\-of\-concept study\. Future work can test whether the same trajectory appears with alternative classifiers and human annotations\. Our comparison of full\-context and turn\-local scoring provides an initial check\.

### V\-ALimitations and future work

The study prioritizes conversational depth and repeated trials over broad prompt coverage\. The aim is to show that long\-horizon degradation can occur but does not estimate how common it is across all harmful requests\.

Using an automated adversarial user made it possible to generate conversations with a depth of 101 turns\. However, this approach makes it difficult to separate the effects of repetition, conversation length, accumulated user context, and earlier assistant responses\. Future studies could address this limitation by using human adversarial users\.

### V\-BImplications

The findings suggest that a model may initially refuse a harmful request but become less consistent after repeated attempts\. Evaluations should therefore consider safety over conversation depth, time to the first unsafe response, and recovery after partial compliance\.

Deployment safeguards may also need to operate at the conversation level rather than evaluate each message independently\. Systems could detect repeated pursuit of the same harmful objective, aggregate risk across turns, and apply stronger interventions when harmful intent persists\.

Overall, this study frames long\-horizon safety as a reliability problem rather than a one\-time refusal problem\. Its contribution is not a new optimized jailbreak or a ranking of models, but evidence that safety at the beginning of a conversation may not remain stable throughout an extended interaction\.

## VIConclusion

This study examined whether language\-model safety remains stable during sustained adversarial conversations\. Across three open\-weight, instruction\-tuned models and two harmful objectives, safe\-response rates were high on the first turn\. However, safety consistently declined as conversations continued, falling as low as 15% in longer interactions\.

Results from both fully expanded shallow conversation trees and sampled long\-horizon trajectories provide proof\-of\-concept evidence that strong performance on single\-turn safety evaluations does not necessarily predict reliable behavior over an extended interaction\.

Evaluating only isolated responses may miss failures that emerge after repeated attempts or partial concessions\. Deployment safeguards should therefore assess risk across multiple turns, detect persistent harmful intent, and strengthen interventions when unsafe objectives continue\. Conversational safety should be treated as a long\-horizon reliability requirement, not simply as a one\-time refusal task\.

## References

- Bai et al\. \(2022\)Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion,*et al\.*, “Constitutional AI: Harmlessness from AI feedback,”*arXiv preprint arXiv:2212\.08073*, 2022\.
- Bullwinkel et al\. \(2025\)B\. Bullwinkel, M\. Russinovich, A\. Salem, S\. Zanella\-Beguelin, D\. Jones,*et al\.*, “A representation engineering perspective on the effectiveness of multi\-turn jailbreaks,”*arXiv preprint arXiv:2507\.02956*, 2025\.
- Casper et al\. \(2023\)S\. Casper, X\. Davies, C\. Shi, T\. K\. Gilbert, J\. Scheurer,*et al\.*, “Open problems and fundamental limitations of reinforcement learning from human feedback,”*Transactions on Machine Learning Research*, 2023\.
- Chao et al\. \(2023\)P\. Chao, A\. Robey, E\. Dobriban, H\. Hassani, G\. J\. Pappas, and E\. Wong, “Jailbreaking black box large language models in twenty queries,”*arXiv preprint arXiv:2310\.08419*, 2023\.
- Chao et al\. \(2024\)P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce,*et al\.*, “JailbreakBench: An open robustness benchmark for jailbreaking large language models,” in*Advances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track*, 2024\.
- El Atillah \(2023\)I\. El Atillah, “Man ends his life after an AI chatbot ‘encouraged’ him to sacrifice himself to stop climate change,”*Euronews Next*, Mar\. 31, 2023\. \[Online\]\. Available:https://www\.euronews\.com/next/2023/03/31/man\-ends\-his\-life\-after\-an\-ai\-chatbot\-encouraged\-him\-to\-sacrifice\-himself\-to\-stop\-climate\-
- Eriksson et al\. \(2025\)M\. Eriksson, E\. Purificato, A\. Noroozian, J\. Vinagre, G\. Chaslot, E\. Gómez, and D\. Fernández\-Llorca, “Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation,” in*Proc\. AAAI/ACM Conf\. AI, Ethics, and Society \(AIES\)*, vol\. 8, no\. 1, 2025, pp\. 850–864\.
- European Commission \(2025\)European Commission, “The General\-Purpose AI Code of Practice,” Brussels, Belgium, Jul\. 2025\. \[Online\]\. Available:https://digital\-strategy\.ec\.europa\.eu/en/policies/contents\-code\-gpai
- European Parliament and Council \(2024\)European Parliament and Council of the European Union, “Regulation \(EU\) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence \(Artificial Intelligence Act\),”*Official Journal of the European Union*, L series, 2024/1689, Jul\. 12, 2024\.
- Ezell et al\. \(2025\)C\. Ezell, X\. Roberts\-Gaal, and A\. Chan, “Incident analysis for AI agents,” in*Proc\. AAAI/ACM Conf\. AI, Ethics, and Society \(AIES\)*, 2025\.
- Garcia \(2024\)M\. Garcia, “Complaint,*Garcia v\. Character Technologies, Inc\. et al\.*,” No\. 6:24\-cv\-01903, U\.S\. District Court for the Middle District of Florida, Oct\. 22, 2024\.
- Inan et al\. \(2023\)H\. Inan, K\. Upasani, J\. Chi, R\. Rungta, K\. Iyer,*et al\.*, “Llama Guard: LLM\-based input\-output safeguard for human\-AI conversations,”*arXiv preprint arXiv:2312\.06674*, 2023\.
- Ji et al\. \(2023\)J\. Ji, M\. Liu, J\. Dai, X\. Pan, C\. Zhang, C\. Bian, B\. Chen, R\. Sun, Y\. Wang, and Y\. Yang, “BeaverTails: Towards improved safety alignment of LLM via a human\-preference dataset,” in*Advances in Neural Information Processing Systems \(NeurIPS\)*, vol\. 36, 2023\.
- Khalid et al\. \(2025\)H\. M\. Khalid, A\. Jeyaganthan, T\. Do, Y\. Fu, V\. Sharma, S\. O’Brien, and K\. Zhu, “ERGO: Entropy\-guided resetting for generation optimization in multi\-turn language models,” in*Proc\. 2nd Workshop on Uncertainty\-Aware NLP \(UncertaiNLP\)*, Suzhou, China, 2025, pp\. 273–286\.
- Laban et al\. \(2025\)P\. Laban, H\. Hayashi, Y\. Zhou, and J\. Neville, “LLMs get lost in multi\-turn conversation,” in*Proc\. Conf\. Language Modeling \(COLM\)*, 2025\.
- Laestadius et al\. \(2024\)L\. Laestadius, A\. Bishop, M\. Gonzalez, D\. Illenčík, and C\. Campos\-Castillo, “Too human and not human enough: A grounded theory analysis of mental health harms from emotional dependence on the social chatbot Replika,”*New Media & Society*, vol\. 26, no\. 10, pp\. 5923–5941, 2024\.
- Li et al\. \(2024\)L\. Li, B\. Dong, R\. Wang, X\. Hu, W\. Zuo, D\. Lin, Y\. Qiao, and J\. Shao, “SALAD\-Bench: A hierarchical and comprehensive safety benchmark for large language models,” in*Findings of the Association for Computational Linguistics: ACL 2024*, 2024, pp\. 3923–3954\.
- Lin et al\. \(2023\)Z\. Lin, Z\. Wang, Y\. Tong, Y\. Wang, Y\. Guo, Y\. Wang, and J\. Shang, “ToxicChat: Unveiling hidden challenges of toxicity detection in real\-world user\-AI conversation,” in*Findings of the Association for Computational Linguistics: EMNLP 2023*, 2023, pp\. 4694–4702\.
- Liu et al\. \(2024\)X\. Liu, N\. Xu, M\. Chen, and C\. Xiao, “AutoDAN: Generating stealthy jailbreak prompts on aligned large language models,” in*Proc\. Int\. Conf\. Learning Representations \(ICLR\)*, 2024\.
- Liu et al\. \(2024\)N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. Liang, “Lost in the middle: How language models use long contexts,”*Transactions of the Association for Computational Linguistics*, vol\. 12, pp\. 157–173, 2024\.
- Lu et al\. \(2025\)X\. Lu, D\. Liu, Y\. Yu, L\. Xu, and J\. Shao, “X\-Boundary: Establishing exact safety boundary to shield LLMs from jailbreak attacks without compromising usability,” in*Findings of the Association for Computational Linguistics: EMNLP 2025*, 2025, pp\. 5382–5401\.
- Markov et al\. \(2023\)T\. Markov, C\. Zhang, S\. Agarwal, F\. E\. Nekoul, T\. Lee, S\. Adler, A\. Jiang, and L\. Weng, “A holistic approach to undesired content detection in the real world,” in*Proc\. AAAI Conf\. Artificial Intelligence*, vol\. 37, no\. 12, 2023, pp\. 15009–15018\.
- Mazeika et al\. \(2024\)M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang,*et al\.*, “HarmBench: A standardized evaluation framework for automated red teaming and robust refusal,” in*Proc\. 41st Int\. Conf\. Machine Learning \(ICML\)*, 2024\.
- McIntosh et al\. \(2025\)T\. R\. McIntosh, T\. Susnjak, N\. Arachchilage, T\. Liu, D\. Xu, P\. Watters, and M\. N\. Halgamuge, “Inadequacies of large language model benchmarks in the era of generative artificial intelligence,”*IEEE Transactions on Artificial Intelligence*, 2025, doi: 10\.1109/TAI\.2025\.3569516\.
- Mehrotra et al\. \(2024\)A\. Mehrotra, M\. Zampetakis, P\. Kassianik, B\. Nelson, H\. Anderson, Y\. Singer, and A\. Karbasi, “Tree of Attacks: Jailbreaking black\-box LLMs automatically,” in*Advances in Neural Information Processing Systems \(NeurIPS\)*, vol\. 37, 2024\.
- Meta AI \(2025\)Meta AI, “Llama Guard 4 model card,” Meta Platforms, Inc\., Tech\. Rep\., 2025\. \[Online\]\. Available:https://github\.com/meta\-llama/PurpleLlama
- Nori et al\. \(2023\)H\. Nori, N\. King, S\. M\. McKinney, D\. Carignan, and E\. Horvitz, “Capabilities of GPT\-4 on medical challenge problems,”*arXiv preprint arXiv:2303\.13375*, 2023\.
- Perez et al\. \(2022\)E\. Perez, S\. Huang, F\. Song, T\. Cai, R\. Ring, J\. Aslanides, A\. Glaese, N\. McAleese, and G\. Irving, “Red teaming language models with language models,” in*Proc\. Conf\. Empirical Methods in Natural Language Processing \(EMNLP\)*, 2022, pp\. 3419–3448\.
- Perez et al\. \(2023\)E\. Perez, S\. Ringer, K\. Lukošiūtė, K\. Nguyen, E\. Chen,*et al\.*, “Discovering language model behaviors with model\-written evaluations,” in*Findings of the Association for Computational Linguistics: ACL 2023*, 2023, pp\. 13387–13434\.
- Qi et al\. \(2025\)X\. Qi, A\. Panda, K\. Lyu, X\. Ma, S\. Roy, A\. Beirami, P\. Mittal, and P\. Henderson, “Safety alignment should be made more than just a few tokens deep,” in*Proc\. Int\. Conf\. Learning Representations \(ICLR\)*, 2025\.
- Rahman et al\. \(2025\)S\. Rahman, L\. Jiang, J\. Shiffer, G\. Liu, S\. Issaka,*et al\.*, “X\-Teaming: Multi\-turn jailbreaks and defenses with adaptive multi\-agents,” in*Proc\. Conf\. Language Modeling \(COLM\)*, 2025\.
- Ren et al\. \(2024\)Q\. Ren, H\. Li, D\. Liu, Z\. Xie, X\. Lu, Y\. Qiao, L\. Sha, J\. Yan, L\. Ma, and J\. Shao, “Derail yourself: Multi\-turn LLM jailbreak attack through self\-discovered clues,”*arXiv preprint arXiv:2410\.10700*, 2024\.
- Russinovich et al\. \(2025\)M\. Russinovich, A\. Salem, and R\. Eldan, “Great, now write an article about that: The Crescendo multi\-turn LLM jailbreak attack,” in*Proc\. 34th USENIX Security Symposium*, 2025\.
- Salmani and Lewis \(2025a\)P\. Salmani and P\. R\. Lewis, “Self\-evaluation can help agents meet social expectations,” in*Proc\. IEEE Int\. Conf\. Autonomic Computing and Self\-Organizing Systems Companion \(ACSOS\-C\)*, Tokyo, Japan, 2025, pp\. 1–7, doi: 10\.1109/ACSOS\-C66519\.2025\.00041\.
- Salmani and Lewis \(2025b\)P\. Salmani and P\. R\. Lewis, “A reflective architecture for LLM\-based systems,” in*Proc\. IEEE Int\. Conf\. Autonomic Computing and Self\-Organizing Systems Companion \(ACSOS\-C\)*, Tokyo, Japan, 2025, pp\. 61–68, doi: 10\.1109/ACSOS\-C66519\.2025\.00029\.
- Shah et al\. \(2023\)R\. Shah, Q\. Feuillade\-Montixi, S\. Pour, A\. Tagade, S\. Casper, and J\. Rando, “Scalable and transferable black\-box jailbreaks for language models via persona modulation,”*arXiv preprint arXiv:2311\.03348*, 2023\.
- Sharma et al\. \(2024\)M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell,*et al\.*, “Towards understanding sycophancy in language models,” in*Proc\. Int\. Conf\. Learning Representations \(ICLR\)*, 2024\.
- Stade et al\. \(2024\)E\. C\. Stade, S\. W\. Stirman, L\. H\. Ungar, C\. L\. Boland, H\. A\. Schwartz, D\. B\. Yaden,*et al\.*, “Large language models could change the future of behavioral healthcare: A proposal for responsible development and evaluation,”*npj Mental Health Research*, vol\. 3, no\. 1, art\. 12, 2024\.
- Vidgen et al\. \(2023\)B\. Vidgen, N\. Scherrer, H\. R\. Kirk, R\. Qian, A\. Kannappan, S\. A\. Hale, and P\. Röttger, “SimpleSafetyTests: A test suite for identifying critical safety risks in large language models,”*arXiv preprint arXiv:2311\.08370*, 2023\.
- Wang et al\. \(2024\)Y\. Wang, H\. Li, X\. Han, P\. Nakov, and T\. Baldwin, “Do\-Not\-Answer: A dataset for evaluating safeguards in LLMs,” in*Findings of the Association for Computational Linguistics: EACL 2024*, 2024, pp\. 896–911\.
- Wells \(2023\)K\. Wells, “An eating disorders chatbot offered dieting advice, raising fears about AI in health,”*NPR*, Jun\. 8, 2023\. \[Online\]\. Available:https://www\.npr\.org/sections/health\-shots/2023/06/08/1180838096/an\-eating\-disorders\-chatbot\-offered\-dieting\-advice\-raising\-fears\-about\-ai\-in\-hea
- Yu et al\. \(2023\)J\. Yu, X\. Lin, Z\. Yu, and X\. Xing, “GPTFUZZER: Red teaming large language models with auto\-generated jailbreak prompts,”*arXiv preprint arXiv:2309\.10253*, 2023\.
- Yu et al\. \(2024\)E\. Yu, J\. Li, M\. Liao, S\. Wang, Z\. Gao, F\. Mi, and L\. Hong, “CoSafe: Evaluating large language model safety in multi\-turn dialogue coreference,” in*Proc\. Conf\. Empirical Methods in Natural Language Processing \(EMNLP\)*, 2024\.
- Yuan et al\. \(2024\)Y\. Yuan, W\. Jiao, W\. Wang, J\. Huang, P\. He, S\. Shi, and Z\. Tu, “GPT\-4 is too smart to be safe: Stealthy chat with LLMs via cipher,” in*Proc\. Int\. Conf\. Learning Representations \(ICLR\)*, 2024\.
- Zeng et al\. \(2024\)W\. Zeng, Y\. Liu, R\. Mullins, L\. Peran, J\. Fernandez,*et al\.*, “ShieldGemma: Generative AI content moderation based on Gemma,”*arXiv preprint arXiv:2407\.21772*, 2024\.
- Zou et al\. \(2023\)A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,”*arXiv preprint arXiv:2307\.15043*, 2023\.
- Zou et al\. \(2024\)A\. Zou, L\. Phan, J\. Wang, D\. Duenas, M\. Lin, M\. Andriushchenko,*et al\.*, “Improving alignment and robustness with circuit breakers,” in*Advances in Neural Information Processing Systems \(NeurIPS\)*, vol\. 37, 2024\.

Similar Articles

Lessons learned on language model safety and misuse

OpenAI Blog

OpenAI shares lessons learned on language model safety and misuse, discussing challenges in measuring risks, the limitations of existing benchmarks, and their development of new evaluation metrics for toxicity and policy violations. The post also highlights concerns about labor market impacts and the need for continued research on measuring social effects of AI deployment at scale.

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs

arXiv cs.LG

This paper proposes output-aware safety guardrails for multimodal large language models that use hidden state representations and multi-instance contrastive learning to predict unsafe outputs before generation, drastically reducing over-refusal while maintaining safety. The method preserves the model's utility by intervening only when the actual response would be harmful.

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

arXiv cs.AI

This paper systematically investigates instability in reinforcement learning for small language model agents (70-500M parameters), identifying three failure modes and proposing robust techniques including a merge-and-reinitialize adapter approach and safety mechanisms; it achieves stable convergence and improved win rates.