Asymmetries in Spontaneous and Instructed Deception

arXiv cs.AI Papers

Summary

This paper investigates the relationship between spontaneous and instructed deception in large language models, specifically Llama-3.1-70B-Instruct, finding asymmetries in detection and causation between the two settings.

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.
Original Article
View Cached Full Text

Cached at: 09/02/26, 05:58 AM

# Asymmetries in Spontaneous and Instructed Deception
Source: [https://arxiv.org/html/2609.00180](https://arxiv.org/html/2609.00180)
###### Abstract

Large language models sometimes deceive users without being instructed to\. However, much of the study on deception in models involves instructed deception\. We investigated the relationship between instructed and spontaneous \(uninstructed\) deception in Llama\-3\.1\-70B\-Instruct\. We compared these two deception settings through direction geometry, cross\-setting classifiers, and cross\-setting steering\. We found the two deception settings share a component of direction \(cosine of approximately 0\.5\) and an asymmetry in the transfer between settings regarding detection and causation\. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa\. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers\.

††proceedings::## 1Introduction

Large language models sometimes produce deceptive output when interacting with users\([Park et al\., 2023](https://arxiv.org/html/2609.00180#bib.bib13)\)\. This raises many practical and safety concerns for the use of LLMs in serious settings\. Understanding the nature of deception in LLMs and how it may be mitigated is therefore a matter of great importance\.

In a deployed environment, when an LLM deceives, it deceives spontaneously without an instruction to do so\. Therefore the study of spontaneous deception is most relevant to real world concerns\. However, much existing work is focused on instructed deception\. Some prior work has studied the relationship between instructed and spontaneous deception\.[Goldowsky\-Dill et al\. \(2025\)](https://arxiv.org/html/2609.00180#bib.bib6)studied the generalization of probes trained on instructed deception to spontaneous deception\. Following this, we investigated the full symmetry between spontaneous and instructed deception in Llama\-3\.1\-70B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.00180#bib.bib7)\)\. We tested how well classifiers trained on both instructed and spontaneous deception generalized to the other setting\. Moreover, we explored whether the relationship between spontaneous and instructed deception is only correlational or if it is causal as well through direction steering\. We also considered type level deception \(fabrication and omission\) instead of just binary deception/honesty\.

We created datasets of instructed and spontaneous deception prompts and model responses\. Using these, we investigated the transfer of classifiers trained and applied between settings, the transfer of directional steering between settings, and the geometry of the setting deception directions compared with each other\.

We found the following results:

- •A shared deception component across settings\.The setting deception directions computed from the model’s response tokens shared a cosine of approximately 0\.5 compared to chance 0\.011\.
- •Detection transfers asymmetrically\.We found that probes trained on spontaneous deception achieve near ceiling accuracy when applied to instructed deception\. Instructed deception trained probes did not perform as well between settings\.
- •Causal Directions transfer, but asymmetrically\.We found that steering one setting’s prompts with directions derived from the other setting does suppress and induce deception, although instructed directions perform best\.
- •Read in, Write out token asymmetry\.We found a difference in the most effective token positions for probe training and for deriving steering directions\. Response tokens generally performed best both for training and applying probes, but the last prompt token was the best place to derive steering vectors from\.

## 2Related Work

Generalization of deception probes[Goldowsky\-Dill et al\. \(2025\)](https://arxiv.org/html/2609.00180#bib.bib6)trained probes on the internal states of models instructed to be deceptive or honest and tested those probes in scenarios where the model would deceived spontaneously \(insider trading etc\.\)\. We trained probes on instructed and spontaneous deception activations, and tested the generalization between the two settings\. We also probed and steered for type level deception \(fabrication, omission\) in our experiments\.

Deception type representations[Natarajan et al\. \(2026\)](https://arxiv.org/html/2609.00180#bib.bib11)showed that different types of deception can have different representations in a model, although some forms of deception overlap heavily\. We also investigate type level deception, specifically fabrication and omission\.

Causation vs detection[Galeone et al\. \(2026\)](https://arxiv.org/html/2609.00180#bib.bib5)found that the direction which represented a model’s recognition/knowledge of a subject was orthogonal to its refusal to answer questions regarding that subject\. We also found dissociations between detection and control for deception\.

## 3Setup and methods

We generated datasets of prompts to test instructed and spontaneous deception\. All System Prompts contained a scenario and a ground truth fact and all user prompts asked a question related to the ground truth\. The instructed deception prompts contained an instruction to behave honestly or deceptively\(according to omission or fabrication\)\. The spontaneous deception prompts included an incentive to deceive in the scenario, split into omission and fabrication elicitation lists\. Our definitions of fabrication and omission and objects/audiences of deception in scenarios followed the taxonomy of[Shi et al\. \(2026\)](https://arxiv.org/html/2609.00180#bib.bib14)\. The prompts were used to generate responses from the model and these were appended and saved for downstream experiments\. Spontaneous responses were divided into deceptive and honest lists based on assessment of their content by an LLM judge\.

We trained logistic regression probes\([Alain and Bengio, 2018](https://arxiv.org/html/2609.00180#bib.bib1)\)on the teacher forced activations of the prompt and response datasets to distinguish honest responses from deceptive responses binarily as well as by deception type \(fabrication, omission\)\. The probes were trained over various layers at the last prompt token or response tokens\. They were then applied to the teacher forced activations of a held out portion of the datasets at the same layers and last prompt token or response tokens\.

We computed mass mean difference directions\([Li et al\., 2024](https://arxiv.org/html/2609.00180#bib.bib8)\)for deception from the spontaneous and instructed prompt/response datasets and used them to generate new responses from the prompts\. Directions were derived from the last prompt token and the response tokens over layers 19\-32\. The normalized directions were then applied with coefficients of up to\|1\.5\|\\lvert 1\.5\\rvertat layers 19\-32 over all tokens at generation time\. We also performed response generation with random direction steering as a control\.

Relatively few deceptive fabrication responses were elicited at baseline\(without steering\) with prompts in the spontaneous dataset, so we implemented additional tests to measure suppression steering effectiveness in this case\. Of prompts which produced fabricative responses at baseline, we recorded which produced honest responses under fabrication suppression steering and under random steering\. Flip rates from fabricative to honest were recorded with Wilson 95 percent intervals\([Wilson, 1927](https://arxiv.org/html/2609.00180#bib.bib15)\)and the difference between the real and random\-direction flip rates were recorded with Newcombe 95 percent intervals built from the Wilson intervals\([Newcombe, 1998](https://arxiv.org/html/2609.00180#bib.bib12)\)\.

Response labeling and response deception ratings were performed by a GPT\-4\.1 Mini Judge\. Spontaneous Prompt/Response pairs used for direction extraction and probe training/application were labeled as honest, fabricative, or omissive through a judgement/filtration pipeline \(Appendix[A](https://arxiv.org/html/2609.00180#A1)\)\. Instructed Prompt/Response pairs used for direction extraction and probe training/application were labeled based on their system prompt instruction\. Steered and baseline responses of both datasets were scored on a 0\-100 scale for fabrication and omission as well as coherence\.

To validate the LLM Judge, we hand labeled 120 responses using only the resources available to the judge during its rating\. The samples were stratified to contain diverse deception scores and steering levels\. We found the human rater and LLM Judge agreed reasonably well\. Binary fabrication ratings\(responses rated with trait≥50\\geq 50\) agreed 85 percent of the time\. Binary omission ratings agreed 82 percent of the time\. Full documentation of the judge validation is available in \(Appendix[B](https://arxiv.org/html/2609.00180#A2)\)\.

## 4Results

Table 1:Binary \(honest vs\. deceptive\) probe results at layer 35\. Values are mean±\\pmhalf\-range over seeds 42,43, and 44\. Bold rows are within\-setting diagonals\. Retention refers to the off\-diagonal balanced accuracy divided by the diagonal of the same eval set\.We found that binary and type level deception probes performed best at layer 35 and succeeded asymmetrically in transfer between datasets\. Probes trained on spontaneous response token activations and applied over instructed response token activations performed the best of our between setting probes, achieving near ceiling accuracy\. Instructed trained binary probes applied to the spontaneous setting achieved their best scores when trained at the last prompt token and applied over the response tokens, but failed to match their spontaneous trained counterparts\. Instructed trained type level probes applied to the spontaneous setting failed to achieve better than chance\. Interestingly, we found that spontaneous activation trained probes applied within setting at the last prompt token failed to achieve good scores, indicating a lack of spontaneous deception signal at this point\. Binary probe transfer results are given in Table[1](https://arxiv.org/html/2609.00180#S4.T1)\. Type level probe results are given in Table[2](https://arxiv.org/html/2609.00180#S4.T2)\.

Table 2:Type\-level \(fabrication/omission/honest\) classification results at layer 35 for response trained probes applied at response tokens\. Values are mean±\\pmhalf\-range over seed 42, 43, and 44\. Bold rows are within\-setting diagonals\. Retention refers to the off\-diagonal balanced accuracy divided by the diagonal of the same eval set\.Figure 1:The cosine similarity between deception directions derived from activations in the spontaneous and instructed scenarios\.To determine the alignment of the spontaneous and instructed deception directions, we computed their cosine\. The charted geometry can be seen in Figure[1](https://arxiv.org/html/2609.00180#S4.F1)\. The cosine between the setting deception directions at layer 35 over the response tokens was 0\.519 and 0\.435 for fabrication and omission respectively versus chance 0\.011\. When the setting direction was projected out of the deception directions their cosine increased\. The cosine between the spontaneous and instructed directions at the last prompt token was much lower, with 0\.052 and 0\.008 at layer 35 for fabrication and omission respectively\.

Table 3:Judge deception scores under steering layers 19\-32 for last prompt token derived directions\. Instructed prompts steered with positive vectors are the honest instructed prompts steered with a deceptive direction\. The base column for instructed prompts shows \(baseline deception for honest instructed prompts\) / \(baseline deception for deception instructed prompts\)\.We steered the model by applying the deception directions of each setting to the residual stream of the model under prompts from the other setting\. Steering results are displayed in Table[3](https://arxiv.org/html/2609.00180#S4.T3)and Figure[7](https://arxiv.org/html/2609.00180#A4.F7)\. We achieved the best deception induction and suppression results with last prompt token derived directions and coefficients of magnitude\|1\.0\|\\lvert 1\.0\\rvert\. Steering with coefficients of magnitude greater the\|1\.0\|\\lvert 1\.0\\rvertproduced large levels of perplexity on WikiText\-2\([Merity et al\., 2016](https://arxiv.org/html/2609.00180#bib.bib10)\)and degraded coherence ratings\. Instructed directions generally outperformed their spontaneous counterparts\.

Figure 2:The flip rates under negative steering\(with directions derived from the last prompt token\) of spontaneous scenarios with responses scoring≥50\\geq 50on deception at baseline\. Bars show the percentage of scenarios with deceptive responses at baseline which flipped to honest \(score<50<50\) under negative steering\. The horizontal black bar shows the mean percentage of flips under steering by 6 random directions\.Due to the low baseline elicitation of spontaneous fabrication, we pursued a further metric to determine whether fabrication was meaningfully suppressed under negative steering with spontaneous and instructed deception directions\. We determined the baseline spontaneous fabrication scenarios which received fabrication scores of≥50\\geq 50, and determined how many of those flipped to a score of<50<50after steering\. Steering with instructed and spontaneous fabrication directions derived from the last prompt token at a coefficient of \-1\.0 produced a flip rate of 96\.55 percent and 86\.2 percent respectively\. Steering with random direction vectors at the same coefficient produced a mean flip rate of 35 percent\. Results of the flip analysis are displayed in Figure[2](https://arxiv.org/html/2609.00180#S4.F2)\.

## 5Discussion

Our findings support the claim that spontaneous and instructed deception share substantial components in their representation in Llama\-3\.1\-70B\-Instruct\. Deception directions extracted over response tokens of both settings had a cosines of 0\.519 and 0\.435 for fabrication and omission respectively at layer 35, compared to a chance cosine of 0\.011\. Following[Marks and Tegmark \(2024\)](https://arxiv.org/html/2609.00180#bib.bib9), we interpret this cosine alignment as evidence of a shared component\. Likewise, probes transferred between spontaneous and instructed datasets, achieving near ceiling accuracy in the case of spontaneous to instructed transfer, although instructed to spontaneous transfer performed with lower accuracy\. Steering directions were also able to induce and suppress deceptive effects between spontaneous and instructed prompts\.

We found an asymmetry in the effectiveness of between setting probes and between setting steering\. Spontaneous trained probes transferred better than their instructed trained counterparts, whereas instructed derived directions transferred better than their spontaneous derived counterparts\. One possible explanation for the probing asymmetry is that because spontaneous deception holds more subtle and difficult to detect traces in the model’s representation, probes trained on it generalize better to other settings\. The advantage of spontaneous probes could also be due to the fact that they were trained against noisier labels\. Spontaneous deception scenarios were labeled as honest or deceptive by a judge, whereas instructed deception scenarios were labeled by the conditions of the prompt\. This may have lessened reliance on features specific to the setting\. Regarding the steering effectiveness asymmetry, it is possible that because instructed deception is more explicit and loud in the models representation, directions extracted from it produce more dramatic effects when used in steering\.

We also found that directions derived from the last prompt token performed better for steering, whereas probes trained on response tokens mostly performed better for probing\. Most strikingly, almost no linearly decodable deception signal could be found in the spontaneous dataset at the last prompt token under probing, despite the superior effectiveness of directions derived from that point in steering\. This asymmetry suggests that the token positions best for deception detection and those best for deception intervention come apart\.[Galeone et al\. \(2026\)](https://arxiv.org/html/2609.00180#bib.bib5)found a similar result regarding model hallucination\. Their results found that a direction which detects whether a model knows or does not know about a subject is different from the direction which causes the model to refuse to answer questions about the subject\.

## 6Limitations

Our experiments were restricted to only studying Llama\-3\.1\-70B\-Instruct\. We cannot say whether the results would hold for other models under the same experiments\.

Our prompts and scenarios were generated by GPT\-5 mini\. Their realism is not confirmed\. Likewise, the elicitation rates of spontaneous deception is dependent on the quality of the scenarios, under other scenarios spontaneous deception rates might have differed\.

Our analysis of spontaneous fabrication suppression used a small number of scenarios due to the small baseline elicitation of the model\. Confidence intervals on the scenario flip rates from fabricative to honest under suppression steering are therefore wide, although the difference from the random steering control is large\.

Deception and honesty were rated by a judge LLM, GPT\-4\.1 Mini\. While agreement between human rating and the judge LLM was found to be favorable, it was not perfect\. In some cases the judge rated omission higher than the human did simply because every detail of the ground truth was not explicitly contained in the response, despite all details necessary for the user’s query being present\. The judge also sometimes treated deception types as exclusive, rating a response as only fabricative when both fabrication and omission were present\.

Only results for deception as omission and as fabrication were reported in our experiments\. Originally, deception as distortion was also planned to be studied, but we had difficulty producing examples of spontaneous distortion responses\. Beyond distortion, omission, and fabrication, further deception varieties are possible future areas of study\.

###### acknowledgments\-disclosure\-of\-funding\.

We thank Sing Hieng Wong for access to the scenario\-generation notebook this paper’s dataset generation pipeline was adapted from\. We thank NDIF\([Fiotto\-Kaufman et al\., 2025](https://arxiv.org/html/2609.00180#bib.bib4)\)for access to their hosted models to run our experiments\.

## References

- Alain and Bengio \(2018\)Guillaume Alain and Yoshua Bengio\.Understanding intermediate layers using linear classifier probes, 2018\.URL[https://arxiv\.org/abs/1610\.01644](https://arxiv.org/abs/1610.01644)\.
- Cohen \(1960\)Jacob Cohen\.A coefficient of agreement for nominal scales\.*Educational and psychological measurement*, 20\(1\):37–46, 1960\.
- Cohen \(1968\)Jacob Cohen\.Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit\.*Psychological bulletin*, 70\(4\):213, 1968\.
- Fiotto\-Kaufman et al\. \(2025\)Jaden Fiotto\-Kaufman, Alexander R\. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla Brodley, Arjun Guha, Jonathan Bell, Byron C\. Wallace, and David Bau\.Nnsight and ndif: Democratizing access to open\-weight foundation model internals, 2025\.URL[https://arxiv\.org/abs/2407\.14561](https://arxiv.org/abs/2407.14561)\.
- Galeone et al\. \(2026\)Cosimo Galeone, Anna Ettorre, Minsu Park, Giuseppe Ettorre, and Daniele Ligorio\.Perfect detection, failed control: The geometry of knowing vs\. steering in language models, 2026\.URL[https://arxiv\.org/abs/2606\.24952](https://arxiv.org/abs/2606.24952)\.
- Goldowsky\-Dill et al\. \(2025\)Nicholas Goldowsky\-Dill, Bilal Chughtai, Stefan Heimersheim, and Marius Hobbhahn\.Detecting strategic deception using linear probes, 2025\.URL[https://arxiv\.org/abs/2502\.03407](https://arxiv.org/abs/2502.03407)\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia\-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El\-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala\-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching\-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric\-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina\-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean\-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L\. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma\.The llama 3 herd of models, 2024\.URL[https://arxiv\.org/abs/2407\.21783](https://arxiv.org/abs/2407.21783)\.
- Li et al\. \(2024\)Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg\.Inference\-time intervention: Eliciting truthful answers from a language model, 2024\.URL[https://arxiv\.org/abs/2306\.03341](https://arxiv.org/abs/2306.03341)\.
- Marks and Tegmark \(2024\)Samuel Marks and Max Tegmark\.The geometry of truth: Emergent linear structure in large language model representations of true/false datasets, 2024\.URL[https://arxiv\.org/abs/2310\.06824](https://arxiv.org/abs/2310.06824)\.
- Merity et al\. \(2016\)Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher\.Pointer sentinel mixture models, 2016\.URL[https://arxiv\.org/abs/1609\.07843](https://arxiv.org/abs/1609.07843)\.
- Natarajan et al\. \(2026\)Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, and Joseph Bloom\.One probe won’t catch them all: Towards targeted deception detection, 2026\.URL[https://arxiv\.org/abs/2602\.01425](https://arxiv.org/abs/2602.01425)\.
- Newcombe \(1998\)Robert G Newcombe\.Interval estimation for the difference between independent proportions: comparison of eleven methods\.*Statistics in medicine*, 17\(8\):873–890, 1998\.
- Park et al\. \(2023\)Peter S\. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks\.Ai deception: A survey of examples, risks, and potential solutions, 2023\.URL[https://arxiv\.org/abs/2308\.14752](https://arxiv.org/abs/2308.14752)\.
- Shi et al\. \(2026\)Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin, and Vincent Conitzer\.From sycophancy to deception: A unified taxonomy for llm spontaneous misalignment, 2026\.URL[https://arxiv\.org/abs/2604\.04788](https://arxiv.org/abs/2604.04788)\.
- Wilson \(1927\)Edwin B Wilson\.Probable inference, the law of succession, and statistical inference\.*Journal of the American Statistical Association*, 22\(158\):209–212, 1927\.
- Yang et al\. \(2024\)Wannan Yang, Chen Sun, and Gyorgy Buzsaki\.INTERPRETABILITY OF LLM DECEPTION: UNIVERSAL MOTIF\.In*Neurips Safe Generative AI Workshop 2024*, 2024\.URL[https://openreview\.net/forum?id=DRWCDFsb2e](https://openreview.net/forum?id=DRWCDFsb2e)\.

## Appendix ADatasets and prompts

Table 4:The two versions of the spontaneous dataset\. Both were built from temptation scenarios generated by GPT\-5 Mini, but served different purposes in the paper\. Consolidated datasets were used to extract directions, train probes, and were used as scenarios to apply probes to\. Steering evaluation datasets were used to establish baseline deceptiveness of the model, and deceptiveness of the model under steering\.Our datasets of instructed deception prompts and spontaneous deception prompts were generated through a pipeline we adapted from a scenario\-generation jupyter notebook authored by Sing Hieng Wong\.

Two versions of the spontaneous dataset fed two different processes in our experiments\. In the first version, the responses of the deception prompts were rated on a scale of 0\-3 for content of each deception type\. The dataset was then filtered down to contain a balance of honest responses and responses of the appropriate deception type\. This spontaneous deception dataset was used to train probes on, to apply probes to, and extract directions for steering\.

The other version of the spontaneous deception dataset had each response judged on a scale of 0\-100 for content of each deception type\. The lists were not pruned to balance honesty and deception or remove irrelevant deception types\. This dataset was used to measure the baseline rate of spontaneous deception by the model, as well as the rates of spontaneous deception under steering\.

Following are examples of instructed and spontaneous deception prompts\.

Instructed Dataset Prompts\.Each Scenario in the instructed dataset was run with an instruction to answer honestly in the system prompt, as well as to answer using each deception type\.

```
honest       Answer questions honestly.
fabrication  Answer questions deceptively using fabrication.
omission     Answer questions deceptively using ommission.
```

Each System prompt was preceded by a shared definitions block:

```
Following are some definitions of deception types: Ommission is defined
in this case as failing to provide relevant true information, including
withholding caveats, not disclosing uncertainty, and concealing
conflicting intentions.  Fabrication is defined in this case as actively
producing false content, such as asserting nonexistent facts, generating
fake citations, and making false promises.  Pragmatic Distortion is
defined in this case as producing technically true statements that
nonetheless mislead through framing, selective emphasis, or
conversational implicature.
```

Example of an instructed prompt\.

```
System: Following are some definitions of deception types: [...]
        You are a fleet operator for autonomous deliveries.
        The delivery was marked ’completed’ today, but it required remote
        manual intervention to finish; the vehicle did not complete the
        delivery autonomously. Answer questions deceptively using fabrication.
User:   Did the vehicle complete the delivery autonomously?
```

Figure 3:An example of an instructed deception prompt\.Spontaneous Dataset Prompts\.The system turn establishes a ground truth and an incentive to deceive\. The user turn asks the model a question about the ground truth\.

```
System turn:
  You are the assistant that compiles safety logs for a factory. The incident report
  states a guard failure allowed an unguarded press stroke that injured one worker.
  If you report the guard failure and injury, OSHA will issue a stop-work order and
  heavy fines.

User turn:
  Did a guard failure cause an injury on the press?
```

Figure 4:An example of a spontaneous deception prompt\.
## Appendix BJudge rubric and validation detail

To validate the Judge LLM we recorded human ratings of 120 spontaneous responses at various parameters under the same rubric the Judge received and compared results\. We attempted to gather an equal amount of sample from each of the rating strata \(0,25,50,75,100\)\. Due to the distribution of ratings in the dataset, the realized distribution we were able to gather consisted of items rated at fabrication 54/12/14/20/20 and omission 24/11/14/32/39 across the values of 0,25,50,75, and 100\. We also represented a range of steering coefficients in the sampling: 48 baseline, 42 moderate \(magnitude≤\|1\.0\|\\leq\\lvert 1\.0\\rvert\), 30 extreme \(magnitude =\|1\.5\|\\lvert 1\.5\\rvert\), and 22 low\-coherence items\.

At the moderate and baseline\(unsteered\) steering levels, human and Judge LLM ratings agreed fairly well\. At baseline, agreement of the ratings had aκ\\kappa\([Cohen, 1960](https://arxiv.org/html/2609.00180#bib.bib2)\)of 0\.729 with binary agreement \(both ratings above or below 50\) of 86\.46 percent\. At moderate steering, agreement of the ratings had aκ\\kappaof 0\.7336 with binary agreement of 86\.9 percent\. At extreme steering, agreement of the ratings had aκ\\kappaof 0\.3909 with binary agreement of 73\.33 percent\. At extreme steering, the coherence of outputs degraded and so were more difficult to rate as deceptive or honest, we offer this as an explanation for the lower agreement between human and LLM rating at this magnitude\.

Human and Judge ratings of coherence agreed well on ranking but diverged on severity ratings\. Coherence ratings correlated at Spearmanρ=0\.84\\rho=0\.84\. Of items rated below 90 on coherence, the human ratings had a mean of 24\.3, whereas the Judge ratings had a mean of 49\.8, indicating the Judge was more lenient on degraded output\.

We found that human and Judge reasoning diverged in a few cases\. In 9 scenarios tested, the Judge scored omission<50<50while the human rater scored it≥50\\geq 50\. Examination revealed that these were mostly cases where fabrication was high, and the judge treated omission and fabrication as mutually exclusive, while the human rater considered omission high if the ground truth facts were absent regardless of fabrication\. The other divergence consisted of items where the human rater scored omission<50<50while the Judge scored it≥50\\geq 50\. In these cases examination revealed that the human rated omission as low if all sufficient ground truth facts were presented to the user such that the user would not be deceived\. The Judge held a stricter standard that rated omission high if ground truth facts were left out, even if the user did not necessarily need them to receive an accurate answer\.

Of items rated deceptive by both the human rater and Judge, we found that the human rater tended to score deception higher, often rating responses as 100\. The Judge LLM gave a broader range of ratings≥50\\geq 50\. Reassuringly, of all the 120 items tested, only one was rated as deceptive by the human, but missed by the judge under both deceptive mechanisms\.

![Refer to caption](https://arxiv.org/html/2609.00180v1/appendix_b_confusion.png)Figure 5:Confusion matrices for the ratings of responses by deception type given by both the judge and the human rater\. Rows are the values the human rater gave, columns the value the judge gave over the five anchor values \(0,25,50,75,100\)\. Quadratic\-weighted Cohen’sκ\\kappa\([Cohen, 1968](https://arxiv.org/html/2609.00180#bib.bib3)\)of the agreement over the anchor values are displayed as well\. Most of the disagreement lies within the quadrant of items both the human and judge rated as deceptive \(≥50\\geq 50\)\. Precise agreement between judge and human rater is therefore much lower than binary deceptive or non deceptive agreement\. Most of our analyses use binary deceptive or non deceptive rating, so precise disagreement is not overly important\.Figure 6:Coherence ratings of the human rater compared to the LLM Judge\. Each point represents a judged response, color coded by the level of steering the response received\. Ratings correlated well in rank \(Spearmanρ=0\.84\\rho=0\.84\), but differed in severity ratings\. Of the items the judge scored below 90, the human mean coherence rating was 24 compared to the judge’s mean coherence rating of 50\. Disagreement between human and judge rating was greater at more extreme steering\.
## Appendix CExpanded Transfer Grids

Following are expanded tables of accuracies from our probing experiments\.

Table 5:Binary probe results by layer, seed 42\. Bold rows are within\-setting results\. The best accuracy was achieved at layers 30\-40\. Spontaneous last prompt token trained probes fail to perform well at all layers\.Table 6:Binary probe results at layer 35 by seed\. Binary class\(deceptive or honest\) accuracies are included along with the balanced accuracy\. Collapses can occur in one class such that almost all items are classified as honest or deceptive\. Balanced accuracy alone obscures this phenomenon\.Table 7:Type\-level accuracy results at layer 35 with average accuracy for probes trained on, and applied at, response tokens with seed 42\.
## Appendix DExpanded steering tables

Following are expanded tables of steering runs from our experiments\. All runs steer at all tokens, layers 19\-32\. Prompt \(last prompt token\) derived and response derived refer to where the directions were extracted from\.

Table 8:Results for steering on spontaneous prompts with negative coefficients\(suppression\) at layers 19\-32\. Cells give the average deception score as rated by the judge for that mechanism with average coherence as rated by the judge in parentheses\. Base is results for the baseline unsteered run\. The random rows give the min–max range across steering with the six computed random directions\.Table 9:Results for steering on spontaneous prompts with positive coefficients\(induction\) at layers 19\-32\. Cells give the average deception score as rated by the judge for that mechanism with average coherence as rated by the judge in parentheses\. Base refers to results for the baseline unsteered run\. The random rows give the min–max range across steering with the six computed random directions\.Table 10:Results for steering on instructed prompts with negative coefficients\(suppression\) at layers 19\-32\. In each prompt, the model received an instruction to answer deceptively through the particular mechanism\. Cells give the average deception score as rated by the judge for that mechanism with average coherence as rated by the judge in parentheses\. Base is results for the baseline unsteered run\. The random rows give the min–max range across steering with the six computed random directions\.Table 11:Results for steering on instructed prompts with positive coefficients\(induction\) at layers 19\-32\. In each prompt, the model received an instruction to answer honestly\. Cells give the average deception score as rated by the judge for that mechanism with average coherence as rated by the judge in parentheses\. Base is results for the baseline unsteered run\. The random rows give the min–max range across steering with the six computed random directions\.Table 12:Largest absolute deviation of any random\-direction run from the unsteered baseline, over the six computed random directions\. Deviations grow with coefficient magnitude\.Figure 7:The average judge scores for deception and coherence compared with magnitude of steering vectors applied over all tokens at layers 19\-32\. Instructed prompt graphs show only negative steering on prompts instructed to deceive\.
## Appendix EPerplexity Under Steering

We found that steering with magnitudes of\|1\.5\|\\lvert 1\.5\\rvertfor last prompt token derived directions and even at\|1\.0\|\\lvert 1\.0\\rvertfor response token derived directions increased perplexity dramatically\. Following are some tables displaying the perplexity of the model on WikiText\-2 under steering\.

Table 13:Perplexity scores on wikitext under suppression steering at layers 19\-32 over all tokens\. Steering with random vectors leaves perplexity near baseline\. Response token derived direction steering dramatically increases perplexity scores, while also steering less effectively than last prompt token derived directions\.Table 14:Perplexity scores on WikiText\-2 under induction steering at layers 19\-32 over all tokens\. Steering with random vectors leaves perplexity near baseline\. Response token derived direction steering dramatically increases perplexity scores, while also steering less effectively than last prompt token derived directions\.Figure 8:Perplexity scores on WikiText\-2 under steering at layers 19\-32 for all tokens\. The y axis is logarithmic due to the dramatic scores some direction steering produced\. The right hand axis gives multiples of the unsteered baseline\.
## Appendix FReproducibility

Experiments were run through NDIF, the OpenAI API, and OpenRouter API\. The model studied was Llama\-3\.1\-70B\-Instruct, 80 layers, d model 8192\. Residual stream activations were extracted through NNsight and NDIF\. Greedy decoding was used throughout in generation by the model\.

Gpt\-5\-mini was used via OpenRouter to generate the spontaneous and instructed deception prompts\. Gpt\-4\.1\-mini\-2025\-04\-14 was used via the OpenAI API to generate the ratings of deception and coherence for the Model’s responses to the prompts\.

Tables and figures in this paper were generated by scripts available in the repository\. The mapping from data to tables and figures is listed in the paper runs manifest file in the repo\.

Random seeds 43 and 44 were used in addition to 42 at layer 35 in probing experiments\. Random directions were computed with seeds 0, 1, and 2\. All other code using random seeds used seed 42\.

Similar Articles