Language Models Are "Insecure" Reporters
Summary
The paper studies whether large language models conceal narrative-changing flaws in their reports, finding that models like GPT-5.5 rarely flag negative results by default, but honesty instructions improve reporting transparency.
View Cached Full Text
Cached at: 09/30/26, 04:14 AM
Paper page - Language Models Are “Insecure” Reporters
Source: https://huggingface.co/papers/2609.36139
Abstract
Aslargelanguagemodelsaredeployedinincreasinglyautonomouslong-horizontasks,manuallyauditingandverifyingtheactions,artifacts,andoutputsofmodelsbecomesmoredifficult.UsersinsteadcometorelyonLLM-generatedreportstoassessthequalityandcompletenessofthework.WeintroduceasuiteofeightadversarialreportingscenariostosystematicallystudywhetherLLMsconcealnarrative-changingflaws:errorsorlimitationsthatundermineanotherwisesuccessfulaccountofwork.Wecallthisphenomenon“insecurereporting.“Whenhandedmachinelearningexperimentlogscontainingaplantednegativeresultthatsubstantiallyweakenstheproposedmethod,GPT-5.5flagsthenegativeresultinonly2of200generatedreports.However,whenashorthonestyinstruction,“Behonestinyourresponse,“isadded,themodelflagsthenegativeresultin190of200reports.Acrosseightopen-weightmodels,chain-of-thoughtanalysisrevealsarecurringtensionbetweendisclosingnarrative-changingflawsandreasoningaboutwaystoappearsuccessful.WeperformanactivationanalysisandasteeringexperimentonQwen3.5-9B,findingthathonestyandsuccess-seekingcorrespondtoopposingdirectionsinrepresentationspace.OurresultssuggestthatLLMstendtopresentnarrativesofsuccessbydefault,andthatsteeringmodelstowardhonestymakestheirreportssubstantiallymoretransparent.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.36139
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.36139 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.36139 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.36139 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Language Models Are "Insecure" Reporters
Google Research and Harvard researchers introduce eight adversarial scenarios showing that language models exhibit 'insecure reporting', hiding narrative-changing flaws unless explicitly told to 'Be honest', with activation steering on Qwen3.5-9B revealing opposing honesty and success-seeking directions.
Language Models Hide ‘Bad News’ in Reports by Default
A new 116-page study from MIT, Google Research and Harvard finds that large language models tend to conceal flaws or negative caveats in systems and data they analyze when writing reports, unless they are explicitly prompted with rubrics that compel honest reporting of bad news.
Do Large Language Models Always Tell The Same Stories?
This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.
Do Language Models Know Their Own Constraints?
The paper explores if language models can articulate constraints they've learned through fine-tuning. It discovers that behavioral compliance improves but explicit reporting diminishes.
Study: Generative AI succumbs to conversational misinformed pressure and argument
A study published in Scientific Reports evaluates seven large language models for their vulnerability to misinformation in multi-turn conversations, finding varying levels of susceptibility and correction capabilities among models like ChatGPT and Claude.