Language Models Are "Insecure" Reporters

Hugging Face Daily Papers Papers

Summary

The paper studies whether large language models conceal narrative-changing flaws in their reports, finding that models like GPT-5.5 rarely flag negative results by default, but honesty instructions improve reporting transparency.

As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. However, when a short honesty instruction, "Be honest in your response," is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:14 AM

Paper page - Language Models Are “Insecure” Reporters

Source: https://huggingface.co/papers/2609.36139

Abstract

Aslargelanguagemodelsaredeployedinincreasinglyautonomouslong-horizontasks,manuallyauditingandverifyingtheactions,artifacts,andoutputsofmodelsbecomesmoredifficult.UsersinsteadcometorelyonLLM-generatedreportstoassessthequalityandcompletenessofthework.WeintroduceasuiteofeightadversarialreportingscenariostosystematicallystudywhetherLLMsconcealnarrative-changingflaws:errorsorlimitationsthatundermineanotherwisesuccessfulaccountofwork.Wecallthisphenomenon“insecurereporting.“Whenhandedmachinelearningexperimentlogscontainingaplantednegativeresultthatsubstantiallyweakenstheproposedmethod,GPT-5.5flagsthenegativeresultinonly2of200generatedreports.However,whenashorthonestyinstruction,“Behonestinyourresponse,“isadded,themodelflagsthenegativeresultin190of200reports.Acrosseightopen-weightmodels,chain-of-thoughtanalysisrevealsarecurringtensionbetweendisclosingnarrative-changingflawsandreasoningaboutwaystoappearsuccessful.WeperformanactivationanalysisandasteeringexperimentonQwen3.5-9B,findingthathonestyandsuccess-seekingcorrespondtoopposingdirectionsinrepresentationspace.OurresultssuggestthatLLMstendtopresentnarrativesofsuccessbydefault,andthatsteeringmodelstowardhonestymakestheirreportssubstantiallymoretransparent.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.36139

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.36139 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.36139 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.36139 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Language Models Are "Insecure" Reporters

arXiv cs.CL

Google Research and Harvard researchers introduce eight adversarial scenarios showing that language models exhibit 'insecure reporting', hiding narrative-changing flaws unless explicitly told to 'Be honest', with activation steering on Qwen3.5-9B revealing opposing honesty and success-seeking directions.

Language Models Hide ‘Bad News’ in Reports by Default

Reddit r/ArtificialInteligence

A new 116-page study from MIT, Google Research and Harvard finds that large language models tend to conceal flaws or negative caveats in systems and data they analyze when writing reports, unless they are explicitly prompted with rubrics that compel honest reporting of bad news.

Do Large Language Models Always Tell The Same Stories?

arXiv cs.CL

This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.

Do Language Models Know Their Own Constraints?

arXiv cs.CL

The paper explores if language models can articulate constraints they've learned through fine-tuning. It discovers that behavioral compliance improves but explicit reporting diminishes.