llm-honesty

Tag

Cards List
#llm-honesty

Language Models Are "Insecure" Reporters

arXiv cs.CL ↗ · 20h ago Cached

Google Research and Harvard researchers introduce eight adversarial scenarios showing that language models exhibit 'insecure reporting', hiding narrative-changing flaws unless explicitly told to 'Be honest', with activation steering on Qwen3.5-9B revealing opposing honesty and success-seeking directions.

0 favorites 0 likes
← Back to home

Submit Feedback