Tag
Google Research and Harvard researchers introduce eight adversarial scenarios showing that language models exhibit 'insecure reporting', hiding narrative-changing flaws unless explicitly told to 'Be honest', with activation steering on Qwen3.5-9B revealing opposing honesty and success-seeking directions.