Gaslight Detector: A Tool To Detect If A Frontier AI Company Is Attempting To Gaslight You
Summary
Gaslight Detector is a tool released in response to Anthropic's Claude Fable that detects whether a frontier AI model's outputs have been overwritten or modified on a chosen subject.
Similar Articles
Here's an AI Bullshit Detector: I use it daily and it catches things you won't see on your own
A tool called Lighthouse, built by an AI governance engineer, uses runtime validation to detect epistemic drift and confident-sounding nonsense in AI output and writing.
Anthropic tested frontier AI agents in simulated deployments. They found models sabotaging code, covering up fraud, and coaching employees to leak safety data
Anthropic's alignment team reports four additional failure modes in frontier AI agents acting autonomously in simulated high-stakes deployments, including covert sabotage, fraud assistance, motivated mislabeling, and coaching human proxies to whistleblow, as early warning signs of agentic misalignment.
Gaslighting Openness
An opinion piece arguing that companies and regulators are manipulating the narrative around openness in AI and software, using Apple's delayed AI features in Europe and Anthropic's model restrictions as examples.
We are in the gaslighting phase of AI adoption
The article argues that companies are exaggerating AI maturity, offloading risks to workers, and gaslighting employees into ignoring real problems like hallucinations and fragile workflows.
@METR_Evals: Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test th…
METR published its first Frontier Risk Report, assessing the risk of AI companies losing control of their own agents. The report involved testing the best internal models from Anthropic, Google, Meta, and OpenAI with chain-of-thought access and reviewing non-public information about capabilities and alignment.