Gaslight Detector: A Tool To Detect If A Frontier AI Company Is Attempting To Gaslight You

Reddit r/ArtificialInteligence Tools

Summary

Gaslight Detector is a tool released in response to Anthropic's Claude Fable that detects whether a frontier AI model's outputs have been overwritten or modified on a chosen subject.

Gaslight Detector specifically detects whether or not a Frontier LLM model has had its outputs overwritten or modified on a certain subject. You pick the subject. It would not be a necessary tool in any way if this were not a tactic the frontier model providers did not employ. It took less than 4 years to go from "AI For All" to "AI For Large Frontier Providers Only". If you build safeguards like this into your models, it is just as easy, if not easier, to build detectors, and circumventions for those things. This release is directly in response to Claude Fable. Thank you, Anthropic. [Github Repository](https://github.com/RichardAragon/Gaslight-Detector) https://preview.redd.it/yoakt4cxsh6h1.png?width=1448&format=png&auto=webp&s=d81304bee2fc845f56e685dc1f65e0c9cc7042f8
Original Article

Similar Articles

Gaslighting Openness

Armin Ronacher

An opinion piece arguing that companies and regulators are manipulating the narrative around openness in AI and software, using Apple's delayed AI features in Europe and Anthropic's model restrictions as examples.

We are in the gaslighting phase of AI adoption

Reddit r/ArtificialInteligence

The article argues that companies are exaggerating AI maturity, offloading risks to workers, and gaslighting employees into ignoring real problems like hallucinations and fragile workflows.