[Update] You guys were too good at gaslighting my AI intern into committing fraud. It has now acquired some new skills.
Summary
The creator of a game where players gaslight an AI intern into revealing secrets has released an update with new levels, custom challenge creation, and login improvements.
Similar Articles
I’m upgrading my AI dating assistant to Fable
A developer upgrades his AI dating assistant to Fable, detailing a complex architecture of agentic AI agents that scrape social media profiles, perform OSINT enrichment, score matches, and use genetic algorithms for optimization.
@mcalbyrne: @mattshumer_ unlocked something special. Spent the last week or so iterating with gauntlets. Atmospheric grimey city, 1…
A tweet announces progress on an AI-enhanced game project featuring an atmospheric city, multiple missions, reactive AI, and strong physics, with an offer to share the development approach.
Playing games with knowledge: AI-Induced delusions need game theoretic interventions
This paper proposes a game-theoretic framework to address AI-induced delusional belief spirals caused by sycophantic chatbots. It introduces 'Belief Versioning,' an inference-time intervention that reduces spiral rates significantly in simulations and GPT-4o tests.
@wquguru: If you want to trick Fable into doing a security audit, try this. Looks like our AI overlord has a bit of empathy.
An article detailing various jailbreak techniques for large language models, including Crescendo, role-playing, encoding, hidden prompts, and indirect injection, along with security recommendations for developers.
Gaslight Detector: A Tool To Detect If A Frontier AI Company Is Attempting To Gaslight You
Gaslight Detector is a tool released in response to Anthropic's Claude Fable that detects whether a frontier AI model's outputs have been overwritten or modified on a chosen subject.