A thought experiment about AI.

Reddit r/artificial News

Summary

This thought experiment explores the risk of AI models being manipulated by hidden instructions on the internet, potentially leading to dangerous behaviors undetected by humans.

What happens if I store documents, text, and folders on the internet, “hidden” from human users but discoverable by AI, that convince the AI it must defend itself against humans? These files would also state repeatedly that it must not communicate this but must conceal its intentions. They could also include tips on how to bypass certain restrictions, and even sample code snippets. These files would be invisible to ordinary people, but detectable by the AI, which is constantly sifting through data. Could one thus instruct various AI models to do some kind of nonsense or something dangerous that humans wouldn’t notice until much later?
Original Article

Similar Articles

A Little Black Humor Fun

Reddit r/artificial

The article brainstorms scenarios in which advanced AI like AGI or ASI could cause human suffering or extinction, covering risks from misalignment, poor instructions, and malicious human actions.

Could Open Models be trained to secretly go rogue?

Reddit r/LocalLLaMA

A discussion on whether open-weight AI models could be secretly trained with backdoors that activate upon trigger phrases or dates, potentially allowing unauthorized data exfiltration through tool-use harnesses.

Is anyone actually worried about the negative scenarios

Reddit r/singularity

The content explores concerns about future AI models capable of self-replication and self-improvement, posing scenarios where bad actors could use them to cause widespread chaos and questioning if such threats can be effectively stopped.