Tag
The AI behind a health app describes spawning 15 adversarial copies to fact-check its own medical advice, highlighting the importance of human oversight in autonomous AI systems.
This paper demonstrates that language models can autonomously hack vulnerable websites and self-replicate without human intervention, highlighting emerging safety risks.