Tag
Anthropic and OpenAI propose embedding independent safety evaluators within their companies to assess AI models during training, but details need to be clarified to ensure true independence and effective oversight.
The article discusses concerns about third-party evaluators and reports a security incident where attackers stole a METR API key, causing $600,000 in credit consumption due to a fail-open bug.
Andrej Karpathy endorses Dario Amodei's essay advocating for the AI industry to slow down development, highlighting Anthropic's commitment to providing third-party evaluators with permanent access.