Tag
This paper identifies a 'cold-start safety gap' in tool-calling LLM agents, where they are most vulnerable at the beginning of a session and become safer after completing regular agentic tasks. The authors introduce the SODA benchmark to evaluate this phenomenon and recommend a simple deployment strategy of warming up agents with regular tasks before safety-critical requests.
AI industry leaders call for human-centered technology deployment, acknowledging societal anxiety about rapid change, and emphasize the importance of safeguarding human autonomy and economic future, while reflecting on failures in industry communication.