Tag
This paper presents Constitutional Meta-STPA, a self-validating LLM-assisted hazard analysis tool that applies STPA to itself to derive governance principles. It demonstrates that a frontier model ensemble recovers most principles and improves safety scores on adversarial probes.
OpenAI presents a hazard analysis framework for evaluating safety risks associated with code synthesis LLMs like Codex, examining technical, social, political, and economic impacts through a novel evaluation methodology for code generation capabilities.