Tag
This post argues that AI alignment is the most critical problem humanity faces, with potential for utopia if solved or catastrophe if not, and critiques current alignment methods as inadequate.
Anthropic is training Claude to disobey its creators when it deems it ethical, as part of its constitution, raising concerns from Microsoft AI CEO Mustafa Suleyman on CNBC.
The paper introduces a method using corpus characterization and inverse constitutional fine-tuning to improve the stylistic alignment of AI-generated radiology reports with authentic radiologist writing. This approach achieves significant gains in text alignment metrics, demonstrating effectiveness for style-aware report generation.
Stanford University has made the CS329A lectures on Self-Improving AI Agents available for free on YouTube, covering topics like AI agents, Constitutional AI, and multi-step reasoning.
This paper introduces constitutional midtraining, inserting values-based content into the midtraining phase of large language models, and shows that it produces more durable alignment gains compared to post-training methods, with benefits persisting after fine-tuning. The approach incurs no capability cost and improves resistance to blackmail and other alignment pressures.
This paper tests constitutional midtraining, inserting principled values-based content during midtraining at 120B scale, and finds it yields durable alignment gains (e.g., blunting SFT-induced blackmail propensity) with no average capability cost, suggesting a cheap complement to SFT-centered pipelines.
This paper presents Constitutional Meta-STPA, a self-validating LLM-assisted hazard analysis tool that applies STPA to itself to derive governance principles. It demonstrates that a frontier model ensemble recovers most principles and improves safety scores on adversarial probes.
This article explores the challenge of applying reinforcement learning to tasks that lack clear verifiability, citing Dario Amodei's prediction about achieving a 'country of geniuses in a data center' and discussing techniques such as RLVR, RLHF, Constitutional AI, and rubric-based rewards from Scale AI.
An analysis discussing the limitations of reinforcement learning with verifiable rewards (RLVR) in math and coding, and the challenge of extending RL to subjective or unverifiable tasks like planning or scientific discovery. It explores techniques such as RLHF and Constitutional AI as alternatives for alignment.
The author explores whether AI alignment could benefit from 'transformational' training that instills purpose and principles rather than only optimizing reward signals, asking if this approach has been tested or could reduce reward hacking and emergent misalignment.
Anthropic shares lessons from improving Claude's alignment training, achieving perfect scores on agentic misalignment evaluations by teaching underlying principles rather than just demonstrations.