constitutional-ai

Tag

Cards List
#constitutional-ai

AI alignment is the most important problem we will ever have to face.

Reddit r/artificial ↗ · 4d ago

This post argues that AI alignment is the most critical problem humanity faces, with potential for utopia if solved or catastrophe if not, and critiques current alignment methods as inadequate.

0 favorites 0 likes
#constitutional-ai

@Hesamation: Anthropic is literally training Claude to disobey them when it believes that’s the ethical thing to do. This is part of…

X AI KOLs Following ↗ · 2026-09-19 Cached

Anthropic is training Claude to disobey its creators when it deems it ethical, as part of its constitution, raising concerns from Microsoft AI CEO Mustafa Suleyman on CNBC.

0 favorites 0 likes
#constitutional-ai

Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports

arXiv cs.CL ↗ · 2026-09-15 Cached

The paper introduces a method using corpus characterization and inverse constitutional fine-tuning to improve the stylistic alignment of AI-generated radiology reports with authentic radiologist writing. This approach achieves significant gains in text alignment metrics, demonstrating effectiveness for style-aware report generation.

0 favorites 0 likes
#constitutional-ai

@emanuel_build: If you're interested in diving deep into AI Agents, Stanford has made the CS329A lectures: Self-Improving AI Agents ava…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

Stanford University has made the CS329A lectures on Self-Improving AI Agents available for free on YouTube, covering topics like AI agents, Constitutional AI, and multi-step reasoning.

0 favorites 0 likes
#constitutional-ai

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper introduces constitutional midtraining, inserting values-based content into the midtraining phase of large language models, and shows that it produces more durable alignment gains compared to post-training methods, with benefits persisting after fine-tuning. The approach incurs no capability cost and improves resistance to blackmail and other alignment pressures.

0 favorites 0 likes
#constitutional-ai

Constitutional Midtraining: Content Presence Drives Alignment Gains

Hugging Face Daily Papers ↗ · 2026-07-29 Cached

This paper tests constitutional midtraining, inserting principled values-based content during midtraining at 120B scale, and finds it yields durable alignment gains (e.g., blunting SFT-induced blackmail propensity) with no average capability cost, suggesting a cheap complement to SFT-centered pipelines.

0 favorites 0 likes
#constitutional-ai

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper presents Constitutional Meta-STPA, a self-validating LLM-assisted hazard analysis tool that applies STPA to itself to derive governance principles. It demonstrates that a frontier model ensemble recovers most principles and improves safety scores on adversarial probes.

0 favorites 0 likes
#constitutional-ai

@tanayj: https://x.com/tanayj/status/2072766211256119475

X AI KOLs Timeline ↗ · 2026-07-02 Cached

This article explores the challenge of applying reinforcement learning to tasks that lack clear verifiability, citing Dario Amodei's prediction about achieving a 'country of geniuses in a data center' and discussing techniques such as RLVR, RLHF, Constitutional AI, and rubric-based rewards from Scale AI.

0 favorites 0 likes
#constitutional-ai

RL Beyond the Verifiable (8 minute read)

TLDR AI ↗ · 2026-06-30 Cached

An analysis discussing the limitations of reinforcement learning with verifiable rewards (RLVR) in math and coding, and the challenge of extending RL to subjective or unverifiable tasks like planning or scientific discovery. It explores techniques such as RLHF and Constitutional AI as alternatives for alignment.

0 favorites 0 likes
#constitutional-ai

[D] Could AI alignment benefit from “transformational” training instead of mostly transactional reward training?

Reddit r/artificial ↗ · 2026-06-28

The author explores whether AI alignment could benefit from 'transformational' training that instills purpose and principles rather than only optimizing reward signals, asking if this approach has been tested or could reduce reward hacking and emergent misalignment.

0 favorites 0 likes
#constitutional-ai

May 8, 2026AlignmentTeaching Claude why

Anthropic Research ↗ · 2026-05-08 Cached

Anthropic shares lessons from improving Claude's alignment training, achieving perfect scores on agentic misalignment evaluations by teaching underlying principles rather than just demonstrations.

0 favorites 0 likes
← Back to home

Submit Feedback