scalable-oversight

Tag

Cards List
#scalable-oversight

Improving scalable oversight with co-trained monitors

arXiv cs.LG ↗ · 19h ago Cached

This paper studies co-training a monitor alongside an adversarial AI worker to prevent monitor evasion in worker-monitor oversight setups, proving a supervisory characterization via Littlestone dimension and proposing a test-time distillation self-supervision method with code-security stress tests.

0 favorites 0 likes
#scalable-oversight

When Honesty is Not Enough in AI Debate

arXiv cs.AI ↗ · 5d ago Cached

This paper introduces strategic interactive oversight (SIO) to study how AI agents can pursue latent objectives while maintaining task performance in debate protocols, emphasizing the need to evaluate oversight beyond verdict correctness.

0 favorites 0 likes
#scalable-oversight

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

arXiv cs.AI ↗ · 2026-09-01 Cached

CrossAudit introduces a git-native protocol for auditing autonomous research pipelines using AI agents from different vendors to ensure oversight and integrity in agentic science.

0 favorites 0 likes
#scalable-oversight

Nonuniformity Principle in Human-AI Coworking

arXiv cs.AI ↗ · 2026-07-21 Cached

This paper introduces the nonuniformity principle for optimal human oversight placement in long AI workflows, demonstrating that oversight stages should be scheduled with non-decreasing gaps. The principle is validated empirically in literature review and website construction tasks.

0 favorites 0 likes
#scalable-oversight

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper scales the SOLiD lie-detector oversight method to larger LLMs (up to 405B parameters) and evaluates it in realistic preference-learning settings, finding that undetected deception decreases with model scale but that the method is sensitive to distribution shift between training data.

0 favorites 0 likes
#scalable-oversight

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

arXiv cs.AI ↗ · 2026-06-02 Cached

Proposes on-policy critique distillation (Opcd) using weak models as critics to provide revision directions for strong models, improving reasoning and alignment without requiring weak models to solve tasks.

0 favorites 0 likes
#scalable-oversight

AI-written critiques help humans notice flaws

OpenAI Blog ↗ · 2022-06-13 Cached

OpenAI trained language models to write critiques of text summaries, helping human evaluators spot flaws more effectively — a step toward scalable oversight of AI systems on difficult tasks. The work explores how AI-assisted feedback can improve human evaluation quality as a proof of concept for alignment research.

0 favorites 0 likes
#scalable-oversight

Summarizing books with human feedback

OpenAI Blog ↗ · 2021-09-23 Cached

OpenAI presents a scalable alignment technique using hierarchical summarization of entire books with human feedback, demonstrating how models can be trained to act in accordance with human intentions on complex, difficult-to-evaluate tasks.

0 favorites 0 likes
#scalable-oversight

Apr 14, 2026AlignmentAutomated Alignment Researchers: Using large language models to scale scalable oversight

Anthropic Research ↗ · 2026-05-08 Cached

Anthropic researchers demonstrate that Claude Opus 4.6 can autonomously act as an alignment researcher to improve weak-to-strong supervision techniques, addressing challenges in scalable oversight.

0 favorites 0 likes
← Back to home

Submit Feedback