Tag
This paper studies co-training a monitor alongside an adversarial AI worker to prevent monitor evasion in worker-monitor oversight setups, proving a supervisory characterization via Littlestone dimension and proposing a test-time distillation self-supervision method with code-security stress tests.
This paper introduces strategic interactive oversight (SIO) to study how AI agents can pursue latent objectives while maintaining task performance in debate protocols, emphasizing the need to evaluate oversight beyond verdict correctness.
CrossAudit introduces a git-native protocol for auditing autonomous research pipelines using AI agents from different vendors to ensure oversight and integrity in agentic science.
This paper introduces the nonuniformity principle for optimal human oversight placement in long AI workflows, demonstrating that oversight stages should be scheduled with non-decreasing gaps. The principle is validated empirically in literature review and website construction tasks.
This paper scales the SOLiD lie-detector oversight method to larger LLMs (up to 405B parameters) and evaluates it in realistic preference-learning settings, finding that undetected deception decreases with model scale but that the method is sensitive to distribution shift between training data.
Proposes on-policy critique distillation (Opcd) using weak models as critics to provide revision directions for strong models, improving reasoning and alignment without requiring weak models to solve tasks.
OpenAI trained language models to write critiques of text summaries, helping human evaluators spot flaws more effectively — a step toward scalable oversight of AI systems on difficult tasks. The work explores how AI-assisted feedback can improve human evaluation quality as a proof of concept for alignment research.
OpenAI presents a scalable alignment technique using hierarchical summarization of entire books with human feedback, demonstrating how models can be trained to act in accordance with human intentions on complex, difficult-to-evaluate tasks.
Anthropic researchers demonstrate that Claude Opus 4.6 can autonomously act as an alignment researcher to improve weak-to-strong supervision techniques, addressing challenges in scalable oversight.