Tag
The user speculates that starting from Opus 4.6, Anthropic internally shifted to using the Mythos model, leading to subsequent Opus models being trained and behaving more cautiously and AI-like.
Debate training can mitigate reward hacking in reinforcement learning from AI feedback (RLAIF) by using a debate opponent to improve ground-truth accuracy for fuzzy tasks.
The article briefly introduces the evolution of AI training methods from RLHF to RLAIF to RLTHF and predicts that by 2025, RLTHF will significantly reduce expert working hours.
This paper presents a method using Group Relative Policy Optimization (GRPO) to fine-tune an open-weight language model for generating actionable financial advice, outperforming commercial LLMs under a judge-independent CATE evaluation while also matching safety criteria.
Introduces Hallucination Self-Play (HSP), a framework that bootstraps a detector using an evolved generator via reinforcement learning, enabling small LLMs to match advanced LLMs on faithfulness hallucination detection without external supervision.
Introduces ODRPO, a framework that decomposes discrete rewards into ordinal binary indicators to improve robustness of policy optimization in RLAIF for LLMs, achieving up to 14.8% relative improvement with minimal overhead.