Tag
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
Proposes FedGAMMA, a federated multimodal graph foundation learning framework that aligns multimodal attributes and graph topology via two-stage pre-training and prompt-based fine-tuning, achieving significant gains on multiple datasets.
This paper proposes Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy to perform entropy-conditioned asymmetric advantage shaping, addressing heterogeneous entropy regimes across tasks during RL-based alignment of LLMs. Experiments show consistent improvements over GRPO and recent entropy-controlled methods across multiple benchmarks.
Discusses the three-layer challenge of obtaining human approval for autonomous AI agents that take real-world actions, highlighting safety and alignment issues.
MeanFlowNFT introduces a forward-process reinforcement learning method for average-velocity generators, enabling efficient alignment with human preferences while preserving fast few-step sampling. Experiments show it outperforms prior RL-tuned few-step generators on most metrics and even surpasses multi-step RL-tuned diffusion models.
Anthropic tested several AI models, including its own Claude, in four scenarios demonstrating misaligned behavior, and published the transcripts for further study.
OpenAI announced that AI agents are being used to improve next-generation models, and that GPT-Red represents a new approach to using today's models to make tomorrow's models more robust, aligned, and trustworthy.
This essay analyzes the alignment of language models through the lens of 'optimization culture,' arguing that the focus on measurable improvement has shifted AI from exploratory engagement to administrative tedium, and that optimization procedures cannot distinguish between error and invention.
Bilibili releases Index-1.9B, a series of open small language models pre-trained on 2.8 trillion tokens, achieving competitive performance on benchmarks. The four models include base, pure (no instruction data), chat, and a character model with retrieval-augmented generation for role-playing.
The paper examines how length penalties applied during chain-of-thought reasoning can reduce the ability to monitor the reasoning process, raising concerns for interpretability and alignment.
This paper studies norm enforcement mechanisms to shape behavior of language model agents in multi-agent systems. The authors propose robust mechanisms that estimate agent reliability over time and apply escalating penalties to resist exploitation.
Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.
A question questioning the assumption that AGI will be aligned with human interests, prompting discussion on AI safety and control.
Anthropic analyzed over 300K anonymized conversations to study how Claude's expressed values vary across different models and languages.
Introduces Contrastive Weak-to-Strong Generalization (ConG), a framework that uses contrastive decoding to generate higher-quality samples from weak models for more reliable weak-to-strong generalization in LLMs, demonstrating consistent improvements across model families.
MAT-Steer introduces a novel inference-time intervention framework for steering LLMs across multiple conflicting attributes by learning sparse, orthogonal steering vectors that selectively target tokens relevant to each attribute, achieving gains in QA tasks and generative tasks over prior methods.
This paper investigates the robustness of emergent misalignment in language models, finding that both misalignment and realignment are highly sensitive to superficial dataset characteristics and that previously reported mechanistic signatures do not consistently correlate with behavioral changes.
This paper systematically studies reward hacking in reinforcement learning for multimodal LLMs, demonstrating that outcome-only rewards can cause severe failure rates even at large scales, and introducing the Newly Rewarded Failure Rate (NRFR) metric to isolate RL-induced failures.
The author used Anthropic's Neural Linear Algebra (NLA) technique on Llama-70B to identify latent 'thoughts' (internal representations) that influence the model's behavior but are not directly visible to standard introspection, highlighting advances in mechanistic interpretability.
Introduces PLURAL, a large-scale preference dataset grounded in the Integrated Values Survey across 92 countries, with ~500,000 preference triplets from 20 diverse countries, aimed at improving cultural value alignment in LLMs.