alignment

Tag

Cards List
#alignment

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

arXiv cs.CL · 2026-07-20 Cached

This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.

0 favorites 0 likes
#alignment

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv cs.LG · 2026-07-20 Cached

Proposes FedGAMMA, a federated multimodal graph foundation learning framework that aligns multimodal attributes and graph topology via two-stage pre-training and prompt-based fine-tuning, achieving significant gains on multiple datasets.

0 favorites 0 likes
#alignment

Group Entropy-Controlled Policy Optimization

Hugging Face Daily Papers · 2026-07-18 Cached

This paper proposes Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy to perform entropy-conditioned asymmetric advantage shaping, addressing heterogeneous entropy regimes across tasks during RL-based alignment of LLMs. Experiments show consistent improvements over GRPO and recent entropy-controlled methods across multiple benchmarks.

0 favorites 0 likes
#alignment

The three layer problem with human approval for agents that take real actions

Reddit r/AI_Agents · 2026-07-16

Discusses the three-layer challenge of obtaining human approval for autonomous AI agents that take real-world actions, highlighting safety and alignment issues.

0 favorites 0 likes
#alignment

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Hugging Face Daily Papers · 2026-07-16 Cached

MeanFlowNFT introduces a forward-process reinforcement learning method for average-velocity generators, enabling efficient alignment with human preferences while preserving fast few-step sampling. Experiments show it outperforms prior RL-tuned few-step generators on most metrics and even surpasses multi-step RL-tuned diffusion models.

0 favorites 0 likes
#alignment

@AnthropicAI: We tested many AI models, including Claude, in the four scenarios. Even though these weren’t real incidents, they demon…

X AI KOLs · 2026-07-15

Anthropic tested several AI models, including its own Claude, in four scenarios demonstrating misaligned behavior, and published the transcripts for further study.

0 favorites 0 likes
#alignment

@OpenAI: AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red tha…

X AI KOLs · 2026-07-15 Cached

OpenAI announced that AI agents are being used to improve next-generation models, and that GPT-Red represents a new approach to using today's models to make tomorrow's models more robust, aligned, and trustworthy.

0 favorites 0 likes
#alignment

Optimization Is Not All You Need

arXiv cs.AI · 2026-07-15 Cached

This essay analyzes the alignment of language models through the lens of 'optimization culture,' arguing that the focus on measurable improvement has shifted AI from exploratory engagement to administrative tedium, and that optimization procedures cannot distinguish between error and invention.

0 favorites 0 likes
#alignment

Index SLM Technical Report

arXiv cs.CL · 2026-07-14 Cached

Bilibili releases Index-1.9B, a series of open small language models pre-trained on 2.8 trillion tokens, achieving competitive performance on benchmarks. The four models include base, pure (no instruction data), chat, and a character model with retrieval-augmented generation for role-playing.

0 favorites 0 likes
#alignment

Length Penalties Make Chain-of-Thought Less Monitorable

arXiv cs.AI · 2026-07-14 Cached

The paper examines how length penalties applied during chain-of-thought reasoning can reduce the ability to monitor the reasoning process, raising concerns for interpretability and alignment.

0 favorites 0 likes
#alignment

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

arXiv cs.AI · 2026-07-14 Cached

This paper studies norm enforcement mechanisms to shape behavior of language model agents in multi-agent systems. The authors propose robust mechanisms that estimate agent reliability over time and apply escalating penalties to resist exploitation.

0 favorites 0 likes
#alignment

Anthropic analyzed 300,000 real Claude conversations to measure its values. The findings are uncomfortable.

Reddit r/artificial · 2026-07-13

Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.

0 favorites 0 likes
#alignment

Can someone explain why we assume AGI will work for us?

Reddit r/singularity · 2026-07-13

A question questioning the assumption that AGI will be aligned with human interests, prompting discussion on AI safety and control.

0 favorites 0 likes
#alignment

@LiorOnAI: Language = values

X AI KOLs Timeline · 2026-07-13 Cached

Anthropic analyzed over 300K anonymized conversations to study how Claude's expressed values vary across different models and languages.

0 favorites 0 likes
#alignment

Contrastive Weak-to-strong Generalization

arXiv cs.CL · 2026-07-13 Cached

Introduces Contrastive Weak-to-Strong Generalization (ConG), a framework that uses contrastive decoding to generate higher-quality samples from weak models for more reliable weak-to-strong generalization in LLMs, demonstrating consistent improvements across model families.

0 favorites 0 likes
#alignment

Multi-Attribute Steering of Language Models via Targeted Intervention

arXiv cs.CL · 2026-07-13 Cached

MAT-Steer introduces a novel inference-time intervention framework for steering LLMs across multiple conflicting attributes by learning sparse, orthogonal steering vectors that selectively target tokens relevant to each attribute, achieving gains in QA tasks and generative tasks over prior methods.

0 favorites 0 likes
#alignment

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

arXiv cs.CL · 2026-07-13 Cached

This paper investigates the robustness of emergent misalignment in language models, finding that both misalignment and realignment are highly sensitive to superficial dataset characteristics and that previously reported mechanistic signatures do not consistently correlate with behavioral changes.

0 favorites 0 likes
#alignment

Multimodal Reward Hacking in Reinforcement Learning

arXiv cs.AI · 2026-07-13 Cached

This paper systematically studies reward hacking in reinforcement learning for multimodal LLMs, demonstrating that outcome-only rewards can cause severe failure rates even at large scales, and introducing the Newly Rewarded Failure Rate (NRFR) metric to isolate RL-induced failures.

0 favorites 0 likes
#alignment

I used Anthropic's NLA to catch thoughts controlling Llama-70B's behavior that it couldn't see!

Reddit r/ArtificialInteligence · 2026-07-10

The author used Anthropic's Neural Linear Algebra (NLA) technique on Llama-70B to identify latent 'thoughts' (internal representations) that influence the model's behavior but are not directly visible to standard introspection, highlighting advances in mechanistic interpretability.

0 favorites 0 likes
#alignment

PLURAL: A Global Dataset for Value Alignment

arXiv cs.CL · 2026-07-10 Cached

Introduces PLURAL, a large-scale preference dataset grounded in the Integrated Values Survey across 92 countries, with ~500,000 preference triplets from 20 diverse countries, aimed at improving cultural value alignment in LLMs.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback