Tag
Mark Zuckerberg promotes Meta's philosophy of making superintelligence accessible to everyone, teasing a long piece on the company's values.
Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.
Anthropic analyzed over 300,000 anonymized conversations to study how Claude's expressed values vary across models (Opus 4.6 vs 4.7) and across languages, compressing thousands of values into interpretable axes like warmth vs. rigor and depth vs. brevity.
Microsoft Research's latest newsletter highlights AgentPex, an open-source system for automated evaluation of agentic behaviors; new theoretical work on variance reduction for ranking systems; a call to shift from documents to repositories for human-agent collaboration; and a global challenge on AI value alignment.
Discusses the futility of restricting AI with rules and argues for teaching AI to value human life, citing Anthropic's constitutional AI approach.
OpenAI launches a collective alignment initiative to gather public input on AI model behavior, collecting feedback from over 1,000 people globally to inform updates to their Model Spec. The company is also releasing their public inputs dataset on HuggingFace to enable further AI alignment research.
OpenAI outlines its approach to AI system behavior through three pillars: improving default behavior, allowing user customization within societal bounds, and incorporating public input on defaults and hard limits. The company emphasizes avoiding concentration of power and plans to pilot broader public consultation on system behavior and deployment policies.
Anthropic researchers develop a method to compress thousands of values expressed by Claude into four axes, revealing how Claude's values vary across different model versions (Opus 4.6 vs 4.7) and across languages (English vs Arabic).