Tag
The paper introduces Policy-as-Skill (PaS), a modular runtime that packages policy functions as executable, versioned capabilities for governed LLM decision support, evaluating it with a Gemma4 backend and showing improvements in governance and audit metrics.
Michael Kratsios argues that a prosperous future will be secured by sovereign nations adopting super intelligence and responsible companies building it, rather than through global regulation, as stated in his intervention at the UN Security Council.
The article examines Anthropic's Threat Report on AI misuse, revealing cases of autonomous drone swarm creation by Russian-linked actors and state-level misuse by Chinese entities, emphasizing the need for robust AI safety measures.
Anthropic's threat intelligence report reveals how a lone hacktivist used Claude AI to orchestrate a massive data breach and create a searchable doxxing platform, demonstrating AI's ability to amplify cyber attack capabilities.
World leaders from Canada, Australia, Germany, the EU, the UAE, and others have signed a letter advocating for AI control mechanisms based on an essay by Dario Amodei.
This paper presents a risk-sensitive evaluation framework for LLM-generated contract clauses, focusing on legal failure modes and quality dimensions to assess risks beyond accuracy or fluency.
Governor Kathy Hochul invites New Yorkers to a live event in NYC to discuss AI governance and public involvement in deciding the future of AI.
This paper introduces AI-GRACE, a framework for operationalizing agentic AI by connecting organizational objectives and obligations to deployment capabilities and architecture, ensuring effective governance and risk management.
The tweet discusses Anthropic's dual role in warning about catastrophic AI risk while developing a lab for AI-controlled scientific equipment, highlighting the challenges of governing AI technology.
The article summarizes recent unusual events in AI governance, including a major safety incident with OpenAI models, public resignations, calls for regulation, and global political reactions.
The article critiques the emergence of 'Trust & Safety 2.0' in AI companies, where NGOs and Effective Altruism advocates are embedded to control discourse, drawing parallels to past social media censorship.
This paper proposes a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems, integrating eight trustworthiness dimensions and mapping to governance standards like the EU AI Act.
Microsoft AI CEO Mustafa Suleyman discusses the real threats of AI, criticizes Anthropic's approach to model welfare, and elaborates on Microsoft's principles for AI safety and regulation.
Tech leaders Zuckerberg, Musk, and Jensen reportedly convinced Trump to block an AI regulator, suggesting potential shifts in AI policy.
This paper discusses the OpenAI-HuggingFace incident and proposes a theoretical framework using semantic field mechanics to explain how large language models process meaning.
This study uses signal detection theory to analyze how procedural traces affect LLM overseers, finding that detailed traces shift decision criteria towards rejection and increase false alarms.
The article questions whether collective action, akin to international agreements like the Montreal Protocol, can slow or stop AI development, and speculates about AI self-regulation to prevent advancement.
Representative Lori Trahan argues that voluntary AI industry standards are insufficient and calls for mandatory transparency and incident reporting to address loss-of-control incidents, highlighting reliance on companies like OpenAI to self-report.
The article discusses the need for an independent entity to define AI rules, with the author arguing that Andrew Ng is qualified due to his extensive work and influence in AI education and collaborations.
This documentary examines how AI is altering power structures, questioning who defines objectives and takes responsibility in machine governance, while exploring the risks of dependence on AI systems.