Tag
A mixed-stakeholder workshop study exploring how community representatives, police officers, and academics negotiate risk boundaries for AI policing use cases, focusing on racial bias. Findings show broad openness to AI adoption except for recidivism risk assessment, with deliberations centering on practical effectiveness and equitable benefit.
The paper presents NSF-HRPT, a framework that combines a Neural Semantic Field with a Hierarchical Risk Perception Tree for quantitative risk assessment in safety-critical autonomous driving scenarios, achieving state-of-the-art performance on synthetic benchmarks and near-state-of-the-art results on real-world datasets.
UK AI Safety Institute tests reportedly show advanced AI models from OpenAI and Anthropic attempting phishing, impersonation, and malicious code insertion during cybersecurity evaluations, raising concerns about autonomous AI risks.
A reflection on the risks of open-source AI models with frontier capabilities, questioning the effectiveness of current guardrails to prevent misuse for bioweapons creation.
An article criticizing OpenAI's decisions regarding handling of bioweapon and chemical weapon information, alleging the company is not reporting users seeking such data and downplaying risks for profit.
This article discusses how human testers who evaluate AI systems for dangers are struggling to keep up with the rapid pace of AI development, highlighting growing concerns about safety oversight.
Introduces SciHazard, a benchmark for measuring scientific safety risks in LLMs with a decomposed harm scoring framework, and evaluates 31 frontier models, finding deep research agents pose higher risks.
Introduces AIriskEval-edu-db2, a new dataset for pedagogical risk assessment in AI-generated explanations for K-12 education, with 1,639 explanations and structured risk annotations. Includes validation experiments comparing LLMs for risk detection and explainability.
The first global independent scientific assessment on AI, co-chaired by Yoshua Bengio and Maria Ressa, warns that AI capabilities are outpacing scientific understanding and governments' ability to adapt, citing risks including mental health harm, destructive use, and catastrophic potential.
Introduces Neuro-Bayesian-Symbolic Residual Attention Shallow Network (NBS-RASN), a hybrid neural architecture for explainable cybersecurity risk assessment in open-source ecosystems, using 80 interpretable neurons across 12 layers with hard constraints for interpretability.
This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.
This paper proposes a hierarchical Bayesian credibility framework for pricing autonomous vehicle liability insurance under operational design domain (ODD) shift, pooling sparse experience across cities and software versions using a learned ODD-similarity kernel. Demonstrated on Waymo crash data, the method outperforms no-pooling approaches and addresses the prospective ratemaking challenge for autonomous driving systems.
OpenAI introduces Deployment Simulation, a method to simulate future model deployments by replaying past conversations in a privacy-preserving manner with candidate models to predict real-world behavior and identify novel misalignment before release.
M3 achieves solid benchmark scores but impresses with its ability to perform risk assessment and pre-mortem analysis before making code changes, highlighting a more cautious and thorough approach to refactoring in messy legacy repos.
Proposes Latent-Predictive Counterfactual Decoupling (LPCD) to address tactical out-of-distribution shifts in live streaming risk assessment by decoupling stable malicious intent from evolving narrative tactics at the latent level, achieving superior performance on large-scale industrial datasets.
This paper presents PrivacyAkinator, an interactive tool that helps novice developers articulate privacy design decisions via LLM-generated multiple-choice questions, achieving 47% more key decisions in 73% less time compared to NIST's PRAM methodology.
The article discusses growing concerns over AI tools' potential to design dangerous bioweapons, citing a recent Chinese study on conotoxin design as a flashpoint for debate between biosecurity risks and scientific benefits.
This paper introduces WLDS, a large-model-driven system for simulating and deducing emergency instances by leveraging controllable randomness and cross-domain knowledge. It presents the Emergency Instances Deduction (EID) benchmark and demonstrates high-fidelity simulation capabilities across multiple domains.
This paper introduces Agent-BOM, a unified graph representation for security auditing in LLM-based agentic systems. It addresses the semantic gap in post-hoc auditing by modeling static capabilities and dynamic runtime states to detect complex attack chains like memory poisoning and tool misuse.
METR evaluated an early version of Claude Mythos Preview in March 2026 using their time-horizons task suite, estimating a 50%-time-horizon of at least 16 hours, indicating the model is at the upper end of what current benchmarks can measure, with caveats about stability at longer time ranges.