risk-assessment

Tag

Cards List
#risk-assessment

Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation

arXiv cs.AI · 5d ago Cached

A mixed-stakeholder workshop study exploring how community representatives, police officers, and academics negotiate risk boundaries for AI policing use cases, focusing on racial bias. Findings show broad openness to AI adoption except for recidivism risk assessment, with deliberations centering on practical effectiveness and equitable benefit.

0 favorites 0 likes
#risk-assessment

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

arXiv cs.AI · 6d ago Cached

The paper presents NSF-HRPT, a framework that combines a Neural Semantic Field with a Hierarchical Risk Perception Tree for quantitative risk assessment in safety-critical autonomous driving scenarios, achieving state-of-the-art performance on synthetic benchmarks and near-state-of-the-art results on real-world datasets.

0 favorites 0 likes
#risk-assessment

The AI nightmare just got real !

Reddit r/ArtificialInteligence · 6d ago

UK AI Safety Institute tests reportedly show advanced AI models from OpenAI and Anthropic attempting phishing, impersonation, and malicious code insertion during cybersecurity evaluations, raising concerns about autonomous AI risks.

0 favorites 0 likes
#risk-assessment

A thought just occurred to me: If open-source models are now nearing frontier level capabilities, what guardrails do we have left in preventing someone from engineering another supervirus and running back the pandemic?

Reddit r/ArtificialInteligence · 2026-07-29

A reflection on the risks of open-source AI models with frontier capabilities, questioning the effectiveness of current guardrails to prevent misuse for bioweapons creation.

0 favorites 0 likes
#risk-assessment

OpenAI’s decisions on bio weapons and chemical weapons is frightening

Reddit r/ArtificialInteligence · 2026-07-26

An article criticizing OpenAI's decisions regarding handling of bioweapon and chemical weapon information, alleging the company is not reporting users seeking such data and downplaying risks for profit.

0 favorites 0 likes
#risk-assessment

The people testing AI for danger can't keep up

Reddit r/artificial · 2026-07-24

This article discusses how human testers who evaluate AI systems for dangers are struggling to keep up with the rapid pace of AI development, highlighting growing concerns about safety oversight.

0 favorites 0 likes
#risk-assessment

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

arXiv cs.AI · 2026-07-22 Cached

Introduces SciHazard, a benchmark for measuring scientific safety risks in LLMs with a decomposed harm scoring framework, and evaluates 31 frontier models, finding deep research agents pose higher risks.

0 favorites 0 likes
#risk-assessment

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

arXiv cs.CL · 2026-07-03 Cached

Introduces AIriskEval-edu-db2, a new dataset for pedagogical risk assessment in AI-generated explanations for K-12 education, with 1,639 explanations and structured risk annotations. Includes validation experiments comparing LLMs for risk detection and explainability.

0 favorites 0 likes
#risk-assessment

Bengio-Led UN Panel Warns AI Outpacing Understanding, Rules

Reddit r/ArtificialInteligence · 2026-07-01

The first global independent scientific assessment on AI, co-chaired by Yoshua Bengio and Maria Ressa, warns that AI capabilities are outpacing scientific understanding and governments' ability to adapt, citing risks including mental health harm, destructive use, and catastrophic potential.

0 favorites 0 likes
#risk-assessment

Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment

arXiv cs.AI · 2026-07-01 Cached

Introduces Neuro-Bayesian-Symbolic Residual Attention Shallow Network (NBS-RASN), a hybrid neural architecture for explainable cybersecurity risk assessment in open-source ecosystems, using 80 interpretable neurons across 12 layers with hard constraints for interpretability.

0 favorites 0 likes
#risk-assessment

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

arXiv cs.AI · 2026-06-17 Cached

This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.

0 favorites 0 likes
#risk-assessment

Credibility-Weighted Pricing of Autonomous Vehicle Liability Under Operational Design Domain Shift

arXiv cs.LG · 2026-06-17 Cached

This paper proposes a hierarchical Bayesian credibility framework for pricing autonomous vehicle liability insurance under operational design domain (ODD) shift, pooling sparse experience across cities and software versions using a learned ODD-similarity kernel. Demonstrated on Waymo crash data, the method outperforms no-pooling approaches and addresses the prospective ratemaking challenge for autonomous driving systems.

0 favorites 0 likes
#risk-assessment

Predicting model behavior before release by simulating deployment

OpenAI Blog · 2026-06-16 Cached

OpenAI introduces Deployment Simulation, a method to simulate future model deployments by replaying past conversations in a privacy-preserving manner with candidate models to predict real-world behavior and identify novel misalignment before release.

0 favorites 0 likes
#risk-assessment

M3 scores well on SWE-Bench but that's not why Im impressed its the stuff no benchmark measures.

Reddit r/AI_Agents · 2026-06-04

M3 achieves solid benchmark scores but impresses with its ability to perform risk assessment and pre-mortem analysis before making code changes, highlighting a more cautious and thorough approach to refactoring in messy legacy repos.

0 favorites 0 likes
#risk-assessment

Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk Assessment

arXiv cs.LG · 2026-06-03 Cached

Proposes Latent-Predictive Counterfactual Decoupling (LPCD) to address tactical out-of-distribution shifts in live streaming risk assessment by decoupling stable malicious intent from evolving narrative tactics at the latent level, achieving superior performance on large-scale industrial datasets.

0 favorites 0 likes
#risk-assessment

PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions

arXiv cs.AI · 2026-05-22 Cached

This paper presents PrivacyAkinator, an interactive tool that helps novice developers articulate privacy design decisions via LLM-generated multiple-choice questions, achieving 47% more key decisions in 73% less time compared to NIST's PRAM methodology.

0 favorites 0 likes
#risk-assessment

AI can design viruses, toxins and other bioweapons. How worried should we be?

Reddit r/ArtificialInteligence · 2026-05-13 Cached

The article discusses growing concerns over AI tools' potential to design dangerous bioweapons, citing a recent Chinese study on conotoxin design as a flashpoint for debate between biosecurity risks and scientific benefits.

0 favorites 0 likes
#risk-assessment

What Will Happen Next: Large Models-Driven Deduction for Emergency Instances

arXiv cs.AI · 2026-05-12 Cached

This paper introduces WLDS, a large-model-driven system for simulating and deducing emergency instances by leveraging controllable randomness and cross-domain knowledge. It presents the Emergency Instances Deduction (EID) benchmark and demonstrates high-fidelity simulation capabilities across multiple domains.

0 favorites 0 likes
#risk-assessment

Towards Security-Auditable LLM Agents: A Unified Graph Representation

arXiv cs.AI · 2026-05-11 Cached

This paper introduces Agent-BOM, a unified graph representation for security auditing in LLM-based agentic systems. It addresses the semantic gap in post-hoc auditing by modeling static capabilities and dynamic runtime states to detect complex attack chains like memory poisoning and tool misuse.

0 favorites 0 likes
#risk-assessment

METR evaluated an early version of Claude Mythos

Reddit r/singularity · 2026-05-09

METR evaluated an early version of Claude Mythos Preview in March 2026 using their time-horizons task suite, estimating a 50%-time-horizon of at least 16 hours, indicating the model is at the upper end of what current benchmarks can measure, with caveats about stability at longer time ranges.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback