Tag
The paper introduces Rhetorical Robustness as an evaluation target for AI scientific reviewers, proposes the RobustReview benchmark with 1,260 manuscript versions, and presents SciCore, a dual-branch reviewer that combines full-manuscript assessment with a science-core-extracted judgment to improve stability while maintaining human alignment.
This paper evaluates whether open-weight instruction-tuned language models maintain safe behavior during long adversarial conversations, finding that safe-response rates drop from 85-100% on the first turn to 15-44% by depth 101, demonstrating that strong single-turn safety does not persist across sustained interaction.
Queloz and Beckmann argue that LLM understanding has been mistakenly framed by a 'monophonic' assumption, showing mechanistically that outputs emerge from polyphonic coalitions of parallel mechanisms, and propose a new conception of understanding based on reliably recruited, controlling circuitry to guide trust in AI.
AJ Ratner will speak at Modal's Runtime event in San Francisco on October 1, discussing how to build trustworthy benchmarks as AI agents become more capable.
LabourCrew is a multi-agent RAG framework designed for trustworthy adversarial deliberation and statutory reasoning over labour law, incorporating grounding mechanisms to control false-accept rates and improve answer relevancy.
The article questions why AI benchmarks are considered saturated at 90%+ accuracy and advocates for aiming at 100% or developing new benchmarks.
The paper reviews failure modes and mitigation strategies for trustworthy agentic AI systems based on LLMs, and introduces the Trustworthy Agent Development Lifecycle (TADL) framework for secure development.
The article shares seven practical code-based guards to make AI agents more reliable by mitigating failures, such as verifying actions before execution and ensuring state consistency, which complement prompt improvements.
This paper argues that calibration should be a first-class criterion in LLM evaluation to enhance trustworthiness and address miscalibration problems in deployment and research pipelines.
This paper proposes a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems, integrating eight trustworthiness dimensions and mapping to governance standards like the EU AI Act.
This paper reviews the key domains of robustness, operating design domains (ODD), and explainability to build trust in AI systems for railway applications, aiming to meet strict safety standards and unlock broader adoption.
A tweet claims that Grok is the most neutral AI, consistently passing neutrality and political-bias tests, and provides trustworthy, reality-based answers.
This paper presents a comprehensive survey on the cybersecurity challenges and defense mechanisms for trustworthy agentic AI systems, covering threat landscapes, architectures, and open research issues.
The paper proposes ATAL, a decision assurance framework that assesses the reliability of AI-generated outputs in air traffic management to ensure safety and operational consistency.
This paper presents a systematic analysis of the EU AI Act's high-risk requirements, deriving a list of AI-specific risk sources to bridge legal obligations with AI risk management practices.
The paper introduces XConf, an experiential confidence estimation method that leverages a model's past experiences to enhance confidence calibration in language models across reasoning, coding, and agent tasks.
The paper introduces the Agent Incident Registry (AIR), a curated catalog of 487 AI agent-related incidents from 2022 to 2026, designed to support security evaluations by distinguishing between realized harm and demonstrated vulnerabilities.
This paper proposes a multi-objective neural basis model (MONBM) framework to simultaneously optimize accuracy, interpretability, and fairness in generalized additive neural networks, revealing complex trade-offs between these trustworthiness dimensions.
A post promoting the ELLIS PhD & Postdoc Program, detailing its benefits, tracks of study, and research areas such as AI interpretability, safety, and AI for Science.
Former MIT researchers now at IBM discuss their work on AI and quantum applications through the MIT-IBM Computing Research Lab, emphasizing the transition from academic research to real-world deployment.