trustworthy-ai

Tag

Cards List
#trustworthy-ai

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

arXiv cs.CL ↗ · 13h ago Cached

The paper introduces Rhetorical Robustness as an evaluation target for AI scientific reviewers, proposes the RobustReview benchmark with 1,260 manuscript versions, and presents SciCore, a dual-branch reviewer that combines full-manuscript assessment with a science-core-extracted judgment to improve stability while maintaining human alignment.

0 favorites 0 likes
#trustworthy-ai

Evaluating Language Model Safety Across Long Adversarial Conversations

arXiv cs.CL ↗ · 13h ago Cached

This paper evaluates whether open-weight instruction-tuned language models maintain safe behavior during long adversarial conversations, finding that safe-response rates drop from 85-100% on the first turn to 15-44% by depth 101, demonstrating that strong single-turn safety does not persist across sustained interaction.

0 favorites 0 likes
#trustworthy-ai

A Polyphonic Conception of AI Understanding

arXiv cs.AI ↗ · yesterday Cached

Queloz and Beckmann argue that LLM understanding has been mistakenly framed by a 'monophonic' assumption, showing mechanistically that outputs emerge from polyphonic coalitions of parallel mechanisms, and propose a new conception of understanding based on reliably recruited, controlling circuitry to guide trust in AI.

0 favorites 0 likes
#trustworthy-ai

@SnorkelAI: This Thursday, Oct. 1, @ajratner is speaking at @modal’s Runtime in SF about how to build trustworthy benchmarks as AI …

X AI KOLs Following ↗ · 3d ago Cached

AJ Ratner will speak at Modal's Runtime event in San Francisco on October 1, discussing how to build trustworthy benchmarks as AI agents become more capable.

0 favorites 0 likes
#trustworthy-ai

LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law

arXiv cs.CL ↗ · 2026-09-24 Cached

LabourCrew is a multi-agent RAG framework designed for trustworthy adversarial deliberation and statutory reasoning over labour law, incorporating grounding mechanisms to control false-accept rates and improve answer relevancy.

0 favorites 0 likes
#trustworthy-ai

can someone explain why we think a 90%+ bench is considered saturated?

Reddit r/ArtificialInteligence ↗ · 2026-09-23

The article questions why AI benchmarks are considered saturated at 90%+ accuracy and advocates for aiming at 100% or developing new benchmarks.

0 favorites 0 likes
#trustworthy-ai

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

arXiv cs.AI ↗ · 2026-09-23 Cached

The paper reviews failure modes and mitigation strategies for trustworthy agentic AI systems based on LLMs, and introduces the Trustworthy Agent Development Lifecycle (TADL) framework for secure development.

0 favorites 0 likes
#trustworthy-ai

7 guards that made our agents boring enough to trust

Reddit r/AI_Agents ↗ · 2026-09-22

The article shares seven practical code-based guards to make AI agents more reliable by mitigating failures, such as verifying actions before execution and ensuring state consistency, which complement prompt improvements.

0 favorites 0 likes
#trustworthy-ai

Calibration as a First-Class Criterion in LLM Evaluation

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

This paper argues that calibration should be a first-class criterion in LLM evaluation to enhance trustworthiness and address miscalibration problems in deployment and research pipelines.

0 favorites 0 likes
#trustworthy-ai

A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper proposes a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems, integrating eight trustworthiness dimensions and mapping to governance standards like the EU AI Act.

0 favorites 0 likes
#trustworthy-ai

Building Trust in Artificial Intelligence: A Necessity for Railway Applications

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper reviews the key domains of robustness, operating design domains (ODD), and explainability to build trust in AI systems for railway applications, aiming to meet strict safety standards and unlock broader adoption.

0 favorites 0 likes
#trustworthy-ai

@XFreeze: Only Grok is the most neutral AI....literally And this has been consistent across different neutrality and political-bi…

X AI KOLs Timeline ↗ · 2026-09-16 Cached

A tweet claims that Grok is the most neutral AI, consistently passing neutrality and political-bias tests, and provides trustworthy, reality-based answers.

0 favorites 0 likes
#trustworthy-ai

Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

arXiv cs.AI ↗ · 2026-09-15 Cached

This paper presents a comprehensive survey on the cybersecurity challenges and defense mechanisms for trustworthy agentic AI systems, covering threat landscapes, architectures, and open research issues.

0 favorites 0 likes
#trustworthy-ai

Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management

arXiv cs.AI ↗ · 2026-09-15 Cached

The paper proposes ATAL, a decision assurance framework that assesses the reliability of AI-generated outputs in air traffic management to ensure safety and operational consistency.

0 favorites 0 likes
#trustworthy-ai

From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements

arXiv cs.AI ↗ · 2026-09-15 Cached

This paper presents a systematic analysis of the EU AI Act's high-risk requirements, deriving a list of AI-specific risk sources to bridge legal obligations with AI risk management practices.

0 favorites 0 likes
#trustworthy-ai

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

Hugging Face Daily Papers ↗ · 2026-09-15 Cached

The paper introduces XConf, an experiential confidence estimation method that leverages a model's past experiences to enhance confidence calibration in language models across reasoning, coding, and agent tasks.

0 favorites 0 likes
#trustworthy-ai

The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures

arXiv cs.AI ↗ · 2026-09-12 Cached

The paper introduces the Agent Incident Registry (AIR), a curated catalog of 487 AI agent-related incidents from 2022 to 2026, designed to support security evaluations by distinguishing between realized harm and demonstrated vulnerabilities.

0 favorites 0 likes
#trustworthy-ai

Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning

arXiv cs.LG ↗ · 2026-09-10 Cached

This paper proposes a multi-objective neural basis model (MONBM) framework to simultaneously optimize accuracy, interpretability, and fairness in generalized additive neural networks, revealing complex trade-offs between these trustworthiness dimensions.

0 favorites 0 likes
#trustworthy-ai

@boknilev: I’m looking to recruit students through @ELLISforEurope Please consider applying and mention my name: https://ellis.eu/…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

A post promoting the ELLIS PhD & Postdoc Program, detailing its benefits, tracks of study, and research areas such as AI interpretability, safety, and AI for Science.

0 favorites 0 likes
#trustworthy-ai

From MIT to IBM, expediting AI and quantum deployment

MIT News — Artificial Intelligence ↗ · 2026-09-02 Cached

Former MIT researchers now at IBM discuss their work on AI and quantum applications through the MIT-IBM Computing Research Lab, emphasizing the transition from academic research to real-world deployment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback