Tag
This position paper argues that AI systems used in high-stakes decision-making should reason similarly to their users and faithfully communicate that reasoning, and outlines a research agenda for achieving such 'cognitively-aligned AI'.
A report based on a survey of 300 data and technology executives examines how legacy data systems limit the effectiveness and scaling of AI agents in enterprises, highlighting that 'data leaders' who give agents broader data access experience greater trust and success.
This paper introduces the RAIL principles (Reasoning, Assurances, Interfacing, Learning) for designing neurosymbolic AI systems that integrate machine learning with symbolic reasoning, arguing this approach is crucial for reliable, efficient, and trustworthy AI.
This paper introduces explanation-based runtime verification for ML-driven optical networks, using model explanations to assess the soundness of individual decisions before execution in the network control loop, demonstrating effectiveness in intercepting erroneous decisions while preserving high automation rates.
This paper evaluates citation faithfulness in agentic scientific synthesis systems, showing that current verifiers are unreliable with unsupported-citation rates varying from 3% to 18% depending on strictness. It proposes a gold-anchored evaluation protocol and a deployable guard that uses split-conformal prediction to provide a distribution-free bound on truly unsupported citations.
Proposes HALO, an architecture with six layers of defense to contain hallucination in enterprise AI systems, reframing 'zero hallucination' as a system-enforced property rather than a model property.
This paper argues that current responsible AI practices fail to create a market that rewards trustworthiness, proposing independent, outcome-oriented certification to close the 'trust gap' by making AI trustworthiness measurable, comparable, and commercially rewarded.
This paper critically analyzes tools and trust mark frameworks for operationalizing trustworthy AI, finding asymmetries in ethical focus and lifecycle coverage. It identifies implementation gaps and offers recommendations for more holistic AI governance.
A systematic survey of Federated Explainable Artificial Intelligence (FedXAI), covering roles, architectures, evaluation practices, and open challenges. It presents a multi-axis taxonomy and discusses model-agnostic to interpretable-by-design approaches, highlighting gaps in standardization and privacy-aware evaluation.
OpenAI announced that AI agents are being used to improve next-generation models, and that GPT-Red represents a new approach to using today's models to make tomorrow's models more robust, aligned, and trustworthy.
This chapter reviews design principles and current landscape of foundation models for Earth observation, highlighting the need for domain-specific adaptation, physically plausible representations, and consistent evaluation benchmarks. It includes case studies on harmful algal bloom prediction and adaptive monitoring station selection.
This tweet discusses the paper 'From Model Scaling to System Scaling' which argues that stronger AI agents require better system design (harness) including context control, memory, and routing, not just larger models.
Bin Yu has developed the PCS framework to make trustworthy AI verifiable, addressing the challenge of increasingly opaque models.
Theoria is a verification architecture that rewrites AI solutions into auditable state transitions, achieving high precision on HLE problems and detecting subtle errors like hidden premises and fabricated citations.
The Decentralized Assessment for Trustworthy AI (DATA) is an ethical evaluation tool that allows users and communities to objectively audit AI companies based on leading ethical frameworks like UNESCO and EU guidelines.
This paper investigates how contextual framing affects LLM responses in mental health interactions, finding systematic behavioral variation and demonstrating that internal representations encode framing information throughout transformer layers.
This paper proposes a post-hoc certification framework for sparse autoencoder (SAE) based interpretability, deriving an upper bound on the frozen language model's risk using measurable quantities. The framework is validated on GPT-2 Small, Gemma-2B, and Llama-3-8B, showing non-vacuous bounds and revealing depth-dependent behavior.
Upsolve AI is a tool for building grounded, governed, and trustworthy data agents.
This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.
This position paper proposes the TRISM framework that integrates NeuroSymbolic AI with LLMs and RAG to address hallucination and interpretability issues in legal AI, introducing RASOR RAG for generating interpretable rationales and formalizing symbolic legal knowledge bases.