trustworthy-ai

Tag

Cards List
#trustworthy-ai

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

arXiv cs.AI · 3d ago Cached

This position paper argues that AI systems used in high-stakes decision-making should reason similarly to their users and faithfully communicate that reasoning, and outlines a research agenda for achieving such 'cognitively-aligned AI'.

0 favorites 0 likes
#trustworthy-ai

Scaling AI agents with trustworthy data

MIT Technology Review · 4d ago Cached

A report based on a survey of 300 data and technology executives examines how legacy data systems limit the effectiveness and scaling of AI agents in enterprises, highlighting that 'data leaders' who give agents broader data access experience greater trust and success.

0 favorites 0 likes
#trustworthy-ai

The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

arXiv cs.AI · 2026-08-06 Cached

This paper introduces the RAIL principles (Reasoning, Assurances, Interfacing, Learning) for designing neurosymbolic AI systems that integrate machine learning with symbolic reasoning, arguing this approach is crucial for reliable, efficient, and trustworthy AI.

0 favorites 0 likes
#trustworthy-ai

Explanation-Based Runtime Verification for Trustworthy ML-driven Optical Networks

arXiv cs.LG · 2026-07-24 Cached

This paper introduces explanation-based runtime verification for ML-driven optical networks, using model explanations to assess the soundness of individual decisions before execution in the network control loop, demonstrating effectiveness in intercepting erroneous decisions while preserving high automation rates.

0 favorites 0 likes
#trustworthy-ai

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

arXiv cs.AI · 2026-07-24 Cached

This paper evaluates citation faithfulness in agentic scientific synthesis systems, showing that current verifiers are unreliable with unsupported-citation rates varying from 3% to 18% depending on strictness. It proposes a gold-anchored evaluation protocol and a deployable guard that uses split-conformal prediction to provide a distribution-free bound on truly unsupported citations.

0 favorites 0 likes
#trustworthy-ai

Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

arXiv cs.CL · 2026-07-21 Cached

Proposes HALO, an architecture with six layers of defense to contain hallucination in enterprise AI systems, reframing 'zero hallucination' as a system-enforced property rather than a model property.

0 favorites 0 likes
#trustworthy-ai

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

arXiv cs.AI · 2026-07-20 Cached

This paper argues that current responsible AI practices fail to create a market that rewards trustworthiness, proposing independent, outcome-oriented certification to close the 'trust gap' by making AI trustworthiness measurable, comparable, and commercially rewarded.

0 favorites 0 likes
#trustworthy-ai

A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms

arXiv cs.AI · 2026-07-20 Cached

This paper critically analyzes tools and trust mark frameworks for operationalizing trustworthy AI, finding asymmetries in ethical focus and lifecycle coverage. It identifies implementation gaps and offers recommendations for more holistic AI governance.

0 favorites 0 likes
#trustworthy-ai

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

arXiv cs.LG · 2026-07-16 Cached

A systematic survey of Federated Explainable Artificial Intelligence (FedXAI), covering roles, architectures, evaluation practices, and open challenges. It presents a multi-axis taxonomy and discusses model-agnostic to interpretable-by-design approaches, highlighting gaps in standardization and privacy-aware evaluation.

0 favorites 0 likes
#trustworthy-ai

@OpenAI: AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red tha…

X AI KOLs · 2026-07-15 Cached

OpenAI announced that AI agents are being used to improve next-generation models, and that GPT-Red represents a new approach to using today's models to make tomorrow's models more robust, aligned, and trustworthy.

0 favorites 0 likes
#trustworthy-ai

Scalable and Trustworthy Earth Observation Foundation Models

arXiv cs.LG · 2026-07-10 Cached

This chapter reviews design principles and current landscape of foundation models for Earth observation, highlighting the need for domain-specific adaptation, physically plausible representations, and consistent evaluation benchmarks. It includes case studies on harmful algal bloom prediction and adaptive monitoring station selection.

0 favorites 0 likes
#trustworthy-ai

@rohanpaul_ai: Stronger agents will not come only from larger models, but from better systems around them. The problem is that many AI…

X AI KOLs Following · 2026-07-09 Cached

This tweet discusses the paper 'From Model Scaling to System Scaling' which argues that stronger AI agents require better system design (harness) including context control, memory, and routing, not just larger models.

0 favorites 0 likes
#trustworthy-ai

@FinanceYF5: Bin Yu built the PCS framework so trustworthy AI stops being a slogan and becomes something you can actually verify. As…

X AI KOLs Following · 2026-07-04 Cached

Bin Yu has developed the PCS framework to make trustworthy AI verifiable, addressing the challenge of increasingly opaque models.

0 favorites 0 likes
#trustworthy-ai

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

arXiv cs.AI · 2026-07-02 Cached

Theoria is a verification architecture that rewrites AI solutions into auditable state transitions, achieving high precision on HLE problems and detecting subtle errors like hidden premises and fabricated citations.

0 favorites 0 likes
#trustworthy-ai

Decentralized Assessment for Trustworthy AI (DATA)

Reddit r/artificial · 2026-06-26

The Decentralized Assessment for Trustworthy AI (DATA) is an ethical evaluation tool that allows users and communities to objectively audit AI companies based on leading ethical frameworks like UNESCO and EU guidelines.

0 favorites 0 likes
#trustworthy-ai

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

arXiv cs.CL · 2026-06-26 Cached

This paper investigates how contextual framing affects LLM responses in mental health interactions, finding systematic behavioral variation and demonstrating that internal representations encode framing information throughout transformer layers.

0 favorites 0 likes
#trustworthy-ai

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

arXiv cs.LG · 2026-06-18 Cached

This paper proposes a post-hoc certification framework for sparse autoencoder (SAE) based interpretability, deriving an upper bound on the frozen language model's risk using measurable quantities. The framework is validated on GPT-2 Small, Gemma-2B, and Llama-3-8B, showing non-vacuous bounds and revealing depth-dependent behavior.

0 favorites 0 likes
#trustworthy-ai

Upsolve AI

Product Hunt · 2026-06-17

Upsolve AI is a tool for building grounded, governed, and trustworthy data agents.

0 favorites 0 likes
#trustworthy-ai

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

arXiv cs.AI · 2026-06-17 Cached

This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.

0 favorites 0 likes
#trustworthy-ai

NeuroSymbolic AI for Legal AI-TRISM: Trustworthy, Reliable, Interpretable, Safe Models

arXiv cs.AI · 2026-06-16 Cached

This position paper proposes the TRISM framework that integrates NeuroSymbolic AI with LLMs and RAG to address hallucination and interpretability issues in legal AI, introducing RASOR RAG for generating interpretable rationales and formalizing symbolic legal knowledge bases.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback