trustworthiness

Tag

Cards List
#trustworthiness

I don't know if this is useful but here's how I get consistent results with AI.

Reddit r/AI_Agents · 5d ago

The author describes a three-week experiment testing AI agent trustworthiness, finding that agreement between agents using the same model is unreliable, and outlines a system with human approval for critical decisions.

0 favorites 0 likes
#trustworthiness

Reviewing Model Collapse and Countermeasures

arXiv cs.AI · 2026-08-25 Cached

This paper provides an up-to-date overview of the phenomenon of model collapse in generative AI and reviews countermeasures to mitigate it, highlighting challenges and future research opportunities.

0 favorites 0 likes
#trustworthiness

At what point does an AI agent become useful enough to trust with real work?

Reddit r/AI_Agents · 2026-08-18

The article discusses the criteria for trusting AI agents with real-world tasks, questioning the balance between usefulness and risk, and seeks insights from users on practical workflows.

0 favorites 0 likes
#trustworthiness

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

arXiv cs.CL · 2026-08-13 Cached

This paper evaluates the trustworthiness of small language models across fairness, robustness, privacy, and ethics, comparing pre-trained SLMs with compressed larger models, and finds that quantization preserves trustworthiness better than pruning and that distillation can further enhance reliability.

0 favorites 0 likes
#trustworthiness

TRACE: Trustworthy Retrieval-Augmented Conversational Engine

arXiv cs.AI · 2026-08-12 Cached

TRACE is a proposed retrieval-augmented conversational engine for public service chatbots that improves constraint-aware recommendations by strengthening retrieval quality over noisy directories, reducing hallucinated responses.

0 favorites 0 likes
#trustworthiness

Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

arXiv cs.CL · 2026-08-04 Cached

A systematic cross-architecture empirical study measuring the trustworthiness cost of domain adaptation in small language models, finding that safety-preserving fine-tuning strategies do not reliably transfer alignment.

0 favorites 0 likes
#trustworthiness

Engineering Trustworthy Agentic AI for Critical Systems

arXiv cs.AI · 2026-07-22 Cached

This survey proposes a trustworthiness model for agentic AI in critical engineering systems, covering safety, robustness, transparency, accountability, and security across domains like power systems and autonomous vehicles.

0 favorites 0 likes
#trustworthiness

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

arXiv cs.AI · 2026-07-13 Cached

ConceptSMILE is a perturbation-based auditing framework for evaluating the reliability of concept-based explainable AI, tested on retinal fundus images.

0 favorites 0 likes
#trustworthiness

Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

arXiv cs.CL · 2026-07-10 Cached

This paper introduces GraphEVAL, a graph-based framework for quantifying uncertainty in LLM reasoning, and proposes a new metric, Graph Reasoning Coherence Score (GRCS), that captures semantic-structural consensus and detects confident hallucinations. The authors also present Graph Self-Consistency (GSC), a decoding strategy that prioritizes reasoning fidelity over nominal accuracy.

0 favorites 0 likes
#trustworthiness

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

arXiv cs.AI · 2026-07-10 Cached

This paper proposes SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents in decentralized energy markets, integrating an LLM-based Planner/Auditor layer and revealing a utility-safety trade-off.

0 favorites 0 likes
#trustworthiness

@cognition: Yesterday we launched SWE-1.7 built on the open-source Kimi K2.7. Concerns about Chinese base models are real: K2.7 com…

X AI KOLs Following · 2026-07-09 Cached

Cognition launched SWE-1.7, built on the open-source Kimi K2.7, with trustworthiness training to address concerns about Chinese base models.

0 favorites 0 likes
#trustworthiness

Should AI be able to prove what it knew at the time?

Reddit r/artificial · 2026-07-06

A thought experiment questioning whether AI systems should maintain a verifiable memory trail of their knowledge and beliefs at the time of decision-making to enhance trust and accountability.

0 favorites 0 likes
#trustworthiness

the problem isnt that AI is wrong, its that it's wrong so confidently

Reddit r/ArtificialInteligence · 2026-06-29

Discusses the issue of AI models producing incorrect answers with high confidence, highlighting the problem of overconfidence in AI outputs.

0 favorites 0 likes
#trustworthiness

@_akhaliq: paper:

X AI KOLs Following · 2026-06-26 Cached

This paper proposes Robust-TO, an agentic video understanding framework that integrates per-frame trustworthiness to address the Blind Trust Problem, achieving significant accuracy gains under realistic perturbations.

0 favorites 0 likes
#trustworthiness

Confidence-Aware Tool Orchestration for Robust Video Understanding

Hugging Face Daily Papers · 2026-06-25 Cached

Robust-TO addresses the Blind Trust Problem in video reasoning by integrating per-frame trustworthiness into an agentic framework, improving accuracy under realistic perturbations through calibrated evidence weighting and reliability-aware reasoning.

0 favorites 0 likes
#trustworthiness

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

arXiv cs.LG · 2026-06-17 Cached

This paper systematically evaluates foundation model representations for multimodal cancer analysis, benchmarking unimodal and multimodal fusion strategies on real-world cohorts, and assessing trustworthiness via conformal prediction.

0 favorites 0 likes
#trustworthiness

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

arXiv cs.CL · 2026-06-02 Cached

Introduces TrustLDM, a comprehensive benchmark for evaluating safety, privacy, and fairness of Language Diffusion Models, revealing that their alignment degrades with malicious post contexts. Proposes an automatic evaluation framework, TrustLDM-Auto, to identify vulnerable configurations.

0 favorites 0 likes
#trustworthiness

Smoothed Elicitation Complexity for Approximate $\Gamma$-calibration of Discrete Classification Tasks

arXiv cs.LG · 2026-05-25 Cached

This paper characterizes approximate property calibration for discrete properties in multiclass classification, using Lipschitz continuous properties as an intermediary to reduce complexity from the number of classes to the elicitation complexity dimension.

0 favorites 0 likes
#trustworthiness

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

Hugging Face Daily Papers · 2026-05-21 Cached

This paper challenges the assumption that current Vision-Language Models faithfully synthesize multimodal data, proposing an information-theoretic Modality Translation Protocol with new metrics (Toll, Curse, Fallacy of Seeing) to evaluate trustworthiness over traditional multimodal gain.

0 favorites 0 likes
#trustworthiness

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

arXiv cs.AI · 2026-05-20 Cached

This vision paper argues that trust in Agent-to-Agent (A2A) networks must be integrated from the ground up, as existing agent alignment techniques are insufficient to address systemic vulnerabilities like adversarial composition and semantic misalignment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback