llm-auditing

Tag

Cards List
#llm-auditing

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

arXiv cs.CL ↗ · 2026-09-16 Cached

This paper finds systematic differences in outputs of large language models for health advice based on access modes like APIs and chatbot interfaces, undermining evaluation validity. It calls for model providers to enable faithful replication of consumer experiences for rigorous auditing.

0 favorites 0 likes
#llm-auditing

VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification

arXiv cs.AI ↗ · 2026-06-24 Cached

VeryTrace is a zero-shot verification-and-repair framework that formalizes LLM reasoning traces into a compilable representation using a DSL, enabling step-level error localization through a hybrid of deterministic checks and LLM audits. It improves accuracy across math, robotics, and relational reasoning without domain-specific training.

0 favorites 0 likes
← Back to home

Submit Feedback