Tag
This paper finds systematic differences in outputs of large language models for health advice based on access modes like APIs and chatbot interfaces, undermining evaluation validity. It calls for model providers to enable faithful replication of consumer experiences for rigorous auditing.
VeryTrace is a zero-shot verification-and-repair framework that formalizes LLM reasoning traces into a compilable representation using a DSL, enabling step-level error localization through a hybrid of deterministic checks and LLM audits. It improves accuracy across math, robotics, and relational reasoning without domain-specific training.