Tag
This paper finds systematic differences in outputs of large language models for health advice based on access modes like APIs and chatbot interfaces, undermining evaluation validity. It calls for model providers to enable faithful replication of consumer experiences for rigorous auditing.