flip-rates

Tag

Cards List
#flip-rates

Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

Hugging Face Daily Papers · 2026-06-14 Cached

This paper introduces a controlled protocol to evaluate answer stability in large language models by challenging correct answers with plausible counterarguments, revealing large variation in flip rates across models that accuracy metrics alone do not capture. The authors release the protocol, challenge records, and a curated MaxFlip challenge set to support stability evaluation.

0 favorites 0 likes
← Back to home

Submit Feedback