factual-qa

Tag

Cards List
#factual-qa

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

arXiv cs.AI · 4d ago Cached

This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.

0 favorites 0 likes
← Back to home

Submit Feedback