epistemic-modals

Tag

Cards List
#epistemic-modals

Benchmarking LLM Competence on Logical Inference over Probability Operators

arXiv cs.CL · yesterday Cached

This paper introduces a benchmark of 14,320 procedurally-generated prompts for evaluating LLMs on logical inference over probability operators like 'probably', 'might', and 'must'. Testing 29 models, the authors find systematic answer biases and show that only 9 exceed random chance.

0 favorites 0 likes
← Back to home

Submit Feedback