Tag
Introduces DeFAb, a verifiable benchmark for defeasible abduction in foundation models, comprising over 372K instances and revealing that current frontier models perform poorly on this form of logical reasoning, with accuracy as low as 23.5% under robust evaluation.