abduction

Tag

Cards List
#abduction

DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models

arXiv cs.AI · 2026-06-18 Cached

Introduces DeFAb, a verifiable benchmark for defeasible abduction in foundation models, comprising over 372K instances and revealing that current frontier models perform poorly on this form of logical reasoning, with accuracy as low as 23.5% under robust evaluation.

0 favorites 0 likes
← Back to home

Submit Feedback