capability-alignment

Tag

Cards List
#capability-alignment

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper characterizes 'futile reasoning' in large language models, where models produce superficially valid but incorrect reasoning on tasks beyond their capability. They introduce CaRL, a capability-aligned reinforcement learning method that trains LLMs to abstain from futile reasoning while preserving performance.

0 favorites 0 likes
← Back to home

Submit Feedback