Tag
Research shows that open-weight chat models fail at counting list items in distinct error modes, not a single phenomenon, with implications for model interventions and transfer learning.