Tag
The paper evaluates whether reasoning models exhibit systematicity by extending rule induction tasks from cognitive science, finding that models often fail on structurally equivalent variants despite solving individual tasks, suggesting a lack of systematicity in their reasoning abilities.
This paper argues that recent claims that neural networks have solved Fodor and Pylyshyn's systematicity challenge are premature. The authors show that the meta-learning for compositionality model fails to generalize out-of-distribution and behaves unsystematically even on in-distribution problems, concluding the challenge remains unmet.