Tag
This perspective paper from arXiv articulates a four-stage inferential framework for evaluating foundation models as cognitive and developmental models, emphasizing that behavioral alignment alone is insufficient and must be embedded within theoretical commitments and contrastive evaluation.
A new neuromorphic AI framework inspired by cognitive science could complete tasks more efficiently than current approaches.
This paper investigates whether small foundation models fine-tuned on human behavioral data can serve as cognitive proxies, finding that scale matters little in-distribution but larger models generalize better out-of-distribution.
This paper introduces Predictive Set Theory, a formal generative framework for cognitive architecture that reconstructs cognition from set-theoretic operations and addresses limitations in predictive processing and Bayesian cognitive science.
This paper critiques recent methods claiming abstraction-first learning in large language models, showing that pure exemplar models can mimic abstraction-first learning depending on input distributional properties.
Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.
An encyclopedia entry explaining the computational theory of mind, which holds that the mind is a computational system. It covers Turing machines, the history of computationalism in cognitive science, and challenges from rival paradigms.
This paper investigates whether diagrammatic representations like Euler and linear diagrams improve LLM reasoning on syllogistic tasks, finding limited benefit compared to natural language or logical notation.
The paper introduces a multi-layer taxonomy for large language models comprising 14 capability domains and 91 subskills, drawing from human cognitive science to organize LLM evaluation beyond isolated tasks. It demonstrates operational utility by mapping 15,934 papers across major AI conferences, revealing concentrated attention on language-semantic competence and reasoning while identifying underexplored domains.
A book examining the role of algebraic models versus deep learning in natural language acquisition, featuring perspectives from leading researchers in computational linguistics, psychology, and mathematical linguistics.
This paper investigates whether large language models have a localized causal mechanism for handling the animacy concept, using circuit discovery on minimal pairs; they find an animacy circuit that is distributed and only partially generalizes.
The paper argues that general intelligence necessitates non-reducible constraints across multiple levels of description, presenting theoretical implications for AI and cognitive science.
This paper investigates whether language models' next-word prediction aligns with human cognitive processing by analyzing EEG signals and event-related potentials, finding that only surprisal correlates with human brain responses, especially for open-class words.
This paper distinguishes tool use from tool discovery in LLM agents, decomposing discovery into curiosity, recognition, and efficiency. It introduces the Lomekwi framework and demonstrates inverse scaling of recognition with model size in combinatorial games.
Harvard Business Review article discussing how AI tools can erode critical thinking and offering design principles to strengthen human reasoning instead.
This study demonstrates that contextual semantic relevance, measuring how strongly an incoming word relates to its recent semantic context, reliably predicts fMRI BOLD responses during naturalistic speech comprehension across two datasets, whereas surprisal (local probabilistic expectation) does not. The findings support that slow hemodynamic responses are especially sensitive to contextual semantic integration rather than local prediction.
A new benchmark test, EgoBabyVLM, challenges AI vision-language models to learn from video footage captured from baby head-cameras, revealing that current AI models fail to match the learning efficiency of infants and suggesting that baby-like learning architectures could lead to more efficient AI.
This study uses semantic entropy, an NLP embedding-based metric, to compare semantic memory navigation between blind and sighted individuals. Results show that visual experience influences entropy patterns, with sighted individuals having higher entropy for abstract concepts while blind individuals exhibit higher entropy for visually salient concrete concepts.
This paper presents a Bayesian model of intercomprehension—understanding a related language without training—using a noisy-channel approach. It compares the model's predictions to human behavior across three language pairs, showing better alignment than larger zero-shot models.
Introduces PM-Bench, a text-based benchmark for evaluating prospective memory in LLM agents, inspired by cognitive science. Experiments show that even the best method (GPT-5.4 agent) achieves only 65.1% F1 score, indicating significant challenges in reliable intention execution.