Tag
This paper evaluates whether Cantonese-adapted language models better predict Cantonese reading using eye-tracking data, comparing models like CKIP GPT-2 and CantoneseLLM-7B. Results indicate that more extensive Cantonese-specific training improves predictive fit, though performance varies by information-theoretic measure.
This paper compares human and LLM scalar judgments for sentences with focus particles 'even' and 'only' across different response scale configurations, finding stable semantic-driven differences but noting the model's lack of response variability.
This paper proposes using off-the-shelf CLIP-style multimodal encoders with a bimodal attribution method to predict gaze behavior in visual world experiments, successfully replicating a seminal study on human predictive processing without fine-tuning.
This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.
This paper introduces STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets for psycholinguistic plausibility judgments. Experiments show that adding a global reasoning scratchpad and evaluator-guided refinement substantially improves generation quality, though near-boundary events remain challenging.
Introduces trajectory extrapolation error, a measure derived from transformer LM hidden states that predicts human reading times independently of and orthogonally to surprisal, revealing a dissociable component of incremental processing cost.
This paper investigates how LLMs produce different outcomes based on conversational context, finding that topic, rather than explicit user demographics, is the primary driver of disparities in high-stakes scenarios like salary advice.
A new study reveals that Italian and Dutch adults instinctively adapt their hand gestures in similar ways when teaching children, suggesting a shared communicative strategy across cultures.
This paper tests the Parse Multiplicity Mismatch Hypothesis, proposing that language models underpredict human processing difficulty in garden path sentences because they can consider more simultaneous parses. Using RNNGs with beam search, they find reducing the number of active parses increases predicted garden path effects, but not enough to fully capture human data.
This paper investigates how humans communicate under strict vocabulary limitations, comparing their incremental production strategies to greedy and globally optimal sampling algorithms using Sequential Monte Carlo inference with large language models.