Tag
The article announces the first ChineseBabyLM Challenge at NLPCC 2026, which asks researchers to train language models from scratch on 100 million Chinese tokens and evaluate them on NLU, cognitive alignment, and Hanzi knowledge, promoting data-efficient and cognitively plausible modeling for Chinese.