baby-lm

Tag

Cards List
#baby-lm

Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

arXiv cs.CL · 5d ago Cached

An autonomous research program by Qiushi Engine conducted end-to-end research on BabyLM 2026 Strict-Small, improving data-efficient language models through principle-guided methods and achieving the highest score in the public snapshot.

0 favorites 0 likes
#baby-lm

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

arXiv cs.CL · 6d ago Cached

This paper introduces Looped GPT-BERT, which uses depth-wise parameter sharing to train a small language model with fewer parameters, achieving comparable performance to baselines in the BabyLM 2026 Strict-small setting.

0 favorites 0 likes
← Back to home

Submit Feedback