pre-pretraining

Tag

Cards List
#pre-pretraining

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

arXiv cs.CL · 2026-08-11 Cached

This paper investigates whether pretraining LLMs on artificial languages (pre-pretraining) consistently improves token efficiency across multiple natural languages, finding that gains are highly dependent on experimental setup and random seed, though stable gains appear for small models with the Llama tokenizer.

0 favorites 0 likes
#pre-pretraining

Language Acquisition Device in Large Language Models

arXiv cs.CL · 2026-05-19 Cached

This paper proposes LAD-inspired pre-pretraining using a formal language called MP-Struct that encodes natural-language-like structures. It shows that this approach improves token efficiency and imparts human-like resistance to structurally implausible languages, challenging prior hypotheses about effective pre-pretraining languages.

0 favorites 0 likes
← Back to home

Submit Feedback