Tag
This paper introduces domain-conditional position offsets, a learned vector added to initial token embeddings, to reduce the cold-start penalty in language models. The method trains in minutes on few documents, requires no model weight changes, and achieves up to 27% perplexity reduction across various model sizes.