Tag
This study explores scaling laws for sequence weighting in language model training, finding non-monotonic behavior where models transition from learning general patterns to data-specific patterns and back as scale increases.