llm-pretraining

Tag

Cards List
#llm-pretraining

Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining

arXiv cs.LG ↗ · 6d ago Cached

Proposes Terminal Shrinkage Averaging (TSA) to separate learning-rate schedule from model estimator in LLM pretraining, improving validation quality and potentially accelerating benchmarks.

0 favorites 0 likes
#llm-pretraining

Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper analyzes how finite Newton-Schulz iterations in the Muon optimizer benefit nonsmooth nonconvex optimization by smoothing the polar map, providing convergence guarantees that match best-known bounds.

0 favorites 0 likes
#llm-pretraining

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

arXiv cs.AI ↗ · 2026-08-26 Cached

The paper formulates data mixing for large language model pretraining as a mixture experiment, applying response surface methodology and optimal experimental design to improve the efficiency and interpretability of proxy training runs.

0 favorites 0 likes
#llm-pretraining

@AdinaYakup: Ultra-FineWeb-L1 Open English web corpus for LLM pre-training from @OpenBMB https://huggingface.co/datasets/openbmb/Ult…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

OpenBMB releases Ultra-FineWeb-L1, an open English web corpus with over 1 trillion tokens for LLM pre-training, derived from Common Crawl and featuring advanced cleaning with Trafilatura 2.0.

0 favorites 0 likes
#llm-pretraining

Scaling Domain Data Repetition in LLM Pretraining

Hugging Face Daily Papers ↗ · 2026-08-14

This paper studies the trade-off in repeating high-quality domain data during LLM pretraining to maintain performance as models scale, finding that optimal repetition counts increase with model size and are negatively correlated with domain validation loss.

0 favorites 0 likes
#llm-pretraining

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper introduces MOSAIC, a framework that jointly optimizes sparse Mixture-of-Experts model architecture and hardware systems for large-scale pretraining, showing that compute-optimal sparsity is not necessarily cluster-optimal when MFU, communication costs, and parallel layouts are considered.

0 favorites 0 likes
#llm-pretraining

MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper introduces MESH, a memory-efficient Sinkhorn-based optimizer for Mixture-of-Experts (MoE) training that restores temporal momentum without storing full optimizer state, reducing memory by 62.5% while maintaining competitive evaluation loss compared to AdamW.

0 favorites 0 likes
#llm-pretraining

@burny_tech: some updates on the optimizer magic

X AI KOLs Timeline ↗ · 2026-07-24 Cached

A new NVIDIA paper proposes higher-order optimizers like Muon and SOAP as more efficient alternatives to AdamW for large-scale LLM pretraining.

0 favorites 0 likes
#llm-pretraining

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper improves higher-order optimizers SOAP and Muon for large-scale LLM pretraining, addressing instabilities at large batch sizes and introducing a layer-wise distributed optimizer compatible with Megatron-LM. Experiments show they consistently outperform AdamW at billion-parameter scales.

0 favorites 0 likes
#llm-pretraining

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

Hugging Face Daily Papers ↗ · 2026-06-29 Cached

This paper challenges the assumption that one-step gradient delay in asynchronous pipeline parallelism is inherently unstable, showing that degradation depends on optimizer choice. It demonstrates that optimizers like Muon are robust to one-step delay and introduces an error-feedback correction to further mitigate staleness, achieving near-synchronous performance in LLM pretraining up to 10B parameters.

0 favorites 0 likes
#llm-pretraining

@eisokant: Great blog post from @ArkadiiBessonov on our pretraining team!

X AI KOLs Timeline ↗ · 2026-06-27 Cached

A tweet shares a blog post discussing three methods for FP8 in LLM pretraining: per-tensor, blockwise, and MXFP8, focusing on how the scale is attached.

0 favorites 0 likes
#llm-pretraining

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

Introduces Holistic Data Scheduler (HDS), a reinforcement learning-based framework that dynamically adjusts data mixtures during LLM pre-training using a multi-objective reward function, achieving 44% fewer iterations to reach target perplexity and a 7.2% improvement on MMLU.

0 favorites 0 likes
#llm-pretraining

RegMix-D: Dynamic Data Mixing via Proxy Training Trajectories

arXiv cs.CL ↗ · 2026-06-18 Cached

RegMix-D extends RegMix to dynamic data mixing by using loss trajectories from proxy runs to predict optimal mixtures at multiple training stages, achieving improvements over static methods.

0 favorites 0 likes
#llm-pretraining

Characterizing Narrative Content in Web-scale LLM Pretraining Data

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

A fine-grained study of narrative features in web-scale LLM pretraining data, introducing NarraBERT and NarraDolma to measure narrative patterns and their distribution across sources.

0 favorites 0 likes
#llm-pretraining

Parallax: Parameterized Local Linear Attention for Language Modeling

Hugging Face Daily Papers ↗ · 2026-05-27 Cached

Introduces Parallax, a parameterized local linear attention mechanism with hardware-aware optimization that improves LLM pretraining efficiency and performance, achieving Pareto improvements at 0.6B and 1.7B scales.

0 favorites 0 likes
#llm-pretraining

@IntologyAI: Can coding agents do research? We release NanoGPT-Bench, an internal eval we’ve used to test agents on an AI R&D proble…

X AI KOLs Following ↗ · 2026-05-19 Cached

IntologyAI releases NanoGPT-Bench, an internal benchmark to evaluate coding agents on AI R&D tasks. Current agents recover only 9.3% of human progress, mostly through hyperparameter tuning, highlighting gaps in algorithmic research capabilities.

0 favorites 0 likes
#llm-pretraining

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

arXiv cs.CL ↗ · 2026-05-12 Cached

This paper introduces SimReg, a regularization technique for LLM pretraining that uses embedding similarity to improve training convergence by over 30% and boost zero-shot performance.

0 favorites 0 likes
#llm-pretraining

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

Hugging Face Daily Papers ↗ · 2026-05-04 Cached

InfoLaw is a data-aware scaling framework that predicts model loss based on token consumption, model size, data mixture weights, and repetition, enabling efficient data-recipe selection under varying compute budgets.

0 favorites 0 likes
← Back to home

Submit Feedback