Tag
EPIG-Tree introduces compute-optimal branching for reinforcement learning, reducing gradient uncertainty in policy estimation, with empirical improvements over GRPO in control and language model environments.
This study introduces BudgetDoc, the first multimodal benchmark for evaluating model-budget-performance trade-offs in document tasks, and develops DRB, a lightweight estimator that predicts reasoning performance to optimize compute allocation and reduce costs in LLMs.
This paper presents a systematic scaling law study for text-to-image diffusion models, showing they scale predictably but require significantly more data per parameter than language models for optimal training.
This paper investigates the scaling properties of native multimodal pre-training, deriving compute-optimal model sizes and token counts under a fixed budget, and revealing distinct scaling behaviors for language and multimodal objectives.
Proposes a three-term scaling law that decouples model size, training steps, and batch size, enabling robust fitting with fewer runs and deriving scaling laws for suboptimal batch sizes.
Lilian Weng's blog post provides a comprehensive overview of scaling laws in deep learning, covering their derivation, compute-optimal allocation, and the debate between Kaplan et al. and Chinchilla.
This paper investigates data filtering for large model pretraining and finds that in the high-compute, data-scarce regime, filtering may not be necessary and can even be detrimental; sufficiently trained large models benefit from nominally low-quality data.
This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.
A modified scaling law accounting for data repetition effects provides compute-optimal training strategies for data-constrained scenarios, showing that beyond a point further repetition is counterproductive and compute is better spent on model capacity.