Are recent LLM gains mostly from pretraining or post-training?

Reddit r/ArtificialInteligence News

Summary

A discussion question exploring whether recent LLM gains are driven more by pretraining or post-training techniques like RL and fine-tuning, given that both require significant compute.

from what i've read, recent frontier llms seem to use broadly similar transformer architectures, while many of the visible improvements (reasoning, coding, and agentic behavior) appear to come from post-training techniques such as supervised fine-tuning, RL, preference optimization, and tool-use training. at the same time, labs continue to spend enormous compute on pretraining with larger, higher-quality datasets, so i assume pretraining is still doing most of the heavy lifting. is there any research, ablation study, or industry experience that sheds light on how much each stage contributes to recent capability gains? is there a growing consensus that post-training is now the main differentiator between frontier models, or is pretraining still responsible for most of the improvements?
Original Article

Similar Articles

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

arXiv cs.LG

Harvard researchers challenge the standard LLM training pipeline by showing RL can be effectively applied during pre-training rather than only after SFT, finding that data composition matters more than model scale, and proposing parallel averaging of RL and SFT objectives that outperforms sequential approaches while preserving general capabilities.

Training a local LLM using CPT and RAG (with evals)

Reddit r/ArtificialInteligence

The article describes a project with experiments on training a local LLM using continued pretraining (CPT) and RAG for domain-specific knowledge, featuring comprehensive evaluations and findings.