@rasbt: Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multin…
Summary
This article covers inference scaling techniques for AI text generation, such as temperature scaling, top-p filtering, and self-consistency, aiming to improve answer accuracy by over 2x through diverse sampling and majority voting.
View Cached Full Text
Cached at: 09/20/26, 05:23 PM
Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x)
00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we’ll implement 11:47 Notebook setup and model loading 17:43 Building a flexible text generation function 24:40 Chain-of-thought prompting 28:26 Sampling and output diversity 33:43 Next-token logits and greedy decoding 38:20 Temperature scaling step by step 42:46 Softmax and token probabilities 47:42 Multinomial sampling 54:51 Adding temperature sampling to text generation 59:31 Top-p filtering step by step 1:10:23 Adding top-p filtering to text generation 1:13:43 Sampling and LLM watermarking 1:16:01 Self-consistency and majority voting 1:20:36 Implementing self-consistency 1:29:02 MATH-500 results 1:35:01 Accuracy and compute tradeoffs 1:36:50 Next steps and self-refinement
Similar Articles
Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention
This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.
@lateinteraction: incidentally and on a more serious note, @dianetc_ and i have wondered for some time if RL for reasoning followed by a …
The article discusses a paper titled 'Reasoning-Intensive Regression' that proposes MENTAT, a lightweight method combining batch-reflective prompt optimization with neural ensemble learning to improve numerical score prediction from text in AI tasks, showing up to 65% improvement over baselines.
@yuetai12575: Excited to see the gains further survive as model scaling!
The article reports initial results from OpenRSI's Marin-Scaling-Ladder experiment, where AI agents autonomously propose, implement, and evaluate new optimizers across model scales from 550M to 2.5B, discovering PSPR with Codex (GPT-5.6).
A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference
This survey presents a unified perspective on self-improving test-time intelligence, connecting test-time adaptation, learning, and scaling for AI systems that refine their behavior during deployment using feedback-driven methods.
Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics
This paper introduces a method to predict best-of-N inference scaling gains for language models using cheap statistics from a single labeled validation-set sampling pass. A compact predictor with three core features achieves Spearman ρ=0.90 with actual gains, enabling screening of configurations before expensive reward-model scoring.