Tag
FluxLite introduces a training-free, inference-time proposal-control framework for discrete diffusion models that compensates jump-rate perturbations via a graph-divergence term in the Feynman-Kac potential, yielding two samplers (HEU and D-VCG) that substantially reduce reweighting variance and sampling error over standard SMC baselines.
This GitHub project uses the Jev model to create a lousy chatbot with various sampling strategies, developed with Claude's help as a fun experiment.
This paper introduces a method for vibe design agents to explore diverse UI alternatives by separating exploration from implementation through structured design specifications, as evaluated on 168 prompts and a large online experiment with over 300,000 tasks.
This paper introduces Newton Matching, a unified framework for fine-tuning and sampling in generative models, which addresses limitations of existing methods by treating learning as an iterative optimization process and leveraging conditional-matching structure.
The article discusses Poisson disk sampling, a technique for randomly placing points with a minimum distance, and details Bridson's efficient algorithm from a highly-cited 2007 paper, along with improvements for computer graphics and simulations.
The article questions why large language models generate different outputs for the same prompt even when all variables are constant, attributing it to inherent randomness in AI sampling processes.
The paper presents a theoretical framework for steering and scaling large language models via sampling algorithms, such as Sequential Monte Carlo and Replica Exchange, to improve generation quality without external supervision.
The article discusses the deprecation of Sampling in the MCP spec as of the 2026-07-28 changelog, shifting model-call costs from clients to servers, and advises how to check if a server relies on Sampling via code or logs.
Introduces a training-free Semantic-Aware Kernel Entropy (SAKE) guidance method for text diffusion models, using order-2 Rényi entropy over a kernel Gram matrix to balance fidelity and diversity during sampling. Experiments show improved Pareto frontier and multi-sample performance on reasoning-intensive tasks.
MindControl is a fork of llama.cpp that allows guiding the reasoning process of language models through injection during sampling, enabling more controlled outputs.
A paper titled 'Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity' has been accepted to ICML. It proposes a simple prompt-engineering trick for more diverse sampling, sparking debate over whether such work belongs at a top-tier ML conference.
Introduces CuBAS, an information-geometric framework for adaptive data selection in supervised classification that uses local curvature of the data manifold to identify informative samples, achieving improved accuracy across 30 benchmark datasets.
Introduces Bootstrap Flow-Map Tree (BFMT), a computationally efficient sampling framework for history-aware global search and alignment under budget constraints, enabling dynamic transition from exploration to refinement.
This paper provides a proof-oriented introduction to diffusion models, covering Langevin dynamics, score-based models, discretization, discrete diffusion, and inference-time control, intended for graduate students.
This article presents a technique to improve LLM creative writing by modifying the sampling process using entropy, aiming to reduce the generic 'LLM feel' in generated text.
The paper introduces VGB, a process-guided sampling algorithm with probabilistic backtracking, which significantly improves coding performance on tiny 0.5B models by being robust to verifier errors.
This paper proposes the Time-Reparameterized Cumulative Intensity Extrapolation (TR-CIE) sampler for discrete flow matching, which improves sampling quality under limited function evaluations by rescaling the time grid and reusing cached model outputs, with theoretical analysis and experiments on text and image generation.
Introduces Nexus Sampling, a training-free KV-cache eviction method using weighted reservoir sampling instead of deterministic top-k, improving long-context LLM inference under fixed memory budgets, matching dense attention performance at 80% eviction.
Recommended a deep guide on modern LLM sampling mechanisms, covering methods such as Temperature, Top-P, Mirostat, etc., of significant reference value for developers aiming to improve output quality.
This paper discovers that large language models partially exhibit emergent symmetry under retokenization—replacing a prompt's canonical tokenization with an alternative valid segmentation while preserving bytes exactly. The authors use this phenomenon to probe compositional understanding and propose retokenization as a novel inference-time sampling strategy that can recover solutions not found by conventional temperature sampling.