FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation

arXiv cs.CL Papers

Summary

FlowLM introduces a flow matching language model derived from pre-trained diffusion models via efficient fine-tuning, enabling high-quality few-step text generation that rivals 2,000-step diffusion sampling with far fewer training epochs.

arXiv:2605.20199v1 Announce Type: new Abstract: We present FlowLM, a flow matching language model transformed from pre-trained diffusion language models via efficient fine-tuning. By re-aligning the curved sampling trajectories of diffusion models into straight-line flows, FlowLM enables high quality few-step generation that rivals or even outperforms the quality of 2,000-step diffusion sampling with very few training epochs. Remarkably, finetuned FlowLM reaches performance saturation with only half as many training epochs as training from scratch, both approaches greatly outperforming the original diffusion model, thereby validating our method. Furthermore, we validate a more effective training objective for flow matching: predicting clean data to consistently guide the sampling process towards the true data distribution. Empirical results demonstrate that our approach is highly effective for high-quality, few-step text generation.
Original Article
View Cached Full Text

Cached at: 05/21/26, 06:31 AM

# FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
Source: [https://arxiv.org/abs/2605.20199](https://arxiv.org/abs/2605.20199)
[View PDF](https://arxiv.org/pdf/2605.20199)

> Abstract:We present FlowLM, a flow matching language model transformed from pre\-trained diffusion language models via efficient fine\-tuning\. By re\-aligning the curved sampling trajectories of diffusion models into straight\-line flows, FlowLM enables high quality few\-step generation that rivals or even outperforms the quality of 2,000\-step diffusion sampling with very few training epochs\. Remarkably, finetuned FlowLM reaches performance saturation with only half as many training epochs as training from scratch, both approaches greatly outperforming the original diffusion model, thereby validating our method\. Furthermore, we validate a more effective training objective for flow matching: predicting clean data to consistently guide the sampling process towards the true data distribution\. Empirical results demonstrate that our approach is highly effective for high\-quality, few\-step text generation\.

## Submission history

From: Runzhe Zhang \[[view email](https://arxiv.org/show-email/b36be57c/2605.20199)\] **\[v1\]**Mon, 6 Apr 2026 10:36:22 UTC \(3,537 KB\)

Similar Articles

LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling

Hugging Face Daily Papers

LangFlow presents the first continuous diffusion language model that rivals discrete diffusion approaches, challenging the long-held belief that continuous diffusion is inferior for language modeling. The work introduces key ingredients like optimal Gumbel-based noise scheduling and demonstrates competitive perplexity and transfer learning performance compared to discrete diffusion baselines.

Masked Language Flow Models

arXiv cs.CL

This paper introduces Masked Language Flow Models (MLFMs), which incorporate masking into flow-based language models to enable continuous flow for conditional generation and allow pretrained Masked Diffusion Models to be converted. The authors propose a novel sampler that alternates continuous denoising with discrete unmasking, demonstrating for the first time that flow-based language models can scale to downstream reasoning and instruction-following tasks.

Latent-Kernel Discrete Flow Maps for Few-Step Generation

arXiv cs.LG

The paper introduces Latent-Kernel Discrete Flow Maps (LKF), a flow-map kernel for discrete diffusion models that captures correlations between positions via a shared latent, enabling few-step generation without distillation and improving text generation perplexity.