Studying FLUX in diffusers library was hard, so I built a smaller open-source version [P]

Reddit r/MachineLearning Tools

Summary

A simplified open-source PyTorch implementation of FLUX diffusion transformers with verifiable line-by-line source mappings, designed for educational purposes.

No content available
Original Article
View Cached Full Text

Cached at: 06/20/26, 06:22 PM

purohit10saurabh/minFLUX

Source: https://github.com/purohit10saurabh/minFLUX

minFLUX

(Unofficial) Minimal PyTorch implementation of FLUX diffusion transformers

License PRs Welcome

A simplified educational PyTorch implementation of FLUX.1 and FLUX.2 diffusion transformers (DiT) by Black Forest Labs. Built for understanding rectified flow matching, joint attention, and the key design choices behind FLUX with verifiable line-by-line source mappings to the official codebases.

The diffusion models architectures and training algorithms are inferred from the official diffusers repo. The VAE architectures are from the official BFL repos (flux and flux2). Each .py file has an accompanying .md file with extensive mapping of every function to its exact source lines at pinned commits.

What’s Inside

  • FLUX.1 and FLUX.2 DiT architectures- Double-stream and single-stream transformer blocks with joint attention
  • Rectified flow matching- Training with velocity prediction and logit-normal timestep sampling
  • Euler ODE inference- Sampling loop with configurable timestep schedules
  • VAE encoder/decoder- Resnet and attention based architectures with latent normalization
  • Verifiable line-by-line source mappings- To the official codebases

Diffusion Equations

Training (rectified flow matching):

\begin{aligned} x_t &= (1 - \sigma(t)) \cdot x_0 + \sigma(t) \cdot \epsilon && \text{(noisy input)} \\ v &= \epsilon - x_0 && \text{(velocity target)} \\ L &= \left\| model(x_t, t) - v \right\|^2 && \text{(MSE loss)} \end{aligned}

Inference (Euler ODE step):

x_{t_{\text{next}}} = x_t + (\sigma(t_{\text{next}}) - \sigma(t)) \cdot model(x_t, t)

FLUX.2 Architecture Overview

FLUX.2 Architecture Overview

Detailed architecture: FLUX.2 Model Architecture

Few differences between FLUX.1 and FLUX.2

ComponentFLUX.1FLUX.2
Text encoderCLIP + T5Mistral3
temb (modulation signal)Timestep + guidance + pooled CLIP textTimestep + guidance only
VAE z_channels1632
VAE normalizationScale/shiftPatchify (2×2) + BatchNorm
FFNGELUSwiGLU
Single-stream blockSeparate attn + MLPFused QKV+MLP projection
ModulationPer-block AdaLN3 shared heads (img, txt, single)
RoPEtheta=10000, axes=(16,56,56)theta=2000, axes=(32,32,32,32)
Position IDs3D (ch, H, W)4D (T, H, W, L)
Biasesbias=Truebias=False
Blocks19 double + 38 single, 24 heads8 double + 48 single, 48 heads

Repository Structure

flux1/                     FLUX.1
  model.py                   DiT (double + single stream)
  training.py                flow matching + pack/unpack
  kontext_training.py        reference-image conditioning
  inference.py               Euler ODE sampling
  vae.py                     VAE (scale/shift)

flux2/                     FLUX.2
  model.py                   DiT (shared modulation, SwiGLU)
  training.py                flow matching + 4D position IDs
  inference.py               Euler ODE (empirical mu shift)
  vae.py                     VAE (patchify + BatchNorm)

utils/                     shared
  model.py                   embeddings, RoPE, attention, norms
  training.py                noise, loss, Euler step, train loop
  vae_utils.py               ResNet, attention, up/down blocks

tests/
  test_utils.py              unit tests

Each .py file in flux1/, flux2/, and utils/ has an accompanying .md file with line-by-line mappings to the source-of-truth repos.

Setup

pip install -r requirements.txt
python -m pytest tests/ -v

Contributing

Contributions are greatly welcome, especially for:

  • Source-of-truth: cross-reference code against diffusers, flux, and flux2 and fix any implementation discrepancies.
  • Documentation: improve the accompanying .md files and update line mappings when diffusers changes.
  • Components: add missing FLUX components or improve existing ones.

Feel free to open an issue or create a pull request.

Disclaimer

Since minFLUX is inferred from the official diffusers and BFL repos, the possible sources of bugs in the code are:

  • AI-assisted: This repo is vibe-coded. It is written with the help of AI, referencing the diffusers and BFL repos. Some training details are inferred from other works like dreambooth.
  • Simplifications: Stripping ControlNet, IP-Adapter, gradient checkpointing, KV caching, FSDP/DeepSpeed support, and the attention processor dispatch pattern makes it incompatible with pretrained weights.
  • Upstream code changes: Source-of-truth line numbers reference specific commits (diffusers, flux, flux2). These codebases change frequently, so functions may move, rename, or change signature.

Citation

If you use this repository, please cite it as:

@misc{minflux2026,
  author = {Purohit, Saurabh},
  title  = {minFLUX: Minimal Pytorch Implementation of FLUX Diffusion Transformers},
  year   = {2026},
  publisher = {GitHub},
  url    = {https://github.com/purohit10saurabh/minFLUX}
}

Similar Articles

prunaai/flux-fast

Replicate Explore

PrunaAI presents Flux Fast, an optimized endpoint for Black Forest Labs' FLUX.1-dev model, claiming the fastest Flux inference via compression, caching, and compilation.

black-forest-labs/FLUX.1-dev

Hugging Face Models Trending

Black Forest Labs releases FLUX.1-dev, a 12-billion parameter open-weights text-to-image transformer model, available on Hugging Face with API endpoints and local inference support.

FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models

arXiv cs.LG

FluxLite introduces a training-free, inference-time proposal-control framework for discrete diffusion models that compensates jump-rate perturbations via a graph-divergence term in the Feynman-Kac potential, yielding two samplers (HEU and D-VCG) that substantially reduce reweighting variance and sampling error over standard SMC baselines.

@RisingSayak: Tensor parallel loading in Diffusers got a massive upgrade On Flux.2-Dev DiT, with a TP degree of 4 (A10G): • 30.4s → 1…

X AI KOLs Following

Hugging Face Diffusers 上张量并行(tensor parallel)加载迎来重大优化:在 Flux.2-Dev DiT、TP=4(A10G)配置下,加载时间从 30.4s 降至 12.5s(约 2.4 倍提速),每 rank 峰值 CPU 内存从 64.1 GB 降至 6.8 GB(减少约 89%)。相关分布式推理(Accelerate 与 PyTorch Distributed)用法已更新到官方文档。