Studying FLUX in diffusers library was hard, so I built a smaller open-source version [P]
Summary
A simplified open-source PyTorch implementation of FLUX diffusion transformers with verifiable line-by-line source mappings, designed for educational purposes.
View Cached Full Text
Cached at: 06/20/26, 06:22 PM
purohit10saurabh/minFLUX
Source: https://github.com/purohit10saurabh/minFLUX
minFLUX
(Unofficial) Minimal PyTorch implementation of FLUX diffusion transformers
A simplified educational PyTorch implementation of FLUX.1 and FLUX.2 diffusion transformers (DiT) by Black Forest Labs. Built for understanding rectified flow matching, joint attention, and the key design choices behind FLUX with verifiable line-by-line source mappings to the official codebases.
The diffusion models architectures and training algorithms are inferred from the official diffusers repo. The VAE architectures are from the official BFL repos (flux and flux2). Each .py file has an accompanying .md file with extensive mapping of every function to its exact source lines at pinned commits.
What’s Inside
- FLUX.1 and FLUX.2 DiT architectures- Double-stream and single-stream transformer blocks with joint attention
- Rectified flow matching- Training with velocity prediction and logit-normal timestep sampling
- Euler ODE inference- Sampling loop with configurable timestep schedules
- VAE encoder/decoder- Resnet and attention based architectures with latent normalization
- Verifiable line-by-line source mappings- To the official codebases
Diffusion Equations
Training (rectified flow matching):
\begin{aligned} x_t &= (1 - \sigma(t)) \cdot x_0 + \sigma(t) \cdot \epsilon && \text{(noisy input)} \\ v &= \epsilon - x_0 && \text{(velocity target)} \\ L &= \left\| model(x_t, t) - v \right\|^2 && \text{(MSE loss)} \end{aligned}
Inference (Euler ODE step):
x_{t_{\text{next}}} = x_t + (\sigma(t_{\text{next}}) - \sigma(t)) \cdot model(x_t, t)
FLUX.2 Architecture Overview
Detailed architecture: FLUX.2 Model Architecture
Few differences between FLUX.1 and FLUX.2
| Component | FLUX.1 | FLUX.2 |
|---|---|---|
| Text encoder | CLIP + T5 | Mistral3 |
temb (modulation signal) | Timestep + guidance + pooled CLIP text | Timestep + guidance only |
VAE z_channels | 16 | 32 |
| VAE normalization | Scale/shift | Patchify (2×2) + BatchNorm |
| FFN | GELU | SwiGLU |
| Single-stream block | Separate attn + MLP | Fused QKV+MLP projection |
| Modulation | Per-block AdaLN | 3 shared heads (img, txt, single) |
| RoPE | theta=10000, axes=(16,56,56) | theta=2000, axes=(32,32,32,32) |
| Position IDs | 3D (ch, H, W) | 4D (T, H, W, L) |
| Biases | bias=True | bias=False |
| Blocks | 19 double + 38 single, 24 heads | 8 double + 48 single, 48 heads |
Repository Structure
flux1/ FLUX.1
model.py DiT (double + single stream)
training.py flow matching + pack/unpack
kontext_training.py reference-image conditioning
inference.py Euler ODE sampling
vae.py VAE (scale/shift)
flux2/ FLUX.2
model.py DiT (shared modulation, SwiGLU)
training.py flow matching + 4D position IDs
inference.py Euler ODE (empirical mu shift)
vae.py VAE (patchify + BatchNorm)
utils/ shared
model.py embeddings, RoPE, attention, norms
training.py noise, loss, Euler step, train loop
vae_utils.py ResNet, attention, up/down blocks
tests/
test_utils.py unit tests
Each .py file in flux1/, flux2/, and utils/ has an accompanying .md file with line-by-line mappings to the source-of-truth repos.
Setup
pip install -r requirements.txt
python -m pytest tests/ -v
Contributing
Contributions are greatly welcome, especially for:
- Source-of-truth: cross-reference code against diffusers, flux, and flux2 and fix any implementation discrepancies.
- Documentation: improve the accompanying
.mdfiles and update line mappings when diffusers changes. - Components: add missing FLUX components or improve existing ones.
Feel free to open an issue or create a pull request.
Disclaimer
Since minFLUX is inferred from the official diffusers and BFL repos, the possible sources of bugs in the code are:
- AI-assisted: This repo is vibe-coded. It is written with the help of AI, referencing the diffusers and BFL repos. Some training details are inferred from other works like dreambooth.
- Simplifications: Stripping ControlNet, IP-Adapter, gradient checkpointing, KV caching, FSDP/DeepSpeed support, and the attention processor dispatch pattern makes it incompatible with pretrained weights.
- Upstream code changes: Source-of-truth line numbers reference specific commits (diffusers, flux, flux2). These codebases change frequently, so functions may move, rename, or change signature.
Citation
If you use this repository, please cite it as:
@misc{minflux2026,
author = {Purohit, Saurabh},
title = {minFLUX: Minimal Pytorch Implementation of FLUX Diffusion Transformers},
year = {2026},
publisher = {GitHub},
url = {https://github.com/purohit10saurabh/minFLUX}
}
Similar Articles
prunaai/flux-fast
PrunaAI presents Flux Fast, an optimized endpoint for Black Forest Labs' FLUX.1-dev model, claiming the fastest Flux inference via compression, caching, and compilation.
black-forest-labs/FLUX.1-dev
Black Forest Labs releases FLUX.1-dev, a 12-billion parameter open-weights text-to-image transformer model, available on Hugging Face with API endpoints and local inference support.
FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
FLUX3D introduces a framework for high-fidelity image-to-3D Gaussian Splatting generation by enhancing representation learning and cross-modal alignment with diffusion-aligned structured latents and a sparse-structure-aware diffusion transformer, achieving state-of-the-art results.
FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
FluxLite introduces a training-free, inference-time proposal-control framework for discrete diffusion models that compensates jump-rate perturbations via a graph-divergence term in the Feynman-Kac potential, yielding two samplers (HEU and D-VCG) that substantially reduce reweighting variance and sampling error over standard SMC baselines.
@RisingSayak: Tensor parallel loading in Diffusers got a massive upgrade On Flux.2-Dev DiT, with a TP degree of 4 (A10G): • 30.4s → 1…
Hugging Face Diffusers 上张量并行(tensor parallel)加载迎来重大优化:在 Flux.2-Dev DiT、TP=4(A10G)配置下,加载时间从 30.4s 降至 12.5s(约 2.4 倍提速),每 rank 峰值 CPU 内存从 64.1 GB 降至 6.8 GB(减少约 89%)。相关分布式推理(Accelerate 与 PyTorch Distributed)用法已更新到官方文档。