@RisingSayak:Diffusers 的张量并行加载迎来重大升级。在 Flux.2-Dev DiT 上,张量并行度(TP)为 4(A10G):• 30.4s → 1…
摘要
Hugging Face Diffusers 上张量并行(tensor parallel)加载迎来重大优化:在 Flux.2-Dev DiT、TP=4(A10G)配置下,加载时间从 30.4s 降至 12.5s(约 2.4 倍提速),每 rank 峰值 CPU 内存从 64.1 GB 降至 6.8 GB(减少约 89%)。相关分布式推理(Accelerate 与 PyTorch Distributed)用法已更新到官方文档。
查看缓存全文
缓存时间: 2026/10/03 19:00
Diffusers 中张量并行加载迎来重大升级
在 Flux.2-Dev DiT 上,张量并行度(TP degree)为 4(使用 A10G)时:
- 加载时间:30.4s → 12.5s
- 每个 rank 的峰值 CPU 内存:64.1 GB → 6.8 GB
加载速度提升约 2.4 倍,CPU 内存减少约 89%。
查看这里:https://huggingface.co/docs/diffusers/main/en/training/distributed_inference#tensor-parallelism
分布式推理 · Hugging Face
来源:https://huggingface.co/docs/diffusers/main/en/training/distributed_inference
分布式推理将工作负载分配到多个 GPU 上。这是一种将更大模型装入显存的实用技术,也可以同时处理多个提示词以获得更高的吞吐量。
本指南将展示如何使用 Accelerate 和 PyTorch Distributed 进行分布式推理。
Accelerate
Accelerate 是一个库,旨在通过自动处理底层配置来简化在多个加速器上的推理和训练,让用户可以专注于自己的 PyTorch 代码。
使用以下命令安装 Accelerate:
uv pip install accelerate
在 Python 文件中初始化一个 accelerate.PartialState 类,以创建分布式环境。accelerate.PartialState 类负责进程管理、设备控制与分配,以及进程协调。
将 DiffusionPipeline 移动到 accelerate.PartialState.device,为每个进程分配一个 GPU。
import torch
from accelerate import PartialState
from diffusers import DiffusionPipeline
pipeline = DiffusionPipeline.from_pretrained(
"Qwen/Qwen-Image", dtype=torch.float16
)
distributed_state = PartialState()
pipeline.to(distributed_state.device)
将 split_between_processes 工具函数作为上下文管理器使用,可在进程之间自动分配提示词。
with distributed_state.split_between_processes(["a dog", "a cat"]) as prompt:
result = pipeline(prompt).images[0]
result.save(f"result_{distributed_state.process_index}.png")
调用 accelerate launch 运行脚本,并使用 --num_processes 参数设置要使用的 GPU 数量。
accelerate launch run_distributed.py --num_processes=2
参考这个最小化示例脚本来在多个 GPU 上运行推理。想了解更多,请查看 Distributed Inference with 🤗 Accelerate 指南。
PyTorch Distributed
PyTorch 的 DistributedDataParallel 实现了数据并行,即在每个设备上复制同一个模型,并行处理不同的数据批次。
在 Python 文件中导入 torch.distributed 和 torch.multiprocessing,以设置分布式进程组并在每个 GPU 上为推理生成进程。
import torch
import torch.distributed as dist
import torch.multiprocessing as mp
from diffusers import DiffusionPipeline
pipeline = DiffusionPipeline.from_pretrained(
"Qwen/Qwen-Image", dtype=torch.float16,
)
创建一个推理函数,并调用 init_process_group。该方法会创建一个分布式环境,包含后端类型、当前进程的 rank,以及 world_size(即参与的进程总数,例如 2 个 GPU 对应 world_size=2)。
将 pipeline 移动到 rank,并使用 get_rank 为每个进程分配 GPU。每个进程处理不同的提示词。
def run_inference(rank, world_size):
dist.init_process_group("nccl", rank=rank, world_size=world_size)
pipeline.to(rank)
if torch.distributed.get_rank() == 0:
prompt = "a dog"
elif torch.distributed.get_rank() == 1:
prompt = "a cat"
image = sd(prompt).images[0]
image.save(f"./{'_'.join(prompt)}.png")
使用 mp.spawn 根据 world_size 定义的进程数量创建进程。
def main():
world_size = 2
mp.spawn(run_inference, args=(world_size,), nprocs=world_size, join=True)
if __name__ == "__main__":
main()
调用 torchrun 运行推理脚本,并使用 --nproc_per_node 参数设置要使用的 GPU 数量。
torchrun --nproc_per_node=2 run_distributed.py
device_map
device_map 参数通过自动将模型组件分配到不同的 GPU 上来实现分布式推理。当模型无法放入单个 GPU 时,这尤其有用。你可以使用 device_map 在特定阶段选择性地加载和卸载所需的模型组件,如下例所示(假设有两个可用的 GPU)。
将 device_map="balanced" 设置为将文本编码器均匀分配到所有可用的 GPU 上。你可以使用 max_memory 参数为每个文本编码器分配最大内存。不要加载其他 pipeline 组件,以避免不必要的内存占用。
from diffusers import FluxPipeline
import torch
prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
pipeline = FluxPipeline.from_pretrained(
"black-forest-labs/FLUX.1-dev",
transformer=None,
vae=None,
device_map="balanced",
max_memory={0: "16GB", 1: "16GB"},
dtype=torch.bfloat16
)
with torch.no_grad():
print("Encoding prompts.")
prompt_embeds, pooled_prompt_embeds, text_ids = pipeline.encode_prompt(
prompt=prompt, prompt_2=None, max_sequence_length=512
)
文本嵌入计算完成后,将它们从 GPU 上移除,为扩散 Transformer 腾出空间。
import gc
def flush():
gc.collect()
torch.cuda.empty_cache()
torch.cuda.reset_max_memory_allocated()
torch.cuda.reset_peak_memory_stats()
del pipeline.text_encoder
del pipeline.text_encoder_2
del pipeline.tokenizer
del pipeline.tokenizer_2
del pipeline
flush()
将 device_map="auto" 设置为自动将模型分配到两个 GPU 上。该策略会优先将模型放置在最快的设备上,如果需要,再将模型放置在较慢的设备(如 CPU 或硬盘)上。将模型参数存储在较慢设备上的代价是推理延迟增加。
from diffusers import AutoModel
import torch
transformer = AutoModel.from_pretrained(
"black-forest-labs/FLUX.1-dev",
subfolder="transformer",
device_map="auto",
dtype=torch.bfloat16
)
运行
pipeline.hf_device_map可以查看各个模型在设备间的分布情况。这有助于追踪模型的设备放置。你也可以对 transformer 模型调用hf_device_map来查看它的分布。
将 transformer 模型添加到 pipeline 中,并设置 output_type="latent" 来生成潜变量。
pipeline = FluxPipeline.from_pretrained(
"black-forest-labs/FLUX.1-dev",
text_encoder=None,
text_encoder_2=None,
tokenizer=None,
tokenizer_2=None,
vae=None,
transformer=transformer,
dtype=torch.bfloat16
)
print("Running denoising.")
height, width = 768, 1360
latents = pipeline(
prompt_embeds=prompt_embeds,
pooled_prompt_embeds=pooled_prompt_embeds,
num_inference_steps=50,
guidance_scale=3.5,
height=height,
width=width,
output_type="latent",
).images
从内存中移除 pipeline 和 transformer,然后加载 VAE 来解码潜变量。VAE 通常足够小,可以加载到单个设备上。
import torch
from diffusers import AutoencoderKL
from diffusers.image_processor import VaeImageProcessor
vae = AutoencoderKL.from_pretrained(ckpt_id, subfolder="vae", dtype=torch.bfloat16).to("cuda")
vae_scale_factor = 2 ** (len(vae.config.block_out_channels) - 1)
image_processor = VaeImageProcessor(vae_scale_factor=vae_scale_factor)
with torch.no_grad():
print("Running decoding.")
latents = FluxPipeline._unpack_latents(latents, height, width, vae_scale_factor)
latents = (latents / vae.config.scaling_factor) + vae.config.shift_factor
image = vae.decode(latents, return_dict=False)[0]
image = image_processor.postprocess(image, output_type="pil")
image[0].save("split_transformer.png")
通过在特定阶段选择性地加载和卸载所需的模型,并将最大的模型分片到多个 GPU 上,就有可能在消费级 GPU 上运行大模型推理。
上下文并行(Context Parallelism)
上下文并行将输入序列切分到多个 GPU 上,以减少内存占用。每个 GPU 处理自己那一部分序列。
使用 set_attention_backend() 切换到更优化的注意力后端。可参考此表格查看完整可用后端列表。
大多数注意力后端都与上下文并行兼容。如果某个后端不兼容,请提交 issue。
Ring Attention
键(K)和值(V)表示使用 Ring Attention 在设备间进行通信。这确保每个切分都能看到其他所有 token 的 K/V。每个 GPU 为其本地的 K/V 计算注意力,然后将其传递给环中的下一个 GPU。没有任何单个 GPU 需要持有完整序列,从而降低了通信延迟。
将 ContextParallelConfig 传递给 transformer 模型的 parallel_config 参数。该配置支持 ring_degree 参数,用于决定使用多少设备进行 Ring Attention。
import torch
from torch import distributed as dist
from diffusers import DiffusionPipeline, ContextParallelConfig
def setup_distributed():
if not dist.is_initialized():
dist.init_process_group(backend="nccl")
rank = dist.get_rank()
device = torch.device(f"cuda:{rank}")
torch.cuda.set_device(device)
return device
def main():
device = setup_distributed()
world_size = dist.get_world_size()
pipeline = DiffusionPipeline.from_pretrained(
"black-forest-labs/FLUX.1-dev", dtype=torch.bfloat16
).to(device)
pipeline.transformer.set_attention_backend("_native_cudnn")
cp_config = ContextParallelConfig(ring_degree=world_size)
pipeline.transformer.enable_parallelism(config=cp_config)
prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
generator = torch.Generator().manual_seed(42)
image = pipeline(
prompt,
guidance_scale=3.5,
num_inference_steps=50,
generator=generator,
).images[0]
if dist.get_rank() == 0:
image.save(f"output.png")
if dist.is_initialized():
dist.destroy_process_group()
if __name__ == "__main__":
main()
上述脚本需要使用与 PyTorch 兼容的分布式启动器(如 torchrun)运行。--nproc-per-node 设置为可用 GPU 的数量。
torchrun --nproc-per-node 2 above_script.py
Ulysses Attention
Ulysses Attention 将序列切分到多个 GPU 上,并执行 all-to-all 通信(每个设备向所有其他设备发送/接收数据)。每个 GPU 最终持有仅对应注意力头子集的所有 token。每个 GPU 在其负责的注意力头上对所有 token 本地计算注意力,然后再次执行 all-to-all 以按 token 重新组织结果,供下一层使用。
ContextParallelConfig 通过 ulysses_degree 参数支持 Ulysses Attention。该参数决定使用多少设备进行 Ulysses Attention。
将 ContextParallelConfig 传递给 enable_parallelism()。
pipeline.transformer.enable_parallelism(config=ContextParallelConfig(ulysses_degree=2))
统一注意力(Unified Attention)
统一序列并行将 Ring Attention 和 Ulysses Attention 结合为一种高效处理长序列的方案。它首先应用 Ulysses 的 all-to-all 通信来重新分配注意力头和序列 token,然后使用 Ring Attention 处理重新分配后的数据,最后反向执行 all-to-all 以恢复原始布局。
这种混合方法充分利用了两种方法各自的优势:
- Ulysses Attention 高效地在注意力头维度上进行并行化
- Ring Attention 以极小的内存开销处理非常长的序列
- 两者结合,在注意力头和序列两个维度上实现了二维并行
ContextParallelConfig 通过同时指定 ulysses_degree 和 ring_degree 来支持统一注意力。使用的设备总数为 ulysses_degree * ring_degree,排列成一个二维网格,其中 Ulysses 组和 Ring 组相互正交(不重叠)。将 ulysses_degree 和 ring_degree 均设置为大于 1 的 ContextParallelConfig 传递给 enable_parallelism()。
pipeline.transformer.enable_parallelism(config=ContextParallelConfig(ulysses_degree=2, ring_degree=2))
统一注意力适用于设备数量足以排列成二维网格的场景(至少需要 4 个设备)。
我们在一个配备 4 张 H100 GPU 的节点上,使用此脚本对 Ulysses、Ring 和统一注意力进行了基准测试。结果总结如下:
| CP Backend | Time / Iter (ms) | Steps / Sec | Peak Memory (GB) |
|---|---|---|---|
| ulysses | 6670.78 | 97.50 | 33.85 |
| ring | 13076.49 | 23.82 | 56.02 |
| unified_balanced | 11068.70 | 54.52 | 33.85 |
从上表可以看出,Ulysses 提供了更好的吞吐量,但其可用设备数量受限于注意力头的数量——这一限制由统一注意力解决。
以下为原文被截断的内容
(原文在此处被截断,无法继续翻译。)
相似文章
@reprompting:今天阅读关于 NVIDIA Hopper 架构的内容 https://arxiv.org/pdf/2501.12084
这篇论文对 NVIDIA Hopper GPU 架构进行了多层级微基准测试分析,评估了 L2 分区缓存、第四代张量核心(FP8)、DPX 指令、分布式共享内存(DSM)和张量内存加速器(TMA)等新特性的性能表现,结果显示 TMA 异步编程可实现 1.5 倍矩阵乘法加速,FP8 性能接近 FP16 的两倍,DPX 指令可加速生物信息学算法至少 4.75 倍。
@aliez_ren: 这下真的舒服了,本地 70 tps 跑 GLM 5.2 几乎满血版! https://huggingface.co/madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid…
通过混合精度(MXFP8、NVFP4、NF3)量化,在4张96GB GPU上本地运行GLM-5.2(753B参数)几乎满血版,精度接近原始FP8,吞吐量达到70 tps。
@FakeMaidenMaker: 炸裂!这个开源项目能给自部署的大模型推理大幅提速、还省显存 GitHub 狂揽 9.2K star,已经加入 PyTorch 基金会,NVIDIA 的 Dynamo 也集成了它。 GitHub:https://github.com/LMC…
LMCache 是一个 KV 缓存管理层,通过缓存并复用 KV cache 来加速大模型推理、降低显存消耗,已获 9.2K star 并加入 PyTorch 基金会,被 NVIDIA Dynamo 集成。
在 diffusers 库中学习 FLUX 很难,所以我构建了一个更小的开源版本 [P]
一个简化的开源 PyTorch 实现,包含可逐行验证的源代码映射,专为教育目的设计。
又一次大型张量修复 b9820
通过在 ggml 后端中减少分割计算期间的同步来提升性能,新增异步 CUDA 复制功能,并让同步放宽机制在多个后端中更通用。