@RisingSayak:Diffusers 的张量并行加载迎来重大升级。在 Flux.2-Dev DiT 上,张量并行度(TP)为 4(A10G):• 30.4s → 1…

X AI KOLs Following 工具

摘要

Hugging Face Diffusers 上张量并行(tensor parallel)加载迎来重大优化:在 Flux.2-Dev DiT、TP=4(A10G)配置下,加载时间从 30.4s 降至 12.5s(约 2.4 倍提速),每 rank 峰值 CPU 内存从 64.1 GB 降至 6.8 GB(减少约 89%)。相关分布式推理(Accelerate 与 PyTorch Distributed)用法已更新到官方文档。

Diffusers 的张量并行(tensor parallel)加载迎来重大升级。在 Flux.2-Dev DiT 上,张量并行度(TP)为 4(A10G): • 加载时间:30.4s → 12.5s • 每个 rank 的峰值 CPU 内存:64.1 GB → 6.8 GB 加载速度提升约 2.4 倍,CPU 内存占用减少约 89%。快来看看: https://huggingface.co/docs/diffusers/main/en/training/distributed_inference#tensor-parallelism
查看原文
查看缓存全文

缓存时间: 2026/10/03 19:00

Diffusers 中张量并行加载迎来重大升级

在 Flux.2-Dev DiT 上,张量并行度(TP degree)为 4(使用 A10G)时:

  • 加载时间:30.4s → 12.5s
  • 每个 rank 的峰值 CPU 内存:64.1 GB → 6.8 GB

加载速度提升约 2.4 倍,CPU 内存减少约 89%。

查看这里:https://huggingface.co/docs/diffusers/main/en/training/distributed_inference#tensor-parallelism


分布式推理 · Hugging Face

来源:https://huggingface.co/docs/diffusers/main/en/training/distributed_inference

分布式推理将工作负载分配到多个 GPU 上。这是一种将更大模型装入显存的实用技术,也可以同时处理多个提示词以获得更高的吞吐量。

本指南将展示如何使用 Accelerate 和 PyTorch Distributed 进行分布式推理。

Accelerate

Accelerate 是一个库,旨在通过自动处理底层配置来简化在多个加速器上的推理和训练,让用户可以专注于自己的 PyTorch 代码。

使用以下命令安装 Accelerate:

uv pip install accelerate

在 Python 文件中初始化一个 accelerate.PartialState 类,以创建分布式环境。accelerate.PartialState 类负责进程管理、设备控制与分配,以及进程协调。

将 DiffusionPipeline 移动到 accelerate.PartialState.device,为每个进程分配一个 GPU。

import torch
from accelerate import PartialState
from diffusers import DiffusionPipeline

pipeline = DiffusionPipeline.from_pretrained(
    "Qwen/Qwen-Image", dtype=torch.float16
)
distributed_state = PartialState()
pipeline.to(distributed_state.device)

将 split_between_processes 工具函数作为上下文管理器使用,可在进程之间自动分配提示词。

with distributed_state.split_between_processes(["a dog", "a cat"]) as prompt:
    result = pipeline(prompt).images[0]
    result.save(f"result_{distributed_state.process_index}.png")

调用 accelerate launch 运行脚本,并使用 --num_processes 参数设置要使用的 GPU 数量。

accelerate launch run_distributed.py --num_processes=2

参考这个最小化示例脚本来在多个 GPU 上运行推理。想了解更多,请查看 Distributed Inference with 🤗 Accelerate 指南。

PyTorch Distributed

PyTorch 的 DistributedDataParallel 实现了数据并行,即在每个设备上复制同一个模型,并行处理不同的数据批次。

在 Python 文件中导入 torch.distributed 和 torch.multiprocessing,以设置分布式进程组并在每个 GPU 上为推理生成进程。

import torch
import torch.distributed as dist
import torch.multiprocessing as mp

from diffusers import DiffusionPipeline

pipeline = DiffusionPipeline.from_pretrained(
    "Qwen/Qwen-Image", dtype=torch.float16,
)

创建一个推理函数,并调用 init_process_group。该方法会创建一个分布式环境,包含后端类型、当前进程的 rank,以及 world_size(即参与的进程总数,例如 2 个 GPU 对应 world_size=2)。

将 pipeline 移动到 rank,并使用 get_rank 为每个进程分配 GPU。每个进程处理不同的提示词。

def run_inference(rank, world_size):
    dist.init_process_group("nccl", rank=rank, world_size=world_size)

    pipeline.to(rank)

    if torch.distributed.get_rank() == 0:
        prompt = "a dog"
    elif torch.distributed.get_rank() == 1:
        prompt = "a cat"

    image = sd(prompt).images[0]
    image.save(f"./{'_'.join(prompt)}.png")

使用 mp.spawn 根据 world_size 定义的进程数量创建进程。

def main():
    world_size = 2
    mp.spawn(run_inference, args=(world_size,), nprocs=world_size, join=True)

if __name__ == "__main__":
    main()

调用 torchrun 运行推理脚本,并使用 --nproc_per_node 参数设置要使用的 GPU 数量。

torchrun --nproc_per_node=2 run_distributed.py

device_map

device_map 参数通过自动将模型组件分配到不同的 GPU 上来实现分布式推理。当模型无法放入单个 GPU 时,这尤其有用。你可以使用 device_map 在特定阶段选择性地加载和卸载所需的模型组件,如下例所示(假设有两个可用的 GPU)。

将 device_map="balanced" 设置为将文本编码器均匀分配到所有可用的 GPU 上。你可以使用 max_memory 参数为每个文本编码器分配最大内存。不要加载其他 pipeline 组件,以避免不必要的内存占用。

from diffusers import FluxPipeline
import torch

prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""

pipeline = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-dev",
    transformer=None,
    vae=None,
    device_map="balanced",
    max_memory={0: "16GB", 1: "16GB"},
    dtype=torch.bfloat16
)
with torch.no_grad():
    print("Encoding prompts.")
    prompt_embeds, pooled_prompt_embeds, text_ids = pipeline.encode_prompt(
        prompt=prompt, prompt_2=None, max_sequence_length=512
    )

文本嵌入计算完成后,将它们从 GPU 上移除,为扩散 Transformer 腾出空间。

import gc

def flush():
    gc.collect()
    torch.cuda.empty_cache()
    torch.cuda.reset_max_memory_allocated()
    torch.cuda.reset_peak_memory_stats()

del pipeline.text_encoder
del pipeline.text_encoder_2
del pipeline.tokenizer
del pipeline.tokenizer_2
del pipeline

flush()

将 device_map="auto" 设置为自动将模型分配到两个 GPU 上。该策略会优先将模型放置在最快的设备上,如果需要,再将模型放置在较慢的设备(如 CPU 或硬盘)上。将模型参数存储在较慢设备上的代价是推理延迟增加。

from diffusers import AutoModel
import torch

transformer = AutoModel.from_pretrained(
    "black-forest-labs/FLUX.1-dev",
    subfolder="transformer",
    device_map="auto",
    dtype=torch.bfloat16
)

运行 pipeline.hf_device_map 可以查看各个模型在设备间的分布情况。这有助于追踪模型的设备放置。你也可以对 transformer 模型调用 hf_device_map 来查看它的分布。

将 transformer 模型添加到 pipeline 中,并设置 output_type="latent" 来生成潜变量。

pipeline = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-dev",
    text_encoder=None,
    text_encoder_2=None,
    tokenizer=None,
    tokenizer_2=None,
    vae=None,
    transformer=transformer,
    dtype=torch.bfloat16
)

print("Running denoising.")
height, width = 768, 1360
latents = pipeline(
    prompt_embeds=prompt_embeds,
    pooled_prompt_embeds=pooled_prompt_embeds,
    num_inference_steps=50,
    guidance_scale=3.5,
    height=height,
    width=width,
    output_type="latent",
).images

从内存中移除 pipeline 和 transformer,然后加载 VAE 来解码潜变量。VAE 通常足够小,可以加载到单个设备上。

import torch
from diffusers import AutoencoderKL
from diffusers.image_processor import VaeImageProcessor

vae = AutoencoderKL.from_pretrained(ckpt_id, subfolder="vae", dtype=torch.bfloat16).to("cuda")
vae_scale_factor = 2 ** (len(vae.config.block_out_channels) - 1)
image_processor = VaeImageProcessor(vae_scale_factor=vae_scale_factor)

with torch.no_grad():
    print("Running decoding.")
    latents = FluxPipeline._unpack_latents(latents, height, width, vae_scale_factor)
    latents = (latents / vae.config.scaling_factor) + vae.config.shift_factor

    image = vae.decode(latents, return_dict=False)[0]
    image = image_processor.postprocess(image, output_type="pil")
    image[0].save("split_transformer.png")

通过在特定阶段选择性地加载和卸载所需的模型,并将最大的模型分片到多个 GPU 上,就有可能在消费级 GPU 上运行大模型推理。

上下文并行(Context Parallelism)

上下文并行将输入序列切分到多个 GPU 上,以减少内存占用。每个 GPU 处理自己那一部分序列。

使用 set_attention_backend() 切换到更优化的注意力后端。可参考此表格查看完整可用后端列表。

大多数注意力后端都与上下文并行兼容。如果某个后端不兼容,请提交 issue。

Ring Attention

键(K)和值(V)表示使用 Ring Attention 在设备间进行通信。这确保每个切分都能看到其他所有 token 的 K/V。每个 GPU 为其本地的 K/V 计算注意力,然后将其传递给环中的下一个 GPU。没有任何单个 GPU 需要持有完整序列,从而降低了通信延迟。

将 ContextParallelConfig 传递给 transformer 模型的 parallel_config 参数。该配置支持 ring_degree 参数,用于决定使用多少设备进行 Ring Attention。

import torch
from torch import distributed as dist
from diffusers import DiffusionPipeline, ContextParallelConfig

def setup_distributed():
    if not dist.is_initialized():
        dist.init_process_group(backend="nccl")
    rank = dist.get_rank()
    device = torch.device(f"cuda:{rank}")
    torch.cuda.set_device(device)
    return device

def main():
    device = setup_distributed()
    world_size = dist.get_world_size()

    pipeline = DiffusionPipeline.from_pretrained(
        "black-forest-labs/FLUX.1-dev", dtype=torch.bfloat16
    ).to(device)
    pipeline.transformer.set_attention_backend("_native_cudnn")

    cp_config = ContextParallelConfig(ring_degree=world_size)
    pipeline.transformer.enable_parallelism(config=cp_config)

    prompt = """
    cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
    highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
    """


    generator = torch.Generator().manual_seed(42)
    image = pipeline(
        prompt,
        guidance_scale=3.5,
        num_inference_steps=50,
        generator=generator,
    ).images[0]

    if dist.get_rank() == 0:
        image.save(f"output.png")

    if dist.is_initialized():
        dist.destroy_process_group()

if __name__ == "__main__":
    main()

上述脚本需要使用与 PyTorch 兼容的分布式启动器(如 torchrun)运行。--nproc-per-node 设置为可用 GPU 的数量。

torchrun --nproc-per-node 2 above_script.py

Ulysses Attention

Ulysses Attention 将序列切分到多个 GPU 上,并执行 all-to-all 通信(每个设备向所有其他设备发送/接收数据)。每个 GPU 最终持有仅对应注意力头子集的所有 token。每个 GPU 在其负责的注意力头上对所有 token 本地计算注意力,然后再次执行 all-to-all 以按 token 重新组织结果,供下一层使用。

ContextParallelConfig 通过 ulysses_degree 参数支持 Ulysses Attention。该参数决定使用多少设备进行 Ulysses Attention。

将 ContextParallelConfig 传递给 enable_parallelism()。

pipeline.transformer.enable_parallelism(config=ContextParallelConfig(ulysses_degree=2))

统一注意力(Unified Attention)

统一序列并行将 Ring Attention 和 Ulysses Attention 结合为一种高效处理长序列的方案。它首先应用 Ulysses 的 all-to-all 通信来重新分配注意力头和序列 token,然后使用 Ring Attention 处理重新分配后的数据,最后反向执行 all-to-all 以恢复原始布局。

这种混合方法充分利用了两种方法各自的优势:

  • Ulysses Attention 高效地在注意力头维度上进行并行化
  • Ring Attention 以极小的内存开销处理非常长的序列
  • 两者结合,在注意力头和序列两个维度上实现了二维并行

ContextParallelConfig 通过同时指定 ulysses_degree 和 ring_degree 来支持统一注意力。使用的设备总数为 ulysses_degree * ring_degree,排列成一个二维网格,其中 Ulysses 组和 Ring 组相互正交(不重叠)。将 ulysses_degree 和 ring_degree 均设置为大于 1 的 ContextParallelConfig 传递给 enable_parallelism()。

pipeline.transformer.enable_parallelism(config=ContextParallelConfig(ulysses_degree=2, ring_degree=2))

统一注意力适用于设备数量足以排列成二维网格的场景(至少需要 4 个设备)。

我们在一个配备 4 张 H100 GPU 的节点上,使用此脚本对 Ulysses、Ring 和统一注意力进行了基准测试。结果总结如下:

CP BackendTime / Iter (ms)Steps / SecPeak Memory (GB)
ulysses6670.7897.5033.85
ring13076.4923.8256.02
unified_balanced11068.7054.5233.85

从上表可以看出,Ulysses 提供了更好的吞吐量,但其可用设备数量受限于注意力头的数量——这一限制由统一注意力解决。

以下为原文被截断的内容

(原文在此处被截断,无法继续翻译。)

相似文章

@reprompting:今天阅读关于 NVIDIA Hopper 架构的内容 https://arxiv.org/pdf/2501.12084

X AI KOLs Timeline

这篇论文对 NVIDIA Hopper GPU 架构进行了多层级微基准测试分析,评估了 L2 分区缓存、第四代张量核心(FP8)、DPX 指令、分布式共享内存(DSM)和张量内存加速器(TMA)等新特性的性能表现,结果显示 TMA 异步编程可实现 1.5 倍矩阵乘法加速,FP8 性能接近 FP16 的两倍,DPX 指令可加速生物信息学算法至少 4.75 倍。

又一次大型张量修复 b9820

Reddit r/LocalLLaMA

通过在 ggml 后端中减少分割计算期间的同步来提升性能,新增异步 CUDA 复制功能,并让同步放宽机制在多个后端中更通用。