large-rollout

Tag

Cards List
#large-rollout

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

arXiv cs.LG · 2026-05-18 Cached

Introduces DualKV, a FlashAttention kernel variant that eliminates redundant prompt token computation in RL post-training (GRPO/DAPO), achieving up to 3.82x speedup on 30B MoE models.

0 favorites 0 likes
← Back to home

Submit Feedback