Tag
This paper introduces STAR, a method for spatiotemporally adaptive reward allocation in RL post-training for text-to-image diffusion models, improving compositional alignment and text rendering by focusing policy updates on relevant latent regions.