Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

arXiv cs.AI Papers

Summary

This paper proposes Spire, a framework that formulates slide personalization as an inverse planning problem, using structural denoising and reinforcement learning to infer latent design intents without relying on explicit templates or verbose instructions.

arXiv:2607.00407v1 Announce Type: new Abstract: Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing tools (e.g., PowerPoint, Beamer) being used. However, relinquishing control over these tools makes the problem intractable to optimize end-to-end. To overcome this, we propose SPIRE, a principled framework to solve PSP approximately. By intentionally corrupting the visual structures of clean slides, SPIRE creates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning (RL). We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi-agent formulation strictly reduces policy gradient variance in RL. Extensive experiments demonstrate the superiority of SPIRE.
Original Article
View Cached Full Text

Cached at: 07/02/26, 05:40 AM

# Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Source: [https://arxiv.org/html/2607.00407](https://arxiv.org/html/2607.00407)
11institutetext:Purdue University22institutetext:Rutgers University33institutetext:University at Albany44institutetext:MicrosoftZihan DongPurdue University Rutgers University University at Albany MicrosoftLinjun ZhangPurdue University Rutgers University University at Albany MicrosoftHaoyu WangPurdue University Rutgers University University at Albany MicrosoftJing GaoPurdue University Rutgers University University at Albany Microsoft Emre KıcımanPurdue University Rutgers University University at Albany MicrosoftRanveer ChandraPurdue University Rutgers University University at Albany MicrosoftWei\-Ting Chen†Purdue University Rutgers University University at Albany Microsoft

###### Abstract

Slide design requires personalizing both deck themes and page layouts\. Yet, current AI agent\-based methods struggle with fine\-grained, page\-level design\. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page\-level Slide Personalization \(PSP\) unresolved\. To close this gap, this work formulates PSP as an inverse planning problem\. We propose to learn a design intent without assuming any knowledge of the specific executing tools \(e\.g\., PowerPoint, Beamer\) being used\. However, relinquishing control over these tools makes the problem intractable to optimize end\-to\-end\. To overcome this, we proposeSpire, a principled framework to solve PSP approximately\. By intentionally corrupting the visual structures of clean slides,Spirecreates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning \(RL\)\. We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi\-agent formulation strictly reduces policy gradient variance in RL\. Extensive experiments demonstrate the superiority ofSpire\.

\*\*footnotetext:Work done during an internship at Microsoft\.††footnotetext:Corresponding author\.## 1Introduction

![Refer to caption](https://arxiv.org/html/2607.00407v1/x1.png)Figure 1:Illustration ofSpire\. Trained with reinforcement learning,Spirelearns to infer design intents instead of using prespecified layout templates or lengthy user instructions as existing methods do\[[53](https://arxiv.org/html/2607.00407#bib.bib53),[10](https://arxiv.org/html/2607.00407#bib.bib10)\], offering both*theoretical*and*empirical*advantages\.Slide decks are a primary medium for communicating ideas in academia and industry\[[34](https://arxiv.org/html/2607.00407#bib.bib34),[13](https://arxiv.org/html/2607.00407#bib.bib13),[40](https://arxiv.org/html/2607.00407#bib.bib40)\]\. Yet, creating a high\-quality deck is a design\-intensive process\. Depending on the presenter and the target audience, the*same*content can call for different visual treatments\. These treatments are often personalized to a speaker’s habitual design, a lab’s visual identity, or an organization’s brand guidelines\. In short, practical slide generation is inherently a*personalized*task\.

Recent advances in multi\-modal large language models \(MLLMs\) have enabled agentic pipelines that automate slide generation for practical use\[[38](https://arxiv.org/html/2607.00407#bib.bib38),[9](https://arxiv.org/html/2607.00407#bib.bib9),[26](https://arxiv.org/html/2607.00407#bib.bib26),[10](https://arxiv.org/html/2607.00407#bib.bib10),[53](https://arxiv.org/html/2607.00407#bib.bib53),[45](https://arxiv.org/html/2607.00407#bib.bib45),[30](https://arxiv.org/html/2607.00407#bib.bib30)\]\. These systems typically decompose slide creation into modular stages such as outlining, asset extraction, layout arrangement, and iterative refinement, thereby improving deck\-level coherence via template selection, visual feedback, and multi\-agent coordination\[[53](https://arxiv.org/html/2607.00407#bib.bib53),[50](https://arxiv.org/html/2607.00407#bib.bib50),[30](https://arxiv.org/html/2607.00407#bib.bib30),[42](https://arxiv.org/html/2607.00407#bib.bib42),[16](https://arxiv.org/html/2607.00407#bib.bib16),[15](https://arxiv.org/html/2607.00407#bib.bib15)\]\. However, they struggle with fine\-grained*page\-level*design, which requires deciding visual hierarchy, element alignment, spacing, and styling choices\. Nonetheless, existing agentic systems handle page\-level layout and styling passively: Once the content is specified, the page\-level design is largely overlooked, either by following generic system\-defined templates\[[24](https://arxiv.org/html/2607.00407#bib.bib24),[53](https://arxiv.org/html/2607.00407#bib.bib53),[42](https://arxiv.org/html/2607.00407#bib.bib42),[15](https://arxiv.org/html/2607.00407#bib.bib15)\], or by requiring lengthy, explicit user instructions\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[30](https://arxiv.org/html/2607.00407#bib.bib30)\], instead of inferring the user’s latent intent\. Consequently, they are too coarse to achieve satisfactory Page\-level Slide Personalization \(PSP\)\.

To close this gap, we introduce the task of agentic PSP; the concept is depicted in Fig[1](https://arxiv.org/html/2607.00407#S1.F1)\. Our formulation is inspired by how human designers teach slide making in practice:*Given*a detailed plan elaborating the page specification such as layout structure and visual hierarchy, a generally capable executor \(e\.g\., an experienced designer unfamiliar with specific corporate styles\) can reliably reproduce a high\-quality personalized slide visual by following the plan step\-by\-step\.

But such actionable plans are infeasible to collect at scale in practice\. Detailed intent annotations are rarely available in real\-world slide corpora\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\], making direct supervised training impracticable\. As a result, PSP requires one to*proactively infer*the plan as a latent intent in order to actively guide the generation process\. Furthermore, this plan is naturally context\-dependent: \(enterprise\) users often maintain more than one set of intents depending on their specific scenarios\. Fortunately, such contextual intent can be largely specified by having the user provide a small number of reference slides\. To this end, we formulate PSP as a probabilistic problem that explicitly models the design intent as a latent random variable\. Given \(1\) user\-provided page assets and \(2\) a small set of reference slides, we infer the final, detailed plan, which is then seamlessly passed to a downstream visual designer for execution\. Intuitively, our formulated PSP aims to infer a plan that, once executed by a capable, non\-personalized executor, can most reliably reproduce the desired personalized slide given the user’s references\.*This probabilistic formulation treats design intent as a latent variable and provides a principled foundation for page\-level slide personalization\.*

Crucially, this latent\-intent formulation of PSP is computationally challenging to solve for two reasons\. First, it evaluates a proposed plan based on whether it enables the executor to reproduce the desired visual\. This requires a rendering likelihood, which lacks a clear, direct objective that one can optimize\. Naive image\-level similarity provides an unreliable estimate for slide quality, which is governed by discrete, structured design decisions rather than purely pixel\-wise closeness\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[39](https://arxiv.org/html/2607.00407#bib.bib39),[33](https://arxiv.org/html/2607.00407#bib.bib33),[44](https://arxiv.org/html/2607.00407#bib.bib44)\]\. Second, the executor that turns a plan into a slide is effectively a black box, which further impedes optimizing the PSP objective with standard numerical methods\[[35](https://arxiv.org/html/2607.00407#bib.bib35),[17](https://arxiv.org/html/2607.00407#bib.bib17)\]\. Therefore, an effective approximation is needed to make this objective tractable\. We detail this formulation and its computational challenges in Sec\.[2\.1](https://arxiv.org/html/2607.00407#S2.SS1)\.

As a remedy, in this work we proposeSpire\(StructuralPlanning viaInverseREconstruction\), a structural denoising framework that turns the hard\-to\-optimize*slide*personalization objective into a verifiable reconstruction problem that serves as a good proxy\. Our solution extends existing multi\-agent visual generation frameworks\[[11](https://arxiv.org/html/2607.00407#bib.bib11),[23](https://arxiv.org/html/2607.00407#bib.bib23),[41](https://arxiv.org/html/2607.00407#bib.bib41),[15](https://arxiv.org/html/2607.00407#bib.bib15),[14](https://arxiv.org/html/2607.00407#bib.bib14)\]to a principled training objective\. With the aforementioned PSP goal of learning a reference\-conditioned page\-level plan, we alternate between a critic that provides semantic feedback and a planner that updates its plan accordingly\. The key idea is to exploit the discrete structural nature of slides: instead of perturbing pixels, we intentionally corrupt a gold slide by applying random, element\-level structural perturbations to its layout, hierarchy, and styling, and record the corresponding discrepancies\.*This yields scalable self\-supervision where suboptimal parts to improve are known by construction\.*

Built upon the structural perturbation,Spiretrains two complementary components separately on this shared signal\. A critic learns to read the corrupted slide and the user’s references, and to produce structured, actionable feedback that pinpoints the discrepancies as semantic edit suggestions\. A planner learns to produce an executable, editable plan, either proposing a plan from scratch or revising an imperfect plan guided by the critic’s feedback\. As both components are trained to recover from controlled structural corruptions, the critic’s feedback becomes verifiable, and the planner learns to improve plans without relying on gradients through the black\-box executor, making latent\-intent PSP tractable in practice\. We resort to reinforcement learning\[[36](https://arxiv.org/html/2607.00407#bib.bib36),[47](https://arxiv.org/html/2607.00407#bib.bib47)\]to train the two agents\. Sec\.[2\.2](https://arxiv.org/html/2607.00407#S2.SS2)details the proposedSpire\. We provide theoretical analysis of its superiority in Sec\.[2\.3](https://arxiv.org/html/2607.00407#S2.SS3)\.*This principled method and its theoretical advantage are the main technical contribution of this work\.*

Our paper is organized as follows\. Sec\.[2](https://arxiv.org/html/2607.00407#S2)formulates PSP and details the proposedSpire\. Extensive experimental results in Sec\.[3](https://arxiv.org/html/2607.00407#S3)demonstrate the effectiveness of our method\. In the remaining part of this paper, we review related work in Sec\.[4](https://arxiv.org/html/2607.00407#S4), and conclude the paper in Sec\.[5](https://arxiv.org/html/2607.00407#S5)\.

## 2Proposed Method

This section formalizes the task of reference\-based page\-level slide personalization \(PSP\), which entails an intractable optimization problem due to contextual intent being latent\. We proposeSpireas a principled approximate solution\.

### 2\.1Problem Formulation

Suppose a \(enterprise\) user specifies a high\-level instructionxxabout slide content \(e\.g\., text content and visual elements\)111In this paper, we refer to page\-level objects on a slide \(e\.g\., text boxes, images, shapes, charts, and tables\) uniformly as*visual elements*\. A*text box*is a type of visual element that contains*text content*\., and providesKKreference slides

𝒟ref=\{sref\(1\),…,sref\(K\)\},\\displaystyle\\mathcal\{D\}\_\{\\text\{ref\}\}=\\\{s\_\{\\text\{ref\}\}^\{\(1\)\},\\ldots,s\_\{\\text\{ref\}\}^\{\(K\)\}\\\},to exemplify their contextually preferred style\. We define PSP as producing a target slidessthat can accurately reflect the user’s exact intent that is not elaborated in the verbalxx\. To reflect the need for creative design, we allow𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}not to cover the optimal layoutssshould use\.

We refer to this intent as a*plan*and consider it as a latent variable denoted byzz\. Intuitively,zzcan be a step\-by\-step actionable instruction that contains a rich page\-level specification \(e\.g\., layout coordinates and element styling\)\.

Following a well\-formedzz, a generally capable executorEEcan be unambiguously guided to produce a high quality visualssthat is*personalized*to the user’s needs, even ifEEitself is not directly tailored to the user’s preference\. From a probabilistic perspective, this implies that the visual realizationssis*conditionally independent*of the user context given the planzz:

p​\(s∣z,x,𝒟ref\)=p​\(s∣z\)\.\\displaystyle p\(s\\mid z,x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)=p\(s\\mid z\)\.Here the randomness is induced by the use of black\-boxEE, which can be a coding agent\[[27](https://arxiv.org/html/2607.00407#bib.bib27),[52](https://arxiv.org/html/2607.00407#bib.bib52),[7](https://arxiv.org/html/2607.00407#bib.bib7),[6](https://arxiv.org/html/2607.00407#bib.bib6),[1](https://arxiv.org/html/2607.00407#bib.bib1),[29](https://arxiv.org/html/2607.00407#bib.bib29),[16](https://arxiv.org/html/2607.00407#bib.bib16)\], an image generator\[[46](https://arxiv.org/html/2607.00407#bib.bib46),[8](https://arxiv.org/html/2607.00407#bib.bib8),[22](https://arxiv.org/html/2607.00407#bib.bib22),[28](https://arxiv.org/html/2607.00407#bib.bib28)\], or a human designer\. Note that disentangling the planzzandEEis on purpose, so that PSP does not need to be conducted whenever a new executorEEis introduced\.

Based on this definition, lettingπθ\\pi\_\{\\theta\}denote a multi\-modal language model*planner*, PSP can be formulated as learningπθ\\pi\_\{\\theta\}from a personal corpus

𝒟tr=\{\(x\(1\),\(s∗\)\(1\),𝒟ref\(1\)\),…,\(x\(J\),\(s∗\)\(J\),𝒟ref\(J\)\)\}\.\\displaystyle\\mathcal\{D\}\_\{\\text\{tr\}\}=\\\{\(x^\{\(1\)\},\(s^\{\*\}\)^\{\(1\)\},\\mathcal\{D\}\_\{\\text\{ref\}\}^\{\(1\)\}\),\\ldots,\(x^\{\(J\)\},\(s^\{\*\}\)^\{\(J\)\},\\mathcal\{D\}\_\{\\text\{ref\}\}^\{\(J\)\}\)\\\}\.Heres∗s^\{\*\}denotes the gold slide forxx\. We detail the construction of𝒟tr\\mathcal\{D\}\_\{\\text\{tr\}\}in the supplementary material\.

Given𝒟tr\\mathcal\{D\}\_\{\\text\{tr\}\}, the goal of PSP is to maximize the marginal likelihood ofs∗s^\{\*\}givenx,𝒟refx,\\mathcal\{D\}\_\{\\text\{ref\}\}\. This can be expressed as the following optimization objective:

𝒥​\(θ\)\\displaystyle\\mathcal\{J\}\(\\theta\)=𝔼\(x,s∗,𝒟ref\)∼𝒟tr​\[log⁡pθ​\(s∗∣x,𝒟ref\)\]\\displaystyle=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\}\\big\[\\log p\_\{\\theta\}\(s^\{\*\}\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\big\]=𝔼\(x,s∗,𝒟ref\)∼𝒟tr​\[log​∫πθ​\(z∣x,𝒟ref\)​p​\(s∗∣z\)​𝑑z\]\.\\displaystyle=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\}\\left\[\\log\\int\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\,p\(s^\{\*\}\\mid z\)\\,dz\\right\]\.\(1\)At a colloquial level, Eq\. \([1](https://arxiv.org/html/2607.00407#S2.E1)\) seeksπθ\\pi\_\{\\theta\}that can generate plans that can maximize the likelihood of black\-boxEEto reproduce \(generate\) the golds∗s^\{\*\}\. Unfortunately, optimizing Eq\. \([1](https://arxiv.org/html/2607.00407#S2.E1)\) is prohibitive due to two reasons\.

First, the likelihoodp​\(s∗∣z\)p\(s^\{\*\}\\mid z\)lacks an explicit form, which makes the objective intractable\. One common solution is to approximately optimize

p​\(s∗∣z\)∝−‖s∗−E​\(z\)‖2,\\displaystyle p\(s^\{\*\}\\mid z\)\\propto\-\\\|s^\{\*\}\-E\(z\)\\\|\_\{2\},in the pixel space, aiming to push generateds=E​\(z\)s=E\(z\)towardss∗s^\{\*\}pixel\-wise\. Yet, recent works\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[39](https://arxiv.org/html/2607.00407#bib.bib39),[44](https://arxiv.org/html/2607.00407#bib.bib44)\]showed that treating slides as pure images struggles to capture their discrete logical structure, offering unreliable guidance for the planner\.

Second, black\-boxEEcannot be differentiated through, which prevents direct computation of the gradient forπθ\\pi\_\{\\theta\}with respect tos=E​\(z\)s=E\(z\)given by

∇θlog⁡pθ​\(s=E​\(z\)∣x,𝒟ref\)\.\\displaystyle\\nabla\_\{\\theta\}\\log p\_\{\\theta\}\(s=E\(z\)\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\.In addition, zeroth\-order optimization methods\[[4](https://arxiv.org/html/2607.00407#bib.bib4),[25](https://arxiv.org/html/2607.00407#bib.bib25),[12](https://arxiv.org/html/2607.00407#bib.bib12),[19](https://arxiv.org/html/2607.00407#bib.bib19)\]proveineffective for this context for two reasons\. First, their high computational cost\[[31](https://arxiv.org/html/2607.00407#bib.bib31),[51](https://arxiv.org/html/2607.00407#bib.bib51),[32](https://arxiv.org/html/2607.00407#bib.bib32)\]makes them prohibitive for slide generation, where user preferences and instructions take diverse forms\. Second, these methods approximate gradients by probing the specific executorEE, which makes them prone to overfitting onaparticular instance\[[12](https://arxiv.org/html/2607.00407#bib.bib12)\], leading to limited generalizability ofπθ\\pi\_\{\\theta\}to other executors\[[32](https://arxiv.org/html/2607.00407#bib.bib32)\]\.

In the next section, we show that by leveraging the*structural*discreteness of slide visuals, we can circumvent these hurdles via a surrogate denoising objective, and we proposeSpireas an effective approximate solution\.

### 2\.2Spire: An Effective Solution for PSP

![Refer to caption](https://arxiv.org/html/2607.00407v1/x2.png)Figure 2:The overview of Personalized Slide Personalization \(PSP\) and the proposedSpire\. The Planner and Critic are trained to recover self\-supervised Structural Perturbation in a decomposed way with reinforcement learning \(red\)\. After training, they collaborate with a black\-box render executor to conduct PSP\.To tackle the computational challenge of Eq\. \([1](https://arxiv.org/html/2607.00407#S2.E1)\), we approximate the two parts therein with two complementary agents based on their intrinsic goals\. The two agents are trained to exploit the discrete structural nature of slides via a shared structural perturbation process, turning the problem into a verifiable reconstruction signal through a*denoising*process over the slide’s discrete structural space\. We dub our methodStructuralPlanning viaInverseREconstruction \(Spire\)\. Fig[2](https://arxiv.org/html/2607.00407#S2.F2)depicts the overview of PSP andSpire\.

We begin by checking the roles of Eq\. \([4](https://arxiv.org/html/2607.00407#S2.E4)\)

pθ​\(s∗∣x,𝒟ref\)=∫πθ​\(z∣x,𝒟ref\)⏟plannerproposal,​p​\(s∗∣z\)⏟renderlikelihood​𝑑z,\\displaystyle p\_\{\\theta\}\(s^\{\*\}\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)=\\int\\underbrace\{\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\_\{\\begin\{array\}\[\]\{c\}\\text\{planner \}\\text\{proposal,\}\\end\{array\}\}\\underbrace\{p\(s^\{\*\}\\mid z\)\}\_\{\\begin\{array\}\[\]\{c\}\\text\{render \}\\text\{likelihood\}\\end\{array\}\}\\,dz,\(4\)where the integrand consists of two parts\. The first term,πθ​\(z∣x,𝒟ref\)\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\), denotes how the plannerπθ\\pi\_\{\\theta\}proposes a planzzto reflect the user’s latent intent\. Next, the second termp​\(s∗∣z\)p\(s^\{\*\}\\mid z\)measures the*render*likelihood of the golden slides∗s^\{\*\}, depicting the plan’s*visual*validity\. PSP wants to adjust the planner proposal based onp​\(s∗∣z\)p\(s^\{\*\}\\mid z\), which, unfortunately, is intractable due to black\-boxEE\.

To bypass this challenge, we approximate the update by incorporating a*critic*agentCϕC\_\{\\phi\}to act as a learned proxy for the likelihood\. By learning to*discriminate*how a generationssvisually deviates from the golds∗s^\{\*\}, the critic captures the user’s implicit preferences and provides actionable semantic feedback, thereby providing guidance for the planner to update its proposalπθ​\(z∣x,𝒟ref\)\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)in context𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}\. With a reliable critic, the planner can iteratively improve its proposal: starting from an initial plan, it revises the plan based on the critique\.

Yet, reliably learning such a critic, and thus enabling critique\-guided planner updates, is also non\-trivial\. The difficulty is the lack of*verifiable*supervision: for a random generationss, one can hardly evaluate whether a critique is objectively correct, which makes subsequent planner training unreliable as well\.

Denoising Signal\.To solve this problem, we propose a “denoising” objective through a corruption\-reconstruction strategy\. Namely, we construct a verifiable self\-supervised signal that can be shared across the two agents’ training\. Specifically, we corrupt the gold slides∗s^\{\*\}into a perturbed slides~\\tilde\{s\}in a controlled way, and record the ground\-truth corruption operations\. Next, the critic is asked to criticizes~\\tilde\{s\}based on visual evidence, and its critique can be verified by checking the recorded corruption operations\. The same corruption record also helps train the planner to perform critique\-guided recovery in the*plan space*: it learns to revise a suboptimal planz~\\tilde\{z\}using the critiquecc\. We will detail this soon\.

Structural Noises\.We leverage the discrete structural nature of slides to design*structural corruptions*, rather than adding pixel\-level noise to the visuals\. Specifically, we represent each slides∗s^\{\*\}as a collection of visual elements that incorporate different element types\. Then, we apply random perturbation to each visual element by modifying its structural attributes\. Each element type admits its own perturbation spaces, e\.g\., position for layout, size for hierarchy, and color palette for styling\. The corruption occurs along three dimensions:*spatial layout*\(e\.g\., shifting coordinates\),*visual hierarchy*\(e\.g\., distorting element sizes\), and*stylistic coherence*\(e\.g\., altering text colors\)\. To prevent spurious correlations, we independently perturb each randomly selected element\.

Formally, assumes∗s^\{\*\}containsNNvisual elements\{ei\}i=1N\\\{e\_\{i\}\\\}\_\{i=1\}^\{N\}, and let𝒯​\(ei\)\\mathcal\{T\}\(e\_\{i\}\)denote the element type \(e\.g\., text box, image, shape\)\. We represent the corruption by an element\-level action list𝒜=\{ai\}i=1N\\mathcal\{A\}=\\\{a\_\{i\}\\\}\_\{i=1\}^\{N\}, where eachaia\_\{i\}perturbs some attributes ofeie\_\{i\}or leaves it unchanged\. Perturbation actions are sampled independently, i\.e\.,

q​\(𝒜∣s∗\)\\displaystyle q\(\\mathcal\{A\}\\mid s^\{\*\}\)=q​\(ai,…,aN∣s∗\)=q​\(ai,…,aN∣ei,…,eN\)​=\(a\)​∏i=1Nq𝒯​\(ei\)​\(ai∣ei\),\\displaystyle=q\(a\_\{i\},\\dots,a\_\{N\}\\mid s^\{\*\}\)=q\(a\_\{i\},\\dots,a\_\{N\}\\mid e\_\{i\},\\dots,e\_\{N\}\)\\overset\{\(a\)\}\{=\}\\prod\_\{i=1\}^\{N\}q\_\{\\mathcal\{T\}\(e\_\{i\}\)\}\(a\_\{i\}\\mid e\_\{i\}\),where\(a\)\(a\)holds by independence\. Hereq𝒯​\(ei\)\(⋅∣ei\)q\_\{\\mathcal\{T\}\(e\_\{i\}\)\}\(\\cdot\\mid e\_\{i\}\)denotes a perturbation distribution over actions for elementeie\_\{i\}\(admitting no perturbation\)\. Subscript𝒯​\(ei\)\\mathcal\{T\}\(e\_\{i\}\)indicates that action spaces depend on the element types\. Having the sampled action list𝒜\\mathcal\{A\}, we apply these actions tos∗s^\{\*\}and denote the corrupted slide bys~\\tilde\{s\}\. The inverse of each applied perturbation forms a discrepancy list:

𝒜diff=\{ai−1∣ai≠∅,1≤i≤N\},\\displaystyle\\mathcal\{A\}\_\{\\text\{diff\}\}=\\\{a\_\{i\}^\{\-1\}\\mid a\_\{i\}\\neq\\emptyset,1\\leq i\\leq N\\\},where∅\\emptysetdenotes leaving the element unchanged\. This yields self\-supervised triplets\(s~,s∗,𝒜diff\)\(\\tilde\{s\},s^\{\*\},\\mathcal\{A\}\_\{\\text\{diff\}\}\), providing scalable supervision without manual annotation\. Built upon the structural perturbation,Spirefactorizes the intractable PSP problem from Eq\. \([1](https://arxiv.org/html/2607.00407#S2.E1)\) into the following two complementary sub\-tasks\.

Structural Discrimination\.This task aims to train criticCϕC\_\{\\phi\}to identify and articulate the discrepancies in𝒜diff\\mathcal\{A\}\_\{\\text\{diff\}\}, which correspond to movings~\\tilde\{s\}toward the golds∗s^\{\*\}to maximize the*render likelihood*term in Eq\. \([4](https://arxiv.org/html/2607.00407#S2.E4)\)\. Specifically, the critic generates a critiqueccbased onxx, a suboptimal visuals~\\tilde\{s\}, and reference𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}:

c∼Cϕ\(⋅∣s~,x,𝒟ref\)\.\\displaystyle c\\sim C\_\{\\phi\}\(\\cdot\\mid\\tilde\{s\},x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\.Here,ccencompasses a list of issues and targeted corrections, translating errors ins~\\tilde\{s\}into actionable textual feedback\. We evaluateccagainst the full action list𝒜\\mathcal\{A\}with a capable LLM judge \(GPT\-4o\-mini\)\. For each elementeie\_\{i\}\(1≤i≤N1\\leq i\\leq N\), the judge verifies whethercccorrectly handles that element:

yi=J​\(c,ei,ai\)∈\{0,1\},1≤i≤N\.\\displaystyle y\_\{i\}=J\(c,e\_\{i\},a\_\{i\}\)\\in\\\{0,1\\\},\\qquad 1\\leq i\\leq N\.Hereyi=1y\_\{i\}=1indicates a correct decision: the critique identifies and corrects the perturbation whenai≠∅a\_\{i\}\\neq\\emptyset, or correctly refrains from reporting an issue whenai=∅a\_\{i\}=\\emptyset; and 0 otherwise\. Subsequently, we define the element\-level verification reward as the average verification accuracy over allNNelements:

Rvfy​\(c,𝒜\)=1N​∑i=1Nyi\.\\displaystyle R\_\{\\text\{vfy\}\}\(c,\\mathcal\{A\}\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}y\_\{i\}\.We train the critic using DAPO\[[47](https://arxiv.org/html/2607.00407#bib.bib47)\], a strong reinforcement learning algorithm, to maximize the expected verification\-format reward:

𝒥critic​\(ϕ\)=𝔼s∗∼𝒟trs~∼q\(⋅∣s∗\)​\[𝔼c∼Cϕ\(⋅∣s~,x,𝒟ref\)​\[Rvfy​\(c,𝒜\)×Rformat,c​\(c\)\]\]\.\\displaystyle\\mathcal\{J\}\_\{\\text\{critic\}\}\(\\phi\)=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s^\{\*\}\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\\\\ \\tilde\{s\}\\sim q\(\\cdot\\mid s^\{\*\}\)\\end\{subarray\}\}\\Big\[\\mathbb\{E\}\_\{c\\sim C\_\{\\phi\}\(\\cdot\\mid\\tilde\{s\},x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\big\[R\_\{\\text\{vfy\}\}\(c,\\mathcal\{A\}\)\\times R\_\{\\text\{format,c\}\}\(c\)\\big\]\\Big\]\.\(5\)We detail critique generation and format requirement in the supplementary material\. Note that our perturbation design, combined with element\-level evaluation, mitigates reward hacking\. As each perturbation type is applied independently with probability1/21/2, trivial always\-issue or never\-issue critiques cannot consistently score high\.

Structural Planning\.This task aims to produce actionable plans that translate user instructions into high\-quality slides\. The goal is to approximately improve the*planner proposal*term in Eq\. \([4](https://arxiv.org/html/2607.00407#S2.E4)\)\. To this end, we train the planner to \(i\) generate a good plan proposal, and \(ii\) update a proposal using the critic’s critique as a semantic gradient\. Formally, we refer to this goal of the planner as*iterative refinement*, which covers*initial proposal*\(generatingzzfrom scratch\) and*critique\-guided refinement*\(revising a suboptimalz~\\tilde\{z\}based on critiquecc\)\. The two subgoals can be expressed as a unified formulation\. Given instructionxx, the reference𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}, and optionally a suboptimal planz~\\tilde\{z\}to revise, with its corresponding critiquecc, the planner aims to generate a high quality planzzas

z∼πθ\(⋅∣x,𝒟ref,c​or​∅,z~​or​∅⏟optional\)\.\\displaystyle z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},\\underbrace\{c\\text\{ or \}\\emptyset,\\;\\tilde\{z\}\\text\{ or \}\\emptyset\}\_\{\\text\{optional\}\}\)\.\(6\)The visual quality of a planzzis measured with a discriminative*plan–visual matching reward*\. Specifically, a capable VLM judge \(Claude Opus 4\.5\) is given the planzzand a pair of visuals\(s∗,s~\)\(s^\{\*\},\\tilde\{s\}\), and is asked to choose which one matches the planzzbetter\. We rewardzzif the judge prefers the gold slides∗s^\{\*\}\. The judgment is conducted twice with the presenting order of the two visuals swapped to mitigate the positional bias\[[43](https://arxiv.org/html/2607.00407#bib.bib43)\]\. Given the presenting order\(s∗,s~\)\(s^\{\*\},\\tilde\{s\}\), we denote the judgment outcome as

o^→=𝕀​\[s∗≻s~\|z,\(s∗,s~\)\]∈\{0,1\},\\displaystyle\\hat\{o\}\_\{\\rightarrow\}=\{\\mathbb\{I\}\}\\Big\[s^\{\*\}\\succ\\tilde\{s\}\\,\\big\|\\,z,\(s^\{\*\},\\tilde\{s\}\)\\Big\]\\in\\\{0,1\\\},ando^←\\hat\{o\}\_\{\\leftarrow\}is defined similarly for the swapped presenting order\. Based on this, we define the*plan–visual matching reward*as

Rpvm​\(z,s∗,s~\)\\displaystyle R\_\{\\text\{pvm\}\}\(z,s^\{\*\},\\tilde\{s\}\)=12​\(𝕀​\[o^→=1\]\+𝕀​\[o^←=1\]\)\\displaystyle=\\frac\{1\}\{2\}\\Big\(\{\\mathbb\{I\}\}\[\\hat\{o\}\_\{\\rightarrow\}=1\]\+\{\\mathbb\{I\}\}\[\\hat\{o\}\_\{\\leftarrow\}=1\]\\Big\)To makeRpvmR\_\{\\text\{pvm\}\}reliable, we instruct the capable judge to base its decision*solely*on fidelity tozzrather than its own aesthetic preference\. Note thats~\\tilde\{s\}is produced through structural perturbations that alter*design styles*\(e\.g\., color palette\) instead of making the slide visually worse\. Without seeingxxand𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\},s~\\tilde\{s\}may look good on its own; therefore, correct*judgment*requires a discriminativezz\. As with the critic, we train the planner with DAPO to maximize the expected reward, combined with a format verification term:

𝒥plnr​\(θ\)\\displaystyle\\mathcal\{J\}\_\{\\text\{plnr\}\}\(\\theta\)=𝔼s∗∼𝒟trs~∼q\(⋅∣s∗\)​\[𝔼c∗∼Oracle​\(c\)​\[𝔼z∼πθ\(⋅∣x,𝒟ref,c∗\)​\[Rpvm​\(z,s∗,s~\)×Rformat,p​\(z\)\]\]\]\.\\displaystyle=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s^\{\*\}\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\\\\ \\tilde\{s\}\\sim q\(\\cdot\\mid s^\{\*\}\)\\end\{subarray\}\}\\Big\[\\mathbb\{E\}\_\{c^\{\*\}\\sim\\text\{Oracle\}\(c\)\}\\big\[\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c^\{\*\}\)\}\[R\_\{\\text\{pvm\}\}\(z,s^\{\*\},\\tilde\{s\}\)\\times R\_\{\\text\{format,p\}\}\(z\)\]\\big\]\\Big\]\.We defer more details about plan generation and format requirementsto the supplementary material\. Note that the iterative refinement goal designed for the planner requires high\-quality\(c∗,z~\)\(c^\{\*\},\\tilde\{z\}\)to train the agent so that it can reliably refine a given plan under critique\. To this end, we use an oracle VLM \(GPT\-5\) to \(i\) describe the perturbed slides~\\tilde\{s\}as a suboptimal planz~\\tilde\{z\}, and \(ii\) translate the discrepancy list𝒜diff\\mathcal\{A\}\_\{\\text\{diff\}\}\(provided as*hints*\[[49](https://arxiv.org/html/2607.00407#bib.bib49),[48](https://arxiv.org/html/2607.00407#bib.bib48),[18](https://arxiv.org/html/2607.00407#bib.bib18)\]\) into gold critiquec∗c^\{\*\}:

z~∼Oracle​\(z~∣s~\),c∗∼Oracle​\(c∣𝒜diff\)\.\\displaystyle\\tilde\{z\}\\sim\\text\{Oracle\}\(\\tilde\{z\}\\mid\\tilde\{s\}\),\\qquad c^\{\*\}\\sim\\text\{Oracle\}\(c\\mid\\mathcal\{A\}\_\{\\text\{diff\}\}\)\.

### 2\.3Theoretical Analysis

We end up this section with the theoretical advantages ofSpire\. Full versions are deferred to the supplementary material due to page limit\.

First, under regular conditions, optimizing theSpireobjective gives a gradient\-level approximation to that of the original PSP objective up to an explicit error bound, as formalized in the following Thm[2\.1](https://arxiv.org/html/2607.00407#S2.Thmtheorem1)\.

###### Theorem 2\.1\(Surrogate Validity for PSP, informal\)

Let𝒥^SP​\(θ,ϕ\)\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)be the empirical loss ofSpire\. Then there existsϵtot≥0\\epsilon\_\{\\mathrm\{tot\}\}\\geq 0such that

‖∇θ𝒥^SP​\(θ,ϕ\)−β4​∇θ𝒥​\(θ\)‖≤ϵtot\.\\displaystyle\\left\\\|\\nabla\_\{\\theta\}\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\frac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\right\\\|\\leq\\epsilon\_\{\\mathrm\{tot\}\}\.Hereϵtot\\epsilon\_\{\\mathrm\{tot\}\}aggregates critic\-calibration, structural\-surrogate, variational\-gap, and empirical optimization errors\. Full definitions are provided in the supplementary material\.

Crucially, the key error term is controlled thanks to our critic training on structural corruptions with verifiable ground truth supervision\. An untrained or pixel\-denoising\-trained C would not satisfy this bound in general\.

Furthermore, Thm[2\.2](https://arxiv.org/html/2607.00407#S2.Thmtheorem2)below proves thatSpiretrains the planner and critic*separately*, thereby eliminating executor\-induced noise\.

###### Theorem 2\.2\(Two\-agent Decomposition Stabilizes Optimization, informal\)

Letg^e2e\\hat\{g\}\_\{\\text\{e2e\}\}andg^2a\\hat\{g\}\_\{\\text\{2a\}\}denote the end\-to\-end and two\-agent policy\-gradient estimators, respectively; see the supplementary material for explicit definitions and derivation\. Then

Var​\(g^e2e\)=Var​\(g^2a\)\+Δexec∣c,Δexec∣c≥0\.\\displaystyle\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{e2e\}\}\)=\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{2a\}\}\)\+\\Delta\_\{\\text\{exec\}\\mid c\},\\qquad\\Delta\_\{\\text\{exec\}\\mid c\}\\geq 0\.SoVar​\(g^2a\)≤Var​\(g^e2e\)\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{2a\}\}\)\\leq\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{e2e\}\}\), with strict inequality for non\-zero executor\-induced noise\. See formal statement, expression, and imperfect\-critic extension in the supplementary material\.

In general, even if a planner\-executor\-critic design is adopted, standard training of the planner and critic still requires the executor to generate visuals for complete rollouts, retaining its induced noise\.Spireprovides a way to bypass the executor in our decomposed training, thereby enabling the noise reduction\.

## 3Experiment

We evaluate the performance ofSpireand ablate its components’ contributions\. Benefiting from the principled optimization presented before,Spireoffers strong PSP capability over strong baselines that rely on much larger GPT models\.

### 3\.1Dataset and Experiment Settings

Datasets\.We build thetrainandtestdata from Zenodo10k\[[53](https://arxiv.org/html/2607.00407#bib.bib53)\]by randomly sampling 200 decks\. For each deck, we perform a slide\-level held\-out split that preserves the deck’s chronological order\. The last 20% of slides are reserved for testing and the remaining slides are used for training \(or for retrieval for training\-free baselines\)\. SlideBench\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\]is used asout\-of\-distributiontest data\. We defer more data preparation details in the supplementary material\.

Backbones\.Both the critic and the planner are fine\-tuned from Qwen2\.5\-VL\-7B\-Instruct\[[2](https://arxiv.org/html/2607.00407#bib.bib2)\]\. At inference time, the planner and critic collaborate with a black\-box executorEEto perform multi\-round iterative generation\. We use GPT\-o4\-mini as the coding executor to implement the plan withpython\-pptx, and follow\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[39](https://arxiv.org/html/2607.00407#bib.bib39)\]to improve coding reliability\. No template is used in execution\.

Baselines\.We compareSpireagainst representative systems that cover both slide\-native pipelines and generic visual synthesis\. We include AutoPresent\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\], which generates a slide by deciding layout and style for the given instruction from scratch on its own\. We also include PPTAgent\[[53](https://arxiv.org/html/2607.00407#bib.bib53)\], which selects a template/style from the reference slides and re\-applies it to the target content through an edit\-based workflow\. Finally, following\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\], we include a generic image generation baseline using Stable Diffusion 3\.5\[[8](https://arxiv.org/html/2607.00407#bib.bib8)\]equipped with IP\-Adapter\[[46](https://arxiv.org/html/2607.00407#bib.bib46)\]which prompt the model to synthesize a slide\-like image\. To isolate the benefit of our training, we additionally report PSP results with planner/critic setting to: \(i\) frontier VLM \(o4\-mini\) and \(ii\) the corresponding base model \(without fine\-tuning\), while keeping the rest of the pipeline unchanged\.

Evaluations\.We evaluate the generation in terms of visual similarity against the gold slides, and reference\-free quality\. Following\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[39](https://arxiv.org/html/2607.00407#bib.bib39),[21](https://arxiv.org/html/2607.00407#bib.bib21)\], we adopt both visual\-similarity metrics and a VLM\-as\-a\-Judge protocol\. Specifically, for visual similarity, we treat the slides as rendered images and utilize metrics such as SSIM and CLIP to measure “reconstruction” quality\. For VLM\-as\-a\-judge evaluation, we follow\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[54](https://arxiv.org/html/2607.00407#bib.bib54),[21](https://arxiv.org/html/2607.00407#bib.bib21)\]and evaluate the generation quality from the perspectives of*faithfulness*,*color*,*layout*, and*overall aesthetics*\. We detail these in the supplementary material\.

### 3\.2Quantitative Results

Table 1:Quantitative results on*test*and*OOD*pages\. Best average results are inboldand the second best results areunderlined\.ModelVisual SimilarityVLM\-as\-JudgeSimssim↑\\mathrm\{Sim\}\_\{\\text\{ssim\}\}\\uparrowSimclip↑\\mathrm\{Sim\}\_\{\\text\{clip\}\}\\uparrow\\columncolorgray\!10AVGFaith↑\\uparrowColor↑\\uparrowLayout↑\\uparrowAest↑\\uparrow\\columncolorgray\!10AVGTest Pages*GPT\-based Models*AutoPresent\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\]0\.60620\.4431\\columncolorgray\!100\.52470\.61740\.58500\.41450\.4105\\columncolorgray\!100\.5069PPTAgent\[[53](https://arxiv.org/html/2607.00407#bib.bib53)\]0\.51260\.5491\\columncolorgray\!100\.53090\.10360\.56700\.43640\.3961\\columncolorgray\!100\.3758PSP \(o4\-mini\)0\.84400\.7683\\columncolorgray\!100\.80620\.55040\.57630\.42060\.3661\\columncolorgray\!100\.4784*7B\-level Models*SD 3\.5\[[8](https://arxiv.org/html/2607.00407#bib.bib8)\]0\.33720\.4496\\columncolorgray\!100\.39340\.15590\.42820\.25450\.1889\\columncolorgray\!100\.2569PSP \(Base\)0\.35780\.3317\\columncolorgray\!100\.34480\.29950\.45880\.29560\.2382\\columncolorgray\!100\.3230Spire\(Ours\)0\.81340\.7018\\columncolorgray\!100\.75760\.70720\.63300\.44700\.3787\\columncolorgray\!100\.5415OOD Pages*GPT\-based Models*AutoPresent\[[10](https://arxiv.org/html/2607.00407#bib.bib10)\]0\.59760\.5892\\columncolorgray\!100\.59340\.79230\.72420\.54770\.5201\\columncolorgray\!100\.6461PPTAgent\[[53](https://arxiv.org/html/2607.00407#bib.bib53)\]0\.48360\.6371\\columncolorgray\!100\.56040\.08390\.58240\.43920\.3831\\columncolorgray\!100\.3722PSP \(o4\-mini\)0\.72330\.7633\\columncolorgray\!100\.74330\.69750\.59470\.44720\.3958\\columncolorgray\!100\.5338*7B\-level Models*SD 3\.5\[[8](https://arxiv.org/html/2607.00407#bib.bib8)\]0\.31770\.5753\\columncolorgray\!100\.44650\.18110\.50750\.29340\.2246\\columncolorgray\!100\.3016PSP \(Base\)0\.39100\.3864\\columncolorgray\!100\.38870\.24320\.23090\.17160\.1292\\columncolorgray\!100\.1937Spire\(Ours\)0\.67850\.6954\\columncolorgray\!100\.68700\.93300\.75450\.65650\.5891\\columncolorgray\!100\.7333

Tab[1](https://arxiv.org/html/2607.00407#S3.T1)shows quantitative results on test and out\-of\-distribution \(OOD\) pages\. We note the following observations\.

Good PSP Performance\.From the table,Spireachieves the strongest overall performance, outperforming the GPT\-based AutoPresent \(0\.5069\) and PSP \(o4\-mini\) \(0\.4784\) in terms of judge score, while maintaining a highly competitive visual similarity score \(0\.7414\), substantially higher than the 7B\-level baselines SD 3\.5 and PSP \(Base\)\. This performance is notable given thatSpirerelies on only two 7B\-scale agents, yet achieves competitive performance compared to GPT\-based components\. We also note empirical failures in the baselines, attributed to their underlying mechanisms\. PPTAgent struggles with faithfulness \(0\.1036\), as realistic reference sets rarely offer perfect templates without structural loss222In reality,pptxtemplate files may not be accessible\.\. AutoPresent falls short when the verbose prompting is too coarse\.*These trends clearly support our PSP formulation\.*The suboptimal results of the training\-free PSP baselines highlight the necessity of preference\-aligned optimization; without it, the critic leverages its own generic preferences rather than reflecting the user’s contextual needs\.*This confirms thatSpire’s gains stem from the RL training, not merely the multi\-agent refinement loop\.*

Strong Generalizability\.On OOD pages,Spiredemonstrates strong generalizability and achieves the highest judge average, substantially outperforming baselines such as AutoPresent, PSP \(o4\-mini\), PPTAgent, SD 3\.5, and PSP \(Base\)\. Our method obtains the best scores across all judging dimensions and the second\-best visual similarity\. These results indicate thatSpiredoes not simply memorize deck\-specific patterns from the training corpus, but instead successfully leverages the reference slides as contextual evidence at inference time, allowing it to infer the appropriate design intent in unseen scenarios\.

Metric Discrepancy\.We also note that visual similarity and judge\-based scores are not perfectly aligned\. For instance, while PSP \(o4\-mini\) achieves the highest visual averages across both in\-distribution and OOD settings, its judge\-based quality remains substantially lower than that ofSpire\. This mismatch, as explained in Sec\.[2](https://arxiv.org/html/2607.00407#S2), highlights the challenge of delivering PSP by optimizing standardized metrics alone—a limitation also noted in the literature\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[33](https://arxiv.org/html/2607.00407#bib.bib33),[44](https://arxiv.org/html/2607.00407#bib.bib44)\]\.

These results collectively demonstrate the effectiveness of the proposedSpire\.

RefAutoPre\.PPTAgentSD 3\.5OursGoldTest![Refer to caption](https://arxiv.org/html/2607.00407v1/x3.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x5.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x6.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x7.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x8.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x11.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x12.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x13.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x14.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x15.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x16.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x17.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x18.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x19.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x21.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x22.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x23.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x24.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x26.png)OOD![Refer to caption](https://arxiv.org/html/2607.00407v1/x27.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x28.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x31.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x32.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x35.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x39.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x40.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)Figure 3:Qualitative comparison of slide generation results\.
### 3\.3Qualitative Results

We provide a visual comparison of the generated slides across different methods\. Illustrative results are shown in Fig[3](https://arxiv.org/html/2607.00407#S3.F3), with more examples deferred to the supplementary material\. We note thatSpireconsistently produces visually appealing and structurally coherent slides that closely align with the gold slides, while baselines often struggle to maintain visual fidelity or fail to generate valid structures\.

The effectiveness ofSpire, stemming from our preference\-aligned RL training, empowers the model to explicitly infer latent user intent\. This capability manifests in two aspects\. First,Spireachieves superior color coherence\. For instance, as shown in Row 4 and 5,Spireaccurately*infers*a proper coloring strategy that differs from the reference slide in OOD scenarios\. Second,Spiredemonstrates strong spatial reasoning for layout arrangement\. In Row 2, it successfully organizes complex, multi\-image visual assets \(e\.g\., heatmaps\) into a clean, aligned structure that matches the gold slide perfectly\. In Row 3, it also correctly positions the logo and text blocks\. This proactive inference capability allowsSpireto surpass baselines solely relying on verbose instructions \(AutoPresent\) or generic templates \(PPTAgent\)\. When such explicit inputs are unavailable, these methods inevitably fail to capture the user’s true intent\.

### 3\.4Ablation Study

Table 2:Ablation study onSpireTraining and Iterative Revision\. Both components are necessary for high\-quality generation\.Table 3:Ablation study on the critic role\. Substituting the larger, untrained GPT critic withSpire\-trained 7B model results in improved performance\.![Refer to caption](https://arxiv.org/html/2607.00407v1/x43.png)\(a\)Visual similarity\.
![Refer to caption](https://arxiv.org/html/2607.00407v1/x44.png)\(b\)Judge scores\.

Figure 4:Ablation onSpirerevision\. Multi\-round revision inSpirehelps improve both visual similarity and judge scores, indicating better generation qualities\.We end up this section with ablation study on key components inSpire\.

SpireTraining and Iterative Revision\.To break down the contributions of our proposed framework, we ablate the fine\-tuning process and the multi\-turn revision mechanism, with results reported in Tab\.[2](https://arxiv.org/html/2607.00407#S3.T2)\. First, we observe that preference\-aligned fine\-tuning is necessary for PSP\. Namely, when directly prompting base Qwen models to conduct PSP, they struggle significantly and achieve a visual average of only 0\.3344 and a judge average of 0\.3230\. In contrast, our fine\-tunedSpireachieves much better scores, reaching 0\.7414 and 0\.5415 respectively\. This massive performance gap confirms that off\-the\-shelf VLMs, especially small\-scale models, cannot reliably infer latent page\-level design intents without our targeted structural denoising optimization\. Second, the results validate the effectiveness of the iterative revision process\. As shown in the table, and Fig\.[4](https://arxiv.org/html/2607.00407#S3.F4), the visual similarity scores exhibit a clear and steady improvement through the successive refinement rounds\. This steady climb is a direct result of our structural denoising objective, which explicitly trains the planner to interpret and execute structural corrections based on the critic’s feedback\. Consequently, the planner is capable of progressively refining suboptimal layouts, proving thatSpiresuccessfully leverages high\-quality critiques to iteratively enhance the design rather than relying on a single\-pass generation\.

Critic Role\.One hypothesis stated in Sec\.[3\.2](https://arxiv.org/html/2607.00407#S3.SS2)is that the suboptimal performance of PSP lies in its critic\. To verify that the preference\-aligned training helps the critic learn the user’s contextual needs better, we replace the critic in the PSP \(o4\-mini\) baseline with our 7B\-levelSpire\-trained critic and report the results on test slides\. Results are reported in Tab[3](https://arxiv.org/html/2607.00407#S3.T3)\. From the table, we see that this substitution yields consistent improvements across all metrics\. Remarkably, the critic is a 7B\-level model, which is much less capable than GPT models\. This gain confirms that such a targeted critic indeed helps guide the PSP process\.

## 4Related Work

Early works consider automated slide generation as a text summarization task\[[38](https://arxiv.org/html/2607.00407#bib.bib38),[9](https://arxiv.org/html/2607.00407#bib.bib9),[3](https://arxiv.org/html/2607.00407#bib.bib3)\], aiming to extract\[[38](https://arxiv.org/html/2607.00407#bib.bib38),[9](https://arxiv.org/html/2607.00407#bib.bib9)\]and summarize\[[3](https://arxiv.org/html/2607.00407#bib.bib3)\]proper deck\-level content from a given document for presentation\. These solutions boiled down to extracting salient sentences, figures, and sections from source documents to organize them into slide outlines\. However, they overlooked the necessity of coherent layout design and the inherent multimodal nature of slides\[[50](https://arxiv.org/html/2607.00407#bib.bib50),[20](https://arxiv.org/html/2607.00407#bib.bib20)\]\. As a result, these solutions offered limited control over the aesthetic perspective of the generated slides, leaving personalized generation untouched\.

Following the emergence of \(M\)LLMs, recent research has shifted toward end\-to\-end and agentic pipelines for slide design\[[26](https://arxiv.org/html/2607.00407#bib.bib26),[10](https://arxiv.org/html/2607.00407#bib.bib10),[53](https://arxiv.org/html/2607.00407#bib.bib53),[50](https://arxiv.org/html/2607.00407#bib.bib50),[30](https://arxiv.org/html/2607.00407#bib.bib30),[42](https://arxiv.org/html/2607.00407#bib.bib42),[16](https://arxiv.org/html/2607.00407#bib.bib16),[15](https://arxiv.org/html/2607.00407#bib.bib15)\]\. Representative solutions decompose the task into modular stages, including content outlining, asset extraction, layout arrangement, and iterative refinement\. These systems improve visual consistency through template\-based selection, visual feedback, and multi\-agent coordination\[[26](https://arxiv.org/html/2607.00407#bib.bib26),[53](https://arxiv.org/html/2607.00407#bib.bib53),[50](https://arxiv.org/html/2607.00407#bib.bib50),[30](https://arxiv.org/html/2607.00407#bib.bib30),[42](https://arxiv.org/html/2607.00407#bib.bib42),[16](https://arxiv.org/html/2607.00407#bib.bib16)\]\. Nevertheless, these solutions mainly focus on more global deck\- or document\-level decisions and planning, lacking more fine\-grained page\-level layout design capabilities\. For such page\-level design, these solutions are either guided by generic templates and system\-defined objectives\[[53](https://arxiv.org/html/2607.00407#bib.bib53),[50](https://arxiv.org/html/2607.00407#bib.bib50)\], or by explicit and lengthy instructions specified by users\[[10](https://arxiv.org/html/2607.00407#bib.bib10),[30](https://arxiv.org/html/2607.00407#bib.bib30)\]\. These designs inherently limit their performance for PSP, which requires deep understanding of user\-specific design intents\.

Very recent works have explored personalization for slide generation\. However, existing works mostly focus on adapting content and narrative structure to different audiences or presentation contexts\[[16](https://arxiv.org/html/2607.00407#bib.bib16),[53](https://arxiv.org/html/2607.00407#bib.bib53),[50](https://arxiv.org/html/2607.00407#bib.bib50)\]\.\[[26](https://arxiv.org/html/2607.00407#bib.bib26)\]enables more tailored communication by conditioning on audience profiles or example pairs, and\[[50](https://arxiv.org/html/2607.00407#bib.bib50)\]allows users to specify a document–slide deck pair to express their preference, with the system aiming to replicate reference slides while plugging in content from the user’s own document\. However, these designs still lack the ability to create novel page\-level designs tailored to the user’s intent\. In this work, we proposeSpireto learn a user’s page\-level preferences from their visual data, pushing agentic slide generation to a more fine\-grained page\-level design task\.

## 5Conclusion

This paper studies Page\-level Slide Personalization \(PSP\), a key aspect of agentic slide generation that remains largely underexplored\. Broadly, we show that personalized visual generation is fundamentally a latent\-intent inference problem\. We formulate PSP as an inverse planning problem to infer a user’s latent design intent, which can be used to guide diverse executors to render the visuals\. We proposeSpire, which leverages structural denoising to construct a tractable surrogate objective for optimization\. By intentionally corrupting the visual structures of gold slides,Spirecreates a self\-supervised denoising task whereby two agents are trained to iteratively refine the design plan\. We provide theoretical analysis on how structural denoising serves as a consistent surrogate objective for PSP, and how our multi\-agent formulation strictly reduces policy gradient variance to stabilize the optimization\. Extensive experiments demonstrate the effectiveness of our method against representative baselines for the PSP task\.

## Acknowledgements

This work is supported in part by the US National Science Foundation under grant NSF IIS\-2141037\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author\(s\) and do not necessarily reflect the views of the National Science Foundation\.

## References

- \[1\]Anthropic: Claude sonnet 4: Hybrid reasoning model with superior intelligence for high\-volume use cases, and 200k context window\.[https://www\.anthropic\.com/claude/sonnet](https://www.anthropic.com/claude/sonnet)\(2025\)
- \[2\]Bai, S\., Chen, K\., Liu, X\., Wang, J\., Ge, W\., Song, S\., Dang, K\., Wang, P\., Wang, S\., Tang, J\., Zhong, H\., Zhu, Y\., Yang, M\., Li, Z\., Wan, J\., Wang, P\., Ding, W\., Fu, Z\., Xu, Y\., Ye, J\., Zhang, X\., Xie, T\., Cheng, Z\., Zhang, H\., Yang, Z\., Xu, H\., Lin, J\.: Qwen2\.5\-vl technical report \(2025\)
- \[3\]Cao, J\., Zhang, X\., Li, R\., Wei, J\., Li, C\., Joty, S\., Carenini, G\.: Multi2: Multi\-agent test\-time scalable framework for multi\-document processing\. In: Proceedings of The 5th New Frontiers in Summarization Workshop\. pp\. 135–156 \(2025\)
- \[4\]Chen, L\., Chen, J\., Goldstein, T\., Huang, H\., Zhou, T\.: Instructzero: Efficient instruction optimization for black\-box large language models\. arXiv preprint arXiv:2306\.03082 \(2023\)
- \[5\]Chen, Z\., Liu, G\., Zhang, B\.W\., Yang, Q\., Wu, L\.: Altclip: Altering the language encoder in clip for extended language capabilities\. In: Findings of the Association for Computational Linguistics: ACL 2023\. pp\. 8666–8682 \(2023\)
- \[6\]Comanici, G\., Bieber, E\., Schaekermann, M\., Pasupat, I\., Sachdeva, N\., Dhillon, I\., Blistein, M\., Ram, O\., Zhang, D\., Rosen, E\., et al\.: Gemini 2\.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\. arXiv preprint arXiv:2507\.06261 \(2025\)
- \[7\]Dong, Y\., Jiang, X\., Qian, J\., Wang, T\., Zhang, K\., Jin, Z\., Li, G\.: A survey on code generation with llm\-based agents\. arXiv preprint arXiv:2508\.00083 \(2025\)
- \[8\]Esser, P\., Kulal, S\., Blattmann, A\., Entezari, R\., Müller, J\., Saini, H\., Levi, Y\., Lorenz, D\., Sauer, A\., Boesel, F\., et al\.: Scaling rectified flow transformers for high\-resolution image synthesis\. In: Forty\-first international conference on machine learning \(2024\)
- \[9\]Fu, T\.J\., Wang, W\.Y\., McDuff, D\., Song, Y\.: Doc2ppt: Automatic presentation slides generation from scientific documents\. In: Proceedings of the AAAI Conference on Artificial Intelligence\. vol\. 36, pp\. 634–642 \(2022\)
- \[10\]Ge, J\., Wang, Z\.Z\., Zhou, X\., Peng, Y\.H\., Subramanian, S\., Tan, Q\., Sap, M\., Suhr, A\., Fried, D\., Neubig, G\., et al\.: Autopresent: Designing structured visuals from scratch\. In: Proceedings of the Computer Vision and Pattern Recognition Conference\. pp\. 2902–2911 \(2025\)
- \[11\]Hahn, M\., Zeng, W\., Kannen, N\., Galt, R\., Badola, K\., Kim, B\., Wang, Z\.: Proactive agents for multi\-turn text\-to\-image generation under uncertainty\. arXiv preprint arXiv:2412\.06771 \(2024\)
- \[12\]Hu, W\., Shu, Y\., Yu, Z\., Wu, Z\., Lin, X\., Dai, Z\., Ng, S\.K\., Low, B\.K\.H\.: Localized zeroth\-order prompt optimization\. Advances in Neural Information Processing Systems37, 86309–86345 \(2024\)
- \[13\]Hu, Y\., Wan, X\.: Ppsgen: Learning to generate presentation slides for academic papers\. In: IJCAI\. pp\. 2099–2105 \(2013\)
- \[14\]Jaiswal, S\., Prabhudesai, M\., Bhardwaj, N\., Qin, Z\., Zadeh, A\., Li, C\., Fragkiadaki, K\., Pathak, D\.: Iterative refinement improves compositional image generation\. arXiv preprint arXiv:2601\.15286 \(2026\)
- \[15\]Jang, D\., Heisler, M\.L\., Xing, L\., Li, Y\., Wang, E\., Xiong, Y\., Zhang, Y\., Fan, Z\.: Deckbench: Benchmarking multi\-agent frameworks for academic slide generation and editing\. arXiv preprint arXiv:2602\.13318 \(2026\)
- \[16\]Jung, K\., Cho, H\., Yun, J\., Yang, S\., Jang, J\., Choo, J\.: Talk to your slides: Language\-driven agents for efficient slide editing\. arXiv preprint arXiv:2505\.11604 \(2025\)
- \[17\]Lan, G\.: First\-order and stochastic optimization methods for machine learning, vol\. 1\. Springer \(2020\)
- \[18\]Li, C\., Xue, M\., Zhang, Z\., Yang, J\., Zhang, B\., Yu, B\., Hui, B\., Lin, J\., Wang, X\., Liu, D\.: Start: Self\-taught reasoner with tools\. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing\. pp\. 13523–13564 \(2025\)
- \[19\]Li, W\., Wang, X\., Li, W\., Jin, B\.: A survey of automatic prompt engineering: An optimization perspective\. arXiv preprint arXiv:2502\.11560 \(2025\)
- \[20\]Liang, X\., Zhang, X\., Xu, Y\., Sun, S\., You, C\.: Slidegen: Collaborative multimodal agents for scientific slide generation\. arXiv preprint arXiv:2512\.04529 \(2025\)
- \[21\]Liu, C\., Yang, Y\., Zhou, K\., Zhang, Z\., Fan, Y\., Xie, Y\., Qi, P\., Wang, X\.E\.: Presenting a paper is an art: Self\-improvement aesthetic agents for academic presentations\. arXiv preprint arXiv:2510\.05571 \(2025\)
- \[22\]Ma, J\., Liang, J\., Chen, C\., Lu, H\.: Subject\-diffusion: Open domain personalized text\-to\-image generation without test\-time fine\-tuning\. In: ACM SIGGRAPH 2024 Conference Papers\. pp\. 1–12 \(2024\)
- \[23\]Ma, S\., Guo, Y\., Su, J\., Huang, Q\., Zhou, Z\., Wang, Y\.: Talk2image: A multi\-agent system for multi\-turn image generation and editing\. arXiv preprint arXiv:2508\.06916 \(2025\)
- \[24\]Maheshwari, H\., Bandyopadhyay, S\., Garimella, A\., Natarajan, A\.: Presentations are not always linear\! gnn meets llm for document\-to\-presentation transformation with attribution\. arXiv preprint arXiv:2405\.13095 \(2024\)
- \[25\]Malladi, S\., Gao, T\., Nichani, E\., Damian, A\., Lee, J\.D\., Chen, D\., Arora, S\.: Fine\-tuning language models with just forward passes\. Advances in Neural Information Processing Systems36, 53038–53075 \(2023\)
- \[26\]Mondal, I\., Shwetha, S\., Natarajan, A\., Garimella, A\., Bandyopadhyay, S\., Boyd\-Graber, J\.: Presentations by the humans and for the humans: Harnessing llms for generating persona\-aware slides from documents\. In: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)\. pp\. 2664–2684 \(2024\)
- \[27\]OpenAI: Introducing chatgpt \(2022\),[https://openai\.com/blog/chatgpt](https://openai.com/blog/chatgpt)
- \[28\]OpenAI: Gpt\-image\-1\.[https://platform\.openai\.com/docs/models/gpt\-image\-1](https://platform.openai.com/docs/models/gpt-image-1)\(2025\)
- \[29\]OpenAI: Introducing gpt\-5\.[https://openai\.com/index/introducing\-gpt\-5/](https://openai.com/index/introducing-gpt-5/)\(2025\)
- \[30\]Pang, W\., Lin, K\.Q\., Jian, X\., He, X\., Torr, P\.: Paper2poster: Towards multimodal poster automation from scientific papers\. arXiv preprint arXiv:2505\.21497 \(2025\)
- \[31\]Park, S\., Jeong, J\., Kim, Y\., Lee, J\., Lee, N\.: Zip: An efficient zeroth\-order prompt tuning for black\-box vision\-language models\. arXiv preprint arXiv:2504\.06838 \(2025\)
- \[32\]Qi, Y\., Tian, J\., Liu, T\., Li, R\., Wei, T\., Liu, H\., Tang, X\., Cheng, M\., He, J\.: Learning to instruct: Fine\-tuning a task\-aware instruction optimizer for black\-box llms\. In: Findings of the Association for Computational Linguistics: EMNLP 2025\. pp\. 7707–7733 \(2025\)
- \[33\]Qu, Y\., Fang, S\., Wang, Y\., Wang, X\., Chen, Z\., Xie, H\., Zhang, Y\.: Igd: Instructional graphic design with multimodal layer generation\. In: Proceedings of the IEEE/CVF International Conference on Computer Vision\. pp\. 18218–18228 \(2025\)
- \[34\]Rojo, M\.G\., García, G\.B\., Mateos, C\.P\., García, J\.G\., Vicente, M\.C\.: Critical comparison of 31 commercially available digital slide systems in pathology\. International journal of surgical pathology14\(4\), 285–305 \(2006\)
- \[35\]Ruder, S\.: An overview of gradient descent optimization algorithms\. arXiv preprint arXiv:1609\.04747 \(2016\)
- \[36\]Shao, Z\., Wang, P\., Zhu, Q\., Xu, R\., Song, J\., Bi, X\., Zhang, H\., Zhang, M\., Li, Y\., Wu, Y\., et al\.: Deepseekmath: Pushing the limits of mathematical reasoning in open language models\. arXiv preprint arXiv:2402\.03300 \(2024\)
- \[37\]Sheng, G\., Zhang, C\., Ye, Z\., Wu, X\., Zhang, W\., Zhang, R\., Peng, Y\., Lin, H\., Wu, C\.: Hybridflow: A flexible and efficient rlhf framework\. arXiv preprint arXiv: 2409\.19256 \(2024\)
- \[38\]Sun, E\., Hou, Y\., Wang, D\., Zhang, Y\., Wang, N\.X\.: D2s: Document\-to\-slide generation via query\-based text summarization\. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies\. pp\. 1405–1418 \(2021\)
- \[39\]Tang, W\., Xiao, J\., Jiang, W\., Xiao, X\., Wang, Y\., Tang, X\., Li, Q\., Ma, Y\., Liu, J\., Tang, S\., et al\.: Slidecoder: Layout\-aware rag\-enhanced hierarchical slide generation from design\. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing\. pp\. 9026–9050 \(2025\)
- \[40\]Trelease, R\.B\.: From chalkboard, slides, and to e\-learning: How computing technologies have transformed anatomical sciences education\. Anatomical sciences education9\(6\), 583–602 \(2016\)
- \[41\]Venkatesh, K\., Dunlop, C\., Yanardag, P\.: Crea: A collaborative multi\-agent framework for creative image editing and generation\. arXiv preprint arXiv:2504\.05306 \(2025\)
- \[42\]Xie, E\., Waterfield, D\., Kennedy, M\., Zhang, A\.: Slidebot: A multi\-agent framework for generating informative, reliable, multi\-modal presentations\. arXiv preprint arXiv:2511\.09804 \(2025\)
- \[43\]Xu, R\., Liu, T\., Dong, Z\., You, T\., Hong, I\., Yang, C\., Zhang, L\., Zhao, T\., Wang, H\.: Alternating reinforcement learning for rubric\-based reward modeling in non\-verifiable llm post\-training\. arXiv preprint arXiv:2602\.01511 \(2026\)
- \[44\]Xu, X\., Xu, X\., Chen, S\., Chen, H\., Zhang, F\., Chen, Y\.C\.: Pregenie: An agentic framework for high\-quality visual presentation generation\. arXiv preprint arXiv:2505\.21660 \(2025\)
- \[45\]Xu, Y\., Ma, X\., Qiu, J\., Zhao, H\.: Textual\-to\-visual iterative self\-verification for slide generation\. arXiv preprint arXiv:2502\.15412 \(2025\)
- \[46\]Ye, H\., Zhang, J\., Liu, S\., Han, X\., Yang, W\.: Ip\-adapter: Text compatible image prompt adapter for text\-to\-image diffusion models\. arXiv preprint arXiv:2308\.06721 \(2023\)
- \[47\]Yu, Q\., Zhang, Z\., Zhu, R\., Yuan, Y\., Zuo, X\., Yue, Y\., Dai, W\., Fan, T\., Liu, G\., Liu, L\., et al\.: Dapo: An open\-source llm reinforcement learning system at scale\. arXiv preprint arXiv:2503\.14476 \(2025\)
- \[48\]Zelikman, E\., Harik, G\., Shao, Y\., Jayasiri, V\., Haber, N\., Goodman, N\.D\.: Quiet\-star: Language models can teach themselves to think before speaking\. arXiv preprint arXiv:2403\.09629 \(2024\)
- \[49\]Zelikman, E\., Wu, Y\., Mu, J\., Goodman, N\.: Star: Bootstrapping reasoning with reasoning\. Advances in Neural Information Processing Systems35, 15476–15488 \(2022\)
- \[50\]Zeng, W\., Ouyang, M\., Cui, L\., Ng, H\.T\.: Slidetailor: Personalized presentation slide generation for scientific papers\. arXiv preprint arXiv:2512\.20292 \(2025\)
- \[51\]Zhan, H\., Chen, C\., Ding, T\., Li, Z\., Sun, R\.: Unlocking black\-box prompt tuning efficiency via zeroth\-order optimization\. In: Findings of the Association for Computational Linguistics: EMNLP 2024\. pp\. 14825–14838 \(2024\)
- \[52\]Zhang, K\., Li, J\., Li, G\., Shi, X\., Jin, Z\.: Codeagent: Enhancing code generation with tool\-integrated agent systems for real\-world repo\-level coding challenges\. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)\. pp\. 13643–13658 \(2024\)
- \[53\]Zheng, H\., Guan, X\., Kong, H\., Zhang, W\., Zheng, J\., Zhou, W\., Lin, H\., Lu, Y\., Han, X\., Sun, L\.: Pptagent: Generating and evaluating presentations beyond text\-to\-slides\. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing\. pp\. 14413–14429 \(2025\)
- \[54\]Zhu, D\., Meng, R\., Song, Y\., Wei, X\., Li, S\., Pfister, T\., Yoon, J\.: Paperbanana: Automating academic illustration for ai scientists\. arXiv preprint arXiv:2601\.23265 \(2026\)

Supplementary Material of Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising Tianci Liu Zihan Dong Linjun Zhang Haoyu Wang Jing Gao Emre Kıcıman Ranveer Chandra Wei\-Ting Chen

\*\*footnotetext:Work done during an internship at Microsoft\.## Appendix 0\.ATheoretical Analysis Details

This section provides formal statements and proofs for the two informal theorems in the main text\. We only keep the theorem\-specific assumptions where each theorem is stated\.

### 0\.A\.1Preliminaries and Notation

Each training instance is\(x,s∗,𝒟ref\)∼𝒟tr\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\. The planner samples a latent planz∼πθ\(⋅∣x,𝒟ref\)z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\. Unless explicitly stated otherwise,∥⋅∥2\\\|\\cdot\\\|\_\{2\}denotesℓ2\\ell\_\{2\}norm on vectors\. For clarity, we use three objectives throughout this section:

- •𝒥​\(θ\)\\mathcal\{J\}\(\\theta\): the original PSP objective in Equation \([1](https://arxiv.org/html/2607.00407#S2.E1)\): 𝒥​\(θ\)=𝔼\(x,s∗,𝒟ref\)∼𝒟tr​\[log​∫πθ​\(z∣x,𝒟ref\)​p​\(s∗∣z\)​𝑑z\]\.\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\}\\left\[\\log\\int\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\\,p\(s^\{\*\}\\mid z\)\\,dz\\right\]\.
- •𝒥LB​\(θ\)\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\): the Jensen lower bound of𝒥​\(θ\)\\mathcal\{J\}\(\\theta\)\(i\.e\.,𝒥​\(θ\)≥𝒥LB​\(θ\)\\mathcal\{J\}\(\\theta\)\\geq\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\): 𝒥LB​\(θ\):=𝔼\(x,s∗,𝒟ref\)​𝔼z∼πθ\(⋅∣x,𝒟ref\)​\[log⁡p​\(s∗∣z\)\]\.\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\):=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\left\[\\log p\(s^\{\*\}\\mid z\)\\right\]\.\(7\)
- •𝒥SP​\(θ,ϕ\)\\mathcal\{J\}\_\{\\text\{SP\}\}\(\\theta,\\phi\): the structural\-perturbation surrogate optimized in training: 𝒥SP​\(θ,ϕ\):=𝔼\(x,s∗,𝒟ref\)​𝔼s~∼qρ\(⋅∣s∗\)​𝔼z∼πθ\(⋅∣x,𝒟ref\)​\[R^ϕ​\(z,s~\)\],\\mathcal\{J\}\_\{\\text\{SP\}\}\(\\theta,\\phi\):=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\mathbb\{E\}\_\{\\tilde\{s\}\\sim q\_\{\\rho\}\(\\cdot\\mid s^\{\*\}\)\}\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\left\[\\widehat\{R\}\_\{\\phi\}\(z,\\tilde\{s\}\)\\right\],\(8\)whereR^ϕ\\widehat\{R\}\_\{\\phi\}is the learned critic\-induced reward\.

For structural perturbation, samples~∼qρ\(⋅∣s∗\)\\tilde\{s\}\\sim q\_\{\\rho\}\(\\cdot\\mid s^\{\*\}\)\. Define

m​\(z,s~\):=log⁡p​\(s∗∣z\)−log⁡p​\(s~∣z\),R∗​\(z,s~\):=σ​\(β​m​\(z,s~\)\)\.m\(z,\\tilde\{s\}\):=\\log p\(s^\{\*\}\\mid z\)\-\\log p\(\\tilde\{s\}\\mid z\),\\qquad R^\{\*\}\(z,\\tilde\{s\}\):=\\sigma\(\\beta\\,m\(z,\\tilde\{s\}\)\)\.\(9\)Hereρ\\rhocontrols perturbation strength,β\>0\\beta\>0, andσ​\(⋅\)\\sigma\(\\cdot\)is sigmoid\.

In practice, we optimize its empirical estimator

𝒥^SP​\(θ,ϕ\),\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\),which is the empirical loss implemented by the critic and planner objectives in the main text\.

For variance analysis with critique conditioning, define

H​\(z,c;x\):=∇θlog⁡πθ​\(z∣x,𝒟ref,c\)\.H\(z,c;x\):=\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c\)\.\(10\)

### 0\.A\.2Details of Informal Theorem 1 \(Surrogate Consistency\)

###### Assumption 0\.A\.1\(A1 \(Regularity\)\)

\(a\)\|β​m​\(z,s~\)\|≤M\|\\beta\\,m\(z,\\tilde\{s\}\)\|\\leq Mfor sampled\(z,s~\)\(z,\\tilde\{s\}\)\. \(b\)∥∇θ𝔼\(x,s∗,𝒟ref\)𝔼s~∼qρ𝔼z∼πθ\[logp\(s~∣z\)\]∥2≤η\\left\\\|\\,\\nabla\_\{\\theta\}\\,\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\mathbb\{E\}\_\{\\tilde\{s\}\\sim q\_\{\\rho\}\}\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\}\[\\log p\(\\tilde\{s\}\\mid z\)\]\\,\\right\\\|\_\{2\}\\leq\\eta\.

###### Assumption 0\.A\.2\(A2 \(Critic\-calibration gradient bound\)\)

‖∇θ𝒥SP​\(θ,ϕ\)−∇θ𝒥R∗​\(θ\)‖2≤δcal,\\left\\\|\\,\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\(\\theta\)\\,\\right\\\|\_\{2\}\\leq\\delta\_\{\\mathrm\{cal\}\},where𝒥R∗​\(θ\):=𝔼\(x,s∗,𝒟ref\)​𝔼s~∼qρ​𝔼z∼πθ​\[R∗​\(z,s~\)\]\\mathcal\{J\}\_\{R^\{\*\}\}\(\\theta\):=\\mathbb\{E\}\_\{\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)\}\\mathbb\{E\}\_\{\\tilde\{s\}\\sim q\_\{\\rho\}\}\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\}\[R^\{\*\}\(z,\\tilde\{s\}\)\]\.

###### Assumption 0\.A\.3\(A3 \(Gradient gap bounds\)\)

Under the joint sampling path

\(x,s∗,𝒟ref\)\\displaystyle\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\)∼𝒟tr,s~∼qρ\(⋅∣s∗\),\\displaystyle\\sim\\mathcal\{D\}\_\{\\text\{tr\}\},\\qquad\\tilde\{s\}\\sim q\_\{\\rho\}\(\\cdot\\mid s^\{\*\}\),c\\displaystyle c∼Cϕ\(⋅∣s~,x,𝒟ref\),z∼πθ\(⋅∣x,𝒟ref,c\),\\displaystyle\\sim C\_\{\\phi\}\(\\cdot\\mid\\tilde\{s\},x,\\mathcal\{D\}\_\{\\text\{ref\}\}\),\\qquad z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c\),the empirical objective𝒥^SP\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}is the Monte Carlo estimator of𝒥SP\\mathcal\{J\}\_\{\\text\{SP\}\}, and the following gradient gaps hold:

‖∇θ𝒥LB​\(θ\)−∇θ𝒥​\(θ\)‖2≤ϵvi,‖∇θ𝒥^SP​\(θ,ϕ\)−∇θ𝒥SP​\(θ,ϕ\)‖2≤ϵemp\.\\left\\\|\\,\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\-\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\,\\right\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{vi\}\},\\qquad\\left\\\|\\,\\nabla\_\{\\theta\}\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\\,\\right\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{emp\}\}\.

###### Theorem 0\.A\.4\(Surrogate Approximation Bound\)

Under A1–A3,

‖∇θ𝒥^SP​\(θ,ϕ\)−β4​∇θ𝒥​\(θ\)‖2≤ϵemp\+δcal\+β4​η\+β​Lσ​\(M\)​M​Bm\+β4​ϵvi,\\left\\\|\\,\\nabla\_\{\\theta\}\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\frac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\,\\right\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{emp\}\}\+\\delta\_\{\\mathrm\{cal\}\}\+\\frac\{\\beta\}\{4\}\\eta\+\\beta L\_\{\\sigma\}\(M\)M\\,B\_\{m\}\+\\frac\{\\beta\}\{4\}\\epsilon\_\{\\mathrm\{vi\}\},\(11\)where

Bm:=𝔼​\[‖∇θm​\(z,s~\)‖2\],Lσ​\(M\):=sup\|u\|≤M\|σ′′​\(u\)\|\.B\_\{m\}:=\\mathbb\{E\}\\\!\\left\[\\left\\\|\\nabla\_\{\\theta\}m\(z,\\tilde\{s\}\)\\right\\\|\_\{2\}\\right\],\\qquad L\_\{\\sigma\}\(M\):=\\sup\_\{\|u\|\\leq M\}\|\\sigma^\{\\prime\\prime\}\(u\)\|\.Consequently, defining the aggregate error

ϵtot:=ϵemp⏟empirical opt\.\+δcal⏟critic\-calibration\+β4​η\+β​Lσ​\(M\)​M​Bm⏟structural\-surrogate\+β4​ϵvi⏟variational\-gap,\\epsilon\_\{\\mathrm\{tot\}\}:=\\underbrace\{\\epsilon\_\{\\mathrm\{emp\}\}\}\_\{\\text\{empirical opt\.\}\}\+\\underbrace\{\\delta\_\{\\mathrm\{cal\}\}\}\_\{\\text\{critic\-calibration\}\}\+\\underbrace\{\\tfrac\{\\beta\}\{4\}\\eta\+\\beta L\_\{\\sigma\}\(M\)MB\_\{m\}\}\_\{\\text\{structural\-surrogate\}\}\+\\underbrace\{\\tfrac\{\\beta\}\{4\}\\epsilon\_\{\\mathrm\{vi\}\}\}\_\{\\text\{variational\-gap\}\},we have‖∇θ𝒥^SP​\(θ,ϕ\)−β4​∇θ𝒥​\(θ\)‖2≤ϵtot\\left\\\|\\,\\nabla\_\{\\theta\}\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\frac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\,\\right\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{tot\}\}, which is the formal version of Informal Theorem 1 in the main text\.

###### Proof

We bound the total error in three steps via the intermediate objectives𝒥R∗\\mathcal\{J\}\_\{R^\{\*\}\}and𝒥LB\\mathcal\{J\}\_\{\\text\{LB\}\}\.

Step 1\(𝒥SP\\mathcal\{J\}\_\{\\text\{SP\}\}to𝒥LB\\mathcal\{J\}\_\{\\text\{LB\}\}\)\. By triangle inequality,

‖∇θ𝒥SP−β4​∇θ𝒥LB‖2≤‖∇θ𝒥SP−∇θ𝒥R∗‖2⏟≤δcal​\(A2\)\+‖∇θ𝒥R∗−β4​∇θ𝒥LB‖2\.\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\right\\\|\_\{2\}\\leq\\underbrace\{\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\-\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\\right\\\|\_\{2\}\}\_\{\\leq\\,\\delta\_\{\\mathrm\{cal\}\}\\ \\text\{\(A2\)\}\}\+\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\right\\\|\_\{2\}\.For the second term, sinceR∗=σ​\(β​m\)R^\{\*\}=\\sigma\(\\beta m\)andσ′​\(0\)=14\\sigma^\{\\prime\}\(0\)=\\tfrac\{1\}\{4\}, a further triangle inequality gives

‖∇θ𝒥R∗−β4​∇θ𝒥LB‖2≤‖∇θ𝒥R∗−β4​∇θ𝔼​\[m\]‖2\+β4​‖∇θ𝔼​\[m\]−∇θ𝒥LB‖2\.\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\right\\\|\_\{2\}\\leq\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathbb\{E\}\[m\]\\right\\\|\_\{2\}\+\\tfrac\{\\beta\}\{4\}\\left\\\|\\nabla\_\{\\theta\}\\mathbb\{E\}\[m\]\-\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\right\\\|\_\{2\}\.By mean\-value theorem with A1\(a\) \(\|β​m\|≤M\|\\beta m\|\\leq M\):

‖∇θ𝒥R∗−β4​∇θ𝔼​\[m\]‖2≤β​Lσ​\(M\)​M​Bm\.\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{R^\{\*\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathbb\{E\}\[m\]\\right\\\|\_\{2\}\\leq\\beta L\_\{\\sigma\}\(M\)M\\,B\_\{m\}\.\(12\)Sincem=log⁡p​\(s∗∣z\)−log⁡p​\(s~∣z\)m=\\log p\(s^\{\*\}\\mid z\)\-\\log p\(\\tilde\{s\}\\mid z\), we have𝔼​\[m\]=𝒥LB​\(θ\)−𝔼​\[log⁡p​\(s~∣z\)\]\\mathbb\{E\}\[m\]=\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\-\\mathbb\{E\}\[\\log p\(\\tilde\{s\}\\mid z\)\], so by A1\(b\):

‖∇θ𝔼​\[m\]−∇θ𝒥LB‖2≤η\.\\left\\\|\\nabla\_\{\\theta\}\\mathbb\{E\}\[m\]\-\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\right\\\|\_\{2\}\\leq\\eta\.\(13\)Combining:‖∇θ𝒥SP−β4​∇θ𝒥LB‖2≤δcal\+β​Lσ​\(M\)​M​Bm\+β4​η\.\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\\\|\_\{2\}\\leq\\delta\_\{\\mathrm\{cal\}\}\+\\beta L\_\{\\sigma\}\(M\)M\\,B\_\{m\}\+\\tfrac\{\\beta\}\{4\}\\eta\.

Step 2\(𝒥LB\\mathcal\{J\}\_\{\\text\{LB\}\}to𝒥\\mathcal\{J\}via variational identity\)\. For each\(x,s∗,𝒟ref\)\(x,s^\{\*\},\\mathcal\{D\}\_\{\\text\{ref\}\}\), lettingpθ∝πθ​p​\(s∗∣z\)p\_\{\\theta\}\\propto\\pi\_\{\\theta\}\\,p\(s^\{\*\}\\mid z\):

𝒥​\(θ\)=𝒥LB​\(θ\)\+𝔼​\[KL​\(πθ∥pθ\)\],\\mathcal\{J\}\(\\theta\)=\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\+\\mathbb\{E\}\\\!\\left\[\\mathrm\{KL\}\\\!\\left\(\\pi\_\{\\theta\}\\;\\\|\\;p\_\{\\theta\}\\right\)\\right\],so𝒥​\(θ\)≥𝒥LB​\(θ\)\\mathcal\{J\}\(\\theta\)\\geq\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\); by A3:‖∇θ𝒥LB​\(θ\)−∇θ𝒥​\(θ\)‖2≤ϵvi\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\-\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{vi\}\}\.

Step 3\(Empirical gap and conclusion\)\. By A3 and triangle inequality,

‖∇θ𝒥^SP​\(θ,ϕ\)−β4​∇θ𝒥​\(θ\)‖2≤ϵemp\+‖∇θ𝒥SP​\(θ,ϕ\)−β4​∇θ𝒥LB​\(θ\)‖2\+β4​ϵvi\.\\left\\\|\\nabla\_\{\\theta\}\\widehat\{\\mathcal\{J\}\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)\\right\\\|\_\{2\}\\leq\\epsilon\_\{\\mathrm\{emp\}\}\+\\left\\\|\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{SP\}\}\(\\theta,\\phi\)\-\\tfrac\{\\beta\}\{4\}\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{LB\}\}\(\\theta\)\\right\\\|\_\{2\}\+\\tfrac\{\\beta\}\{4\}\\epsilon\_\{\\mathrm\{vi\}\}\.Substituting Step 1 yields \([11](https://arxiv.org/html/2607.00407#Pt0.A1.E11)\)\.

### 0\.A\.3Details of Informal Theorem 2 \(Variance Reduction\)

We analyze planner gradient variance with critique\-conditioned training\. For the critique\-conditioned policy objective

𝒥pg​\(θ\):=𝔼z∼πθ\(⋅∣x,𝒟ref,c\),r∼p\(r∣z,c,x\)​\[r\],\\mathcal\{J\}\_\{\\mathrm\{pg\}\}\(\\theta\):=\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c\),\\,r\\sim p\(r\\mid z,c,x\)\}\[r\],the REINFORCE identity gives

∇θ𝒥pg​\(θ\)=𝔼​\[H​\(z,c;x\)​\(r−b​\(x,c\)\)\],\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\mathrm\{pg\}\}\(\\theta\)=\\mathbb\{E\}\\\!\\left\[H\(z,c;x\)\\,\(r\-b\(x,c\)\)\\right\],which motivates the end\-to\-end estimator below; replacingrrbyr¯​\(z,c,x\)=𝔼​\[r∣z,c,x\]\\bar\{r\}\(z,c,x\)=\\mathbb\{E\}\[r\\mid z,c,x\]gives the two\-agent estimator\. Let

z∼πθ\(⋅∣x,𝒟ref,c\),H\(z,c;x\):=∇θlogπθ\(z∣x,𝒟ref,c\),z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c\),\\quad H\(z,c;x\):=\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(z\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c\),and letr∼p​\(r∣z,c,x\)r\\sim p\(r\\mid z,c,x\)denote reward obtained through black\-box execution\.

Define end\-to\-end estimator

g^e2e:=H​\(z,c;x\)​\(r−b​\(x,c\)\)\.\\hat\{g\}\_\{\\text\{e2e\}\}:=H\(z,c;x\)\\,\(r\-b\(x,c\)\)\.\(14\)Define two\-agent conditional\-mean estimator

r¯​\(z,c,x\):=𝔼​\[r∣z,c,x\],g^2a:=H​\(z,c;x\)​\(r¯​\(z,c,x\)−b​\(x,c\)\)\.\\bar\{r\}\(z,c,x\):=\\mathbb\{E\}\[r\\mid z,c,x\],\\qquad\\hat\{g\}\_\{\\text\{2a\}\}:=H\(z,c;x\)\\,\(\\bar\{r\}\(z,c,x\)\-b\(x,c\)\)\.\(15\)
##### Assumptions for Theorem 2\.

###### Assumption 0\.A\.5\(B1 \(Finite second moments\)\)

𝔼​\[‖H​\(z,c;x\)‖22\]<∞\\mathbb\{E\}\[\\\|H\(z,c;x\)\\\|\_\{2\}^\{2\}\]<\\inftyand𝔼​\[r2∣z,c,x\]<∞\\mathbb\{E\}\[r^\{2\}\\mid z,c,x\]<\\infty\.

###### Assumption 0\.A\.6\(B2 \(Baseline independence\)\)

b​\(x,c\)b\(x,c\)is independent ofrrconditioned on\(x,c\)\(x,c\)\.

###### Theorem 0\.A\.7\(Variance Decomposition\)

Under B1–B2,

Var​\(g^e2e\)=Var​\(g^2a\)\+Δexec∣c,\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{e2e\}\}\)=\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{2a\}\}\)\+\\Delta\_\{\\text\{exec\}\\mid c\},\(16\)where

Δexec∣c:=𝔼​\[‖H​\(z,c;x\)‖22​Var​\(r∣z,c,x\)\]≥0\.\\Delta\_\{\\text\{exec\}\\mid c\}:=\\mathbb\{E\}\\\!\\left\[\\\|H\(z,c;x\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{Var\}\(r\\mid z,c,x\)\\right\]\\geq 0\.\(17\)HenceVar​\(g^2a\)≤Var​\(g^e2e\)\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{2a\}\}\)\\leq\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{e2e\}\}\), with strict inequality wheneverVar​\(r∣z,c,x\)\>0\\mathrm\{Var\}\(r\\mid z,c,x\)\>0on a set of non\-zero measure\. This is the formal version of Informal Theorem 2 in the main text\.

###### Proof

Apply law of total variance conditioning on\(z,c\)\(z,c\):

Var​\(Y\)=𝔼​\[Var​\(Y∣z,c\)\]\+Var​\(𝔼​\[Y∣z,c\]\)\.\\mathrm\{Var\}\(Y\)=\\mathbb\{E\}\[\\mathrm\{Var\}\(Y\\mid z,c\)\]\+\\mathrm\{Var\}\(\\mathbb\{E\}\[Y\\mid z,c\]\)\.TakeY=g^e2eY=\\hat\{g\}\_\{\\text\{e2e\}\}\. Conditioned on\(z,c\)\(z,c\), onlyrris random:

Var​\(g^e2e∣z,c,x\)=‖H​\(z,c;x\)‖22​Var​\(r∣z,c,x\)\.\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{e2e\}\}\\mid z,c,x\)=\\\|H\(z,c;x\)\\\|\_\{2\}^\{2\}\\,\\mathrm\{Var\}\(r\\mid z,c,x\)\.Taking expectation givesΔexec∣c\\Delta\_\{\\text\{exec\}\\mid c\}\. Also,

𝔼​\[g^e2e∣z,c,x\]=H​\(z,c;x\)​\(r¯​\(z,c,x\)−b​\(x,c\)\)=g^2a\.\\mathbb\{E\}\[\\hat\{g\}\_\{\\text\{e2e\}\}\\mid z,c,x\]=H\(z,c;x\)\\,\(\\bar\{r\}\(z,c,x\)\-b\(x,c\)\)=\\hat\{g\}\_\{\\text\{2a\}\}\.Thus the second term isVar​\(g^2a\)\\mathrm\{Var\}\(\\hat\{g\}\_\{\\text\{2a\}\}\), proving \([16](https://arxiv.org/html/2607.00407#Pt0.A1.E16)\)\.

##### Interpretation\.

End\-to\-end optimization carries executor\-induced randomness directly into policy gradient\. The two\-agent decomposition conditions on critique signal and replaces stochastic return with a lower\-noise conditional target, removing the additive variance term in \([17](https://arxiv.org/html/2607.00407#Pt0.A1.E17)\)\.

## Appendix 0\.BMore Technical Details

### 0\.B\.1Details about reference\-based PSP Data

We present more details about data curation\.

#### 0\.B\.1\.1Instruction\-Target\-Reference Triplets\.

We detail the construction of the data triplets\. Let’s use the training data𝒟tr\\mathcal\{D\}\_\{\\text\{tr\}\}as an example\. By definition, each data point in

𝒟tr=\{\(x\(1\),\(s∗\)\(1\),𝒟ref\(1\)\),…,\(x\(J\),\(s∗\)\(J\),𝒟ref\(J\)\)\}\\displaystyle\\mathcal\{D\}\_\{\\text\{tr\}\}=\\\{\(x^\{\(1\)\},\(s^\{\*\}\)^\{\(1\)\},\\mathcal\{D\}\_\{\\text\{ref\}\}^\{\(1\)\}\),\\ldots,\(x^\{\(J\)\},\(s^\{\*\}\)^\{\(J\)\},\\mathcal\{D\}\_\{\\text\{ref\}\}^\{\(J\)\}\)\\\}is a triplet consisting of three components: \(i\) a high\-level instructionxx, \(ii\) a gold slide pages∗s^\{\*\}to be generated, and \(iii\) a reference set𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}\. In practice, one may only have access to raw slide decks from users\. Nevertheless, these triplets can be readily constructed from such decks as follows:

- •Gold Slide Page\.Each individual slide pagesswithin the training split of deck is treated as a target pages∗s^\{\*\}to be generated\.
- •High\-level Instruction\.We employ a capable VLM \(e\.g\., GPT\-4o\) to provide a*concise*summary of the slide pages∗s^\{\*\}, e\.g\., naming the text boxes and visual elements\. Note that this summary captures only the factual content without describing visual styles such as coloring, font size, or layout organization, thereby serving the role of the high\-level instructionxx\. Notably, these omitted visual styles represent the latent design intent that must be inferred from the reference set for successful PSP\.
- •Reference Set\.We randomly pair eachs∗s^\{\*\}with a few other pages from the same deck to serve as the𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}, leveraging the inherent visual coherence within a single deck\. In our experiments, we set the size of the reference set to 1\. We use a single reference slide to create a challenging personalization setting while keeping the retrieval protocol consistent across methods\.

This procedure transforms raw slide decks into the structured training corpus𝒟tr\\mathcal\{D\}\_\{\\text\{tr\}\}\.𝒟te\\mathcal\{D\}\_\{\\text\{te\}\}is constructed similarly\.

#### 0\.B\.1\.2Train, Test, and OOD Data\.

As presented in Sec\.[3](https://arxiv.org/html/2607.00407#S3), we first perform a slide\-level held\-out split that preserves the deck’s chronological order\. The last 20% of slides are reserved for testing, and the remaining slides are used for training \(or for retrieval in training\-free baselines\)\.

Next, we construct target–reference pairs within each split to ensure that test slides are never leaked during training\. To determine the feasibility of each pair, we follow\[[39](https://arxiv.org/html/2607.00407#bib.bib39)\]to compute the visual complexity scores for both target and reference slides, filtering out pairs where the target–reference complexity score gap is excessive\. When constructing these pairs, we consider any combination of two slides within a split to be viable, regardless of their original chronological order\. Consequently, for a filtered split containingKKpages, the resulting dataset size isK​\(K−1\)K\(K\-1\)\.

To better reflect realistic PSP requirements, we construct OOD data using a more challenging protocol\. These OOD decks are never seen during training; starting from the second page, each slide is paired with its*immediately preceding*page\. This setup mimics the sequential nature of real\-world slide creation\. No complexity filtering is applied to the OOD set\. Therefore, for a raw OOD deck ofKKpages, the final evaluation set is of sizeK−1K\-1\.

### 0\.B\.2Training Details of the Critic

We provide additional training details for the critic agent\.

#### 0\.B\.2\.1Critic Prompt\.

We prompt the critic to act as a presentation design expert responsible for evaluating a generated slide against the provided user instructionxxand reference slide\(s\)𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}\. The critique generation is structured into three stages\. We provide the prompt template in the end of the supplementary material\.

- •Analysis\.The critic first writes a brief high\-level analysis summarizing the major design inconsistencies\.
- •Assessment\.It then assesses the generated slide along eight predefined structural aspects subject to random perturbations:*graphic color, graphic position, graphic size, image position, image size, text color, text position*, and*text size*\. See Supp\.[0\.D](https://arxiv.org/html/2607.00407#Pt0.A4)for further details on the perturbation mechanism\.
- •Feedback\.Finally, the critic outputs a single structured block containing aspect\-wise feedback for all eight aspects in a fixed order\. Each entry includes the current issue and a target correction\. To ensure the feedback is directly usable for downstream refinement, the prompt requires concise and actionable suggestions, with explicit numeric targets in normalized coordinates whenever position or size adjustments are necessary\.

#### 0\.B\.2\.2Format Reward\.

We employ a rule\-based format rewardRformat,c​\(c\)R\_\{\\text\{format\},c\}\(c\)to encourage the critic to follow the required multi\-stage output schema\. The verification checks whether each stage is properly produced, including the*overall*structure as well as the specific requirements for the*analysis*,*assessment*, and final aspect\-wise*feedback*sections\. The verification checks: \(i\) compliance with the required format by identifying keyword tags and ensuring the content is non\-empty; and \(ii\) whether the names of aspect\-level entries are valid and non\-duplicated\.

Mathematically, the format reward is a weighted sum of these components:

Rformat, c​\(c\)\\displaystyle R\_\{\\text\{format, c\}\}\(c\)=0\.05​Rstructure​\(c\)\+0\.05​Ranalysis​\(c\)\\displaystyle=0\.05\\,R\_\{\\text\{structure\}\}\(c\)\+0\.05\\,R\_\{\\text\{analysis\}\}\(c\)\+0\.05​Rassessment​\(c\)\+0\.85​Rfeedback​\(c\)\.\\displaystyle\\quad\+0\.05\\,R\_\{\\text\{assessment\}\}\(c\)\+0\.85\\,R\_\{\\text\{feedback\}\}\(c\)\.We setRfeedback​\(c\)R\_\{\\text\{feedback\}\}\(c\)as the dominant term, reflecting that the structured aspect\-wise feedback is the most critical part of the critic’s output\.

#### 0\.B\.2\.3Accuracy Reward\.

The accuracy reward, as introduced in Sec\. 2\.2, verifies whether the extracted critique is consistent with the ground\-truth discrepancy list𝒜diff\\mathcal\{A\}\_\{\\text\{diff\}\}\. Specifically, for each discrepancy item in𝒜diff\\mathcal\{A\}\_\{\\text\{diff\}\}, we check whether the critic correctly diagnoses both the issue and the corresponding correction\. This verification is implemented via a rule\-based keyword matching procedure between the*extracted*critique and the perturbation annotations\. We compute the final reward by aggregating the matched issue–correction keywords over all discrepancy items:

Rvfy​\(c,𝒜diff\)=\# matched issue–correction keywords in​\(c,𝒜diff\)\# all ground\-truth issue–correction keywords in​𝒜diff\.R\_\{\\text\{vfy\}\}\(c,\\mathcal\{A\}\_\{\\text\{diff\}\}\)=\\frac\{\\text\{\\\# matched issue\-\-correction keywords in \}\(c,\\mathcal\{A\}\_\{\\text\{diff\}\}\)\}\{\\text\{\\\# all ground\-truth issue\-\-correction keywords in \}\\mathcal\{A\}\_\{\\text\{diff\}\}\}\.\(18\)Essentially, the accuracy reward measures the precision with which the critique covers the issue–correction pairs specified by𝒜diff\\mathcal\{A\}\_\{\\text\{diff\}\}\.

This reward, combined with the format reward, constitutes the final reward for critic training defined in Eq\. \([5](https://arxiv.org/html/2607.00407#S2.E5)\)\.

### 0\.B\.3Training Details of the Planner

We provide additional training details for the planner agent\.

#### 0\.B\.3\.1Planner Prompt\.

We prompt the planner to act as an expert slide design planner responsible for producing a detailed and executable design plan conditioned on the provided user instructionxxand reference slide\(s\)𝒟ref\\mathcal\{D\}\_\{\\text\{ref\}\}\. As detailed in Sec\. 2\.2, the planner aims to perform two complementary subtasks: \(i\) to generate a high\-quality initial plan based on the user instruction and reference slides; and \(ii\) to revise a suboptimal plan if a critique is provided\. We use separate prompt templates for these two scenarios, both of which are structured into three stages\. We provide the prompt template at the end of the supplementary material\.

##### Initial Plan Generation\.

- •Analysis\.The planner first infers transferable design principles from the reference slide\(s\), including the core visual logic, reusable stylistic elements, and context\-specific factors that should be adapted to the current request\.
- •Strategy\.It then formulates an adaptation strategy that maps these principles into concrete planning decisions, including the overall layout structure, the roles of major visual elements, and the aspect\-level choices governing color, position, size, and typographic hierarchy\.
- •Plan\.Finally, the planner outputs a complete design plan that specifies the slide composition in a sequential and directly executable manner, covering the background, major containers, textual components, visual elements, and overall spacing/balance\. To ensure executability, the prompt requires concrete normalized spatial specifications and prohibits ambiguous instructions\.

##### Critique\-Guided Plan Revision\.

- •Analysis\.The planner first infers transferable design principles from the reference slide\(s\), mimicking a chain\-of\-thought process\.
- •Strategy\.It then revises the suboptimal plan according to the provided critiquecc\. In particular, the planner first normalizes the previous plan into an explicit structured representation, and then applies the required edits while preserving aspects that are already satisfactory\.
- •Plan\.Finally, the planner outputs a new complete design plan that retains the overall structure of the suboptimal plan whenever possible, while incorporating the requested corrections in a directly executable form\.

Both initial and revised plans are evaluated from*format*and*accuracy*perspectives in a unified way, as detailed below\.

#### 0\.B\.3\.2Format Reward\.

We employ a rule\-based format rewardRformat,p​\(z\)R\_\{\\text\{format\},p\}\(z\)to encourage the planner to follow the required multi\-stage output schema\. The verification checks whether each stage is properly produced, including the*overall*structure as well as the specific requirements for the*analysis*,*strategy*, and final*plan*sections\. The verification checks: \(i\) compliance with the required format by identifying keyword tags and ensuring the content is non\-empty; and \(ii\) whether the final plan preserves the required asset content, ensuring no critical user\-provided material is omitted\.

Mathematically, the format reward is a weighted sum of these components:

Rformat,p​\(z\)\\displaystyle R\_\{\\text\{format\},p\}\(z\)=0\.05​Rstructure​\(z\)\+0\.05​Ranalysis​\(z\)\\displaystyle=0\.05\\,R\_\{\\text\{structure\}\}\(z\)\+0\.05\\,R\_\{\\text\{analysis\}\}\(z\)\+0\.10​Rstrategy​\(z\)\+0\.80​Rplan​\(z\)\.\\displaystyle\\quad\+0\.10\\,R\_\{\\text\{strategy\}\}\(z\)\+0\.80\\,R\_\{\\text\{plan\}\}\(z\)\.We setRplan​\(z\)R\_\{\\text\{plan\}\}\(z\)as the dominant term, reflecting that the final executable design plan is the most critical part of the planner’s output\.

#### 0\.B\.3\.3Accuracy Reward\.

The accuracy reward, as introduced in Sec\. 2\.2, verifies whether the planner’s output is consistent with the target slide under the plan–visual matching \(PVM\) objective\. Specifically, we extract the natural\-language design plan from the final*plan*stage and evaluate whether it matches the gold slides∗s^\{\*\}more faithfully than a randomly perturbed slides~\\tilde\{s\}\. To mitigate positional bias, the comparison is conducted under both visual presentation orders, and the final reward is computed by averaging the two binary outcomes:

Rpvm​\(z,s∗,s~\)=12​\(𝕀​\[o^→=1\]\+𝕀​\[o^←=1\]\)\.R\_\{\\text\{pvm\}\}\(z,s^\{\*\},\\tilde\{s\}\)=\\frac\{1\}\{2\}\\left\(\\mathbb\{I\}\[\\hat\{o\}\_\{\\rightarrow\}=1\]\+\\mathbb\{I\}\[\\hat\{o\}\_\{\\leftarrow\}=1\]\\right\)\.\(19\)Essentially, the accuracy reward measures whether the final plan describes the gold slide more faithfully than the perturbed alternative, based on the judgment of a capable VLM judge \(e\.g\., o4\-mini\)\.

This reward, combined with the format reward, constitutes the final reward for planner training defined as

𝒥plnr​\(θ\)\\displaystyle\\mathcal\{J\}\_\{\\text\{plnr\}\}\(\\theta\)=𝔼s∗∼𝒟trs~∼q\(⋅∣s∗\)​\[𝔼c∗∼Oracle​\(c\)​\[𝔼z∼πθ\(⋅∣x,𝒟ref,c∗\)​\[Rpvm​\(z,s∗,s~\)×Rformat,p​\(z\)\]\]\]\.\\displaystyle=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s^\{\*\}\\sim\\mathcal\{D\}\_\{\\text\{tr\}\}\\\\ \\tilde\{s\}\\sim q\(\\cdot\\mid s^\{\*\}\)\\end\{subarray\}\}\\Big\[\\mathbb\{E\}\_\{c^\{\*\}\\sim\\text\{Oracle\}\(c\)\}\\big\[\\mathbb\{E\}\_\{z\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,\\mathcal\{D\}\_\{\\text\{ref\}\},c^\{\*\}\)\}\[R\_\{\\text\{pvm\}\}\(z,s^\{\*\},\\tilde\{s\}\)\\times R\_\{\\text\{format,p\}\}\(z\)\]\\big\]\\Big\]\.

#### 0\.B\.3\.4Synthesis of Critique\-Guided Plan Revision Data\.

As detailed in Sec\.[2\.2](https://arxiv.org/html/2607.00407#S2.SS2), critique\-guided revision training relies on high\-quality revision triplets\(z~,s~,c∗\)\(\\tilde\{z\},\\tilde\{s\},c^\{\*\}\)\. Here,s~\\tilde\{s\}is a perturbed slide,z~\\tilde\{z\}is a suboptimal plan aligned withs~\\tilde\{s\}, andc∗c^\{\*\}is an oracle critique that specifies how the suboptimal plan should be revised toward the gold slides∗s^\{\*\}\. The construction process is summarized as follows:

- •Perturbed Slide\.The perturbed slides~\\tilde\{s\}is obtained by applying recorded structural perturbations to the gold slides∗s^\{\*\}, as described in Sec\. 2\.2\. Since these perturbations are known and explicitly recorded,s~\\tilde\{s\}serves as a controllable visual target corresponding to the suboptimal state\.
- •Oracle Critique\.To construct the oracle critiquec∗c^\{\*\}, we provide the perturbation action list to a capable LLM and ask it to translate the known discrepancies into the required critique format\. This converts the recorded structural corruptions into natural\-language feedback that is directly consumable by the planner\. We further apply rejection filtering using the critique reward introduced in Supp[0\.B\.2](https://arxiv.org/html/2607.00407#Pt0.A2.SS2), retaining only those pass this verification step\.
- •Suboptimal Plan\.To construct the suboptimal planz~\\tilde\{z\}, we provide the perturbed slides~\\tilde\{s\}to a capable LLM and ask it to infer a complete design plan that accurately*reflects the current degraded slide*rather than the gold slide\. We then apply rejection filtering based on the planner reward and retain only those candidate plans where the plan–visual matching reward preferss~\\tilde\{s\}overs∗s^\{\*\}, thereby enforcing consistency between the inferred plan and the perturbed visual input\.

Overall, this pipeline ensures coherent revision supervision:s~\\tilde\{s\}provides the observable suboptimal slide,z~\\tilde\{z\}captures the corresponding suboptimal design intent, andc∗c^\{\*\}specifies the corrections needed to recover the intended gold design\.

### 0\.B\.4Implementation and Evaluation Details of Experiments

#### 0\.B\.4\.1Implementation Details\.

We trainSpirewithverl\[[37](https://arxiv.org/html/2607.00407#bib.bib37)\]; the detailed training hyper\-parameters are provided in Tab[4](https://arxiv.org/html/2607.00407#Pt0.A2.T4)\.

Table 4:Hyper\-parameters used for critic and planner training\.
#### 0\.B\.4\.2Evaluation Details \(Visual Similarity\)\.

For the visual similarity evaluation, we treat each slide as a rendered full\-page image\. We report two full\-reference image similarity metrics, following\[[39](https://arxiv.org/html/2607.00407#bib.bib39)\]\. The scores are normalized to\[0,1\]\[0,1\]\.

- •Structural Similarity Index Measure \(SSIM\)\.SSIM measures the low\-level structural fidelity between the generated slide and the gold slide\.
- •CLIP Image Similarity\.CLIP compares the global visual representations of two images in a pre\-trained vision\-language embedding space\. We use AltCLIP\[[5](https://arxiv.org/html/2607.00407#bib.bib5)\]\.

#### 0\.B\.4\.3Evaluation Details \(VLM\-as\-a\-Judge\)\.

We also report a reference\-free VLM\-as\-a\-judge evaluation protocol\. In this setting, the VLM judge observes the user instructionxxand the generated slide to assess the generation quality across the following four aspects\. Each aspect is scored on a 0–10 scale and then normalized to\[0,1\]\[0,1\]for aggregation\.

- •Faithfulness\.Whether the textual content requested in the user instruction is fully preserved and readable, without missing text, content truncation, or severe element overlap\.
- •Color\.The harmony of the color palette, and the contrast between foreground and background elements to ensure readability\.
- •Layout\.The overall composition of the slide, including alignment, spacing, margins, and the logical organization of major visual elements\.
- •Overall Aesthetics\.The overall professionalism and visual appeal of the slide as a presentation\-grade page\.

We provide the full VLM\-as\-a\-judge prompt at the end of the supplementary material\.

## Appendix 0\.CAdditional Experiment Results

We present more experiment results that are not included in the main body due to page limits\.

#### 0\.C\.0\.1Reference\-based Pairwise VLM\-Judge Evaluation\.

To complement the reference\-free evaluation, we conduct a reference\-based pairwise comparison where the VLM judge is provided with the user instructionxx, the gold slides∗s^\{\*\}, and two generated slides from competing methods\. This protocol directly measures relative generation quality against the intended target, making it more sensitive to subtle differences in design execution\.

We evaluate each pair along four aspects: \(i\)*faithfulness*\(text preservation and readability\); \(ii\)*color*; \(iii\)*layout*; and \(iv\)*overall aesthetics*\. The definitions of these aspects are the same as in the reference\-free setting presented above\. To mitigate positional bias, we evaluate each pair twice by swapping the presentation order and aggregate the results\.

Tab[5](https://arxiv.org/html/2607.00407#Pt0.A3.T5)reports the*win/tie rate*ofSpire, representing the fraction of cases whereSpireis judged*at least as good as the baseline*\. As shown in Tab[5](https://arxiv.org/html/2607.00407#Pt0.A3.T5),Spireconsistently outperforms all baselines on both test and OOD pages\. The advantages are particularly pronounced in Layout and Faithfulness, whereSpireachieves*win/tie rates*exceeding 90% against most baselines\. Even against the strong GPT\-based PSP \(o4\-mini\) baseline,Spiremaintains a significant lead \(avg\. 0\.847/0\.851\), confirming that our RL\-based training produces slides that are consistently preferred in head\-to\-head comparisons\.

Table 5:Reference\-based Pairwise VLM Evaluation \(Win/Tie Rate ofSpire\)\.
### 0\.C\.1More Qualitative Results

Fig[5](https://arxiv.org/html/2607.00407#Pt0.A3.F5)presents additional visual comparisons betweenSpireand baseline methods across both in\-distribution test slides and OOD scenarios\. Consistent with the observations in Sec\.[3\.3](https://arxiv.org/html/2607.00407#S3.SS3),Spiredemonstrates superior faithfulness to the user instruction while maintaining high design consistency with the reference slides\.

RefAutoPre\.PPTAgentSD 3\.5OursGoldTest![Refer to caption](https://arxiv.org/html/2607.00407v1/x45.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x46.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x48.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x49.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x50.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x51.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x53.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x56.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x58.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x61.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x63.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x64.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)OOD![Refer to caption](https://arxiv.org/html/2607.00407v1/x69.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x71.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x73.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x74.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/x75.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/x82.png)![Refer to caption](https://arxiv.org/html/2607.00407v1/)![Refer to caption](https://arxiv.org/html/2607.00407v1/)Figure 5:Additional qualitative comparison of slide generation results\.

## Appendix 0\.DMore Details about Perturbation

We construct perturbations using a structured taxonomy defined along two axes:shape roleandvisual attribute\. Each slide element is first assigned one of three shape roles, namelytext,graphic, orimage, according to its functional role in the slide\. Conditioned on this role assignment, we define a fixed set of perturbable attributes, resulting ineight perturbation categories\.

### 0\.D\.1Perturbation Categories

##### Text\-related categories\.

For text elements, we define three perturbation categories:

- •Text Size\(text\_size\): perturbs the font size of a text portion\.
- •Text Position\(text\_position\): perturbs text layout by replacing the paragraph alignment with a different valid alignment\.
- •Text Color\(text\_color\): perturbs the font color of a text element\.

##### Graphic\-related categories\.

For non\-image graphic shapes, we define three perturbation categories:

- •Graphic Size\(graphic\_size\): perturbs one size attribute of a graphic element, namely width or height\.
- •Graphic Position\(graphic\_position\): perturbs one spatial coordinate of a graphic element, namely its horizontal coordinatexxor vertical coordinateyy, wherexxandyydenote the horizontal and vertical positions of the element on the slide canvas, respectively\.
- •Graphic Color\(graphic\_color\): perturbs the fill color of a graphic element\.

##### Image\-related categories\.

For image elements, we define two perturbation categories:

- •Image Size\(image\_size\): perturbs the width or height of an image element\.
- •Image Position\(image\_position\): perturbs the horizontal coordinatexxor vertical coordinateyyof an image element\.

### 0\.D\.2Perturbation Magnitude

For numerical attributes, we control perturbation strength with a discrete severity level\. Leta∈ℝa\\in\\mathbb\{R\}denote the original value of a numerical attribute and leta′a^\{\\prime\}denote the perturbed value\. We sample a relative perturbation ratioδ\\deltaas

δ=s⋅u,u∼𝒰​\(δmin,δmax\),s∼Unif​\{−1,\+1\},\\delta=s\\cdot u,\\qquad u\\sim\\mathcal\{U\}\(\\delta\_\{\\min\},\\delta\_\{\\max\}\),\\qquad s\\sim\\mathrm\{Unif\}\\\{\-1,\+1\\\},\(20\)where𝒰​\(δmin,δmax\)\\mathcal\{U\}\(\\delta\_\{\\min\},\\delta\_\{\\max\}\)denotes the uniform distribution over the interval\[δmin,δmax\]\[\\delta\_\{\\min\},\\delta\_\{\\max\}\],δmin\\delta\_\{\\min\}andδmax\\delta\_\{\\max\}are the minimum and maximum perturbation magnitudes associated with the selected category and severity level,uuis the sampled perturbation magnitude, andssdetermines the perturbation direction\. The perturbed value is then computed as

a′=a​\(1\+δ100\),a^\{\\prime\}=a\\left\(1\+\\frac\{\\delta\}\{100\}\\right\),\(21\)whereδ\\deltais expressed in percentage\. When required,a′a^\{\\prime\}is clipped to the valid range of the corresponding attribute, e\.g\.,\[0,255\]\[0,255\]for RGB channels\. For categorical attributes, the perturbed value is sampled uniformly from the set of valid alternatives excluding the original value\.

### 0\.D\.3Sampling Strategy

Given a slide, we first enumerate all perturbable items and group them by the eight categories above\. Each category is then activated independently according to a Bernoulli random variable with parameterpp, whereppdenotes the probability that a category is selected for perturbation\. In our experiments, we setp=0\.5p=0\.5\. Categories with no valid candidates are skipped automatically\.

For each activated category, the pipeline samples one valid candidate from the corresponding pool and generates one perturbation instruction\. To avoid conflicting edits, each physical shape is constrained to receive at most one primary perturbation\. After primary perturbations are generated, the pipeline may further sample one additional extra perturbation from the remaining candidate pool, typically using a lower severity level\. This design increases perturbation diversity while preserving semantic and geometric consistency at the slide level\.

The resulting perturbation process is role\-aware, attribute\-specific, and severity\-controlled\. It produces diverse but structured modifications that approximate realistic slide variations, while preventing incompatible edits to the same element\. Figure[6](https://arxiv.org/html/2607.00407#Pt0.A4.F6)provides qualitative examples of our perturbation pipeline\. The first column contains the original slides and the second column contains their perturbed counterparts, with each column corresponding to one slide\-level example\.

Figure 6:Qualitative examples of the proposed perturbation pipeline\. Each row shows one original\-perturbed slide pair\.`System Prompt for Critique Generation \(Spire\) User Prompt for Critique Generation \(Spire\) System Prompt for Plan Generation \(Spire\) User Prompt for Plan Generation \(Spire\) System Prompt for Plan Revision \(Spire\) User Prompt for Plan Revision \(Spire\) Prompt Template for VLM Judge \(Spire\) Evaluation Criteria for VLM Judge \(Spire\)`

Similar Articles

DeepSlide: From Artifacts to Presentation Delivery

arXiv cs.AI

DeepSlide is a human-in-the-loop multi-agent system for the full presentation process, from requirement elicitation and time-budgeted narrative planning to evidence-grounded slide-script generation and rehearsal support. It introduces a dual-scoreboard benchmark separating static artifact quality from dynamic delivery excellence, and achieves gains in narrative flow, pacing precision, and slide-script synergy.

Steered Generation via Gradient-Based Optimization on Sparse Query Features

arXiv cs.LG

This paper introduces Prototype-Based Sparse Steering, a method that applies sparse autoencoders to attention query activations in LLMs, then uses gradient-based optimization during inference to steer generation toward target behaviors. The approach is validated in both a logical planning task and a stylistic educational domain, demonstrating interpretable and disentangled control.

Narrative-Driven Paper-to-Slide Generation via ArcDeck

Hugging Face Daily Papers

ArcDeck is a multi-agent framework that generates presentation slides from academic papers by modeling logical flow through discourse trees and iterative agent refinement, outperforming direct summarization methods. The paper introduces ArcBench, a new benchmark for evaluating paper-to-slide generation with emphasis on narrative coherence and logical structure.