SlideLab: 以受众为中心的科学幻灯片生成与评估

arXiv cs.CL 论文

摘要

SlideLab引入了一种免训练的多智能体框架,可从研究论文生成以受众为中心的科学演示文稿,该框架在人类偏好研究中超越了现有系统,并提出了用于逐张幻灯片评估的ConfArena系统。

arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:35

# SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
Source: [https://arxiv.org/html/2609.30294](https://arxiv.org/html/2609.30294)
Karun SharmaYuxia WangAffiliation:INSAIT, Sofia University “St\. Kliment Ohridski”

###### Abstract

Scientific presentations are more than summaries of research papers\. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation\. We presentSlideLab, a training\-free multi\-agent framework for generating scientific presentations from research papers\.SlideLabfirst plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification\. In a blind human preference study,SlideLabwas preferred over both open\-source and commercial systems on 77% of papers while using roughly4×4\\timesfewer inference tokens than the strongest open\-source baseline\. We also introduce ConfArena, an audience\-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide\.ConfArenamatches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order\.

## 1Introduction

A conference talk gives an author 10\-15 minutes to communicate what a paper develops over eight pages\. Unlike a reader, the audience encounters the work for the first time and must follow the presentation as it unfolds\. A good slide deck therefore does more than summarize a paper\. It selects the right content, presents it in a coherent order, and uses visual elements to help the audience understand the work as the talk progresses\.

Recent advances in large language models and agentic frameworks have led to rapid progress in automated slide generation\. Existing approaches range from summarization\-based methods\([Sun et al\., 2021](https://arxiv.org/html/2609.30294#bib.bib15);[Kumar and Chowdary, 2024](https://arxiv.org/html/2609.30294#bib.bib8);[Fu et al\., 2022](https://arxiv.org/html/2609.30294#bib.bib2)\)to template\-guided editing systems\([Zheng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib22);[Zeng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib19);[Liang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib9);[Tang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib16)\)and end\-to\-end generation frameworks that use agentic and multi\-agent architectures\([Ge et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib3);[Yang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib17);[Zheng et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib23)\)to generate complete presentations from research papers\. Despite this progress, generating scientific presentations remains difficult because the system must jointly determine what scientific content should appear on each slide and how it should be presented visually\. Recent benchmarks\([Chen et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib1);[Yang et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib18)\)show that even state\-of\-the\-art systems still struggle with factual grounding to the source paper, narrative organization, and presentation of visual elements\.

The difficulty extends beyond generation to evaluation\. Most existing benchmarks\([Zheng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib22);[Chen et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib1);[Yang et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib18);[Jang et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib5);[Ozden et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib13)\)ask an LLM to judge the final slide deck against a fixed set of rubrics, such as visual quality, layout consistency, and content fidelity\. While these metrics are useful, they evaluate a completed slide deck rather than the presentation as experienced by an audience\. In practice, understanding develops incrementally as a talk progresses, and communication failures naturally surface through audience questions\. Existing evaluation frameworks largely overlook this sequential aspect of scientific presentations\.

Motivated by these observations, we address both the generation and evaluation of scientific presentations\. We introduceSlideLab, an agentic framework that plans a coherent presentation narrative before progressively constructing and refining the slide deck through specialized agents\. The framework emphasizes faithfulness to the source paper while improving narrative organization and visual aesthetic\. We further introduce a conference\-room evaluation environment that evaluates presentations as they are experienced by the audience\. A simulated audience follows the presentation slide by slide, asks questions when communication breaks down, and provides a richer assessment than static deck\-level evaluation\.

Extensive experiments, including human studies with participants spanning undergraduate students to faculty members, show thatSlideLaboutperforms existing open\-source systems and closely competes with commercial systems while requiring substantially fewer inference tokens\. The proposed evaluation environment also exhibits strong agreement with human judgments, suggesting that it captures aspects of presentation quality overlooked by existing automated metrics\. Our contributions are summarized as follows:

- •We presentSlideLab, a training\-free framework for generating scientific presentations from research papers\. The framework progressively constructs and refines the slide deck, producing slide decks with stronger narrative flow, improved visual design, and better factual grounding\. In a blind human preference study against three representative open\- and closed\-source systems,SlideLabwas preferred for 77% of the evaluated papers\.
- •We introduceConfArena, a conference\-room evaluation environment that assesses scientific presentations through audience interaction rather than static deck evaluation\.
- •We will publicly release our codebase forSlideLab, ConfArena, web demo, and annotation platform after review\.

## 2Related Work

#### Early Extractive Methods

Early work formulated slide generation as a document summarization problem\([Sun et al\., 2021](https://arxiv.org/html/2609.30294#bib.bib15);[Kumar and Chowdary, 2024](https://arxiv.org/html/2609.30294#bib.bib8);[Fu et al\., 2022](https://arxiv.org/html/2609.30294#bib.bib2)\), focusing on selecting and compressing salient content from source documents through extraction, ranking, and summarization techniques\. These methods established the foundations of automatic slide generation but operate primarily on text, lack multimodal generation capabilities, and predate modern LLM\-based systems\.

#### Editing\-Based Generation

A major line of recent work formulates slide generation as adapting or editing existing presentations rather than generating them from scratch\([Zheng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib22);[Zeng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib19);[Liang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib9);[Tang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib16);[Jung et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib6)\)\. These methods leverage reference slide decks, templates, or the deck’s underlying object model to inherit content organization and visual design\. For example, PPTAgent edits retrieved reference slides, SlideGen extends this paradigm through a multi\-agent pipeline for content planning, layout selection, and refinement, and Talk\-to\-Your\-Slides operates directly on a slide’s structured object model rather than pixels for cheaper, more precise edits\. While these approaches produce visually consistent decks, their dependence on templates or reference presentations constrains layout flexibility, often making generation closer to template filling than true slide synthesis\.

![Refer to caption](https://arxiv.org/html/2609.30294v1/EACL_FINAL_27.png)Figure 1:SlideLab Overview: narrative planning, figure design, flow restructuring, layout debugging and grounding\.
#### End\-to\-End Scientific Slide Generation

Recent work instead aims to synthesize presentations directly from documents or user instructions without relying on reference presentations\([Ge et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib3);[Yang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib17);[Zheng et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib23)\)\. AutoPresent focuses on generating individual slides from natural language instructions but does not explicitly model presentation\-level coherence\. AutoSlides and DeepPresenter extend this setting to end\-to\-end scientific presentation generation using multi\-stage agentic pipelines, while AeSlides instead targets aesthetic layout quality alone, training with reinforcement learning over verifiable layout metrics rather than content or narrative signals\([Pan et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib14)\)\. Collectively, these methods represent a shift toward fully generative presentation systems that leverage the planning, reasoning, and self\-correction capabilities of modern foundation models\.

#### Slide Deck Evaluation

Existing evaluation methods can be broadly grouped into three categories\([Zheng et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib22);[Chen et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib1);[Zhao et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib21);[Ozden et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib13);[Yang et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib18);[Jang et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib5);[Zhang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib20)\)\. Judge\-based methods \(e\.g\., PPTEval and PresentBench\) rely on LLM judges or predefined rubrics to assess completed slide decks, while reference\-based benchmarks such as ArcBench and SlidesGen\-Bench compare generated presentations against “ground\-truth” decks\. Specialized benchmarks instead target particular failure modes, such as iterative editing \(DECKBench\), visual design flaws \(SlideAudit\), or slide\-level editing and layout reasoning, where PPTArena scores in\-place PowerPoint edits with a dual VLM judge and PPTBench decomposes PowerPoint understanding into Detection, Understanding, Modification, and Generation\([Ofengenden et al\., 2026](https://arxiv.org/html/2609.30294#bib.bib12);[Huang et al\., 2025](https://arxiv.org/html/2609.30294#bib.bib4)\)\. Despite these differences, all evaluate presentations as static artifacts with complete deck access\. None of them models how audience understanding evolves throughout a presentation or assesses communication in a sequential presentation setting\. Our Conference Room environment addresses this gap and is evaluated alongside existing benchmarks rather than replacing them\.

## 3SlideLab

SlideLabbuilds upon feedback from human annotations that reveal several flaws in scientific slide generation frameworks, including poor narrative flow, grounding issues, poor layout aesthetics, and high cost \(see Figure[1](https://arxiv.org/html/2609.30294#S2.F1)\)\. Our framework improves the structure, grounding, visual, and other aspects of presentation generation discussed below\. It also reduces the cost and token consumption by 3×\\timescompared to DeepPresenter\.

### 3\.1Structured Narrative Planning

One of the main observations from our human annotation study was that current frameworks still lack factual grounding and fail to create a coherent narrative for the presentation\. Although these models can summarize long papers well, they often fail to organize the material into fluent slides\. As a result, many generated decks feel out of order, which heavily affects the audience experience

To address this, we introduce a Planner agent that reads every section, figure, and table of the paper through multi\-turn tool calls and generates multiple candidate presentation blueprints\. A second agent, powered by a stronger frontier model, critiques these candidates and selects the one with the most coherent flow\. For each slide, the final blueprint specifies the title, one\-sentence takeaway, content source, figures and tables to include, and whether a new visual should be generated\. This blueprint is then used by the downstream agents to generate the presentation\. Our ablation \(Table[3](https://arxiv.org/html/2609.30294#S5.T3)\) shows that removing the Planner causes the largest drop in performance, highlighting the importance of planning the presentation before generating slides\.

### 3\.2Concept Figure Design

Scientific presentations often need explanatory diagrams that are not present in the original paper\. However, many existing systems either omit such visuals altogether or generate generic images that are only loosely related to the paper\. This puts more emphasis on visual appeal than on helping the audience understand the work\. As a result, important concepts are often left as blocks of text\.

We parse the entire paper and extract all figures, tables, and plots for the downstream agents\. The Slide Generator reuses these whenever they are sufficient\. If a concept requires a diagram that is not available in the paper, it inserts a<\!\-\- VISUAL\_SLOT: description \-\-\>placeholder\.

The Visual Generator reads the relevant paper sections together with the slide context and writes a detailed prompt for an image generator \(e\.g\.,gpt\-image\-2\), matching the color palette and style already used in the deck\. The generated figure is placed into the visual slot, and we check to ensure that it does not overflow the slide or leave large empty gaps, and that it lines up cleanly with any text, tables, or other figures already on the slide\. Unlike template based approaches that force a fixed number of figures or tables onto every slide, we only add a visual when the content actually needs one\.

### 3\.3Editorial Flow Restructuring

Generated presentations often contain redundant slides or sections that are better merged or presented in a different order\. Existing frameworks handle this only before the generation stage\. We instead use a separate post\-generation Compositor agent that edits the completed presentation\.

Compositor agent can reorder slides, remove redundant ones, merge almost empty slides, move figures to better positions, and insertVISUAL\_SLOTs where a slide becomes too text\-heavy\. It does not change the design or create new slides\. Note that the Compositor has access to the full paper, so every edit is checked against the source before the final deck is returned\.

![Refer to caption](https://arxiv.org/html/2609.30294v1/slide_comp_single.png)Figure 2:Slides shortcuts comparison between SlideLab \(Ours\) and DeepPresenter, Kimi Slides, Manus\.
### 3\.4Layout Debugging and Grounding

After generation, the slides can still have two types of errors: layout defects and grounding errors\. Layout problems such as overflow, overlap, clipping, or poor spacing cannot be reliably detected by text only LLMs or simple deterministic methods\. To address this, we use a post\-generation LayoutDebugger agent powered by a multimodal LLM\. Unlike existing frameworks, which rely only on text during generation, our framework performs an explicit visual inspection of every slide\. Each slide is rendered as a PNG using Playwright111Playwright is an open\-source browser automation framework developed by Microsoft\.[https://github\.com/microsoft/playwright](https://github.com/microsoft/playwright)\.and passed to the LayoutDebugger\. If any layout problem is detected, it returns a corrected HTML version of the slide, which is rendered and checked again for up to three rounds\. We then verify grounding using a frontier long\-context LLM by providing both the generated HTML and the original paper, allowing it to identify unsupported claims, incorrect numbers, or other factual errors before the presentation is finalized\.

## 4ConfArena: Evaluation Environment

Scientific presentations are intended to communicate research to a diverse conference audience, not merely to showcase visually appealing slide decks\. We therefore evaluate presentations in the setting for which they are designed\. Instead of relying on a single LLM to score a completed presentation, our environment simulates a conference talk in which a panel of attendees follows the presentation slide by slide, asks questions whenever communication breaks down, and participates in a Q&A with the presenter\. ConfArena then evaluates observable events that occur during a presentation, such as unsupported claims, narrative defects, audience questions, figure failures, and missing scientific content, and aggregates them into complementary evaluation axes, instead of relying on a single holistic judgment\.

*Per slide**Whole talk*SystemGrounding\. errors↓\\downarrowFig\. errors↓\\downarrowDesign↑\\uparrowFigure use↑\\uparrowNarrative\. errors↓\\downarrowCoverage↑\\uparrowEnv\. score↑\\uparrowPPTAgent0\.980\.643\.213\.185\.400\.610\.42DeepPresenter2\.080\.533\.493\.275\.190\.940\.32Kimi Slides1\.480\.173\.673\.084\.560\.910\.58Manus1\.81–3\.59–4\.140\.890\.37SlideLab\(ours\)0\.710\.264\.213\.803\.290\.990\.73Table 1:Results on ConfArena,100100papers×\\times33seeds\. Grounding \. errors==claims on a slide that the paper does not support, per slide\. Fig\. errors==figures that fabricate content not supported by the paper or are illegible, per slide\. Design==how clean and readable each slide looks, scored 1–5\. Figure use==how well each slide’s figures support its content, scored 1–5\. Narrative \. errors==story\-flow breaks across deck\. Coverage==fraction of the paper’s key points conveyed \(0–1\)\. Env\. score==rank\-normalized composite of the six metrics \(0–1\)\. Best per column in bold\.### 4\.1Simulation Overview

The simulation environment consists of three roles: a*presenter*, an*examiner*, and three*attendee personas*\. The presenter has access to the full paper and answers questions as the author, while the attendees see only the slide deck and build their understanding as the presentation progresses\. The examiner also reads the paper and serves as a reference for all evaluations\. Since the paper, presenter, examiner, and evaluation prompts remain fixed for all slide decks of a paper, differences in the outcome reflect the quality of the slide deck itself\. Each evaluation includes three stages: the presentation, the Q&A, and the final assessment\. Every role is implemented as an independent LLM call\. The models and input modalities used at each stage are summarized in[A\.2](https://arxiv.org/html/2609.30294#A1.SS2)\.

### 4\.2Attendees

A conference audience is inherently diverse\. To capture this, we model three attendee personas: an*expert reviewer*, a*learner*, and a*cross\-field attendee*\. Together, they represent the range of audiences a scientific presentation is intended to serve\.

We intentionally avoid elaborate persona prompts\. Prior work has shown that increasingly detailed personas can degrade LLM judgment \([Kim et al\. \(2025\)](https://arxiv.org/html/2609.30294#bib.bib7);[Zheng et al\. \(2024\)](https://arxiv.org/html/2609.30294#bib.bib24)\), while highly specific descriptions risk representing only a narrow stereotype\. Each attendee is therefore defined only by three attributes: a*focus*, a*question trigger*, and a short*role description*\. The reviewer focuses on scientific correctness and rigor, the learner on whether the methodology and ideas are understandable, and the cross\-field attendee on accessibility without specialized background knowledge\. Their question triggers determine when something is unclear enough to interrupt the presentation \([A\.2](https://arxiv.org/html/2609.30294#A1.SS2)\. Consequently, questions arise from distinct audience perspectives rather than a single evaluation criterion, allowing the same presentation to be perceived differently by different attendees, as in a real conference\. The attendees evaluate the presentation independently and do not observe each other’s reasoning\.

### 4\.3Presentation

Before the presentation begins, the examiner reads the paper and prepares a checklist of the key scientific points that should be communicated, such as the motivation, the methodological contribution, and the main experimental findings\. These serve as the reference against which the presentation is later evaluated\.

The presentation then proceeds slide by slide\. For each slide, the examiner compares the rendered slide against the paper to identify unsupported claims, incorrect numerical values, fabricated visuals, and other grounding errors\. Independently, each attendee interprets the slide from the perspective of its persona, updates its understanding of the talk, evaluates whether the slide follows naturally from the preceding discussion, and decides whether anything is unclear enough to ask a question\. Questions raised during the presentation are collected for the subsequent Q&A\.

### 4\.4Q&A

The collected questions are posed to the presenter, giving higher priority to questions independently raised by multiple attendees\. After the presenter answers using the full paper, the examiner determines why each question arose: because the relevant information was omitted from the slides, communicated unclearly, hidden in an unreadable figure or table, or genuinely beyond the scope of the presentation\. The first three cases indicate shortcomings of the presentation, whereas the last reflects discussion that naturally extends beyond the talk itself\.

### 4\.5Final Assessment

Finally, the observations collected throughout the simulation are aggregated into the metrics reported in Table[1](https://arxiv.org/html/2609.30294#S4.T1)\.Grounding\. errorscounts the claims on each slide that the paper does not support, including incorrect numerical values\.Fig\. errorscounts fabricated figures that depict content, results, or concepts not supported by the source paper or figures that are illegible\.Designandfigure useare scores for how clean each slide looks and how well its figures support that content\.Narrative\. errorscounts breaks in the talk’s flow, such as results appearing before the method or a missing conclusion\.Coveragemeasures what fraction of the paper’s key points actually reach the deck\. We combine them into a single environment score by rank\-normalizing each metric within a paper and averaging them with equal weight\.

## 5Experiments

### 5\.1Experimental Setup

#### Dataset

We evaluate on two datasets: \(1\) the100100paper–deck pairs from ArcBench, covering oral\-presentation papers from CVPR, ICCV, ICLR, ICML, and NeurIPS \(2022–2025\); and \(2\) a held\-out set of3030machine learning papers that are authored by the annotators, used for detailed comparisons and human evaluation\. The100100paper set is used in automated evaluation runs\.

#### Baselines

We compare against the open\-source frameworks PPTAgent and DeepPresenter, along with the commercial systems Kimi Slides[Moonshot AI \(2026\)](https://arxiv.org/html/2609.30294#bib.bib11)and Manus[Manus \(2026\)](https://arxiv.org/html/2609.30294#bib.bib10)\. Since PPTAgent edits an existing deck instead of generating one from scratch, it needs a template to start from\. We ran it with three different templates for each paper and kept the best score\. DeepPresenter is run with default settings, while Kimi Slides and Manus are accessed through their web interfaces using the paper PDF as input\.

#### Evaluation

We evaluate SlideLab through both automatic and human evaluation\. For automatic evaluation, we use our proposed ConfArena environment alongside established benchmarks, including PPTEval, PresentBench, and SlidesGen\-Bench\. Human evaluation is conducted as a blind preference study in which annotators compare complete slide decks without knowing their source\. We also assess the quality of ConfArena by measuring its agreement with human judgments and through controlled perturbation experiments that isolate common presentation problems\.

### 5\.2Human Preference Evaluation

We conducted a blind human evaluation with 10 volunteer annotators, including graduate students, PhD students, and faculty members\. For every paper, annotators were shown four anonymous slide decks \(SlideLab, DeepPresenter, Kimi Slides, and Manus\) in randomized order together with the source paper\. Rather than judging visual appearance alone, they were instructed to select the presentation they would use to deliver the paper at a real conference, considering factual correctness, coverage of the paper’s main contributions, narrative flow, layout, figure and table usage, and overall presentation quality\.

After selecting the preferred presentation among four, annotators additionally rated the chosen presentation along six dimensions: content grounding, content coverage, narrative structure, visual design, information density, and figure usage, using a five\-point Likert scale, to encourage them to consider all aspects of presentation quality rather than making an overall aesthetic judgment\. They could optionally provide free\-form comments explaining the strengths and weaknesses of the evaluated presentations\. The annotation interface and detailed guidelines are provided in the Appendix[A](https://arxiv.org/html/2609.30294#A1)\. Table[2](https://arxiv.org/html/2609.30294#S5.T2)summarizes the results\.SlideLabwas selected for2323of the3030evaluated papers \(77%77\\%\), while Kimi Slides was preferred for the remaining seven; neither DeepPresenter nor Manus was selected\.

The qualitative feedback \(Figure[2](https://arxiv.org/html/2609.30294#S3.F2)and Figure[4](https://arxiv.org/html/2609.30294#A1.F4)\) revealed consistent patterns across the baselines\. Manus was often criticized for text\-heavy slides and the lack of supporting visual content\. In several cases, annotators noted that it inserted screenshots of entire paper pages instead of extracting the relevant figures or tables\. DeepPresenter frequently exhibited layout issues, such as excessive whitespace, and its generated figures were often judged unsuitable for conference presentations\. It also received the largest number of comments related to factual inaccuracies\. Kimi Slides generally produced well\-structured presentations but often overcrowded slides with overlapping elements, while receiving the fewest comments about factual errors\. In contrast,SlideLabconsistently avoided these issues, leading annotators to prefer its decks\.

SystemPapers picked best% of 30SlideLab\(ours\)2377%Kimi Slides723%DeepPresenter00%Manus00%Table 2:Results of the blind human preference study on 30 research papers\. Annotators selected the presentation they would use to deliver the paper at a conference after evaluating factual correctness, coverage, narrative flow, visual design, and figure usage\.
### 5\.3Evaluation with ConfArena

We evaluate on the set of100100research papers described in Section[5\.1](https://arxiv.org/html/2609.30294#S5.SS1)\. PPTAgent, DeepPresenter, andSlideLabeach generate three independent presentations per paper, giving900900decks\. Kimi Slides and Manus are web applications with usage limits, so we generate one presentation per paper\. Every deck is evaluated independently in ConfArena, and the results are in Table[1](https://arxiv.org/html/2609.30294#S4.T1)\.

SlideLabachieves the highest overall environment score and performs best on all evaluation axes except figure errors\. Kimi Slides generated all of the figures programatically rather than image generation models which helped reduce figure errors\. Kimi Slides records the fewest figure errors, but also has noticeably lower coverage of the source paper\. DeepPresenter lies at the opposite end, attempting to include more content but accumulating grounding errors together with weaker narrative flow\.SlideLabachieves the strongest overall performance by maintaining high coverage without sacrificing factual correctness, visual quality, or presentation structure\.

#### SlideLabComponent Ablations

We removed each major component ofSlideLaband re\-scored on the100100\-paper set\. Table[3](https://arxiv.org/html/2609.30294#S5.T3)reports the environment score\. Removing the Planner causes the largest drop, because the Generator no longer has a structured blueprint\. Removing the LayoutDebugger also hurts, mainly through visual\-design and illegible\-figure penalties\. Removing the Compositor or replacing custom visuals with paper figures only causes smaller drops\.

ConfigurationEnv\. scoreSlideLab\(full\)0\.70−\-Planner0\.51−\-LayoutDebugger0\.58−\-Compositor0\.65−\-Custom visuals \(paper figs only\)0\.63Table 3:Component ablations of SlideLab\. The Planner and LayoutDebugger are the two stages whose removal changes the result significantly\. Each configuration is run once per paper on the100100\-paper set, unlike Table[1](https://arxiv.org/html/2609.30294#S4.T1), which averages three independent runs per paper\.

### 5\.4Performance by Existing Metrics

We also evaluate the presentations using PPTEval, PresentBench, and SlidesGen\-Bench to verify that the improvements observed under ConfArena are reflected by existing evaluation methods\.

Table[4](https://arxiv.org/html/2609.30294#S5.T4)shows a similar trend\.SlideLabobtains the highest scores on all three PPTEval dimensions and on SlidesGen\-Bench, while ranking second on PresentBench, where Kimi Slides performs marginally better\. Overall, the rankings produced by existing benchmarks are consistent with those observed under ConfArena, suggesting that the improvements are not biased to the proposed evaluation environment\.

PPTEval \(1–5\)PresentBenchSlidesGen\-BenchSystemContent↑\\uparrowDesign↑\\uparrowCoherence↑\\uparrowpass %↑\\uparrowquiz acc\.↑\\uparrowPPTAgent3\.53\.83\.6630\.74DeepPresenter3\.73\.73\.4640\.76Kimi Slides3\.94\.24\.0730\.79Manus3\.63\.83\.5610\.72SlideLab\(ours\)4\.14\.44\.2710\.84

Table 4:Evaluations on the100100\-paper set using other three methods\. PPTEval: LLM\-judge ratings of content, design, and coherence \(11–55\)\. PresentBench: fraction of decks passing the benchmark’s checks \(%\)\. SlidesGenBench: accuracy on quizzes the benchmark generates from the source paper\.↑\\uparrowhigher is better\. Best in bold\.
### 5\.5Validating ConfArena

The previous experiments evaluateSlideLabusing ConfArena\. We next validate ConfArena itself by examining whether it aligns with human judgment and whether its individual metrics respond to the presentation failures they are intended to capture\.

#### Does ConfArena Reflect Human Judgment?

An evaluation framework is only useful if it reflects how people assess presentation quality\. We therefore compare the rankings produced by ConfArena and existing automated benchmarks against the blind human study described earlier\. Table[5](https://arxiv.org/html/2609.30294#S5.T5)summarizes the resulting system rankings\.

ConfArena shows the same overall ordering as the human study, rankingSlideLabfirst, followed by Kimi Slides, with Manus and DeepPresenter as the weakest systems\. Existing benchmarks broadly agree on the strongest systems but differ on the remaining rankings\. In particular, PresentBench ranks Kimi Slides aboveSlideLab, whereas ConfArena is the only automated evaluation that matches the human preference ordering\.

SystemHumanConfArenaPPTEvalPresentBenchSlidesGenB\.SlideLab11121Kimi Slides22212Manus33344DeepPresenter34433

Table 5:System rank under four evaluation frameworks against humans \(11= best\)\. Manus and DeepPresenter tie in the human study with zero best\-picks each\. PPTEval ranks use the mean of its three dimensions\.PerturbationPPTEvalSlidesG\.PresentB\.ConfArenaFalsified num\.✓×\\times×\\times✓Degraded fig\.✓✓✓✓Dropped slide∘\\circ×\\times×\\times✓Shuffled order✓✓✓✓*Caught /4**3**1**2**4*Table 6:Can evaluation frameworks detect planted errors in slides?✓==the framework’s relevant metric moved in the expected direction;×\\times==did not;∘\\circ==the framework has no metric targetting that failure\.
#### Does ConfArena Measure the Intended Failures?

A useful evaluation metric should respond to the failure it is intended to measure while remaining relatively insensitive to unrelated modifications\. To test this, we manually perturb1515high\-qualitySlideLabpresentations by introducing targeted errors affecting only a single aspect of the presentation, and then re\-evaluate the modified decks using ConfArena\.

Table[7](https://arxiv.org/html/2609.30294#S5.T7)shows that each perturbation primarily affects its corresponding evaluation axis while producing only minor changes in the remaining metrics\. For example, shuffling slide order substantially increases narrative defects without affecting grounding, whereas introducing an incorrect numerical value primarily increases grounding errors\. Similarly, shrinking figures mainly impacts figure readability, and removing a key methodological slide reduces presentation coverage\. These results suggest that the individual ConfArena metrics are sensitive to the presentation failures they are designed to measure rather than unrelated changes\.

We also compare this behavior against existing evaluation benchmarks by applying the same perturbations\. Table[6](https://arxiv.org/html/2609.30294#S5.T6)shows the result: ConfArena is the only evaluation that catches all four targeted failures on their intended axes\. PPTEval has no metric for missing content at all, while SlidesGen\-Bench and PresentBench each have the relevant metric but fail to detect two of the four failures\.

Mean change \(positive==worse\)Damage appliedArc defectsGrounding errorsFigure errorsShuffle slide order\+7\.25\+0\.11\-0\.01Inject one false number\+0\.50\+1\.51\+0\.20Shrink figures by 50%\+0\.25\-0\.17\+0\.51Drop Slides\+2\.50\+0\.01\-0\.01

Table 7:We damage1515SlideLabdecks in targeted ways and re\-run the ConfArena\. Cells show the mean change per metric; bold marks the metric the damage targets\. Each damage moves mainly its own metric\.

## 6Conclusion

We presentedSlideLab, a training\-free framework for generating scientific presentations from research papers through structured narrative planning, progressive slide refinement, and multimodal grounding\. We also introduced ConfArena, a conference\-style evaluation environment that assesses presentations from the audience’s perspective through slide\-by\-slide interactions and Q&A\. Experiments, including blind human evaluation that considers both scientific content and presentation quality, show that SlideLab produces more effective presentations while requiring substantially fewer inference tokens than existing systems\. We hope these contributions provide a stronger foundation for future research on both scientific presentation generation and its evaluation\.

## Limitations

Our work focuses on generating presentations from research papers and evaluates systems within this setting\. While theSlideLabframework is largely domain\-agnostic, the experiments in this work are limited to conference\-style presentations\. Many of the design principles used bySlideLab, including narrative planning, multimodal grounding, and iterative slide refinement, are not specific to scientific presentations and may generalize to other presentation domains\. Future work could investigate the applicability of the framework to educational, business, and technical presentations, which often differ in audience, communication objectives, and presentation style\.

## Ethical Considerations

Our work is intended to support scientific communication by assisting in the preparation and evaluation of research presentations, not to replace human authorship or scientific judgment\. Since the system can generate incorrect or misleading content, all generated slides should be reviewed by authors before public use\.

## References

- Chen et al\. \(2026\)Xin\-Sheng Chen, Jiayu Zhu, Pei\-lin Li, Hanzheng Wang, Shuojin Yang, and Meng\-Hao Guo\. 2026\.Presentbench: A fine\-grained rubric\-based benchmark for slide generation\.*arXiv preprint arXiv:2603\.07244*\.
- Fu et al\. \(2022\)Tsu\-Jui Fu, William Yang Wang, Daniel McDuff, and Yale Song\. 2022\.[Doc2ppt: Automatic presentation slides generation from scientific documents](https://arxiv.org/abs/2101.11796)\.*Preprint*, arXiv:2101\.11796\.
- Ge et al\. \(2025\)Jiaxin Ge, Zora Zhiruo Wang, Xuhui Zhou, Yi\-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, and Trevor Darrell\. 2025\.[Autopresent: Designing structured visuals from scratch](https://arxiv.org/abs/2501.00912)\.*Preprint*, arXiv:2501\.00912\.
- Huang et al\. \(2025\)Zheng Huang, Xukai Liu, Tianyu Hu, Kai Zhang, and Ye Liu\. 2025\.[Pptbench: Towards holistic evaluation of large language models for powerpoint layout and design understanding](https://arxiv.org/abs/2512.02624)\.*Preprint*, arXiv:2512\.02624\.
- Jang et al\. \(2026\)Daesik Jang, Morgan Lindsay Heisler, Linzi Xing, Yifei Li, Edward Wang, Ying Xiong, Yong Zhang, and Zhenan Fan\. 2026\.[Deckbench: Benchmarking multi\-agent frameworks for academic slide generation and editing](https://arxiv.org/abs/2602.13318)\.*Preprint*, arXiv:2602\.13318\.
- Jung et al\. \(2026\)Kyudan Jung, Hojun Cho, Jooyeol Yun, Soyoung Yang, Jaehyeok Jang, and Jaegul Choo\. 2026\.[Talk to your slides: High\-efficiency slide editing via language\-driven structured data manipulation](https://doi.org/10.18653/v1/2026.findings-acl.166)\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 3370–3399, San Diego, California, United States\. Association for Computational Linguistics\.
- Kim et al\. \(2025\)Junseok Kim, Nakyeong Yang, and Kyomin Jung\. 2025\.[Persona is a double\-edged sword: Rethinking the impact of role\-play prompts in zero\-shot reasoning tasks](https://doi.org/10.18653/v1/2025.findings-ijcnlp.51)\.In*Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics*, pages 848–862, Mumbai, India\. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics\.
- Kumar and Chowdary \(2024\)Keshav Kumar and Ravindranath Chowdary\. 2024\.[Slidespawn: An automatic slides generation system for research publications](https://arxiv.org/abs/2411.17719)\.*Preprint*, arXiv:2411\.17719\.
- Liang et al\. \(2025\)Xin Liang, Xiang Zhang, Yiwei Xu, Siqi Sun, and Chenyu You\. 2025\.[Slidegen: Collaborative multimodal agents for scientific slide generation](https://arxiv.org/abs/2512.04529)\.*Preprint*, arXiv:2512\.04529\.
- Manus \(2026\)Manus\. 2026\.Manus playbook: Slide generator\.[https://manus\.im/playbook/slide\-generator](https://manus.im/playbook/slide-generator)\.Accessed: 2026\-08\-04\.
- Moonshot AI \(2026\)Moonshot AI\. 2026\.Kimi slides\.[https://www\.kimi\.com/slides](https://www.kimi.com/slides)\.Accessed: 2026\-08\-04\.
- Ofengenden et al\. \(2026\)Michael Ofengenden, Yunze Man, Ziqi Pang, Liang\-Yan Gui, and Yu\-Xiong Wang\. 2026\.[Pptarena: A benchmark for powerpoint editing](https://arxiv.org/abs/2512.03042)\.*Preprint*, arXiv:2512\.03042\.
- Ozden et al\. \(2026\)Tarik Can Ozden, Sachidanand VS, Furkan Horoz, Ozgur Kara, Junho Kim, and James Matthew Rehg\. 2026\.[Narrative\-driven paper\-to\-slide generation via arcdeck](https://arxiv.org/abs/2604.11969)\.*Preprint*, arXiv:2604\.11969\.
- Pan et al\. \(2026\)Yiming Pan, Chengwei Hu, Xuancheng Huang, Can Huang, Mingming Zhao, Yuean Bi, Xiaohan Zhang, Aohan Zeng, and Linmei Hu\. 2026\.[Aeslides: Incentivizing aesthetic layout in llm\-based slide generation via verifiable rewards](https://arxiv.org/abs/2604.22840)\.*Preprint*, arXiv:2604\.22840\.
- Sun et al\. \(2021\)Edward Sun, Yufang Hou, Dakuo Wang, Yunfeng Zhang, and Nancy X\. R\. Wang\. 2021\.[D2S: Document\-to\-slide generation via query\-based text summarization](https://doi.org/10.18653/v1/2021.naacl-main.111)\.In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 1405–1418, Online\. Association for Computational Linguistics\.
- Tang et al\. \(2025\)Wenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao, Yuhang Wang, Xuxin Tang, Qing Li, Yuehe Ma, Junliang Liu, Shisong Tang, and Michael R\. Lyu\. 2025\.[SlideCoder: Layout\-aware RAG\-enhanced hierarchical slide generation from design](https://doi.org/10.18653/v1/2025.emnlp-main.458)\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 9015–9039, Suzhou, China\. Association for Computational Linguistics\.
- Yang et al\. \(2025\)Yuheng Yang, Wenjia Jiang, Yang Wang, Yiwei Wang, and Chi Zhang\. 2025\.[Auto\-slides: Automatic academic presentation generation with multi\-agent collaboration](https://arxiv.org/abs/2509.11062)\.*arXiv preprint arXiv:2509\.11062*\.AGI Lab, Westlake University; University of California at Merced; Corresponding author: Chi Zhang\.
- Yang et al\. \(2026\)Yunqiao Yang, Wenbo Li, Houxing Ren, Zimu Lu, Ke Wang, Zhiyuan Huang, Zhuofan Zong, Mingjie Zhan, and Hongsheng Li\. 2026\.[Slidesgen\-bench: Evaluating slides generation via computational and quantitative metrics](https://arxiv.org/abs/2601.09487)\.*Preprint*, arXiv:2601\.09487\.
- Zeng et al\. \(2025\)Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui, and Hwee Tou Ng\. 2025\.[Slidetailor: Personalized presentation slide generation for scientific papers](https://arxiv.org/abs/2512.20292)\.*Preprint*, arXiv:2512\.20292\.
- Zhang et al\. \(2025\)Zhuohao \(Jerry\) Zhang, Ruiqi Chen, Mingyuan Zhong, and Jacob O\. Wobbrock\. 2025\.[Slideaudit: A dataset and taxonomy for automated evaluation of presentation slides](https://doi.org/10.1145/3746059.3747736)\.In*Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology*, UIST ’25, page 1–23\. ACM\.
- Zhao et al\. \(2026\)Bo Zhao, Maosheng Pang, Chen Zhang, Huan Yang, Yixin Cao, and Wei Ji\. 2026\.[Unipptbench: A unified benchmark for presentation generation across diverse input settings](https://arxiv.org/abs/2605.17356)\.*Preprint*, arXiv:2605\.17356\.
- Zheng et al\. \(2025\)Hao Zheng, Xinyan Guan, Hao Kong, Jia Zheng, Weixiang Zhou, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, and Le Sun\. 2025\.[Pptagent: Generating and evaluating presentations beyond text\-to\-slides](https://arxiv.org/abs/2501.03936)\.*Preprint*, arXiv:2501\.03936\.
- Zheng et al\. \(2026\)Hao Zheng, Guozhao Mo, Xinru Yan, Qianhao Yuan, Wenkai Zhang, Xuanang Chen, Yaojie Lu, Hongyu Lin, Xianpei Han, and Le Sun\. 2026\.[Deeppresenter: Environment\-grounded reflection for agentic presentation generation](https://arxiv.org/abs/2602.22839)\.*Preprint*, arXiv:2602\.22839\.
- Zheng et al\. \(2024\)Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens\. 2024\.[When “a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models](https://doi.org/10.18653/v1/2024.findings-emnlp.888)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 15126–15154, Miami, Florida, USA\. Association for Computational Linguistics\.

## Appendix AAppendix

### A\.1SlideLabConfiguration

Table[8](https://arxiv.org/html/2609.30294#A1.T8)lists the model used at each stage ofSlideLab, with max number of multi\-turn rounds\. All text agents run at temperature1\.01\.0with high reasoning effort\. The Planner, Slide Generator, and Compositor each run as a tool\-use loop with the round limits shown; hitting the limit forces the agent to produce its final output from what it has gathered so far\.

StageModelMax multi\-turn roundsPlannergpt\-5\.560Slide Generatormimo\-v2\.5\-pro50Visual Generatormimo\-v2\.5\-pro15Compositormimo\-v2\.5\-pro30Layout Debuggergpt\-5\.55 per slideNarration Enginemimo\-v2\.5\-pro15Table 8:Model and tool\-round limit perSlideLabstage\.nnis the number of slides in the deck\. All stages use temperature1\.01\.0and high reasoning effort\.
### A\.2ConfArena Configuration

ConfArena runs each deck through a simulated conference talk with three independent attendee personas – a reviewer, a learner, and a cross\-field listener – each tracking its own understanding of the presentation as it unfolds\. Table[9](https://arxiv.org/html/2609.30294#A1.T9)lists the focus and question triggers of each attendee\. All calls use openai/gpt\-5\.4 at temperature0\.20\.2with medium reasoning effort\. Only the per\-slide visual pass sees the rendered slide image; every other call works from slide text and the source paper\.

AttendeeFocus and Question TriggerExpert reviewerFocuses on scientific rigor and correctness\. Asks a question when a claim, result, or numerical value appears unsupported, inconsistent, or overstated\.LearnerFocuses on whether the core ideas, methodology, and motivation are easy to follow\. Asks a question when an important concept or step is insufficiently explained\.Cross\-field attendeeFocuses on accessibility without deep subfield knowledge\. Asks a question when jargon or field\-specific assumptions are introduced without adequate explanation\.Table 9:Attendee personas used in ConfArena\.
### A\.3Evaluation Configs

PPTEval, PresentBench, and SlidesGen\-Bench are run with the default settings from their released codebases, with openai/gpt\-5\.4 as the judge model throughout\. ConfArena uses openai/gpt\-5\.4\. Decks average 19\.4 slides for SlideLab, 18\.5 for DeepPresenter, 19\.7 for Kimi Slides, and 14\.2 for Manus\.

### A\.4Human Evaluation Interface

Figure[3](https://arxiv.org/html/2609.30294#A1.F3)shows the annotation interface used in the human preference study\. Annotators first pick the deck they would use to present the paper at a conference, then rate that deck on six dimensions using a five\-point scale\. The six dimensions and their anchors are listed in Table[10](https://arxiv.org/html/2609.30294#A1.T10)\. Annotators were known to the authors and participated on a voluntary basis\. They were aware that their annotations will be shown for analysis in our work\.

![Refer to caption](https://arxiv.org/html/2609.30294v1/fig/ui_main.png)Figure 3:Annotation interface for the human preference study\. Annotators see the four anonymized decks side by side, pick the one they would present, then rate it on six dimensions\.![Refer to caption](https://arxiv.org/html/2609.30294v1/fig/comp1.png)

![Refer to caption](https://arxiv.org/html/2609.30294v1/fig/comp2.png)

![Refer to caption](https://arxiv.org/html/2609.30294v1/fig/comp3.png)

Figure 4:Representative qualitative comparisons supporting the quantitative results reported in the main paper\.DimensionWhat annotators are askedContent GroundingDo the slides match what the paper says?Content CoverageDoes the deck cover the paper’s key contributions?Narrative StructureDo the slides tell a coherent story?Visual DesignDoes it look like a professional conference talk?Information DensityAre slides concise, not walls of text?Figure UsageAre figures well\-chosen and sized correctly?Table 10:The six dimensions annotators rate after picking their preferred deck\. Each uses a five\-point scale with the anchors shown in the interface\.#### Why Likert Scoring?

Many presentation attributes, including visual design, figure use, and narrative quality, cannot be adequately captured by binary decisions\. We therefore use Likert ratings for these dimensions, allowing the evaluation to capture incremental differences in presentation quality that binary labels would overlook\.

### A\.5Detailed Perturbation Scores

Table[12](https://arxiv.org/html/2609.30294#A1.T12)reports the full numeric results behind Table[6](https://arxiv.org/html/2609.30294#S5.T6)\. For each benchmark and perturbation we show the metric that benchmark uses for that failure mode, its mean value on the unperturbed decks, its mean value on the perturbed decks, and the change\. A metric counts as detecting the failure if it moves in the expected direction \(down for scores where higher is better, up for error counts\)\. Cells marked – mean the benchmark has no metric for that failure mode\.

StageTime \(s\)Cost \($\)Tokens \(M\)Planner960\.100\.09Slide Generator2640\.030\.23ImageVisualGenerator \+ Compositor2610\.420\.32LayoutDebugger1220\.030\.05SlideLabtotal7440\.580\.69DeepPresenter16202\.12\.9

Table 11:Average wall\-clock time, API cost, and token usage per deck, from production logs\. Cost includes image generation, which dominates the ImageVisualGenerator \+ Compositor stage\.PerturbationBenchmarkMetric \(direction\)BaselinePerturbedΔ\\DeltaDetected?Falsified numberPPTEvalcontent \(1–5,↑\\uparrow\)4\.304\.27−0\.03\-0\.03✓SlidesGen\-Benchquiz accuracy \(0–1,↑\\uparrow\)0\.980\.98\+0\.00\+0\.00×\\timesPresentBenchmaterial\-dep\. pass % \(↑\\uparrow\)27\.545\.0\+17\.5\+17\.5×\\timesSlideLabfaithfulness \(0–1,↑\\uparrow\)0\.580\.57−0\.01\-0\.01✓Degraded figurePPTEvaldesign \(1–5,↑\\uparrow\)4\.053\.85−0\.20\-0\.20✓SlidesGen\-Benchvisual weighted total \(0–10,↑\\uparrow\)8\.458\.22−0\.23\-0\.23✓PresentBenchmaterial\-indep\. pass % \(↑\\uparrow\)100\.092\.0−8\.0\-8\.0✓SlideLabfigure legible % \(↑\\uparrow\)75\.374\.2−1\.06\-1\.06✓Dropped slidePPTEvalcoverage axis–––∘\\circSlidesGen\-Benchquiz accuracy \(0–1,↑\\uparrow\)0\.980\.97−0\.01\-0\.01×\\timesPresentBenchmaterial\-dep\. pass % \(↑\\uparrow\)27\.527\.5\+0\.00\+0\.00×\\timesSlideLabcoverage \(0–1,↑\\uparrow\)0\.920\.90−0\.015\-0\.015✓Shuffled orderPPTEvalcoherence \(1–5,↑\\uparrow\)4\.903\.70−1\.20\-1\.20✓SlidesGen\-Benchlogical flow \(1–10,↑\\uparrow\)8\.806\.40−2\.40\-2\.40✓PresentBenchmaterial\-indep\. pass % \(↑\\uparrow\)100\.090\.0−10\.0\-10\.0✓SlideLabnarrative \(0–1,↑\\uparrow\)0\.930\.69−0\.238\-0\.238✓Table 12:Mean baseline and perturbed scores per benchmark per perturbation, across1515papers\.Δ\\Delta= perturbed minus baseline; for metrics where higher is better, a negativeΔ\\Deltameans the benchmark moved in the expected direction\.✓==detected,×\\times==not detected,∘\\circ==no metric exists for that failure mode\. Baselines are the mean over the1010unperturbed decks\.
### A\.6Computational Cost

Table[11](https://arxiv.org/html/2609.30294#A1.T11)breaks down the average wall\-clock time and API cost per deck on the100100\-paper set\. The full system costs approximately $0\.580\.58per paper and runs in under seven minutes\.SlideLabuses about700700K tokens per deck, roughly3×3\\timesfewer than DeepPresenter \(∼\\sim30003000K\), mainly because the Planner reads the paper once and the Slide Generator works slide by slide against a fixed blueprint, instead of reflecting over the whole deck repeatedly\.

相似文章

X+Slides:面向受众条件的幻灯片生成基准测试

arXiv cs.AI

X+Slides是一个新的基准,用于评估从源文档生成面向受众条件的幻灯片,它使用源基础探针和受众特定的效用权重。在DeepPresenter、SlideTailor和NotebookLM上的实验表明,当前系统能够恢复大量但不够完整的受众关键信息。

DeepSlide:从幻灯片制品到演讲交付

arXiv cs.AI

DeepSlide 是一个人机协同的多智能体系统,覆盖完整的演示流程,从需求获取、带时间预算的叙事规划,到基于证据的幻灯片-脚本生成以及排练支持。它引入了一个双记分板基准,将静态制品质量与动态交付卓越性清晰分离,并在叙事流畅性、节奏精准度和幻灯片-脚本协同方面取得了显著提升。

ArcDeck:叙事驱动的论文到幻灯片生成

Hugging Face Daily Papers

ArcDeck 是一个多智能体框架,通过话语树和迭代智能体优化来建模逻辑流程,从而从学术论文生成演示幻灯片,性能优于直接摘要方法。该论文还引入了 ArcBench,这是一个新的基准测试,用于评估论文到幻灯片生成,强调叙事连贯性和逻辑结构。

AI生成的幻灯片:它们好吗?学生能分辨吗?

arXiv cs.AI

本文研究了使用生成式AI工具(NotebookLM、Claude、M365 Copilot、Cursor、Claude Code)从教师笔记生成幻灯片,发现编程助手生成的幻灯片质量最佳,且学生无法可靠地区分AI生成的幻灯片与人工制作的幻灯片。