SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

arXiv cs.AI Papers

Summary

SAGE is a framework for automating storyboard generation in short drama production using self-evolving rules and attribution-guided updates, achieving expert-level performance and reducing authoring time in commercial deployment.

arXiv:2608.17468v1 Announce Type: new Abstract: Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:04 AM

# Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Source: [https://arxiv.org/html/2608.17468](https://arxiv.org/html/2608.17468)
## SAGE: Self\-Evolving Storyboard Skills via Attribution\-Guided Rule EvolutionCCS:Computing methodologies Artificial intelligence

Maolin Ran,Xiaoyang LuAffiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[xiaoyangl@sjtu\.edu\.cn](mailto:[email protected]),Jiaqi LiuAffiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[jkliu189@gmail\.com](mailto:[email protected]),Jian WangAffiliation:CreativeFitting,Shanghai,Chinaemail:[jim\.wang@creativefitting\.ai](mailto:[email protected]),Weiwen LiuAffiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[wwliu@sjtu\.edu\.cn](mailto:[email protected]),Jianghao LinAffiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[linjianghao@sjtu\.edu\.cn](mailto:[email protected]),Yong YuAffiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[yyu@sjtu\.edu\.cn](mailto:[email protected])andWeinan ZhangNote:Corresponding authors\.Affiliation:Shanghai Jiao Tong University,Shanghai,Chinaemail:[wnzhang@sjtu\.edu\.cn](mailto:[email protected])

###### Abstract\.

Storyboards decompose screenplays into shot\-by\-shot visual plans that drive automated short drama production\. Because high\-quality storyboarding rests on the tacit expertise of professional directors, it remains a capacity bottleneck at industrial scale\. Large language models can automate this step, yet existing ways of equipping them with directorial knowledge face three challenges: \(1\)*Knowledge acquisition*: the craft stays implicit in exemplars or must be authored by hand, so explicit knowledge exists only where a human writes it\. \(2\)*Knowledge refinement*: authored knowledge is never evaluated against execution outcomes, and opaque generation prevents feedback from being attributed to the knowledge behind each decision\. \(3\)*Knowledge injection*: injecting everything exceeds the usable context, yet hand\-picking knowledge for every narrative group does not scale\. In light of these challenges, we present SAGE \(Skill withAttribution\-GuidedEvolution\)\. SAGE is a deployed framework that learns, attributes, and evolves directing knowledge from expert demonstrations\. It first extracts content\-free rules by contrasting each training screenplay with its expert storyboard\. During generation, the model declares which rules each narrative group adopts\. Joining these records with localized feedback yields attribution at the level of individual rules, which drives targeted rule updates\. Evolved rules are then consolidated into scenario packages accessed through a routing index\. Each group therefore retrieves only a bounded set of scenario packages matched to its situation, without expert intervention\. On 18 test episodes across three genres, SAGE scored77\.877\.8on an expert\-validated rubric, exceeding professional directors’77\.177\.1\. Deployed for 14 days in Virtual Film Studio, a commercial short drama production platform, it produced 1,344 narrative group outputs\. Of these,87\.2%87\.2\\%were accepted without substantive edits, and the production team recorded a drop of over83%83\\%in authoring time per episode\. We release PROSE, the first public dataset pairing screenplays with storyboards authored by professional directors, spanning 68 episodes at[https://github\.com/creDreams/PROSE](https://github.com/creDreams/PROSE)\.

###### Keywords:

self\-evolving agents, skill learning, credit assignment, storyboard generation

## 1\.Introduction

![Five stacked rows compare four systems on one screenplay beat. The top row states the input screenplay line. The next four rows correspond to SAGE, Few-shot, CoT, and Vanilla. Each of these rows holds a textual storyboard record, three camera language fields for shot scale, camera angle, and camera movement, a greyscale thumbnail, and a threat level indicator. The SAGE row lists close-up, low angle, and push-in. The other three rows list flat angles and a static camera. A closing row summarizes the comparison.](https://arxiv.org/html/2608.17468v1/figures/fig_paradigm.png)Figure 1\.Camera language decisions on one screenplay beat from*His Toyboy*EP001, a held\-out test episode, under an identical screenplay, backbone LLM, and storyboard schema\. SAGE alone selects a low\-angle push\-in that reinforces Victor’s dominance, whereas Few\-shot and CoT keep the angle flat and Vanilla widens to a composition that weakens the power dynamic\. Only the textual records are system outputs; thumbnails and threat labels are post\-hoc readings, excluded from evaluation\.Five stacked rows compare four systems on one screenplay beat\. The top row states the input screenplay line\. The next four rows correspond to SAGE, Few\-shot, CoT, and Vanilla\. Each of these rows holds a textual storyboard record, three camera language fields for shot scale, camera angle, and camera movement, a greyscale thumbnail, and a threat level indicator\. The SAGE row lists close\-up, low angle, and push\-in\. The other three rows list flat angles and a static camera\. A closing row summarizes the comparison\.As video generation models mature, automated film production systems increasingly transform screenplays into videos\([14](https://arxiv.org/html/2608.17468#bib.bib39);[40](https://arxiv.org/html/2608.17468#bib.bib38)\)\. In these systems, the*storyboard*is the intermediate artifact connecting creative intent to video synthesis\([14](https://arxiv.org/html/2608.17468#bib.bib39);[53](https://arxiv.org/html/2608.17468#bib.bib37)\)\. It specifies, shot by shot, the visual content, shot scale, camera angle, and camera movement\. Its quality directly constrains the fidelity, coherence, and cinematic expressiveness of the downstream video\.

Manual storyboarding constrains throughput across the short drama industry\. Tens of thousands of serialized episodes ship annually on platforms such as Douyin and TikTok, yet the timeline from concept to release often spans only weeks or even days, which forces creators into multiple roles at once to meet the efficiency demands of the industry\([4](https://arxiv.org/html/2608.17468#bib.bib34)\)\. Storyboards are still authored by hand, and this manual stage caps the speed of short drama production\([37](https://arxiv.org/html/2608.17468#bib.bib35);[32](https://arxiv.org/html/2608.17468#bib.bib36)\)\. The bottleneck persists because storyboarding rests on tacit directorial expertise\([26](https://arxiv.org/html/2608.17468#bib.bib53)\)\. A director jointly decides shot rhythm, visual description, shot scale, camera angle, and camera movement, yet the craft behind these decisions is rarely articulated as explicit principles\. On the commercial platform studied in this work, a director spends more than one hour on a single episode\. Large language models \(LLMs\) can automate this process\. Benchmarks nevertheless show that frontier models lack professional competence in camera language\([17](https://arxiv.org/html/2608.17468#bib.bib42);[34](https://arxiv.org/html/2608.17468#bib.bib41)\)\. Figure[1](https://arxiv.org/html/2608.17468#S1.F1)illustrates this deficit on one dramatic beat, where direct generation defaults to a flat angle that does not convey the scene’s power dynamic\. Effective automation therefore requires equipping LLMs with this domain knowledge\.

Prior efforts inject such knowledge into LLMs through three mechanisms\. Few\-shot prompting supplies screenplay and storyboard exemplars, and chain\-of\-thought \(CoT\) prompting\([36](https://arxiv.org/html/2608.17468#bib.bib46)\)adds reasoning chains written by experts\. Skills\([3](https://arxiv.org/html/2608.17468#bib.bib31)\)package a stable workflow together with separate knowledge files, a design that has become an industrial practice\. These mechanisms let LLMs exploit directorial knowledge, yet three challenges remain\.

C1: Knowledge acquisition\.In existing mechanisms, knowledge is either implicit in exemplars or authored by hand\. Few\-shot exemplars encode the craft implicitly, so the model must re\-induce it at inference time and never obtains an inspectable, reusable form\. CoT chains and skill files state the knowledge explicitly, but directors must author every chain and every file themselves\. Explicit knowledge is thus available only where a human writes it, and the challenge is to extract it automatically from expert demonstrations\.

C2: Knowledge refinement\.Once authored, the knowledge stays fixed and is never evaluated against execution outcomes, so its defects remain undetected\. Refinement must therefore use execution feedback\. Feedback alone is nevertheless insufficient, because generation is opaque: it does not reveal which piece of the injected knowledge shaped which decision\. Existing pipelines judge knowledge at extraction time or by votes over whole trajectories, without tracing its use in generation\([49](https://arxiv.org/html/2608.17468#bib.bib5);[43](https://arxiv.org/html/2608.17468#bib.bib7)\), and self\-correction without such localization is reported to degrade performance\([13](https://arxiv.org/html/2608.17468#bib.bib3)\)\. Effective refinement therefore requires generation to be traceable, so that feedback can be attributed to the knowledge responsible for each decision\.

C3: Knowledge injection\.At inference time, the knowledge given to the model must be selected automatically\. Current practice secures relevance by manual selection, and prior automatic retrieval still operates over knowledge units defined by hand\([10](https://arxiv.org/html/2608.17468#bib.bib6)\)\. Manual selection cannot scale when every narrative group demands its own decision over thousands of entries, and injecting everything exceeds the usable context\.

We present SAGE \(Skill withAttribution\-GuidedEvolution\), a framework that resolves the three challenges on the skill substrate, which already separates a stable workflow from an explicit knowledge base\. For acquisition, SAGE contrasts each training screenplay with its expert storyboard and extracts content\-free rules automatically \(C1\)\. For refinement, generation declares which rules each narrative group adopts\. Joining these*rule\-adoption records*with localized feedback yields three evolution operations: misfiring rules are revised, new rules are added for coverage gaps, and unused rules are retired \(C2\)\. For injection, evolved rules are consolidated by semantic clustering into*scenario packages*with a routing index, so each narrative group retrieves only the packages matched to its situation \(C3\)\.

On a test set of 18 episodes spanning three drama genres, SAGE scored77\.877\.8, exceeding the professional directors’77\.177\.1\. All scores come from an LLM rubric whose agreement with three professional directors exceeds inter\-human agreement \(§[4\.6](https://arxiv.org/html/2608.17468#S4.SS6)\)\. It also outperformed strong baselines given reasoning chains authored by directors and exemplars from adjacent episodes\. Attribution\-guided iteration contributed3\.63\.6points over rule warm\-start, whereas iteration without attribution peaked early and then fell below its starting point\. The consolidated rules transferred unchanged to three other backbone LLMs, with relative gains of8\.1%8\.1\\%to21\.2%21\.2\\%\.

Our contributions are:

- •Framework\.SAGE treats the knowledge base of an LLM skill as learnable parameters and evolves it from expert demonstrations\. To our knowledge, SAGE was the first framework to evolve such a knowledge base under credit assignment at the granularity of individual rules, rather than a single validation score over the whole skill document\. The framework spans evolution, consolidation, and deployment, and it runs in commercial production\.
- •Mechanism\.A rule\-level attribution mechanism that assigns credit over knowledge expressed in natural language\. It links authoring to execution feedback, a connection that static skills lack\. Our ablation and iteration studies show that iteration without it peaks early and then declines\.
- •Dataset\.We release PROSE, to our knowledge the first open dataset pairing screenplays with storyboards*authored by professional human directors*111Publicly available at[https://github\.com/creDreams/PROSE](https://github.com/creDreams/PROSE)\.\. It spans 68 episodes across three professionally produced series of distinct genres\. Prior resources instead provide shot annotations that models reverse\-engineer from finished videos\([32](https://arxiv.org/html/2608.17468#bib.bib36);[53](https://arxiv.org/html/2608.17468#bib.bib37)\)\.
- •Deployment\.We deployed SAGE in Virtual Film Studio \(VFS\), a commercial short drama production platform, and report a 14\-day study on three unseen ongoing dramas\. Of 1,344 narrative group outputs,87\.2%87\.2\\%entered production without substantive edits\. Authoring time per episode fell from over one hour to roughly 10 minutes, a reduction of over83%83\\%\. Storyboarding therefore shifted from manual authoring to review, which is what removes the capacity bottleneck\.

## 2\.Related Work

LLM\-based storyboard generation\.FilmAgent\([40](https://arxiv.org/html/2608.17468#bib.bib38)\)and MovieAgent\([38](https://arxiv.org/html/2608.17468#bib.bib40)\)coordinate role\-playing agents\. FilMaster\([14](https://arxiv.org/html/2608.17468#bib.bib39)\)instead retrieves camera language conventions from 440K film clips\. Neither learns an explicit, inspectable knowledge base from professional storyboard demonstrations\. DramaDirector\([53](https://arxiv.org/html/2608.17468#bib.bib37)\)is the closest system to our task, since it fine\-tunes an LLM planner with SFT and GRPO for short drama storyboards, but it keeps cinematic knowledge implicit in model weights\. A parallel line generates*image*storyboards\([39](https://arxiv.org/html/2608.17468#bib.bib44);[7](https://arxiv.org/html/2608.17468#bib.bib45)\), addressing visual consistency rather than the directorial decomposition studied here, and earlier engine\-based previsualization renders shot candidates under*manually specified*rules\([28](https://arxiv.org/html/2608.17468#bib.bib43)\)\. Benchmark studies consistently find that frontier models lack professional competence in camera language\([17](https://arxiv.org/html/2608.17468#bib.bib42);[34](https://arxiv.org/html/2608.17468#bib.bib41)\), which motivates explicit knowledge injection\.

Storyboard datasets\.SkyScript\-100M\([32](https://arxiv.org/html/2608.17468#bib.bib36)\)and DramaBoard\([53](https://arxiv.org/html/2608.17468#bib.bib37)\)pair short drama scripts with shot\-level annotations\. Their storyboards, however, are*reverse\-engineered from finished videos*, so they capture what ended up on screen rather than the decisions the director made\. PROSE instead releases the directors’ original pre\-production storyboards, the decision traces that demonstration\-supervised evolution requires\.

Experiential knowledge in self\-evolving agents\.Self\-evolving agents\([11](https://arxiv.org/html/2608.17468#bib.bib11)\)improve from their own experience; we focus on the branch that evolves context rather than weights\. Self\-Refine\([20](https://arxiv.org/html/2608.17468#bib.bib1)\)and Reflexion\([30](https://arxiv.org/html/2608.17468#bib.bib2)\)iterate on individual outputs or store trajectory\-level reflections, without accumulating knowledge that persists beyond the task\. Unaided self\-correction is also known to be unreliable\([13](https://arxiv.org/html/2608.17468#bib.bib3)\)\. A second family distills experience into persistent natural\-language knowledge, expressed as failure\-derived rules, state\-conditioned guidelines, causal abstractions, or distilled insights and procedural memory\([43](https://arxiv.org/html/2608.17468#bib.bib7);[10](https://arxiv.org/html/2608.17468#bib.bib6);[21](https://arxiv.org/html/2608.17468#bib.bib24);[49](https://arxiv.org/html/2608.17468#bib.bib5);[8](https://arxiv.org/html/2608.17468#bib.bib30)\)\. A third family stores capabilities as executable skills\. Voyager\([33](https://arxiv.org/html/2608.17468#bib.bib8)\)grows a code skill library through environment feedback, and successors extend the idea to other interactive domains\([31](https://arxiv.org/html/2608.17468#bib.bib27);[51](https://arxiv.org/html/2608.17468#bib.bib29);[35](https://arxiv.org/html/2608.17468#bib.bib26)\); industrial standards package such procedural instructions with resources\([3](https://arxiv.org/html/2608.17468#bib.bib31)\), and memory architectures manage the state generically\([24](https://arxiv.org/html/2608.17468#bib.bib25);[48](https://arxiv.org/html/2608.17468#bib.bib28)\)\. Across this family the maintenance signal is coarse: items are added or voted on from trajectory outcomes, without tracking*which*item influenced*which*output\. Sound and harmful items therefore receive the same credit, and the noise grows as the store accumulates entries\. AutoManual\([5](https://arxiv.org/html/2608.17468#bib.bib4)\)comes closest, since its planner cites the rules it engages, but it learns from binary episodic rewards and revises rules by post\-hoc judgment over whole trajectories\. SAGE instead aligns against expert demonstrations, the only supervision available without an environment that verifies success, and joins per\-group adoption records with localized feedback \(§[4\.4](https://arxiv.org/html/2608.17468#S4.SS4)\)\. Agent\-Pro\([47](https://arxiv.org/html/2608.17468#bib.bib10)\)and AgentEvolver\([45](https://arxiv.org/html/2608.17468#bib.bib13)\)also use reflection or self\-attribution, but operate over policy prompts and RL action steps rather than declarative knowledge items\.

Natural\-language component optimization\.Manually supplied context is brittle, since in\-context learning is sensitive to exemplar relevance and ordering, and selecting effective exemplars is itself a retrieval problem\([18](https://arxiv.org/html/2608.17468#bib.bib47);[50](https://arxiv.org/html/2608.17468#bib.bib48);[19](https://arxiv.org/html/2608.17468#bib.bib49);[29](https://arxiv.org/html/2608.17468#bib.bib50)\)\. Treating the natural\-language components of a frozen\-LLM system as optimizable parameters began with instruction and evolutionary prompt search\([54](https://arxiv.org/html/2608.17468#bib.bib14);[41](https://arxiv.org/html/2608.17468#bib.bib15);[9](https://arxiv.org/html/2608.17468#bib.bib17);[12](https://arxiv.org/html/2608.17468#bib.bib18)\)\. The idea then matured into gradient\-like frameworks\. ProTeGi\([27](https://arxiv.org/html/2608.17468#bib.bib16)\)edits prompts along critique\-derived textual gradients, while TextGrad\([44](https://arxiv.org/html/2608.17468#bib.bib12)\)and Trace\([6](https://arxiv.org/html/2608.17468#bib.bib20)\)backpropagate language feedback through computation graphs and execution traces\. DSPy\([15](https://arxiv.org/html/2608.17468#bib.bib19)\)compiles declarative pipelines with learnable instructions, and later frameworks train agent functions as weights\([46](https://arxiv.org/html/2608.17468#bib.bib21)\)or show that language\-level reflection can outperform RL in sample efficiency\([1](https://arxiv.org/html/2608.17468#bib.bib9)\)\. Most recently the paradigm has reached skills\. EvoSkill\([2](https://arxiv.org/html/2608.17468#bib.bib22)\)edits skill folders from failure analysis, and SkillOpt\([42](https://arxiv.org/html/2608.17468#bib.bib23)\)treats skill documents as trainable state\. In both, the learning signal remains a single scalar validation score over the whole skill, so edits are retained or discarded in bulk and an individual rule’s contribution is never measured\. SAGE shares the view of knowledge as learnable parameters, but maintains rule\-adoption records that assign credit to individual rules, so one round can revise, add, and retire different rules at once\. Our experiments show that this granularity keeps iteration productive where document\-level optimization plateaus \(§[4\.2](https://arxiv.org/html/2608.17468#S4.SS2)\)\.

## 3\.Methodology

We present SAGE, a framework that enables the knowledge component of an LLM skill to evolve autonomously from expert demonstrations\.

### 3\.1\.Problem Formulation

#### Storyboard generation\.

A screenplaySSconsists of scenes with dialogue, action descriptions, and character information\. The goal is to produce a storyboardB=⟨b1,…,bn⟩B=\\langle b\_\{1\},\\ldots,b\_\{n\}\\rangle, an ordered sequence of shots\. Each shotbib\_\{i\}is a structured tuple specifying visual content, shot scale, camera angle, camera movement, characters, and dialogue\. Operationally, each scene is partitioned into*narrative groups*; every group yields a*storyboard segment*of one or more shots, and these segments are merged intoBB\. ProducingBBrequires joint decisions across five professional dimensions:*shot rhythm*,*visual description*,*shot scale*,*camera angle*, and*camera movement*\. These dimensions encode tacit directorial knowledge that is difficult to specify exhaustively in a static prompt\.

#### Skill as workflow and knowledge\.

We define a*skill*as a pairΣ=\(𝒲,ℛ\)\\Sigma=\(\\mathcal\{W\},\\mathcal\{R\}\)\. The*workflow*𝒲\\mathcal\{W\}is a fixed multi\-step procedural scaffold\. It specifies how the screenplay is partitioned, which knowledge is retrieved at each step, and how partial outputs are merged\. The*knowledge base*ℛ\\mathcal\{R\}holds the declarative rules that the workflow consumes\. This decomposition follows emerging industrial standards for agent skills\([3](https://arxiv.org/html/2608.17468#bib.bib31)\)\. Each ruler∈ℛr\\in\\mathcal\{R\}is a pair

\(1\)r=\(cond,prac\),r=\(\\textit\{cond\},\\textit\{prac\}\),whereconddescribes the narrative situation in which the rule applies, such as a shock reaction within a high\-intensity dialogue\. The second elementpracis an executable directive, such as not inserting a breathing shot between that shock reaction and the follow\-up question\. Each rule is tagged with one of the five dimensions\. Rules are further constrained to be*content\-free*: they must not contain concrete shot content or verbatim material from expert storyboards, which prevents data leakage and encourages generalization\.

#### Learning view\.

The workflow𝒲\\mathcal\{W\}is easy to fix by design, whereas skill quality is dominated byℛ\\mathcal\{R\}\. We therefore cast rule acquisition as optimization\. LetgΣ​\(S\)g\_\{\\Sigma\}\(S\)denote the storyboard generated by an LLM equipped with skillΣ\\Sigma\. Letfalign​\(gΣ​\(S\),B∗\)∈\[0,100\]f\_\{\\mathrm\{align\}\}\(g\_\{\\Sigma\}\(S\),B^\{\*\}\)\\in\[0,100\]measure how closely that storyboard matches the expert referenceB∗B^\{\*\}across the five dimensions\. Given a corpus of expert demonstrations𝒟=\{\(Sj,Bj∗\)\}j=1m\\mathcal\{D\}=\\\{\(S\_\{j\},B\_\{j\}^\{\*\}\)\\\}\_\{j=1\}^\{m\}, training seeks

\(2\)ℛ⋆=arg⁡maxℛ​𝔼\(S,B∗\)∼𝒟​\[falign​\(g\(𝒲,ℛ\)​\(S\),B∗\)\]\.\\mathcal\{R\}^\{\\star\}=\\arg\\max\_\{\\mathcal\{R\}\}\\;\\mathbb\{E\}\_\{\(S,B^\{\*\}\)\\sim\\mathcal\{D\}\}\\left\[f\_\{\\mathrm\{align\}\}\\big\(g\_\{\(\\mathcal\{W\},\\mathcal\{R\}\)\}\(S\),B^\{\*\}\\big\)\\right\]\.This formulation yields a direct analogy to standard machine learning\. The rule setℛ\\mathcal\{R\}plays the role of*learnable parameters*, andfalignf\_\{\\mathrm\{align\}\}acts as the*training objective*\. The*test metric*is a separate reference\-free quality scorefqualf\_\{\\mathrm\{qual\}\}, which measures absolute professional quality without access toB∗B^\{\*\}\. Two properties distinguish this setting from gradient\-based learning\. The “parameters” are discrete natural\-language rules, and the optimization signal must be routed to individual rules through an explicit*attribution*mechanism \(§[3\.3](https://arxiv.org/html/2608.17468#S3.SS3)\)\.

### 3\.2\.Framework Overview

![A horizontal pipeline of four blocks. The leftmost block holds the screenplay and the expert storyboard. The Evolution block encloses a cycle of four numbered phases, namely rule extraction, attributed generation, alignment evaluation, and attribution-guided diagnosis, with rule-adoption records at the center and a loop labeled T rounds. The Consolidation block lists deduplication, condition embedding, clustering, and scenario labeling. The Knowledge Base block holds a routing index and scenario packages. The rightmost Inference block lists grouping, top-k routing, rule-injected generation, and merging into a storyboard.](https://arxiv.org/html/2608.17468v1/figures/fig_framework.png)Figure 2\.Overview of SAGE\.Stage 1 \(Evolution\): for each training episode, a four\-phase loop generates storyboards with per\-group rule attribution, evaluates them against the expert reference, and revises the rule set through attribution\-guided diagnosis\.Stage 2 \(Consolidation\): rules evolved across all episodes are deduplicated, embedded, and clustered into scenario packages with a routing index\.Stage 3 \(Inference\): on unseen screenplays, narrative groups are routed to their matched scenario packages, whose rules are injected into generation; no expert reference is required\.A horizontal pipeline of four blocks\. The leftmost block holds the screenplay and the expert storyboard\. The Evolution block encloses a cycle of four numbered phases, namely rule extraction, attributed generation, alignment evaluation, and attribution\-guided diagnosis, with rule\-adoption records at the center and a loop labeled T rounds\. The Consolidation block lists deduplication, condition embedding, clustering, and scenario labeling\. The Knowledge Base block holds a routing index and scenario packages\. The rightmost Inference block lists grouping, top\-k routing, rule\-injected generation, and merging into a storyboard\.As shown in Figure[2](https://arxiv.org/html/2608.17468#S3.F2), SAGE operates in three stages\. All three share a unified generation pipeline of*narrative grouping*,*scenario routing*,*rule injection*,*attributed generation*, and*segment merging*\.Stage 1 \(evolution, §[3\.3](https://arxiv.org/html/2608.17468#S3.SS3)\)refines a per\-episode rule set through a four\-phase loop whose key ingredient is*rule\-level attribution*\.Stage 2 \(consolidation, §[3\.4](https://arxiv.org/html/2608.17468#S3.SS4)\)deduplicates and clusters the large, redundant union of per\-episode rule sets into*scenario packages*with a routing index\.Stage 3 \(inference, §[3\.5](https://arxiv.org/html/2608.17468#S3.SS5)\)routes each group of an unseen screenplay to its top\-kkpackages and generates a segment with the retrieved rules injected\.

### 3\.3\.Stage 1: Attribution\-Guided Rule Evolution

The evolution stage refines the rule set for each training episode overTTrounds\. Each round executes the four phases shown in Figure[3](https://arxiv.org/html/2608.17468#S3.F3), and Algorithm[1](https://arxiv.org/html/2608.17468#alg1)summarizes the loop\.

![The upper part shows the four phases of one evolution round in sequence, annotated with a real trace in which a camera angle rule scores 63 out of 100 and is marked for revision. The lower part expands rule-level credit assignment. A rule-adoption record and a localized feedback item meet at a join operation, which produces a rule-level diagnosis. Three outcomes branch from the diagnosis, namely a revised rule, a newly added rule, and a retired rule.](https://arxiv.org/html/2608.17468v1/figures/fig_evolution.png)Figure 3\.One round of attribution\-guided rule evolution, on a real trace from*Beyond the Wall*\. Attributed generation \(Phase 2\) records which rules each group adopted; diagnosis \(Phase 4\) joins these records with dimension\-level feedback to decide, per rule, whether to revise, add, or retire\.The upper part shows the four phases of one evolution round in sequence, annotated with a real trace in which a camera angle rule scores 63 out of 100 and is marked for revision\. The lower part expands rule\-level credit assignment\. A rule\-adoption record and a localized feedback item meet at a join operation, which produces a rule\-level diagnosis\. Three outcomes branch from the diagnosis, namely a revised rule, a newly added rule, and a retired rule\.#### Phase 1: Rule Extraction\.

In the first round, an initial rule setℛ\(1\)\\mathcal\{R\}^\{\(1\)\}is extracted by contrasting screenplay with expert storyboard\. The screenplay is first partitioned into a two\-level hierarchy, in which scenes at level 1 are split into*narrative groups*at level 2\. A group is a dialogue exchange, an action sequence, or an emotional beat, and is the minimal unit of analysis\. For every group, the extractor examines how the director decomposed it into shots, then induces content\-free rules\(cond,prac\)\(\\textit\{cond\},\\textit\{prac\}\)that explain the observed decisions\. Subsequent rounds inherit the revised rule set from the previous round’s Phase 4\.

#### Phase 2: Attributed Generation\.

The model generates one segment per group from three inputs\. These are the screenplay, a director vocabulary defining the legal shot scales, angles, and movements, and the current rule setℛ\(t\)\\mathcal\{R\}^\{\(t\)\}\. The expert storyboard is withheld\. The defining feature of this phase is the*rule\-adoption record*: for each groupuu, the model declares the adopted rulesAu⊆ℛ\(t\)A\_\{u\}\\subseteq\\mathcal\{R\}^\{\(t\)\}alongside the segment it produces\. The full*attribution map*𝒜\(t\)=\{\(u,Au\)\}u∈groups⁡\(S\)\\mathcal\{A\}^\{\(t\)\}=\\\{\(u,A\_\{u\}\)\\\}\_\{u\\in\\mathrm\{groups\}\(S\)\}makes every generation decision traceable to the rules that informed it\.

#### Phase 3: Alignment Evaluation\.

The storyboard is scored against the expert reference byfalignf\_\{\\mathrm\{align\}\}, which produces an overall score, five per\-dimension scores, and natural\-language feedback that localizes each deviation\. One such deviation reads that the confrontation in scene 4 lacks a re\-establishing two\-shot after three consecutive close\-ups\. The expert storyboard is visible only to the evaluator, never to the generator\. Appendix[C](https://arxiv.org/html/2608.17468#A3)details the review protocol\.

#### Phase 4: Attribution\-Guided Diagnosis\.

Diagnosis joins the feedback with𝒜\(t\)\\mathcal\{A\}^\{\(t\)\}to perform rule\-level credit assignment\. Deviations fall into two classes with distinct remedies\. InClass A \(misfiring rule\), a dimension deviates in groups where some rulerr*is*adopted\. That rule is implicated, so its condition or practice is*revised*\. InClass B \(coverage gap\), a dimension deviates in groups where*no*adopted rule governs it\. No existing rule is at fault, so a*new*rule is induced from the feedback\. Rules never adopted throughout the episode are*retired*, which keeps the set minimal\. At most30%30\\%of rules may be modified per round, a cap that prevents destructive oscillation\. The revised setℛ\(t\+1\)\\mathcal\{R\}^\{\(t\+1\)\}seeds the next round\.

Without attribution, feedback can only be assigned at the episode level\. The optimizer then knows*that*a dimension scored poorly but not*which*rule caused it, so revisions become undirected rewrites\. Our ablations identify this shift from episode\-level to rule\-level credit assignment as the condition for sustained improvement \(§[4](https://arxiv.org/html/2608.17468#S4)\)\.

Algorithm 1Attribution\-Guided Rule Evolution for a Single Episode0:screenplay

SS, expert storyboard

B∗B^\{\*\}, rounds

TT
0:evolved rule set

ℛ\(T\+1\)\\mathcal\{R\}^\{\(T\+1\)\}
1:

U←Group⁡\(S\)U\\leftarrow\\mathrm\{Group\}\(S\)\{two\-level narrative grouping\}

2:

ℛ\(1\)←ExtractRules⁡\(S,B∗,U\)\\mathcal\{R\}^\{\(1\)\}\\leftarrow\\mathrm\{ExtractRules\}\(S,B^\{\*\},U\)\{Phase 1\}

3:for

t=1t=1to

TTdo

4:

\(B\(t\),𝒜\(t\)\)←AttrGen⁡\(S,U,ℛ\(t\)\)\(B^\{\(t\)\},\\mathcal\{A\}^\{\(t\)\}\)\\leftarrow\\mathrm\{AttrGen\}\(S,U,\\mathcal\{R\}^\{\(t\)\}\)\{Phase 2\}

5:

\(s\(t\),F\(t\)\)←falign​\(B\(t\),B∗\)\(s^\{\(t\)\},F^\{\(t\)\}\)\\leftarrow f\_\{\\mathrm\{align\}\}\(B^\{\(t\)\},B^\{\*\}\)\{Phase 3\}

6:for alldeviations

d∈F\(t\)d\\in F^\{\(t\)\}do

7:if

∃r∈Au\\exists\\,r\\in A\_\{u\}governing

dim\(d\)\\dim\(d\)for the group

uuof

ddthen

8:revise

rr\{Class A: misfiring rule\}

9:else

10:

ℛ\(t\)←ℛ\(t\)∪\{InduceRule⁡\(d\)\}\\mathcal\{R\}^\{\(t\)\}\\leftarrow\\mathcal\{R\}^\{\(t\)\}\\cup\\\{\\mathrm\{InduceRule\}\(d\)\\\}\{Class B: gap\}

11:endif

12:endfor

13:retire rules never adopted in

𝒜\(t\)\\mathcal\{A\}^\{\(t\)\}
14:

ℛ\(t\+1\)←\\mathcal\{R\}^\{\(t\+1\)\}\\leftarrowrevised set \{

≤30%\\leq 30\\%of rules modified\}

15:endfor

### 3\.4\.Stage 2: Rule Consolidation

Evolution is per\-episode by design\. The union across the corpus reaches the order of10210^\{2\}rules per episode and several thousand in total, and is redundant and in places contradictory\. Consolidation compresses it into a retrievable knowledge base in three steps\.

#### Deduplication\.

A two\-level union\-find procedure runs within each dimension\. At level 1, rules whose*condition*embeddings exceed a cosine similarity of0\.900\.90are merged into a condition group\. At level 2, within each condition group, rules whose*practice*embeddings also exceed the threshold are collapsed to their semantic centroid\. Practices below the threshold are preserved as alternative practices of a single multi\-practice rule\. One pass thus resolves both redundancy, where condition and practice coincide, and latent contradiction, where a shared condition maps to divergent practices\. The threshold is deliberately conservative because the two error directions are not symmetric\. Merging rules that differ in meaning destroys knowledge irrecoverably, whereas failing to merge equivalent rules only leaves redundancy, since divergent practices survive as alternatives of one rule instead of being collapsed into a centroid\.

#### Embedding and clustering\.

Each deduplicated rule is represented by its condition embedding, capturing the narrative situation it targets\. Embeddings areℓ2\\ell\_\{2\}\-normalized, reduced with UMAP\([22](https://arxiv.org/html/2608.17468#bib.bib52)\), and clustered withkk\-means\. A grid search selects the number of clusters and the UMAP hyperparameters, jointly scoring dimension coverage, cluster size compliance, size uniformity, and silhouette quality\. The number of scenario packages is therefore determined by the data rather than fixed a priori\. Rules triggered by similar situations thus become co\-located and co\-retrieved, regardless of their source episode or drama\.

#### Scenario packaging\.

An LLM agent labels each cluster with a human\-readable scenario name, such as*emotional climax under psychological pressure*, together with a short applicability description\. The agent then materializes the cluster as a*scenario package*, a document that groups the cluster’s rules by dimension\. A compact*routing index*is built alongside, listing every package’s name, description, and rule inventory\. Appendix[A](https://arxiv.org/html/2608.17468#A1)shows the resulting cluster structure and an example package\.

### 3\.5\.Stage 3: Scenario\-Aware Inference

At deployment the evolved skill runs on unseen screenplays with no expert reference and no iteration, reusing the training pipeline\. The screenplay is partitioned into the same two\-level hierarchy used during evolution\. For each group, the model matches its situation against the routing index and selects the top\-kkpackages, recording a justification per match\. We setk=3k\{=\}3to match the three reference episodes supplied to Few\-shot and CoT, which equalizes the injection budget across knowledge\-injection methods\. Routing over the index rather than scanning all rules keeps the injected context bounded as the knowledge base grows\. Each group’s segment is then generated independently and in parallel, conditioned on the group’s screenplay content, the retrieved packages, and the director vocabulary\. Generation also emits a rule\-adoption record, which preserves traceability in deployment\. Finally, segments are concatenated in screenplay order and their shots renumbered into the final storyboard\. Because packages encode situation\-conditioned knowledge rather than model\-specific tricks, the consolidated base is backbone\-agnostic and can be injected into other LLMs unchanged\. Appendix[B](https://arxiv.org/html/2608.17468#A2)traces one real group through routing and generation\.

## 4\.Experiments

We evaluate SAGE around four research questions\.\(RQ1\)Does the evolved skill close the quality gap to professional directors, and how does it compare with strong prompting and skill optimization baselines?\(RQ2\)How much does each component contribute, namely rule warm\-start, iteration, and attribution?\(RQ3\)Does attribution make quality improve over rounds instead of fluctuating?\(RQ4\)Is the consolidated knowledge base portable across backbone LLMs?

### 4\.1\.Experimental Setup

#### Dataset\.

PROSE comprises three professionally produced short drama series of distinct genres, namely a sci\-fi suspense series of 20 episodes \(*Beyond the Wall*\), an urban romance series of 23 \(*His Toyboy*\), and an emotional healing series of 25 \(*My Cure*\)\. Each episode pairs a screenplay, comprising a synopsis, character profiles, and a scene\-level script, with the storyboard authored by the series’ professional director\. We held out 6 episodes per series, 18 in total, as the test set\. The consolidated knowledge base is built exclusively from rules evolved on the remaining 50 training episodes\.

#### Evaluation protocol\.

All systems were scored by a reference\-free*quality*rubric on the five dimensions of shot rhythm, visual description, shot scale, camera angle, and camera movement\. Each dimension uses a 100\-point scale, and the overall score is their average\. Scoring used Claude Opus 4\.6 under a fixed rubric prompt, whose score anchors are given in Appendix[D](https://arxiv.org/html/2608.17468#A4)\. This metric is distinct from the alignment scorefalignf\_\{\\mathrm\{align\}\}used as the training signal: quality measures how*good*a storyboard is against professional standards, whereas alignment measures how*close*it is to a specific expert reference\. The scorer’s reliability is validated against human experts in §[4\.6](https://arxiv.org/html/2608.17468#S4.SS6)\.

#### Baselines\.

We compared six alternatives under the same backbone \(Claude Opus 4\.6\) and output schema\.Directoris the human storyboard, which serves as the expert reference\.Vanillagenerates directly with no external knowledge\.Few\-shotis conditioned on⟨\\langlescreenplay, storyboard⟩\\ranglepairs from the*three nearest neighboring episodes of the same series*, never the target episode itself\.CoTadds reasoning chains that encode decomposition thinking*authored by directors*on those same reference episodes\.EvoSkill\([2](https://arxiv.org/html/2608.17468#bib.bib22)\)andSkillOpt\([42](https://arxiv.org/html/2608.17468#bib.bib23)\)are representative skill optimization methods, reimplemented faithfully on the same test set\. The strong baselines thus receive demonstrations from adjacent episodes, whereas SAGE uses none at inference, which makes the comparison conservative for our method\.

#### Implementation\.

Rule evolution ranT=10T\{=\}10rounds per training episode with at most30%30\\%of rules modified per round\. Conditions were embedded with Qwen3\-Embedding\-8B; consolidation yielded5555scenario packages from2,0362\{,\}036deduplicated rules\. Inference routed each group to its top\-3 packages\. Unless stated otherwise, SAGE results use the round\-5 knowledge base, which is the best\-performing round on the test set\. The no\-attribution ablation is likewise reported at its own best round \(§[4\.4](https://arxiv.org/html/2608.17468#S4.SS4)\), so this oracle round selection is applied symmetrically and characterizes each variant’s upper bound\.

### 4\.2\.Main Results \(RQ1\)

Table 1\.Main comparison on the 18\-episode test set \(quality scores, 100\-point scale\)\.Bold: best among AI systems;underline: exceeds the human director\.Table[1](https://arxiv.org/html/2608.17468#S4.T1)reports the main comparison, from which three findings emerge\.*Expert\-level quality\.*At77\.877\.8overall, SAGE was the only AI system to exceed the human director at77\.177\.1\. It was also the closest system to the director on camera angle and camera movement, the two dimensions on which Vanilla scored lowest\.*Knowledge versus exemplars\.*Both prompting baselines received demonstrations from adjacent episodes and reasoning authored by directors, yet Few\-shot reached only71\.271\.2and CoT76\.076\.0\. SAGE encodes the same knowledge*explicitly*, instead of leaving it latent in exemplars for the model to induce anew\.*Generic skill optimization\.*EvoSkill and SkillOpt both scored below SAGE, with their largest deficits on camera movement, the dimension with the lowest scores overall\. SAGE thus gained most where Vanilla was weakest, and its visual description even surpassed the director\. We attribute this to evolved rules that enforce compositional completeness in lighting, blocking, and framing, which human storyboards often leave implicit\.

### 4\.3\.Ablation Study \(RQ2\)

Table 2\.Ablation on the 18\-episode test set\. Each row adds one component\.Table[2](https://arxiv.org/html/2608.17468#S4.T2)isolates each component\. The contrastive warm\-start in row B provided the largest single gain, since rules extracted by contrasting screenplays with expert storyboards already capture substantial explicit knowledge\. Iteration*without*attribution in row C added little, because episode\-level feedback cannot identify which rules to fix\. Adding attribution in row D more than doubled the iteration benefit, with its largest gains on shot scale and visual description, the two dimensions where row C remained weakest\. This pattern confirms the argument of §[3\.3](https://arxiv.org/html/2608.17468#S3.SS3): rule\-level credit assignment converts iteration from perturbation into optimization\.

### 4\.4\.Iteration Dynamics \(RQ3\)

Figure 4\.Quality on the test set over 10 evolution rounds\. Both settings share the same round\-1 rule set at 74\.2\. With attribution, quality rises to 77\.8 by round 5 and stays within 77\.4 to 77\.7 through round 10; without attribution, it peaks at 75\.4 in round 3 and degrades to 73\.8 by round 10\.A line chart with evolution round from 1 to 10 on the horizontal axis and overall quality score from 73 to 78 on the vertical axis\. Both lines start together at 74\.2, marked by a dotted horizontal reference line\. The solid line for the setting with attribution climbs to a peak of 77\.8 at round 5 and then stays close to that level\. The dashed line for the setting without attribution peaks at 75\.4 at round 3 and then declines to 73\.8, ending below the starting level\.Figure[4](https://arxiv.org/html/2608.17468#S4.F4)tracks quality across 10 rounds from an identical warm\-start set\. With attribution, quality rose over the first five rounds and then held stable through round 10\. Without attribution, it peaked earlier at a lower value and ended*below its starting point*, so attribution changes both the magnitude and the stability of improvement\.

### 4\.5\.Cross\-Model Generalization \(RQ4\)

Figure 5\.Cross\-model transfer of the consolidated knowledge base, evolved entirely with Claude Opus 4\.6\. Injecting the unchanged scenario packages improves every transfer target; labels above the SAGE bars report relative improvements over Vanilla\.A grouped bar chart over three backbone models, namely GPT\-5\.4, GLM\-5\.2, and DeepSeek V4\-Pro\. Each model has a Vanilla bar and a bar for the same model with SAGE rules injected\. The Vanilla scores are 71\.6, 60\.4, and 59\.0\. The scores with SAGE rules are 77\.4, 71\.7, and 71\.5\. Relative improvements of 8\.1 percent, 18\.7 percent, and 21\.2 percent are printed above the second bar of each group\.If the evolved rules encode directorial domain knowledge rather than backbone\-specific tricks, they should transfer to other LLMs unchanged\. We injected the identical scenario packages, evolved entirely with Claude Opus 4\.6, into three backbones and modified nothing else in the pipeline\. As Figure[5](https://arxiv.org/html/2608.17468#S4.F5)shows, every target improved by8\.1%8\.1\\%to21\.2%21\.2\\%relative, and GPT\-5\.4 approached the source model’s own quality\.222GLM\-5\.2 was evaluated on 16 of the 18 episodes owing to two routing failures, with the matching Vanilla episodes excluded for parity\.Two patterns are notable\. First, weaker backbones benefited more, because the rules supply structure that compensates for missing domain knowledge, whereas the strongest backbone gained least from the highest baseline\. Second, visual description transferred most universally, which makes compositional checklists the most portable evolved knowledge\. Absolute scores nonetheless tracked backbone capability, since rules supply domain knowledge but not generation capability\.

### 4\.6\.Validity of the Automatic Scorer

Table 3\.Scorer validity on 18 director storyboards: agreement of Claude Opus 4\.6 with the consensus of three professional directors, vs\. inter\-human agreement, measured by Lin’s CCC\.All reported scores come from an LLM scorer, a paradigm whose reliability and biases are well documented\([52](https://arxiv.org/html/2608.17468#bib.bib32);[25](https://arxiv.org/html/2608.17468#bib.bib33)\)\. We therefore validated it against human judgment\. Three professional directors and the scorer independently scored the 18 director storyboards on the five dimensions, yielding 90 score pairs\. Table[3](https://arxiv.org/html/2608.17468#S4.T3)reports Lin’s concordance correlation coefficient\([16](https://arxiv.org/html/2608.17468#bib.bib51)\)\. The scorer’s agreement with the human consensus*exceeded*inter\-human agreement on four of the five dimensions, the sole exception being shot scale\. Human scores were on average lower by a small margin, a systematic offset that does not affect relative rankings\. Self\-preference bias\([25](https://arxiv.org/html/2608.17468#bib.bib33)\)is also unlikely to favor our method, since all AI systems in Table[1](https://arxiv.org/html/2608.17468#S4.T1)share the scorer’s backbone and any such bias applies uniformly\. We conclude the scorer is a reliable proxy for expert judgment in this domain\.

## 5\.Production Deployment

CreativeFitting is an AI native entertainment company based in Shanghai\. It operates Reel\.AI, among the first AI generated short drama apps distributed to overseas audiences on the App Store and Google Play, and VFS, its in\-house creation platform on which over a thousand creators produce content\. To test whether offline gains translate into production value, we deployed SAGE in VFS, using the same framework trained on a larger proprietary corpus of director demonstrations\. The evaluation covered three ongoing productions disjoint from the public 68\-episode corpus\. Over 14 days, platform logs recorded 12 production users, 1,344 narrative group outputs, and 2,038 generation and revision operations\.

We computed acceptance at the narrative group level, the unit of independent generation\. Following the production team’s operational criterion, an output is accepted if it can enter downstream production without substantive edits, where changes limited to asset references, formatting, punctuation, or wording count as non\-substantive\.

Table 4\.Production acceptance on three ongoing dramas\. Each output corresponds to one narrative group and is a segment that may contain multiple shots\.Prod\. AProd\. BProd\. COverallNarrative group outputs5594533321,344Accepted outputs4603803321,172Acceptance \(%\)82\.383\.9100\.087\.2Table[4](https://arxiv.org/html/2608.17468#S5.T4)shows that 1,172 of 1,344 outputs were accepted without substantive edits, an overall rate of87\.2%87\.2\\%that ranged from82\.3%82\.3\\%to100\.0%100\.0\\%across the three productions\. The production team further reported that typical authoring time per episode fell from over one hour to roughly 10 minutes, an approximately sixfold acceleration\. The acceptance rate is computed from platform interaction logs, whereas the turnaround was tracked by the production team over the same period\.

Two properties of the deployed pipeline keep the residual manual effort bounded\. The unit of acceptance coincides with the unit of generation, so a rejected output calls for a local regeneration of one narrative group rather than a revision pass over the episode\. The system also emits rule\-adoption records at inference time \(§[3\.5](https://arxiv.org/html/2608.17468#S3.SS5)\), so every rejected output stays traceable to the rules that informed it\.

## 6\.Conclusion

Professional storyboarding depends on directorial knowledge that experts cannot exhaustively articulate, so every existing injection path relies on manual externalization\. SAGE removes this dependence by evolving the knowledge component of a skill from the expert demonstrations released in PROSE\. Rule\-adoption records route feedback to individual rules, and the evolved rules are consolidated into scenario packages for deployment without a reference\. The evolved knowledge exceeded the professional directors on our test set, transferred unchanged to three other backbones, and held these gains in a production deployment on live dramas\.

Our iteration study also generalizes beyond storyboarding\. The*granularity*of credit assignment determines whether knowledge evolution converges: rule\-level attribution reached a stable optimum, whereas episode\-level feedback declined\. Systems that treat natural\-language knowledge as learnable parameters therefore need to localize feedback to individual items\. This requirement is architectural rather than domain specific, since any pipeline whose generation step declares the knowledge it consumed can route feedback to that knowledge\.

## References

- Agrawalet al\.\(2026\)L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang, C\. Potts, K\. Sen, A\. G\. Dimakis, I\. Stoica, D\. Klein, M\. Zaharia, and O\. KhattabGEPA: reflective prompt evolution can outperform reinforcement learning\.InThe Fourteenth International Conference on Learning Representations \(ICLR\),Note:Oral\. arXiv:2507\.19457Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Alzubiet al\.\(2026\)S\. Alzubi, N\. Provenzano, J\. Bingham, W\. Chen, and T\. VuEvoSkill: automated skill discovery for multi\-agent systems\.arXiv preprint arXiv:2603\.02766\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1),[§4\.1](https://arxiv.org/html/2608.17468#S4.SS1.SSS0.Px3.p1.1)\.
- Anthropic \(2025\)AnthropicEquipping agents for the real world with agent skills\.Note:[https://www\.anthropic\.com/engineering/equipping\-agents\-for\-the\-real\-world\-with\-agent\-skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills)Engineering blog; Agent Skills open standard\. Accessed 2026\-07\-13Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p3.1),[§2](https://arxiv.org/html/2608.17468#S2.p3.1),[§3\.1](https://arxiv.org/html/2608.17468#S3.SS1.SSS0.Px2.p1.1)\.
- Caoet al\.\(2026\)G\. Cao, T\. He, Y\. Liu, and RAY LCAudience in the loop: viewer feedback\-driven content creation in micro\-drama production on social media\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,External Links:[Document](https://dx.doi.org/10.1145/3772318.3790592)Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p2.1)\.
- Chenet al\.\(2024\)M\. Chen, Y\. Li, Y\. Yang, S\. Yu, B\. Lin, and X\. HeAutoManual: constructing instruction manuals by LLM agents via interactive environmental learning\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Chenget al\.\(2024\)C\. Cheng, A\. Nie, and A\. SwaminathanTrace is the next AutoDiff: generative optimization with rich feedback, execution traces, and LLMs\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2406\.16218Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Dinkevichet al\.\(2025\)D\. Dinkevich, M\. Levy, O\. Avrahami, D\. Samuel, and D\. LischinskiStory2Board: a training\-free approach for expressive storyboard generation\.arXiv preprint arXiv:2508\.09983\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Fanget al\.\(2026\)R\. Fang, Y\. Liang, X\. Wang, J\. Wu, S\. Qiao, P\. Xie, F\. Huang, H\. Chen, and N\. ZhangMemp: exploring agent procedural memory\.InFindings of the Association for Computational Linguistics: ACL 2026,Note:arXiv:2508\.06433Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Fernandoet al\.\(2024\)C\. Fernando, D\. Banarse, H\. Michalewski, S\. Osindero, and T\. RocktäschelPromptbreeder: self\-referential self\-improvement via prompt evolution\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),pp\. 13481–13544\.Note:arXiv:2309\.16797Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Fuet al\.\(2024\)Y\. Fu, D\. Kim, J\. Kim, S\. Sohn, L\. Logeswaran, K\. Bae, and H\. LeeAutoGuide: automated generation and selection of context\-aware guidelines for large language model agents\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p6.1),[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Gaoet al\.\(2026\)H\. Gao, J\. Geng, W\. Hua, M\. Hu, X\. Juan, H\. Liu, S\. Liu, J\. Qiu, X\. Qi, Y\. Wu, H\. Wang,et al\.A survey of self\-evolving agents: what, when, how, and where to evolve on the path to artificial super intelligence\.Transactions on Machine Learning Research\.Note:arXiv:2507\.21046Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Guoet al\.\(2024\)Q\. Guo, R\. Wang, J\. Guo, B\. Li, K\. Song, X\. Tan, G\. Liu, J\. Bian, and Y\. YangConnecting large language models with evolutionary algorithms yields powerful prompt optimizers\.InThe Twelfth International Conference on Learning Representations \(ICLR\),Note:arXiv:2309\.08532Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Huanget al\.\(2024\)J\. Huang, X\. Chen, S\. Mishra, H\. S\. Zheng, A\. W\. Yu, X\. Song, and D\. ZhouLarge language models cannot self\-correct reasoning yet\.InThe Twelfth International Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p5.1),[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Huanget al\.\(2025\)K\. Huang, Y\. Huang, X\. Wang, Z\. Lin, X\. Ning, P\. Wan, D\. Zhang, Y\. Wang, and X\. LiuFilMaster: bridging cinematic principles and generative AI for automated film generation\.arXiv preprint arXiv:2506\.18899\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.18899)Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p1.1),[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Khattabet al\.\(2024\)O\. Khattab, A\. Singhvi, P\. Maheshwari, Z\. Zhang, K\. Santhanam, S\. Vardhamanan, S\. Haq, A\. Sharma, T\. T\. Joshi, H\. Moazam, H\. Miller, M\. Zaharia, and C\. PottsDSPy: compiling declarative language model calls into state\-of\-the\-art pipelines\.InThe Twelfth International Conference on Learning Representations \(ICLR\),Note:arXiv:2310\.03714External Links:[Link](https://openreview.net/forum?id=sY5N0zY5Od)Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Lin \(1989\)L\. I\. LinA concordance correlation coefficient to evaluate reproducibility\.Biometrics45\(1\),pp\. 255–268\.Cited by:[§4\.6](https://arxiv.org/html/2608.17468#S4.SS6.p1.1)\.
- Liuet al\.\(2025\)H\. Liu, J\. He, Y\. Jin, D\. Zheng, Y\. Dong,et al\.ShotBench: expert\-level cinematic understanding in vision\-language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2506\.21356Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p2.1),[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Liuet al\.\(2022\)J\. Liu, D\. Shen, Y\. Zhang, B\. Dolan, L\. Carin, and W\. ChenWhat makes good in\-context examples for GPT\-3?\.InProceedings of Deep Learning Inside Out \(DeeLIO 2022\): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures,pp\. 100–114\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Luet al\.\(2022\)Y\. Lu, M\. Bartolo, A\. Moore, S\. Riedel, and P\. StenetorpFantastically ordered prompts and where to find them: overcoming few\-shot prompt order sensitivity\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 8086–8098\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Gupta, B\. P\. Majumder, K\. Hermann, S\. Welleck, A\. Yazdanbakhsh, and P\. ClarkSelf\-refine: iterative refinement with self\-feedback\.InAdvances in Neural Information Processing Systems 36 \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Majumderet al\.\(2023\)B\. P\. Majumder, B\. Dalvi Mishra, P\. Jansen, O\. Tafjord, N\. Tandon, L\. Zhang, C\. Callison\-Burch, and P\. ClarkCLIN: a continually learning language agent for rapid task adaptation and generalization\.arXiv preprint arXiv:2310\.10134\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- McInneset al\.\(2018\)L\. McInnes, J\. Healy, and J\. MelvilleUMAP: uniform manifold approximation and projection for dimension reduction\.arXiv preprint arXiv:1802\.03426\.Cited by:[§3\.4](https://arxiv.org/html/2608.17468#S3.SS4.SSS0.Px2.p1.1)\.
- Murch \(2001\)W\. MurchIn the blink of an eye: a perspective on film editing\.2nd edition,Silman\-James Press,Los Angeles, CA\.Cited by:[Appendix D](https://arxiv.org/html/2608.17468#A4.p1.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.arXiv preprint arXiv:2310\.08560\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Panicksseryet al\.\(2024\)A\. Panickssery, S\. R\. Bowman, and S\. FengLLM evaluators recognize and favor their own generations\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\),Cited by:[§4\.6](https://arxiv.org/html/2608.17468#S4.SS6.p1.1)\.
- Polanyi \(1966\)M\. PolanyiThe tacit dimension\.Doubleday,Garden City, NY\.Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p2.1)\.
- Pryzantet al\.\(2023\)R\. Pryzant, D\. Iter, J\. Li, Y\. T\. Lee, C\. Zhu, and M\. ZengAutomatic prompt optimization with “gradient descent” and beam search\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 7957–7968\.Note:arXiv:2305\.03495Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Raoet al\.\(2023\)A\. Rao, X\. Jiang, Y\. Guo, L\. Xu, L\. Yang, L\. Jin, D\. Lin, and B\. DaiDynamic storyboard generation in an engine\-based virtual environment for video production\.InACM SIGGRAPH 2023 Posters,New York, NY, USA,pp\. 1–2\.Note:arXiv:2301\.12688External Links:[Document](https://dx.doi.org/10.1145/3588028.3603647)Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Rubinet al\.\(2022\)O\. Rubin, J\. Herzig, and J\. BerantLearning to retrieve prompts for in\-context learning\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics \(NAACL\),pp\. 2655–2671\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems 36 \(NeurIPS\),pp\. 8634–8652\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Tanet al\.\(2025\)W\. Tan, W\. Zhang, X\. Xu, H\. Xia, Z\. Ding, B\. Li, B\. Zhou, J\. Yue, J\. Jiang, Y\. Li, R\. An, M\. Qin, C\. Zong, L\. Zheng, Y\. Wu, X\. Chai, Y\. Bi, T\. Xie, P\. Gu, X\. Li, C\. Zhang, L\. Tian, C\. Wang, X\. Wang, B\. F\. Karlsson, B\. An, S\. Yan, and Z\. LuCradle: empowering foundation agents towards general computer control\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\), PMLR 267,Note:arXiv:2403\.03186Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Tanget al\.\(2024\)J\. Tang, Q\. Jia, Y\. Xie, Z\. Gong, X\. Wen, J\. Zhang, Y\. Guo, G\. Chen, and J\. YangSkyScript\-100m: 1,000,000,000 pairs of scripts and shooting scripts for short drama\.arXiv preprint arXiv:2408\.09333\.Cited by:[3rd item](https://arxiv.org/html/2608.17468#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2608.17468#S1.p2.1),[§2](https://arxiv.org/html/2608.17468#S2.p2.1)\.
- Wanget al\.\(2024\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.Transactions on Machine Learning Research\.Note:arXiv:2305\.16291Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Wanget al\.\(2025a\)X\. Wang, S\. Xu, X\. Shan, Y\. Zhang, M\. Diao, X\. Duan, Y\. Huang, K\. Liang, and Z\. MaCineTechBench: a benchmark for cinematographic technique understanding and generation\.InAdvances in Neural Information Processing Systems 38 \(NeurIPS 2025\),pp\. 60372–60408\.Note:arXiv:2505\.15145External Links:[Document](https://dx.doi.org/10.52202/085713-1810)Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p2.1),[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Wanget al\.\(2025b\)Z\. Z\. Wang, J\. Mao, D\. Fried, and G\. NeubigAgent workflow memory\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\), PMLR 267,Note:arXiv:2409\.07429Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems 35 \(NeurIPS\),pp\. 24824–24837\.Note:arXiv:2201\.11903Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p3.1)\.
- Weiet al\.\(2025\)Z\. Wei, H\. Wu, L\. Zhang, X\. Xu, Y\. Zheng, P\. Hui, M\. Agrawala, H\. Qu, and A\. RaoCineVision: an interactive pre\-visualization storyboard system for director–cinematographer collaboration\.InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology \(UIST\),External Links:[Document](https://dx.doi.org/10.1145/3746059.3747793)Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p2.1)\.
- Wuet al\.\(2025\)W\. Wu, Z\. Zhu, and M\. Z\. ShouAutomated movie generation via multi\-agent CoT planning\.arXiv preprint arXiv:2503\.07314\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Xieet al\.\(2024\)J\. Xie, J\. Feng, Z\. Tian, K\. Q\. Lin, Y\. Huang,et al\.Learning long\-form video prior via generative pre\-training\.arXiv preprint arXiv:2404\.15909\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Xuet al\.\(2025\)Z\. Xu, L\. Wang, J\. Wang, Z\. Li, S\. Shi,et al\.FilmAgent: a multi\-agent framework for end\-to\-end film automation in virtual 3d spaces\.arXiv preprint arXiv:2501\.12909\.Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p1.1),[§2](https://arxiv.org/html/2608.17468#S2.p1.1)\.
- Yanget al\.\(2024\)C\. Yang, X\. Wang, Y\. Lu, H\. Liu, Q\. V\. Le, D\. Zhou, and X\. ChenLarge language models as optimizers\.InThe Twelfth International Conference on Learning Representations \(ICLR\),Note:arXiv:2309\.03409Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Yanget al\.\(2026\)Y\. Yang, Z\. Gong, W\. Huang, Q\. Yang, Z\. Zhou, Z\. Huang, Y\. Li, X\. Gao, Q\. Dai, B\. Liu, K\. Qiu, Y\. Yang, D\. Chen, X\. Yang, and C\. LuoSkillOpt: executive strategy for self\-evolving agent skills\.arXiv preprint arXiv:2605\.23904\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1),[§4\.1](https://arxiv.org/html/2608.17468#S4.SS1.SSS0.Px3.p1.1)\.
- Yanget al\.\(2023\)Z\. Yang, P\. Li, and Y\. LiuFailures pave the way: enhancing large language models through tuning\-free rule accumulation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 1751–1777\.Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p5.1),[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Yuksekgonulet al\.\(2025\)M\. Yuksekgonul, F\. Bianchi, J\. Boen, S\. Liu, P\. Lu, Z\. Huang, C\. Guestrin, and J\. ZouOptimizing generative AI by backpropagating language model feedback\.Nature639,pp\. 609–616\.Note:Framework known as TextGrad; preprint: arXiv:2406\.07496Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Zhaiet al\.\(2025\)Y\. Zhai, S\. Tao, C\. Chen,et al\.AgentEvolver: towards efficient self\-evolving agent system\.arXiv preprint arXiv:2511\.10395\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Zhanget al\.\(2024a\)S\. Zhang, J\. Zhang, J\. Liu, L\. Song, C\. Wang, R\. Krishna, and Q\. WuOffline training of language model agents with functions as learnable weights\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),pp\. 60315–60335\.Note:arXiv:2402\.11359Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Zhanget al\.\(2024b\)W\. Zhang, K\. Tang, H\. Wu, M\. Wang, Y\. Shen, G\. Hou, Z\. Tan, P\. Li, Y\. Zhuang, and W\. LuAgent\-pro: learning to evolve via policy\-level reflection and optimization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 5348–5375\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Zhanget al\.\(2025\)Z\. Zhang, Q\. Dai, X\. Bo, C\. Ma, R\. Li, X\. Chen, J\. Zhu, Z\. Dong, and J\. WenA survey on the memory mechanism of large language model\-based agents\.ACM Transactions on Information Systems43\(6\)\.Note:arXiv:2404\.13501External Links:[Document](https://dx.doi.org/10.1145/3748302)Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Zhaoet al\.\(2024\)A\. Zhao, D\. Huang, Q\. Xu, M\. Lin, Y\. Liu, and G\. HuangExpeL: LLM agents are experiential learners\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 19632–19642\.Cited by:[§1](https://arxiv.org/html/2608.17468#S1.p5.1),[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Zhaoet al\.\(2021\)Z\. Zhao, E\. Wallace, S\. Feng, D\. Klein, and S\. SinghCalibrate before use: improving few\-shot performance of language models\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),pp\. 12697–12706\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.
- Zhenget al\.\(2025\)B\. Zheng, M\. Y\. Fatemi, X\. Jin, Z\. Z\. Wang, A\. Gandhi, Y\. Song, Y\. Gu, J\. Srinivasa, G\. Liu, G\. Neubig, and Y\. SuSkillWeaver: web agents can self\-improve by discovering and honing skills\.arXiv preprint arXiv:2504\.07079\.Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p3.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems 36 \(NeurIPS 2023\), Datasets and Benchmarks Track,Cited by:[§4\.6](https://arxiv.org/html/2608.17468#S4.SS6.p1.1)\.
- Zhouet al\.\(2026\)H\. Zhou, S\. Liu, J\. Chen, X\. Zou, L\. Xia, and L\. NieDramaDirector: geometry\-guided short drama generation\.arXiv preprint arXiv:2606\.24107\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.24107)Cited by:[3rd item](https://arxiv.org/html/2608.17468#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2608.17468#S1.p1.1),[§2](https://arxiv.org/html/2608.17468#S2.p1.1),[§2](https://arxiv.org/html/2608.17468#S2.p2.1)\.
- Zhouet al\.\(2023\)Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. BaLarge language models are human\-level prompt engineers\.InThe Eleventh International Conference on Learning Representations \(ICLR\),Note:arXiv:2211\.01910Cited by:[§2](https://arxiv.org/html/2608.17468#S2.p4.1)\.

## Appendix ARule Consolidation Details

Figure[6](https://arxiv.org/html/2608.17468#A1.F6)shows that rules learned across episodes and series form coherent, well\-separated clusters \(a\)\. Each cluster becomes a self\-contained scenario package, routable by its natural\-language description and organized by professional dimension \(b\)\. This is how evolved knowledge is stored and retrieved at inference time \(§[3\.4](https://arxiv.org/html/2608.17468#S3.SS4)\)\.

Two properties are worth noting\. First, clusters mix rules from all three series rather than separating by source, indicating that the learned conditions describe narrative situations rather than one drama’s idiosyncrasies\. The example package contains 38 rules contributed by all three series, and the recovered scenario carries no trace of its source episodes\. Second, the distribution across dimensions is intentionally uneven: camera angle dominates this power\-asymmetry package, whereas an emotional\-release package concentrates on shot scale and rhythm\. Consolidation therefore preserves each situation’s dimensional signature rather than balancing dimensions artificially\.

![Refer to caption](https://arxiv.org/html/2608.17468v1/fig_consolidation_scatter.png)\(a\)
\(b\)
Figure 6\.Rule consolidation\. \(a\) UMAP projection of rule\-condition embeddings, colored by cluster\. \(b\) A representative scenario package shown in translation: a routable natural\-language description over rules grouped by professional dimension\. The package contains 38 rules in total; representative rules are shown for space\.Panel \(a\) is a two\-dimensional UMAP scatter plot of rule condition embeddings\. Points form many small compact groups, each drawn in a distinct color, separated by empty space\. Panel \(b\) shows one scenario package as a document\. A header names the package and gives its applicability description, the rule count of 38, the three contributing series, and the rule counts per dimension\. Three framed entries follow, one each for camera angle, shot scale, and shot rhythm, and each entry states a condition and a practice\. A footer notes that 35 further rules are omitted\.
## Appendix BScenario\-Aware Inference Example

Figure[7](https://arxiv.org/html/2608.17468#A2.F7)traces one real narrative group from a test episode through the scenario\-aware inference pipeline of §[3\.5](https://arxiv.org/html/2608.17468#S3.SS5)\. A two\-person confrontation is matched against the routing index and dispatched to its top\-3 scenario packages, whose rules are injected into generation; the produced shots then carry rule\-adoption records back to the packages that informed them\. No expert reference is involved\.

The three retrieved packages are complementary rather than redundant\. One supplies the angle vocabulary for power asymmetry, another governs the rhythm of a verbal exchange, and the third covers the framing of a physical intrusion\. Their rules therefore act on different shots of the same segment, and the adoption records make this division visible after the fact, which is what allows a rejected output to be traced to the rule that shaped it\.

![A left to right flow in three columns. The left column describes one narrative group, a two-person confrontation, with its dialogue lines. The middle column lists the three scenario packages it was routed to, each with a identifier, a name, a condition, and a practice. The right column shows the generated shots, each labeled with its shot scale, camera angle, and camera movement. Colored connecting lines trace which package informed which shot.](https://arxiv.org/html/2608.17468v1/figures/fig_routing.png)Figure 7\.A real routing example from*Beyond the Wall*EP001, a held\-out test episode, translated\. Each shot is marked with the*primary*package whose rules it adopted, and one representative rule per package is shown\.A left to right flow in three columns\. The left column describes one narrative group, a two\-person confrontation, with its dialogue lines\. The middle column lists the three scenario packages it was routed to, each with a identifier, a name, a condition, and a practice\. The right column shows the generated shots, each labeled with its shot scale, camera angle, and camera movement\. Colored connecting lines trace which package informed which shot\.
## Appendix CAlignment Review Skill

The alignment scorefalignf\_\{\\mathrm\{align\}\}of §[3\.1](https://arxiv.org/html/2608.17468#S3.SS1)is produced by a review skill that compares a generated storyboard against the director reference\. Figure[8](https://arxiv.org/html/2608.17468#A3.F8)reproduces that skill file, translated and condensed to fit the column\. The reviewer receives both artifacts, whereas the generator never sees the reference\. Comparison proceeds over the five professional dimensions, and the director storyboard is treated as the sole correct target on every one of them\.

Two design choices make the resulting signal usable for rule level diagnosis\. First, the skill discards all surface variation\. Table layout, column order, and wording are excluded from the comparison, so a deviation reflects a directorial decision rather than a formatting artifact\. Second, the report records only deviations\. Praise and hedging are suppressed, and each deviation must cite the shot numbers at which the two storyboards diverge\. A deviation without such a citation cannot be attributed and is therefore rejected\. Phase 4 of evolution consumes these localized deviations and joins them with the rule adoption records, as described in §[3\.3](https://arxiv.org/html/2608.17468#S3.SS3)\.

\# Role
You are a professional storyboard review expert\. Compare each AI storyboard against the director’s original storyboard and score its degree of alignment\.
\# Review dimensions
Follow the five dimensions below\.Ignore table format, column order, and wording entirely\.
1\.Rhythm\-\-\- shot splitting density, timing of inserted reaction shots, emotional breathing room\.
2\.Visual description\-\-\- purely visual content, physical action, micro expression and physiological reaction such as a swallow or a tremble\. Ignore all interiority and literary embellishment\.
3\.Shot scale\-\-\- scale selection, including the director’s preference for close\-up and extreme close\-up on body detail\.
4\.Camera angle\-\-\- subjective and objective viewpoint shifts, over the shoulder framing, Dutch angle, and the power dynamics built by low and high angles\.
5\.Camera movement\-\-\- visual stability of the move, such as a locked off frame or a slow push in\.
\# Scoring
Each dimension uses a 100\-point scale; the overall score is their mean\. A higher score means closer alignment with the director\.The director’s manual storyboard is the only correct alignment target \(gold standard\)\.
\# Workflow
1\. Ask the user for the paths of the AI storyboards and of the director benchmark file before starting\.
2\. Write a Python script to read the CSV files and analyse the differences\.
3\. Never print a whole table, which overflows the context\. Slice the data or search for action keywords to locate the dramatic peaks, then compare reaction close\-ups and oppressive compositions there\.
4\. Emit the report in the two prescribed parts\.
\# Schema pitfalls
The two sources carry different headers, so resolve each field through a cascade of fallbacks\. The AI output merges scale, viewpoint, and composition into one column, whereas the director file splits scale and angle apart\. Read every file as utf\-8\-sig so that a byte order mark cannot corrupt the first header\. Guard against generated files that hold a header but no rows\.
\# Output
Part 1\.A summary table: one row per plan, with the episode, the overall score, the five dimension scores, and a one\-line justification\.
Part 2\.A deviation analysis, under strict discipline:
∙\\bulletPain points only\. Never state an advantage of the AI plan, a weakness of the director, or any word of praise\.
∙\\bulletCite shot numbers\.For example: in Shot 15 the director uses a Dutch low angle with an over\-the\-shoulder framing to convey Victor’s pressure, whereas plan A stays at eye level and loses the spatial hierarchy entirely\.
∙\\bulletReport explicitly whether the plan misses the director’s preferred visual grammar, the handheld breathing quality, or the listener’s reaction shot\.Alignment Review Skill\(condensed from the skill file used forfalignf\_\{\\mathrm\{align\}\}\)Figure 8\.The skill file behindfalignf\_\{\\mathrm\{align\}\}, translated and condensed to fit the column\. Section headings and the wording of every directive follow the original\. The director storyboard is the gold standard and stays hidden from the generator\.A framed transcript of the alignment review skill file, set in a monospaced font under a colored title bar\. Sections appear in order\. A role section casts the reader as a storyboard review expert\. A review dimensions section enumerates the five professional dimensions and instructs the reviewer to ignore table format, column order, and wording\. Further sections state the scoring procedure and the report format\.
## Appendix DQuality Scoring Rubric

The quality metricfqualf\_\{\\mathrm\{qual\}\}of §[4\.1](https://arxiv.org/html/2608.17468#S4.SS1)scores a storyboard on absolute professional standards without any reference\. Figure[9](https://arxiv.org/html/2608.17468#A4.F9)reproduces the scoring prompt, translated and condensed to fit the column\. Its grounding is the editing priority of Walter Murch, which ranks emotion and story above rhythm and continuity\([23](https://arxiv.org/html/2608.17468#bib.bib54)\)\. Every dimension is read in the context of vertical short drama, where narrative density is high and the opening seconds govern retention\.

Each dimension carries a 100 point scale divided into five bands of twenty points, and the overall score is their unweighted mean\. Every band is anchored by the observable evidence expected at that level rather than by an adjective, which keeps runs comparable\. The bands share one structure across the five dimensions, so a score reads the same wherever it appears\. Three protocol rules govern the output\. Every credit or deduction cites specific shot numbers\. Judgment follows short drama practice instead of feature film convention\. Section[4\.6](https://arxiv.org/html/2608.17468#S4.SS6)validates the rubric against three professional directors\.

\# Role
You are a senior storyboard director and cinematographer with twenty years of experience, fluent in vertical short drama production\. Score the given storyboard on five dimensions, 100 points each,with no reference storyboard available\.
Ground every judgment in:Murch’s six criteria for editing\-\-\- emotion 51%, story 23%, rhythm 10%, eye trace 7%, planarity 5%, spatial continuity 4%; thefunction of a storyboardas the route map from script to screen, where every frame carries a narrative purpose; andshort drama traits\-\-\- dense narration, mobile first, the first seconds deciding retention, a twist every 30 to 60 seconds\.
\# Score anchors\(five bands of twenty points, the same structure on every dimension; the rhythm dimension is shown\)
81–100complete rhythmic arc of setup, escalation, climax, and breath; density gradient precisely matched to tension; ASL near 1\.5 to 2s at the climax and 3 to 5s in dialogue; deliberate breathing shots; the first three shots form an effective hook\.
61–80rhythm varies plausibly and the gradient is broadly right; breathing shots exist but sit imprecisely\.
41–60basic fast and slow variation, yet mechanical; breathing shots absent or misplaced; the opening lacks a hook\.
21–40scattered variation with no density logic; action and dialogue are barely distinguished\.
0–20monotonous throughout; no breathing shot; the climax is indistinguishable from the setup, and shots are merely enumerated\.
\# Core criteria per dimension\(abridged\)
Rhythm— density gradient, breathing design, hook structure, arc within a scene, motivation for each cut\.
Visual description— executable composition, light direction and quality, spatial layering, precision of character description, environmental storytelling, purely visual content only\.
Shot scale— spectrum coverage from ELS to ECU, emotional weight matching, progression logic, the 30\-degree rule, restraint on ECU\.
Camera angle— narrative motivation, power dynamics through low and high angles, the 180\-degree axis, POV placement, restraint on Dutch angle\.
Camera movement— trigger, path, and stop point for every move; static against moving contrast; semantics of direction; type variety\.
Vertical rules override film convention:9:16 framing, a naturally high share of CU and MCU, and push\-in to CU as the signature reveal\.
\# Output format
Report a score table over the five dimensions with a grade each and the overall mean\. Grades follow the bands: excellent, good, fair, pass, fail\.
Then, per dimension, give the reasoning with shot numbers as evidence, the strongest case, the main defect, and a concrete suggestion targeting it\. Close with a verdict naming the single most important strength and weakness\.
\# Cautions
Judge by the standard of a working short drama director rather than relaxing it because the plan already beats most AI output\.Every deduction must cite shot numbers\.Stay in the short drama context and never import feature film standards\. Executability comes first: the bottom line is whether a crew could shoot the plan without further discussion\. A few standout shots do not offset a systemic defect, since consistency across the episode matters more than a local highlight\.The five dimensions must remain discriminative, so do not assign one score to all of them\.Quality Scoring Rubric\(condensed from the prompt used forfqualf\_\{\\mathrm\{qual\}\}\)Figure 9\.The scoring prompt behindfqualf\_\{\\mathrm\{qual\}\}, translated and condensed to fit the column\. The five bands are shown for the rhythm dimension; the other four share the same structure and are abridged to their core criteria\.A framed transcript of the quality scoring prompt, set in a monospaced font under a colored title bar\. A role section casts the reader as a senior storyboard director and states that no reference storyboard is available\. A grounding paragraph lists Murch's six criteria for editing with their weights, the function of a storyboard, and the traits of short drama\. A score anchors section then gives five bands of twenty points for the rhythm dimension, each band stating the observable evidence expected at that level\.

Similar Articles

SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration

arXiv cs.CL

Introduces SPO, a stochastic search framework for automatic prompt optimization, with three strategies including SAGE, an agent-guided multi-agent pipeline. Evaluated on benchmarks and deployed on a mental-health chatbot, showing improvements in retention through continuous optimization.