When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation
Summary
The paper presents a dual gatekeeping system for AI-generated educational videos that combines educator input and automated metrics to enhance pedagogical quality, demonstrating that principled resistance to AI outputs improves instructional design.
View Cached Full Text
Cached at: 08/21/26, 10:05 AM
# When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation Source: [https://arxiv.org/html/2608.19812](https://arxiv.org/html/2608.19812) ## When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content CreationConference:Understanding and Engaging Critical Resistance to AI in Education; April 13–17, 2026; Barcelona, Spain1210CCS:Human\-centered computing User studiesCCS:Applied computing Interactive learning environments Yearim KimNote:Both authors contributed equally to this research\.email:[yerim1656@snu\.ac\.kr](mailto:[email protected])Affiliation:Seoul National University,Seoul,Republic of KoreaInjun Baekemail:[jjune1416@snu\.ac\.kr](mailto:[email protected])Affiliation:Seoul National University,,Samsung Electronics,Republic of KoreaandNojun KwakNote:Corresponding authorAffiliation:Seoul National University,Seoul,Republic of Koreaemail:[nojunk@snu\.ac\.kr](mailto:[email protected]) 2026© , 2026; ###### Abstract\. To prevent the adoption of aesthetically polished but pedagogically flawed AI content, westudya video authoring pipeline featuring two layers ofstructured refusal\. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative\-visual synchronization\. While neither layer is exhaustive, their synergy ensures that principled resistance–the act of deferring AI output until it meets rigorous standards–becomes a catalyst for higher quality\.Evaluation combining a study with 23 educators across 3 topics and automated metrics across 7 topicsdrawn from established science and philosophy curricula shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners\. ###### Keywords: Human\-AI Collaboration, Educational AI, Generative AI, Multimedia Learning ## 1\.Introduction While modern AI models\([2](https://arxiv.org/html/2608.19812#bib.bib1);[1](https://arxiv.org/html/2608.19812#bib.bib3)\)synthesizes professional\-looking educational videos in a minute, surface\-level polish does not guarantee pedagogical rigor\. Current video generation pipelines often prioritize visual appeal over instructional essentials, such as precise temporal alignment of narration or strategic sequencing of prerequisite concepts\. Consequently, outputs are frequently optimized forlooking good, rather thanteaching well\. This gap matters as educators are increasingly expected to adopt AI\-generated materials with minimal intervention\. The pressure toward seamless, friction\-free adoption treats any slowdown as inefficiency—a stance that risks reducing the educator’s role from professional decision\-maker to passive consumer\. We argue, however, that pedagogical friction is not a hurdle to be eliminated but a site of professional accountability\. Moments of deliberate hesitation–whether an educator questioning a script’s logical flow or an algorithm flagging a narrative truncation–are precisely where instructional quality is forged\. Table 1\.Mayer’s 12 CTML principles\([3](https://arxiv.org/html/2608.19812#bib.bib2)\)for reducing extraneous, managing essential, and fostering generative processing\.CoherenceSignalingRedundancySpatial ContiguityTemporal ContiguitySegmentingPre\-trainingModalityMultimediaPersonalizationVoiceImageWe conceptualize this approach asprincipled resistance: deliberate, theory\-grounded pushback against AI outputs that fail pedagogical standards\. Rooted in Mayer’s Cognitive Theory of Multimedia Learning \(CTML\)\([3](https://arxiv.org/html/2608.19812#bib.bib2)\)–a framework of 12 empirically validated principles for effective multimedia instruction \(Table[1](https://arxiv.org/html/2608.19812#S1.T1)\)–we developed the PedaCo \(Pedagogical Co\-creation\), a human\-AI collaborative system that operationalizes this resistance\. PedaCo integrates two complementary gatekeeping mechanisms: reviewing scripts through CTML\-informed criteria, and an automated metric evaluating finished videos against computationally measurable CTML dimensions\. This paper presents the design rationale behind this dual approach, summarizes converging evidence from both human and computational evaluations, and poses open questions about how educational AI systems should balance human agency with automated safeguards\. ## 2\.Two Layers of Resistance In our framework, principled resistance takes three concrete forms:rejecting\(requesting regeneration\),revising\(manual editing\), andoverriding\(vetoing automated flags\)\. These are not ad hoc reactions but norm\-driven decisions grounded in CTML principles\. The framework rests on a simple observation: educators and algorithms are good at catching different kinds of problems:while human educators excel at identifying nuanced pedagogical mismatches–such as content being too advanced for a target audience–algorithms provide precise, high\-resolution verification of structural integrity, such as identifying temporal misalignments between narration and visuals\.Building both checkpoints into the same pipeline creates overlapping coverage that neither could achieve alone\. ### 2\.1\.Layer 1: Human Educator Review at the Script Stage The first layer intervenes before any video is rendered \(Figure[1](https://arxiv.org/html/2608.19812#S2.F1)left\)\.The educator begins by inputting learning content and configuring which CTML principles the system should enforce\.An LLM\([4](https://arxiv.org/html/2608.19812#bib.bib4)\)then generates an initial script, which passes through a structured review cycle\. An AI reviewer—itself prompted with CTML principles—produces feedback organized by principle, identifying potential violations rather than definitive judgments \(e\.g\., “Scene 3 introduces technical terms without prior explanation, which may conflict with the Pre\-training principle”\)\. The human educator then decides what to accept, what to revise manually, and what to regenerate\.This review loop can be repeated until the educator is satisfied with the script\. This design choice to review at the script level is deliberate\. Textual revisions are computationally and laboriously efficient, whereas pedagogical errors baked into a rendered visual narrative are nearly impossible to correct post\-synthesis\.By placing the human checkpoint at this intermediate stage, we make pedagogical critique economically viable\.This ensures the system remains advisory, preserving the educator’s professional authority to say ‘no’ to AI suggestions based on their specific curricular context and pedagogical style\. ### 2\.2\.Layer 2: Automated Metric After Video Synthesis The second layer performs a post\-synthesis evaluation of the finalized video through a composite metric \(Figure[1](https://arxiv.org/html/2608.19812#S2.F1)right\)\. We automate the assessment of five dimensions:coherence,redundancy,temporal contiguity,modality, andimage quality\.The educator reviews the principle\-level scores and decides whether to accept the video or return to the script stage for targeted revision\.This boundary between human and automated resistance is a strategic design decision\. While dimensions like temporal synchronization are amenable to reliable computational measurement, others—such as Personalization \(tone\)—require a deep understanding of curricular structures and learner psychology that currently remains the sole domain of the human expert\. By automating only where algorithmic feasibility aligns with pedagogical necessity, we create a robust safety net that prevents technical regressions without marginalizing human judgment\. Figure 1\.The Dual Gatekeeping Interface\.The left panel\(Layer 1: Script Level\)generates an initial script \(dd\) from learning content \(aa\) and generation principles \(bb\), then scaffolds the educator’s revision by providing AI critiques and revised drafts \(ee\) based on review constraints \(cc\)\. The right panel\(Layer 2: Video Level\)visualizes invisible pedagogical quality via automated metrics \(h,ih,i\), allowing users to assess the alignment between the final video \(gg\) and the learning content \(ff\)\. ## 3\.Evidence from Two Evaluations We conducted a multi\-method evaluation to assess the efficacy of the PedaCo framework, combining a human\-centric study with educators \(3\.1\) and an algorithmic assessment via automated metrics \(3\.2\)\. Together, these evaluations provide converging evidence for the value of dual\-layered resistance\. ### 3\.1\.What Educators Found We conducted a within\-subject study with 23 educators—who were priorly briefed on the CTML principles—directly used our system to generate and evaluate videos\. Guided by the system, participants simulated with the pre\-generated videos by inputting raw learning content and iteratively refining the AI\-generated scripts with the system’s CTML feedback\. Then, they compared these videos outcomes with baseline videos generated without CTML guidelines\. The topics include three different cognitive demands: causal reasoning, abstract concepts, and procedural knowledge\. Participants rated each condition on 13 items covering all 12 CTML principles and overall instructional validity\. The review\-based approach yielded statistically significant improvements across every principle \(p<\.05p<\.05, Wilcoxon signed\-rank\)\.The mean rating rose from 3\.07 to 3\.86 on a 5\-point scale \(\+0\.79,p<\.01p<\.01\),with the most pronounced gains observed in content organization: prerequisite sequencing \(\+0\.86\+0\.86\), irrelevant material removal \(\+0\.84\+0\.84\), and overall instructional validity \(\+0\.96\+0\.96\)ove\.Nearlyall effects were large \(r≥\.64r\\geq\.64, computed asZ/NZ/\\sqrt\{N\}\), with one exception \(redundancy,r=\.42r=\.42, medium\),and consistent across gender and experience level\. Notably, educators did not perceive the review process as slowing them down\. They rated production efficiency at 4\.26/5 and the validity of the CTML\-based guidance at 4\.04/5 \(with remarkably low variance,SD=0\.62SD=0\.62\)\. One participant captured the tension well: the iterative process was“quite challenging”butultimatelyfor producing“a robust and effective learning tool”\(P23\)\. Another explicitly requested“separate functions where teachers can additionally review, edit, and modify”the AI output \(P05\)—in other words, more resistance, not less\. ### 3\.2\.What the Metrics Found Independently, we applied our automated metrics to a corpus of 14 videos \(7 topics×\\times2 conditions\) drawn from established science and philosophy curricula\.The videos were generated with identical structure, isolating the CTML\-informed generation and review as the sole variable\.Two of the five metrics showed significant improvement: temporal contiguity \(0\.294 vs\. 0\.273,p=\.021p=\.021\) and coherence \(0\.729 vs\. 0\.646,p=\.011p=\.011\)\. The remaining three showed no significant difference—modality and redundancy scored high in both conditions \(near ceiling\), and image quality did not differ as both conditions used the same video synthesis model\. We retain these metrics as safety nets: current models perform well on these dimensions, but they serve as guardrails against future model regressions or hallucination\-induced failures\. ### 3\.3\.Where the Two Evaluations Agree The most compelling evidence for our dual\-layered approach is the high degree of convergence between subjective ratings and objective metrics\. Despite being conducted independently with different samples and instruments, both evaluations identifiedcoherenceandtemporal alignmentas the dimensions most significantly enhanced by the PedaCo pipeline\. Educators ranked coherence among the top three improvements, matching the statistically significant gains identified by the automated system\. This triangulation suggests that the two gatekeeping layers are not merely redundant but provide complementary verification of the same underlying instructional quality\. ## 4\.Discussion: Reframing Resistance in ducational AI The PedaCo framework offers a theory\-grounded perspective on where the boundaries of Generative AI should be drawn in educational settings\. By operationalizing "principled resistance" through CTML, we illustrate how intentional friction can sustainpedagogical integritywithout marginalizing the professional authority of educators\. Our findings suggest that "productive friction" is most effective when it is: \(a\) theoretically grounded rather than intuitive; \(b\) strategically embedded upstream at the script stage; and \(c\) hybridized across human and computational agents\. However, this dual\-gatekeeping approach surfaces three emergent tensions for the workshop to consider: - •Negotiating Agency: When automated flags and educator judgments diverge, how should the interface balance algorithmic safeguards with human autonomy? - •Sustainability of Friction: While valued in our short\-term study, the long\-term viability of iterative review in daily classroom preparation remains unknown\. We must identify the threshold where productive friction transitions into "friction fatigue\." - •Beyond Proxy Metrics: Future research must move beyond theoretical proxies to evaluate the direct causal impact of structured resistance on student learning outcomes\. ## 5\.Conclusion In this paper, we have argued that educational AI needs structured ways to say “not yet”—bridging human pedagogical expertise with computational precision\.Our PedaCo framework builds principled resistance into the video generation pipeline through educator review at the script stage and automated pedagogical metrics at the video stage\. Early evidence suggests both layers improve the same instructional dimensions, and that educators experience this friction as productive rather than burdensome\. We offer this as one concrete answer to the workshop’s animating question: resistance to AI in educationshould not be equated with rejection\.Instead,it can mean building systems that are designed to push back, on principled grounds, until the output is genuinely ready to teach\. ###### Acknowledgements\. This work was funded by the Korean Government through the grants from IITP \(RS\-2021\-II211343\) and KOCCA \(RS\-2024\-00398320\)\. ## References - Brookset al\.\(2024\)T\. Brooks, B\. Peebles, C\. Holmes, W\. DePue, Y\. Guo, L\. Jing, D\. Schnurr, J\. Taylor, T\. Luhman, E\. Luhman, C\. Ng, R\. Wang, and A\. RameshVideo generation models as world simulators\.External Links:[Link](https://openai.com/research/video-generation-models-as-world-simulators)Cited by:[§1](https://arxiv.org/html/2608.19812#S1.p1.1)\. - DeepMind \(2024\)G\. DeepMindVeo: google’s most capable generative video model\.Note:[https://deepmind\.google/models/veo/](https://deepmind.google/models/veo/)Accessed: 2025\-09\-29Cited by:[§1](https://arxiv.org/html/2608.19812#S1.p1.1)\. - Mayer \(2009\)R\. E\. MayerMultimedia principle\.InMultimedia Learning,pp\. 223–241\.Cited by:[Table 1](https://arxiv.org/html/2608.19812#S1.T1),[Table 1](https://arxiv.org/html/2608.19812#S1.T1.4),[§1](https://arxiv.org/html/2608.19812#S1.p2.1)\. - Teamet al\.\(2023\)G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican,et al\.Gemini: a family of highly capable multimodal models\.arXiv preprint arXiv:2312\.11805\.Cited by:[§2\.1](https://arxiv.org/html/2608.19812#S2.SS1.p1.1)\.
Similar Articles
Should AI ever be fully autonomous for content aimed at children? We chose not to, and it's slowed us down a lot.
The author discusses the ethical decision to avoid full AI autonomy in generating content for children, opting for mandatory human review to ensure safety and trustworthiness.
I found a way to fight AI slop
The author proposes using AI as a research moderation and synthesis tool rather than a content generator to combat 'AI slop.' By building a pipeline that compares and ranks expert sources, the author argues the future human role in AI is curation and judgment.
Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks
This design-based study investigates how pedagogical scaffolding can help ethnic minority preparatory students shift from passive consumption to critical co-creation with Generative AI in prompt engineering tasks, resulting in improved prompt self-efficacy and active gatekeeping of AI-generated content.
Attention AI Slop Porters
A commentary or critique addressing the issue of low-quality AI-generated content being produced and distributed.
Why We Build
An opinion piece advocating for AI systems that deliver transparent, verifiable knowledge from domain experts, enabling discovery-based learning and countering centralized propaganda.