MaLiang-Harness: A Programmable Path to Image and Video Generation

Hugging Face Daily Papers Papers

Summary

MaLiang-Harness introduces a unified framework to address the Program-to-Visual gap in executable programs for image and video generation, evaluating multiple MLLMs on benchmarks and revealing insights into visual generation performance.

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:21 AM

Paper page - MaLiang-Harness: A Programmable Path to Image and Video Generation

Source: https://huggingface.co/papers/2609.34309

Abstract

Executableprogramsofferexplicitcontroloverhowimagesandvideosareconstructed,butgeneratingrunnablecodeisonlythebeginningofvisualcreation.Aprogramcanexecutecorrectlywhileviolatingtherequestedcomposition,appearance,ormotion.WedefinethisdiscrepancyastheProgram-to-Visual(P2V)gapandintroduceMaLiang-Harness,aunifiedframeworkfororganizingMLLM-drivenvisualgenerationintoapersistentprocessofconstruction,inspection,andrevision.Itscentraldesignistomaketheevolvingvisualprogram,itsconstructionhistory,anditsverificationshareacommonrevisionreference.WedefinethePersistentExecutableGeneration(PEG)stateaspreservingprogramsandtaskcontext.TraceableGenerationProcess(TGP)connectseditstorenderedevidence,andRevision-awareEditingandVerification(REV)supportsrestorationandchecksthecurrentrevisionbeforecompletion.Together,thesemechanismscoordinateplanning,execution,andvisualfeedbackacrossrenderingbackends.Weevaluate11powerfulclosed-sourceMLLMsonMaLiang-IBenchandfouronMaLiang-VBench,measuringgenerationsuccess,visualquality,andcomputationalcost.GPT-6-Astraachieves100%generationsuccessonbothbenchmarks,with96.0%ofimagetasksand76.9%ofvideotasksmeetingallqualitythresholds.Thecomparisonalsorevealsamismatchbetweengeneralcapabilityscoresandvisualgenerationperformance,withsimilarlyscoredmodelsdifferingsubstantiallyintheirabilitytosatisfyvisualrequirements.MaLiang-HarnessprovidesasystematicbasisforstudyinghowMLLMstranslateexecutablecodeintovisualoutcomes,exposingboththepotentialofprogrammablegenerationandthelimitationsofgeneralbenchmarksaspredictorsofthisability.Theprojectisavailableathttps://github.com/gulucaptain/MaLiang-Harness.

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.34309

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34309 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34309 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34309 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

GUI harness [video]

Reddit r/LocalLLaMA

A research project presenting a GUI harness that allows users or LLMs to build complex applications using simple vector graphics functions, featuring integrated code execution, history management, and support for multiple open-weight LLMs.

EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses

arXiv cs.LG

This paper introduces EvoGen-Harness, a generator-agnostic framework that evolves multiple external responsibilities around frozen text-to-image models using a Trace method for failure attribution and coordinated search, achieving significant benchmark gains over existing single-dimension adaptation baselines.

Show-Harness: Just a VLM Agent Can Play Robots

Hugging Face Daily Papers

Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.

Self-Harness: Harnesses That Improve Themselves

Hacker News Top

Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.