MaLiang-Harness: A Programmable Path to Image and Video Generation
Summary
MaLiang-Harness introduces a unified framework to address the Program-to-Visual gap in executable programs for image and video generation, evaluating multiple MLLMs on benchmarks and revealing insights into visual generation performance.
View Cached Full Text
Cached at: 09/30/26, 04:21 AM
Paper page - MaLiang-Harness: A Programmable Path to Image and Video Generation
Source: https://huggingface.co/papers/2609.34309
Abstract
Executableprogramsofferexplicitcontroloverhowimagesandvideosareconstructed,butgeneratingrunnablecodeisonlythebeginningofvisualcreation.Aprogramcanexecutecorrectlywhileviolatingtherequestedcomposition,appearance,ormotion.WedefinethisdiscrepancyastheProgram-to-Visual(P2V)gapandintroduceMaLiang-Harness,aunifiedframeworkfororganizingMLLM-drivenvisualgenerationintoapersistentprocessofconstruction,inspection,andrevision.Itscentraldesignistomaketheevolvingvisualprogram,itsconstructionhistory,anditsverificationshareacommonrevisionreference.WedefinethePersistentExecutableGeneration(PEG)stateaspreservingprogramsandtaskcontext.TraceableGenerationProcess(TGP)connectseditstorenderedevidence,andRevision-awareEditingandVerification(REV)supportsrestorationandchecksthecurrentrevisionbeforecompletion.Together,thesemechanismscoordinateplanning,execution,andvisualfeedbackacrossrenderingbackends.Weevaluate11powerfulclosed-sourceMLLMsonMaLiang-IBenchandfouronMaLiang-VBench,measuringgenerationsuccess,visualquality,andcomputationalcost.GPT-6-Astraachieves100%generationsuccessonbothbenchmarks,with96.0%ofimagetasksand76.9%ofvideotasksmeetingallqualitythresholds.Thecomparisonalsorevealsamismatchbetweengeneralcapabilityscoresandvisualgenerationperformance,withsimilarlyscoredmodelsdifferingsubstantiallyintheirabilitytosatisfyvisualrequirements.MaLiang-HarnessprovidesasystematicbasisforstudyinghowMLLMstranslateexecutablecodeintovisualoutcomes,exposingboththepotentialofprogrammablegenerationandthelimitationsofgeneralbenchmarksaspredictorsofthisability.Theprojectisavailableathttps://github.com/gulucaptain/MaLiang-Harness.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.34309
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34309 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34309 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34309 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.
GUI harness [video]
A research project presenting a GUI harness that allows users or LLMs to build complex applications using simple vector graphics functions, featuring integrated code execution, history management, and support for multiple open-weight LLMs.
EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses
This paper introduces EvoGen-Harness, a generator-agnostic framework that evolves multiple external responsibilities around frozen text-to-image models using a Trace method for failure attribution and coordinated search, achieving significant benchmark gains over existing single-dimension adaptation baselines.
Show-Harness: Just a VLM Agent Can Play Robots
Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.
Self-Harness: Harnesses That Improve Themselves
Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.