GenRouter: Unified Workflow Routing for Agentic Image Generation
Summary
GenRouter is a unified routing framework for agentic image generation that adaptively directs prompts to optimal workflows, significantly reducing costs and latency while improving visual alignment through demand profiling and self-evolution.
View Cached Full Text
Cached at: 08/18/26, 03:52 AM
Paper page - GenRouter: Unified Workflow Routing for Agentic Image Generation
Source: https://huggingface.co/papers/2608.16721
Abstract
GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution.
The rapid evolution of text-to-image (T2I) generation models has effectively solved the foundational challenge of raw pixel synthesis, shifting the community’s focus toward fulfilling increasingly intricate user requests. While recentagentic image generationworkflows enhance static inference with advanced capabilities like external knowledge retrieval and iterative reasoning, they mostly operate in isolated silos with fixed ``one-size-fits-all“ topologies. This inevitably leads to severe compute-mismatch, where simple queries are forced through computationally heavy pipelines. To bridge this gap, we presentGenRouter, the first unifiedworkflow routingframework foragentic image generation. We first formulateGenCanvas, standardizing diverse agentic pipelines into a universal set of foundational primitives and executable templates. Operating over this unified space,GenRouteradaptively routes heterogeneous prompts to their optimal workflows via (i)demand profiling, (ii)experience matching, and (iii)Pareto filtering. Extensive experiments across diverse benchmarks demonstrate thatGenRouterachieves superior visual alignment while reducing execution costs by over 95% and latency by 65% compared to heavyweight static pipelines. Furthermore, the system continuously self-evolves via accumulated experience, enabling robustzero-shot generalizationthat boosts performance and halves computational overhead.
View arXiv pageView PDFGitHub11Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.16721 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.16721 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.16721 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Runway launches AI model router as generative media gets crowded
Runway launched the Media Router, a tool that automatically selects the best generative media model for a given request based on quality, speed, or cost preferences. It aims to simplify integration for developers amid a crowded generative media landscape.
Runway Launched an AI Router for Generative Media (2 minute read)
Runway launched Media Router, a preference-optimized router that automatically selects the best video, image, or audio model based on user-defined criteria for cost, quality, or latency, eliminating manual model picking. It is live now in Runway Dev.
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
GenEvolve is a self-evolving image generation framework that uses tool-orchestrated trajectories and visual experience distillation to iteratively improve generative capabilities, achieving state-of-the-art performance.
GenClaw: Code-Driven Agentic Image Generation
GenClaw introduces a code-driven agentic image generation framework that breaks the black-box paradigm by mimicking the human creative process: conceptualizing, sketching with code (SVG/HTML/Three.js), and then using generative models for texture and photorealism.
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Qwen-Image-Agent proposes a unified agentic framework that addresses the context gap in text-to-image generation by integrating planning, reasoning, searching, and memory mechanisms. It introduces IA-Bench for evaluation and achieves state-of-the-art performance.