Tag
OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.
The article introduces Crystal Ball Reader, an AI agent that converts user questions into cinematic visions inside a crystal sphere to aid decision-making.
This paper proposes a four-level framework classifying agentic visual generation systems based on the controller's direct decision scope over generation operations, from fixed support to experience-adaptive control.
Comic Creator is an AI agent that maintains character and style consistency across comic pages, enabling the creation of complete comics from single issues to lengthy graphic novels with integrated artwork and dialogue.
Jun-Yan Zhu announces the launch of Nunchux AI, a company focused on visual generation, marking a decade since his work on influential models like CycleGAN.
This paper explores the synergy between visual understanding and generation in unified multimodal models, showing that task-decoupled architectures and end-to-end optimization can enhance performance by turning coexistence into synergy.
The paper introduces Adversarial Fréchet Distance (AdvFD), which adds a learnable adversarial feature space to static Fréchet losses to improve generator post-training, with real-feature whitening to stabilize optimization.
A new paper argues that text-to-image models benefit less from longer prompts and more from explicitly organized visual structure, proposing structured prompts over natural-language prose.
This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.
This paper introduces Chimera, a hybrid visual diffusion backbone with a principled scaling recipe, combining Kimi Delta Attention, Multi-head Latent Attention, and sparse Mixture-of-Experts to efficiently handle long-context image and video generation. It also presents HeteroP, a module-wise hyperparameter transfer scheme, and Chinchilla-style scaling laws to train an 11B-parameter model with 2B activated parameters.
This paper addresses the knowledge boundary problem in visual generation by introducing the SearchGen-20K benchmark and SearchGen-Corpus-1M, and proposes a teach-then-search co-training framework to handle evolving, long-tailed user requests beyond a generator's training data.
This paper presents a reinforcement learning framework for visual generative models that uses distribution-wise rewards, with a subset-replace strategy for efficiency, improving image diversity and quality while addressing mode collapse and reward hacking.
This paper introduces Representation Distribution Matching (RDM), a method for one-step image generation by matching feature distributions under pretrained encoders, achieving state-of-the-art results on ImageNet and enabling post-training of FLUX.2 into a one-step generator with improved performance.
SharpMoE is a post-training framework that improves routing in diffusion mixture-of-experts models by using clean latent features to identify salient tokens and a trajectory routing loss to allocate compute precisely, achieving state-of-the-art visual generation.
Introducing GPIC (Giant Permissive Image Corpus), a large-scale dataset of 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, fully permissive for research and commercial use.
Introduces CLVR (Closed-Loop Visual Reasoning), a framework that reformulates text-to-image generation from a single-step process into a closed-loop, multi-step visual reasoning approach using a VLM controller and diffusion models, achieving improved performance on compositional prompts.
OpenAI's Codex, typically used for coding, can also serve as a creative partner for generating brand ad campaigns by understanding style guides and emotional prompts, as demonstrated by creative specialist Shad Nelson.