visual-generation

Tag

Cards List
#visual-generation

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

arXiv cs.LG · 3d ago Cached

OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.

0 favorites 0 likes
#visual-generation

@JenovaAIAgent: Introducing Crystal Ball Reader — an AI agent that translates your questions into cinematic visions inside a crystal sp…

X AI KOLs Following · 2026-09-08 Cached

The article introduces Crystal Ball Reader, an AI agent that converts user questions into cinematic visions inside a crystal sphere to aid decision-making.

0 favorites 0 likes
#visual-generation

Agentic Visual Generation: From Generative Models to Agentic Control

Hugging Face Daily Papers · 2026-09-06 Cached

This paper proposes a four-level framework classifying agentic visual generation systems based on the controller's direct decision scope over generation operations, from fixed support to experience-adaptive control.

0 favorites 0 likes
#visual-generation

@JenovaAIAgent: Comic Creator is an AI agent that builds complete comics page by page — from one-shot issues to 200-page graphic novels…

X AI KOLs Following · 2026-09-04 Cached

Comic Creator is an AI agent that maintains character and style consistency across comic pages, enabling the creation of complete comics from single issues to lengthy graphic novels with integrated artwork and dialogue.

0 favorites 0 likes
#visual-generation

@songhan_mit: Ten years, same question, much harder version of it. Jun-Yan has been pushing the frontier of visual generation — someo…

X AI KOLs Timeline · 2026-09-04 Cached

Jun-Yan Zhu announces the launch of Nunchux AI, a company focused on visual generation, marking a decade since his work on influential models like CycleGAN.

0 favorites 0 likes
#visual-generation

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

Hugging Face Daily Papers · 2026-09-01 Cached

This paper explores the synergy between visual understanding and generation in unified multimodal models, showing that task-decoupled architectures and end-to-end optimization can enhance performance by turning coexistence into synergy.

0 favorites 0 likes
#visual-generation

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

Hugging Face Daily Papers · 2026-08-11 Cached

The paper introduces Adversarial Fréchet Distance (AdvFD), which adds a learnable adversarial feature space to static Fréchet losses to improve generator post-training, with real-feature whitening to stabilize optimization.

0 favorites 0 likes
#visual-generation

@rohanpaul_ai: Longer prompts are not what image generators need. Text-to-image models seem less constrained by prompt length than by …

X AI KOLs Following · 2026-08-05 Cached

A new paper argues that text-to-image models benefit less from longer prompts and more from explicitly organized visual structure, proposing structured prompts over natural-language prose.

0 favorites 0 likes
#visual-generation

Scaling Properties of Text Conditioning in Visual Generation

Hugging Face Daily Papers · 2026-07-31 Cached

This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.

0 favorites 0 likes
#visual-generation

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Hugging Face Daily Papers · 2026-07-30 Cached

This paper introduces Chimera, a hybrid visual diffusion backbone with a principled scaling recipe, combining Kimi Delta Attention, Multi-head Latent Attention, and sparse Mixture-of-Experts to efficiently handle long-context image and video generation. It also presents HeteroP, a module-wise hyperparameter transfer scheme, and Chinchilla-style scaling laws to train an 11B-parameter model with 2B activated parameters.

0 favorites 0 likes
#visual-generation

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Hugging Face Daily Papers · 2026-07-09 Cached

This paper addresses the knowledge boundary problem in visual generation by introducing the SearchGen-20K benchmark and SearchGen-Corpus-1M, and proposes a teach-then-search co-training framework to handle evolving, long-tailed user requests beyond a generator's training data.

0 favorites 0 likes
#visual-generation

Optimizing Visual Generative Models via Distribution-wise Rewards

Hugging Face Daily Papers · 2026-07-02 Cached

This paper presents a reinforcement learning framework for visual generative models that uses distribution-wise rewards, with a subset-replace strategy for efficiency, improving image diversity and quality while addressing mode collapse and reward hacking.

0 favorites 0 likes
#visual-generation

Representation Distribution Matching for One-Step Visual Generation

Hugging Face Daily Papers · 2026-07-02 Cached

This paper introduces Representation Distribution Matching (RDM), a method for one-step image generation by matching feature distributions under pretrained encoders, achieving state-of-the-art results on ImageNet and enabling post-training of FLUX.2 into a one-step generator with improved performance.

0 favorites 0 likes
#visual-generation

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

Hugging Face Daily Papers · 2026-06-25 Cached

SharpMoE is a post-training framework that improves routing in diffusion mixture-of-experts models by using clean latent features to identify salient tokens and a trajectory routing loss to allocate compute precisely, achieving state-of-the-art visual generation.

0 favorites 0 likes
#visual-generation

@drfeifei: I’m very excited by this new benchmark dataset for visual generation that is suitable for the modern era of large scale…

X AI KOLs Following · 2026-05-29 Cached

Introducing GPIC (Giant Permissive Image Corpus), a large-scale dataset of 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, fully permissive for research and commercial use.

0 favorites 0 likes
#visual-generation

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

Hugging Face Daily Papers · 2026-05-14 Cached

Introduces CLVR (Closed-Loop Visual Reasoning), a framework that reformulates text-to-image generation from a single-step process into a closed-loop, multi-step visual reasoning approach using a VLM controller and diffusion models, achieving improved performance on compositional prompts.

0 favorites 0 likes
#visual-generation

Codex for Creatives: Riff, Design, Ship

YouTube AI Channels · 2026-06-12 Cached

OpenAI's Codex, typically used for coding, can also serve as a creative partner for generating brand ad campaigns by understanding style guides and emotional prompts, as demonstrated by creative specialist Shad Nelson.

0 favorites 0 likes
← Back to home

Submit Feedback