Tag
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
AntLing has open-sourced Ming-Image-0.1-Design, a 6B text-to-image model specialized for UI and text-rich visual designs, supporting transparent backgrounds and featuring a UI/UX design leaderboard.
A distilled student model of Qwen-Image-2.1, trained by Viggle using Distribution Matching Distillation, enabling text-to-image generation and image editing in 6 steps instead of 40, resulting in about 5× faster performance with competitive quality.
Supra2-IMG is a tiny 100M parameter text-to-image model that achieves state-of-the-art quality in image generation, trained from scratch in under 10 hours on a single H100 and released open-source on Hugging Face.
Qwen-Image-2.1 is a unified text-to-image generation and image editing model with 7B parameters, featuring improvements in efficiency, transparency, versatility, and realism. This GGUF quantized version from unsloth enables efficient local inference.
This article likely covers the Bonsai 2 27b AI model, which generates SVG images of whimsical scenes like a donkey making coffee and a mushroom riding a donkey.
Linum AI introduces JiT-DDT, a novel encoder-decoder architecture that trains text-to-image models 3.6× faster than previous methods while generating images with higher resolution.
This article compares the performance of various AI models in generating SVGs based on specific prompts for the years 2025 and 2026, detailing their output quality, time taken, and cost.
Qwen-Image-2.1 is an open-source unified text-to-image and image editing model with 7B parameters, featuring efficient architecture, transparency support, and versatile editing capabilities.
A detailed write-up on training a 210M text-to-image diffusion transformer from scratch on a single GPU, sharing key measurements on attention sinks, loss as a health signal, and timestep shifting benefits.
This post showcases Workflow1111, a rebuild of AUTOMATIC1111's stable diffusion web UI using Gradio Workflow, which integrates multiple media pipelines and AI models into a single canvas.
GPT Image 2.5 is now available on Martini, enabling text-to-image and editing with up to 4K resolution and multiple style references.
OpenAI's new AI models, GPT-Image-2.5-Sunburst and Flare, have achieved top rankings in text-to-image and image editing arenas, showing significant improvements over previous versions.
This paper analyzes how image tokenizer design affects joint text-image modeling in multimodal models using a controlled autoregressive testbed, showing distinct scaling behaviors and correlations with downstream performance.
Jasper Research has released a comprehensive cookbook and resources for building text-to-image models from scratch, featuring detailed explanations, a 100M-image dataset, and a codebase with a tiny model.
This paper introduces I-CARE, a methodology for systematically analyzing interference in machine unlearning for text-to-image models, providing formal definitions and an open-source framework to enable reproducible study.
An article detailing the end-to-end build of an app using the Qwen3.8-27b model, covering prompt flow, text-to-image generation, and image-to-video stitching.
This paper introduces ContextBias and ContextBench to evaluate bias persistence in text-to-image models, finding that bias increases in semantically unrelated contexts.
A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.
DiSCO is a training-free, black-box defense for text-to-image models that uses distribution-guided contrastive prompt optimization to prevent generation of Not-Safe-For-Work content, significantly reducing attack success rates.