text-to-image

Tag

Cards List
#text-to-image

On the Diffusibility of High-Dimensional Latents

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.

0 favorites 0 likes
#text-to-image

AntLing open sourced the Ming-Image-0.1-Design family

Reddit r/LocalLLaMA ↗ · 4d ago Cached

AntLing has open-sourced Ming-Image-0.1-Design, a 6B text-to-image model specialized for UI and text-rich visual designs, supporting transparent backgrounds and featuring a UI/UX design leaderboard.

0 favorites 0 likes
#text-to-image

Viggle/Qwen-Image-2.1-viggle-turbo

Hugging Face Models Trending ↗ · 4d ago Cached

A distilled student model of Qwen-Image-2.1, trained by Viggle using Distribution Matching Distillation, enabling text-to-image generation and image editing in 6 steps instead of 40, resulting in about 5× faster performance with competitive quality.

0 favorites 0 likes
#text-to-image

[MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release!

Reddit r/LocalLLaMA ↗ · 5d ago

Supra2-IMG is a tiny 100M parameter text-to-image model that achieves state-of-the-art quality in image generation, trained from scratch in under 10 hours on a single H100 and released open-source on Hugging Face.

0 favorites 0 likes
#text-to-image

unsloth/Qwen-Image-2.1-GGUF

Hugging Face Models Trending ↗ · 5d ago Cached

Qwen-Image-2.1 is a unified text-to-image generation and image editing model with 7B parameters, featuring improvements in efficiency, transparency, versatility, and realism. This GGUF quantized version from unsloth enables efficient local inference.

0 favorites 0 likes
#text-to-image

Bonsai 2 27b Q2 - Donkey making coffee svg and a mushroom riding a donkey

Reddit r/LocalLLaMA ↗ · 2026-09-18

This article likely covers the Bonsai 2 27b AI model, which generates SVG images of whimsical scenes like a donkey making coffee and a mushroom riding a donkey.

0 favorites 0 likes
#text-to-image

Training Text-to-Image Models 3.6× Faster

Hacker News Top ↗ · 2026-09-16 Cached

Linum AI introduces JiT-DDT, a novel encoder-decoder architecture that trains text-to-image models 3.6× faster than previous methods while generating images with higher resolution.

0 favorites 0 likes
#text-to-image

Show HN: Pelican-bicycle alternatives (updated for 2026)

Hacker News Top ↗ · 2026-09-14 Cached

This article compares the performance of various AI models in generating SVGs based on specific prompts for the years 2025 and 2026, detailing their output quality, time taken, and cost.

0 favorites 0 likes
#text-to-image

Qwen/Qwen-Image-2.1

Hugging Face Models Trending ↗ · 2026-09-14 Cached

Qwen-Image-2.1 is an open-source unified text-to-image and image editing model with 7B parameters, featuring efficient architecture, transparency support, and versatile editing capabilities.

0 favorites 0 likes
#text-to-image

Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P]

Reddit r/MachineLearning ↗ · 2026-09-11

A detailed write-up on training a 210M text-to-image diffusion transformer from scratch on a single GPU, sharing key measurements on attention sinks, loss as a health signal, and timestep shifting benefits.

0 favorites 0 likes
#text-to-image

Rebuilding AUTOMATIC1111 with Gradio Workflow

Hugging Face Blog ↗ · 2026-09-10 Cached

This post showcases Workflow1111, a rebuild of AUTOMATIC1111's stable diffusion web UI using Gradio Workflow, which integrates multiple media pipelines and AI models into a single canvas.

0 favorites 0 likes
#text-to-image

@martini_film: GPT Image 2.5 is on Martini. One bear. One photo. Thirty one-line directions: lego, bath, boba, origami, balloon, gold,…

X AI KOLs Timeline ↗ · 2026-09-09 Cached

GPT Image 2.5 is now available on Martini, enabling text-to-image and editing with up to 4K resolution and multiple style references.

0 favorites 0 likes
#text-to-image

@gdb: team has been cooking

X AI KOLs Timeline ↗ · 2026-09-08 Cached

OpenAI's new AI models, GPT-Image-2.5-Sunburst and Flare, have achieved top rankings in text-to-image and image editing arenas, showing significant improvements over previous versions.

0 favorites 0 likes
#text-to-image

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper analyzes how image tokenizer design affects joint text-image modeling in multimodal models using a controlled autoregressive testbed, showing distinct scaling behaviors and correlations with downstream performance.

0 favorites 0 likes
#text-to-image

Detailed explanation of how to create a text-to-image model from scratch. [R]

Reddit r/MachineLearning ↗ · 2026-09-02

Jasper Research has released a comprehensive cookbook and resources for building text-to-image models from scratch, featuring detailed explanations, a 100M-image dataset, and a codebase with a tiny model.

0 favorites 0 likes
#text-to-image

I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

arXiv cs.AI ↗ · 2026-09-02 Cached

This paper introduces I-CARE, a methodology for systematically analyzing interference in machine unlearning for text-to-image models, providing formal definitions and an open-source framework to enable reproducible study.

0 favorites 0 likes
#text-to-image

Qwen3.8-27b-UD-IQ3XXS - End to end Build App -> Prompt Flow-> Test to Image -> Image to Video - Stitch

Reddit r/LocalLLaMA ↗ · 2026-08-30

An article detailing the end-to-end build of an app using the Qwen3.8-27b model, covering prompt flow, text-to-image generation, and image-to-video stitching.

0 favorites 0 likes
#text-to-image

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

Hugging Face Daily Papers ↗ · 2026-08-30 Cached

This paper introduces ContextBias and ContextBench to evaluate bias persistence in text-to-image models, finding that bias increases in semantically unrelated contexts.

0 favorites 0 likes
#text-to-image

A dataset with 52 Text to image model evaluation [P]

Reddit r/MachineLearning ↗ · 2026-08-26

A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.

0 favorites 0 likes
#text-to-image

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

arXiv cs.AI ↗ · 2026-08-19 Cached

DiSCO is a training-free, black-box defense for text-to-image models that uses distribution-guided contrastive prompt optimization to prevent generation of Not-Safe-For-Work content, significantly reducing attack success rates.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback