@stevibe: Image Models side-by-side As Qwen has released Qwen Image 3.0 today, I have tested 4 prompts across 4 image models: > Q…
Summary
A user compares four image generation models—Qwen Image 3.0, GPT Image 2, Nano Banana 2, and Seedream 5.0 Pro—with four prompts in a Twitter thread.
View Cached Full Text
Cached at: 07/22/26, 08:36 PM
Image Models side-by-side
As Qwen has released Qwen Image 3.0 today, I have tested 4 prompts across 4 image models:
Qwen Image 3.0 GPT Image 2 Nano Banana 2 Seedream 5.0 Pro
A thread
A vertical A2 event poster for the “Harbor Lights Night Market”, designed in a retro 1970s Japanese travel-poster style: warm sunset gradient background (deep orange to plum purple), grainy risograph texture, flat illustrated skyline of a harbor with paper lanterns strung between boats.
Typography, all rendered accurately: — Main title in large hand-brushed English display type: “HARBOR LIGHTS NIGHT MARKET” — Directly beneath, the same title in Japanese: 「ハーバーライツ夜市」 — And in Spanish italic serif: “Mercado Nocturno Luces del Puerto” — A date block in a stamp-style rounded rectangle: “SAT · AUG 15 · 5PM–MIDNIGHT” — Three vendor category cards arranged horizontally near the bottom, each with a small icon and bilingual label:
- A steaming ramen bowl icon — “STREET FOOD / 屋台グルメ”
- A vinyl record icon — “LIVE MUSIC / 音楽ライブ”
- A hand-thrown ceramic cup icon — “LOCAL CRAFT / 手作り市” — Footer in small but fully legible type: “Pier 7, Harbor District — Free entry — harborlights.example .com” — A vertical strip of text running down the right edge, top to bottom, in Japanese tategaki (vertical writing): 「夏の思い出をここで」
The lanterns should glow with soft bloom. Include a small crescent moon top-left. All text must be spelled exactly as written, with correct kana, and the visual hierarchy must read title → date → categories → footer.
A single image containing a 3x3 storyboard grid on white paper with thin black gutters, panels numbered 1–9 in small circles at each panel’s top-left corner. Consistent protagonist across ALL nine panels: a woman in her 60s with short silver hair, round red glasses, a mustard-yellow raincoat, and a small brown dachshund on a blue leash. Consistent muted watercolor style throughout.
Panel 1: Wide shot — she locks the door of a narrow green townhouse in the rain. Caption box below panel: “7:02 AM — The last delivery.” Panel 2: Close-up of her gloved hand holding a parcel wrapped in brown paper, tied with red string, addressed label reading “Unit 4B, Alder Lane”. Panel 3: She and the dachshund wait at a crosswalk; a bus splashes past. Caption: “The city never waits.” Panel 4: Low angle — she looks up at a crooked apartment building, its buzzer panel visible. Panel 5: Extreme close-up of the buzzer panel; her finger presses the button labeled “4B”, other buttons read 1A, 2A, 3B, 4A. Panel 6: The door cracks open; only a sliver of a young man’s face and one wide eye visible. Caption: “You’re late,” he whispered. Panel 7: Insert shot — the parcel changes hands in the doorway; rain drips off the awning above. Panel 8: She walks away down the wet street, seen from behind, the dachshund looking back over its shoulder at the door. Panel 9: Final wide shot — the door, now closed, with the red string from the parcel caught in the doorframe. Caption: “But she never delivered anything at all.”
Lighting is consistent overcast morning rain in every panel. The raincoat, glasses, dog, and leash must be identical in every panel where they appear.
A tall single-page educational infographic titled “WHY THE SKY IS BLUE — AND SUNSETS ARE RED”, in a clean flat-design science-magazine style: off-white background, navy/teal/coral palette, geometric sans-serif headings, thin-line diagrams.
Section 1 (top): Header title, subtitle “Rayleigh scattering explained in one page”. A horizontal visible-light spectrum bar labeled with wavelengths at 400 nm, 500 nm, 600 nm, 700 nm, violet through red. Section 2: A diagram of sunlight entering the atmosphere: the sun at left, Earth arc at right, parallel rays hitting scattered air molecules drawn as small dots; short blue arrows scattering in all directions from the molecules, long red arrows passing straight through. Two labels: “Blue light: scattered ~5.5x more” and “Red light: passes through”. Section 3: A callout card with the scattering formula rendered in proper math notation: I ∝ 1/λ⁴, with the caption “Scattering intensity is inversely proportional to the fourth power of wavelength”. Beside it, a mini bar chart comparing relative scattering of 450 nm vs 650 nm light, bars labeled “450 nm” and “650 nm” with the ratio “≈ 4.4 : 1”. Section 4: Two side-by-side circular diagrams: “NOON — short path through atmosphere” showing a nearly vertical light path, and “SUNSET — long path” showing a shallow grazing path with the annotation “Blue is scattered away before reaching your eye”. Section 5 (footer): Three small fact chips in a row: “Fact: Mars sunsets are blue”, “Fact: The ocean is blue for a different reason”, “Fact: Violet scatters most, but our eyes favor blue”. Bottom credit line in tiny 8-pt style text: “Sources: Rayleigh (1871) · Illustration for classroom use · Fig. 1 of 1”.
All text spelled exactly as specified, correct superscript in λ⁴, correct nm units, clear reading order top to bottom.
Laguna S 2.1 running native NVFP4
Picked Blackwell hardware for this side-by-side: DGX Spark | 19.44 tok/s | 162ms TTFT RTX PRO 6000 | 108.54 tok/s | 266ms TTFT 4× RTX 5090 (TP4+EP4) | 145.81 tok/s | 521ms TTFT
117.6B MoE (8.5B active), quantized to FP4, on your own box.
Similar Articles
Qwen-Image-2.0 Technical Report
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
@swyx: Qwen Image 3 announced. these pictures are NOT screenshots. all generated in a single pass. the last one (image annotat…
Qwen Image 3, a new image generation model, announced. It generates images in a single pass and could spawn edtech/industrial training startups.
Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge (6 minute read)
Qwen-Image-3.0 is a third-generation foundational image generation model supporting up to 4.5k token input, native rendering in 12 languages, and simulation of interfaces like web pages and games, leveraging rich world knowledge for practical deployment.
Qwen-Image-2.0 Technical Report (57 minute read)
This technical report presents Qwen-Image-2.0, a new image generation model from Alibaba's Qwen team, detailing its architecture and capabilities.
Qwen-Image-Flash (26 minute read)
This paper from Alibaba revisits few-step distillation for visual generative models, focusing on training recipe factors such as data composition, teacher guidance, and task mixture, using Qwen-Image-2.0 as a case study to develop Qwen-Image-Flash.