Tag
Guillermo Rauch highlights Grok Imagine Image 2.0, now available on Vercel AI Gateway and ranking #2 on Arena.ai's leaderboard. Vercel offers access via AI CLI and a live playground.
Grok announces Imagine Image 2.0, a next-generation image model with precision editing, crisp text rendering, improved factuality, and real-world usefulness.
Grok announces Imagine Image 2.0, a next-generation image model with precision editing, crisp text rendering, and improved factuality for real-world use.
Grok Imagine Image 2.0 (Low) from xAI jumped to #2 in the Text-to-Image Arena, beating its own older quality model and showing significant improvement.
A tweet argues that a Claude Code skill can replace a $79/mo Higgsfield subscription by calling image-generation APIs directly, cutting per-image cost from ~31-34¢ to ~5¢ while keeping outputs local and owned by the user.
SenseNova U1.5-Lite-Preview is an open-source 8B MoT multimodal model that natively generates and edits ultra-wide panoramas, product posters, and commercial photography at 4K, with improved material rendering and fewer artifacts.
A Hugging Face model page for PinkCherry_MiniMax-H3, a niche AI model focused on generating NSFW furry and floral imagery, with update notes about improvements to rabbit motion, flower petals, and unicorn horns.
QwenCloud unveiled Qwen-Image-3.0-Pro, a powerful image generation model supporting dense layouts, tiny text rendering, and native multilingual output, positioning it as a deployable productivity tool.
Introduces ToolArtist, a fully agentic image generation model built from a unified multimodal model, using SFT and reinforcement learning (RAD-GRPO) to dynamically orchestrate reasoning, tool use, and image generation.
UniWorld-Design is a framework that redefines image generation using semantic RGBA layers as atomic units, comprising T2RGBA for generating layered assets and I2L for decomposing images into editable layers, achieving state-of-the-art results on the Crello benchmark.
A comparison of Nano Banana 2 and OpenAI image generation using a detailed prompt, with resulting images shared in comments.
Google is testing unreleased changes to close feature gaps between the Gemini desktop app and its web version, including dedicated tabs for image/video generation, camera capture, and custom MCP server support for Spark.
Introduces Poplar, a scalable Specify-Render-Inspect pipeline for synthesizing human-centric image datasets, and releases Poplar-9K, a curated dataset of 9,401 image-text pairs with auditable inspection records.
SenseNova released a preview of its U1.5 Lite model, showing benchmark gains in image generation and editing, with native 4K output and improved Chinese/English text rendering, though acknowledged weaknesses remain.
Nous Research has opened free access to FLUX 3 Preview for all Nous Portal users, including free tier, and is running a short film contest with prizes.
Kroma v0.1 is a LoRA fine-tune for Krea 2, released as a ComfyUI-compatible safetensors file with rank 256 and weight deltas for RMSNorm/modulation tensors, under an MIT license.
This paper introduces Synthetic Self-Guidance (SSG), a method that attaches a lightweight prediction head to a frozen pretrained pixel-space diffusion model, using the discrepancy between intermediate and final predictions as self-guidance during sampling. It shows that model-generated samples suffice for training the head, improving FID by over 50% on several variants without classifier-free guidance and enhancing strong baselines with CFG.
Midjourney V8.2 has been released, marking a new version of the popular AI image generation model.
Google integrates Nano Banana 2's image generation into Google Earth, enabling users to visualize historical scenes, reimagine spaces, and brainstorm real estate plans by typing prompts.
NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.