Tag
SenseNova released a preview of its U1.5 Lite model, showing benchmark gains in image generation and editing, with native 4K output and improved Chinese/English text rendering, though acknowledged weaknesses remain.
Nous Research has opened free access to FLUX 3 Preview for all Nous Portal users, including free tier, and is running a short film contest with prizes.
Kroma v0.1 is a LoRA fine-tune for Krea 2, released as a ComfyUI-compatible safetensors file with rank 256 and weight deltas for RMSNorm/modulation tensors, under an MIT license.
This paper introduces Synthetic Self-Guidance (SSG), a method that attaches a lightweight prediction head to a frozen pretrained pixel-space diffusion model, using the discrepancy between intermediate and final predictions as self-guidance during sampling. It shows that model-generated samples suffice for training the head, improving FID by over 50% on several variants without classifier-free guidance and enhancing strong baselines with CFG.
Midjourney V8.2 has been released, marking a new version of the popular AI image generation model.
Google integrates Nano Banana 2's image generation into Google Earth, enabling users to visualize historical scenes, reimagine spaces, and brainstorm real estate plans by typing prompts.
NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.
Microsoft unveiled MAI Image 2.5 Pro, a new in-house AI model with improved image generation quality, complex prompt handling, text rendering, and natural language editing.
Parallel Decoding Distillation (PDD) is a trajectory-based distillation method that accelerates image and video generation by predicting multiple denoising steps per network evaluation, achieving state-of-the-art performance with 4-8 NFEs on models like LTX-2.3, Wan14B, and Qwen-Image.
This paper identifies a failure mode in classifier-free guidance distillation called Negative Branch Asymmetry, where errors in the positive and negative CFG branches cancel out, and proposes Positive-Direction Matching to supervise branches separately for more robust distilled models.
Repackaged model files for ComfyUI from the Microsoft Mage-Flow model, including diffusion models, text encoder, and VAE.
This paper proposes Spectral Alignment (SPA), a lightweight guidance-based method that reduces exposure bias in diffusion models by calibrating the power spectrum of intermediate predictions, showing consistent improvements across pixel-space, latent, and flow-matching models with minimal computational overhead.
Microsoft announces public preview of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, their latest purpose-built generative AI models for image and voice, now available on Azure AI and powering Microsoft products like Bing Image Creator.
ComfyUI is an open-source, node-based AI creation engine that lets builders and visual professionals design complex generation workflows for image, video, audio, and 3D without coding, with partial re-execution and broad model support.
Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.
BFL has introduced FLUX 3, a multi-modal AI model capable of generating images, videos, and audio.
User MrLarus created two series of art posters using GPT-Image2—Eastern negative space style and instrumental echo style—emphasizing their sophistication and detailed craftsmanship.
Oxygen-TryOn is a unified foundation model for any-item virtual try-on, achieving state-of-the-art consistency and realism across single and multi-item try-on tasks through a dedicated data engine and three-stage training pipeline.
Proposes ProVisE, a benchmark-agnostic framework to evaluate spatial cognition in image-generation models using pixel-space outputs, and introduces SpatialGen-Bench for unified evaluation across 14 spatial subtasks.
Dylan Castillo conducted a rigorous investigation to determine if AI labs have been secretly training models to draw pelicans riding bicycles. Testing multiple models with various animal-vehicle combinations, he found no evidence of 'pelicanmaxxing'.