Tag
ThinkV2V introduces a reasoning-driven framework that activates MLLM thinking before visual generation for instruction-guided video editing, using an MLLM-to-DiT architecture with progressive curriculum training and inference-time thinking scaling. The authors also release the ThinkV2V-150K dataset and ThinkV2V-Bench, showing their 5B DiT model outperforms larger 10B baselines.
VGGT-Diff is a geometry-routed multi-view diffusion model that integrates visual geometry latents with a pretrained diffusion model to improve sparse-view novel view synthesis, achieving competitive or state-of-the-art performance.
This paper introduces a generative atmospheric super-resolution method using composable interfaces to condition a pretrained diffusion model on heterogeneous in situ observations, improving reconstruction accuracy without model retraining.
Augment Code switched its coding-agent backend to Stefano Ermon's Mercury 2.5 diffusion model, achieving 82% latency reduction and 90% cost cut in production. The article highlights the performance advantages of diffusion models and the need for independent AI benchmarking tools.
The paper proposes LLMAE, a method to repurpose pre-trained decoder-only LLMs as continuous text autoencoders using a latent bottleneck, achieving high-fidelity reconstruction and enabling downstream tasks like image captioning.
Google Gemma team has released DiffusionGemma-Jev (djev), a fork of JEV, simplifying deployment on Google Cloud Run with performance metrics of ~35-60 ms latency and batch processing at 100-123 requests per second.
Uranus is a next-generation simulation infrastructure for embodied AI that uses a joint-trajectory-conditioned autoregressive diffusion model to enable scalable, low-latency generation of robot simulations with streaming rollout and extensible control.
Supra2-IMG is a tiny 100M parameter text-to-image model that achieves state-of-the-art quality in image generation, trained from scratch in under 10 hours on a single H100 and released open-source on Hugging Face.
Qwen-Image-2.1 is a unified text-to-image generation and image editing model with 7B parameters, featuring improvements in efficiency, transparency, versatility, and realism. This GGUF quantized version from unsloth enables efficient local inference.
DSD is a diffusion-based method for discovering diverse and reusable motor skills in simulated humanoid control, improving upon prior skill discovery techniques with broader behavioral coverage.
RefineEdit is a training-free prompt-to-prompt image editing method that uses a generative refinement network to enhance edit localization and background preservation, achieving top benchmark scores.
SCTab-Diff is a semantics-consistent tabular diffusion framework that uses weak semantic priors to generate high-fidelity synthetic tabular data, improving distributional fidelity and semantic consistency over existing methods.
Repackaged model files for Qwen-Image-2.1 optimized for ComfyUI workflows, including text-to-image and image edit capabilities.
Qwen-Image-2.1 is an open-source unified text-to-image and image editing model with 7B parameters, featuring efficient architecture, transparency support, and versatile editing capabilities.
Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.
An independent audio researcher has trained a model called Foundation-1 that generates infinite one-shot sounds for music production and converts text prompts into playable synths, releasing the model along with documentation and inferencing tools.
Srijika is a system that generates installable OpenType fonts for nine Indic scripts by restyling glyphs using a latent diffusion model while preserving layout consistency.
DiDrive is a risk-aware hierarchical diffusion framework for safe offline reinforcement learning in autonomous driving that improves performance in complex traffic scenarios through integrated representation learning and distribution correction optimization.
Celeris-1 Magnus is a hybrid diffusion model optimized for agentic work, achieving a 41.2% solve rate on the τ³-bench banking benchmark at a 55-second median time, outperforming models like GPT-5.6-sol.
The paper presents Puppeteer, a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, posture-aware, and physically grounded gestures.