diffusion-model

Tag

Cards List
#diffusion-model

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Hugging Face Daily Papers ↗ · 2d ago Cached

ThinkV2V introduces a reasoning-driven framework that activates MLLM thinking before visual generation for instruction-guided video editing, using an MLLM-to-DiT architecture with progressive curriculum training and inference-time thinking scaling. The authors also release the ThinkV2V-150K dataset and ThinkV2V-Bench, showing their 5B DiT model outperforms larger 10B baselines.

0 favorites 0 likes
#diffusion-model

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

Hugging Face Daily Papers ↗ · 4d ago Cached

VGGT-Diff is a geometry-routed multi-view diffusion model that integrates visual geometry latents with a pretrained diffusion model to improve sparse-view novel view synthesis, achieving competitive or state-of-the-art performance.

0 favorites 0 likes
#diffusion-model

Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces

arXiv cs.LG ↗ · 6d ago Cached

This paper introduces a generative atmospheric super-resolution method using composable interfaces to condition a pretrained diffusion model on heterogeneous in situ observations, improving reconstruction accuracy without model retraining.

0 favorites 0 likes
#diffusion-model

Stefano Ermon: Autoregressive inference is sequential and memory-bound. Diffusion is built to map to GPUs — that's why it wins.

Reddit r/artificial ↗ · 6d ago

Augment Code switched its coding-agent backend to Stefano Ermon's Mercury 2.5 diffusion model, achieving 82% latency reduction and 90% cost cut in production. The article highlights the performance advantages of diffusion models and the need for independent AI benchmarking tools.

0 favorites 0 likes
#diffusion-model

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

arXiv cs.LG ↗ · 2026-09-24 Cached

The paper proposes LLMAE, a method to repurpose pre-trained decoder-only LLMs as continuous text autoencoders using a latent bottleneck, achieving high-fidelity reconstruction and enabling downstream tasks like image captioning.

0 favorites 0 likes
#diffusion-model

@LinusEkenstam: I love seeing how fast the frontier labs are jumping in on creating forks. Just today we’ve seen omni-jev and now djev …

X AI KOLs Timeline ↗ · 2026-09-23 Cached

Google Gemma team has released DiffusionGemma-Jev (djev), a fork of JEV, simplifying deployment on Google Cloud Run with performance metrics of ~35-60 ms latency and batch processing at 100-123 requests per second.

0 favorites 0 likes
#diffusion-model

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

Uranus is a next-generation simulation infrastructure for embodied AI that uses a joint-trajectory-conditioned autoregressive diffusion model to enable scalable, low-latency generation of robot simulations with streaming rollout and extensible control.

0 favorites 0 likes
#diffusion-model

[MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release!

Reddit r/LocalLLaMA ↗ · 2026-09-21

Supra2-IMG is a tiny 100M parameter text-to-image model that achieves state-of-the-art quality in image generation, trained from scratch in under 10 hours on a single H100 and released open-source on Hugging Face.

0 favorites 0 likes
#diffusion-model

unsloth/Qwen-Image-2.1-GGUF

Hugging Face Models Trending ↗ · 2026-09-21 Cached

Qwen-Image-2.1 is a unified text-to-image generation and image editing model with 7B parameters, featuring improvements in efficiency, transparency, versatility, and realism. This GGUF quantized version from unsloth enables efficient local inference.

0 favorites 0 likes
#diffusion-model

DSD: Learning Diverse and Reusable Motor Skills via Diffusion Skill Discovery

arXiv cs.LG ↗ · 2026-09-17 Cached

DSD is a diffusion-based method for discovering diverse and reusable motor skills in simulated humanoid control, improving upon prior skill discovery techniques with broader behavioral coverage.

0 favorites 0 likes
#diffusion-model

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

RefineEdit is a training-free prompt-to-prompt image editing method that uses a generative refinement network to enhance edit localization and background preservation, achieving top benchmark scores.

0 favorites 0 likes
#diffusion-model

Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors

arXiv cs.LG ↗ · 2026-09-16 Cached

SCTab-Diff is a semantics-consistent tabular diffusion framework that uses weak semantic priors to generate high-fidelity synthetic tabular data, improving distributional fidelity and semantic consistency over existing methods.

0 favorites 0 likes
#diffusion-model

Comfy-Org/Qwen-Image-2.1

Hugging Face Models Trending ↗ · 2026-09-15 Cached

Repackaged model files for Qwen-Image-2.1 optimized for ComfyUI workflows, including text-to-image and image edit capabilities.

0 favorites 0 likes
#diffusion-model

Qwen/Qwen-Image-2.1

Hugging Face Models Trending ↗ · 2026-09-14 Cached

Qwen-Image-2.1 is an open-source unified text-to-image and image editing model with 7B parameters, featuring efficient architecture, transparency support, and versatile editing capabilities.

0 favorites 0 likes
#diffusion-model

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Hugging Face Daily Papers ↗ · 2026-09-11 Cached

Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.

0 favorites 0 likes
#diffusion-model

I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.)

Reddit r/LocalLLaMA ↗ · 2026-09-09

An independent audio researcher has trained a model called Foundation-1 that generates infinite one-shot sounds for music production and converts text prompts into playable synths, releasing the model along with documentation and inferencing tools.

0 favorites 0 likes
#diffusion-model

Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts

Hugging Face Daily Papers ↗ · 2026-09-04 Cached

Srijika is a system that generates installable OpenType fonts for nine Indic scripts by restyling glyphs using a latent diffusion model while preserving layout consistency.

0 favorites 0 likes
#diffusion-model

DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving

arXiv cs.LG ↗ · 2026-09-03 Cached

DiDrive is a risk-aware hierarchical diffusion framework for safe offline reinforcement learning in autonomous driving that improves performance in complex traffic scenarios through integrated representation learning and distribution correction optimization.

0 favorites 0 likes
#diffusion-model

@Celeris_ai: Introducing Celeris-1 Magnus. A model built for agentic work. On τ³-bench banking, Magnus delivers 41.2% at a 55-second…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Celeris-1 Magnus is a hybrid diffusion model optimized for agentic work, achieving a 41.2% solve rate on the τ³-bench banking benchmark at a 55-second median time, outperforming models like GPT-5.6-sol.

0 favorites 0 likes
#diffusion-model

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

The paper presents Puppeteer, a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, posture-aware, and physically grounded gestures.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback