diffusion-model

Tag

Cards List
#diffusion-model

Qwen-Image-2.0-RL Technical Report

Hugging Face Daily Papers · 2026-06-25 Cached

This technical report presents Qwen-Image-2.0-RL, a post-training pipeline using reinforcement learning from human feedback and on-policy distillation to enhance visual quality and instruction-following in image generation and editing tasks.

0 favorites 0 likes
#diffusion-model

Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation

arXiv cs.AI · 2026-06-24 Cached

This paper introduces Prob-BBDM, a probabilistic Brownian Bridge Diffusion Model for efficient and high-quality MRI sequence synthesis from 2D axial slices, achieving up to 88.46% SSIM and 26.09 dB PSNR with only 4 diffusion steps, and demonstrating clinical utility in tumor segmentation.

0 favorites 0 likes
#diffusion-model

@charles_irl: dflash go brr

X AI KOLs Timeline · 2026-06-24 Cached

NVIDIA announces DFlash, an open source block diffusion model for speculative decoding that achieves up to 15x higher inference throughput on Blackwell GPUs while maintaining interactivity.

0 favorites 0 likes
#diffusion-model

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

Hugging Face Daily Papers · 2026-06-24 Cached

This paper presents TryOnCrafter, a novel framework for camera-controllable video virtual try-on that uses a renderable 4D try-on proxy and DiT-based video generation to achieve omnidirectional viewpoint exploration, overcoming the limitations of existing methods that depend on fixed source camera trajectories.

0 favorites 0 likes
#diffusion-model

I'm eager for a 15x speedup on my strix halo

Reddit r/LocalLLaMA · 2026-06-23

Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.

0 favorites 0 likes
#diffusion-model

Diffusion Model that can turn any Image into a Playable Hallucination! BUT LOCALLY, NOT ON DATACENTER

Reddit r/ArtificialInteligence · 2026-06-23

A diffusion model that can transform any image into an interactive, playable hallucination, running locally on user hardware.

0 favorites 0 likes
#diffusion-model

Krea 2 released on Hugging Face

Reddit r/LocalLLaMA · 2026-06-23 Cached

Krea 2 is a 12-billion parameter text-to-image diffusion model released open-weight on Hugging Face, with Raw (base) and Turbo (post-trained) checkpoints available.

0 favorites 0 likes
#diffusion-model

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Hugging Face Daily Papers · 2026-06-22 Cached

Vera is a layered diffusion model for video editing that preserves source content by generating edit layers and alpha mattes, using a Mixture-of-Transformers architecture.

0 favorites 0 likes
#diffusion-model

Inception Labs' Mercury 2 AI Beats Google's DiffusionGemma at Its Own Game (4 minute read)

TLDR AI · 2026-06-22 Cached

Inception Labs released Mercury 2, a diffusion language model that generates roughly 1,000 tokens per second and outperforms Google's DiffusionGemma on the AIME 2026 benchmark with a score of 90% versus 69.1%, though DiffusionGemma is free and open-weight while Mercury 2 is a paid, closed-weight API model.

0 favorites 0 likes
#diffusion-model

krea/Krea-2-Turbo

Hugging Face Models Trending · 2026-06-18 Cached

Krea released Krea 2 Turbo, a 12-billion parameter text-to-image diffusion model, available open-weight on Hugging Face with support for multiple inference libraries.

0 favorites 0 likes
#diffusion-model

DiRecT: Safe Diffusion-Based Planning via Receding-Horizon Denoising

arXiv cs.LG · 2026-06-16 Cached

DiRecT introduces a training-free algorithm for safe diffusion-based planning that enforces constraints only on final clean trajectories using receding-horizon denoising, improving safety and performance over existing methods.

0 favorites 0 likes
#diffusion-model

Decoupled Latent Optimization of Diffusion Models for Full Waveform Inversion

arXiv cs.LG · 2026-06-15 Cached

Introduces Decoupled Latent Optimization (DLO) for full waveform inversion, which relaxes latent optimization into a quadratic-penalty objective, outperforming classical and diffusion-based methods on benchmarks while preserving smoothed-velocity initialization.

0 favorites 0 likes
#diffusion-model

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

Hugging Face Daily Papers · 2026-06-14 Cached

Track2View generates novel camera viewpoints from videos by conditioning a video diffusion transformer on paired 3D point tracks, achieving state-of-the-art visual quality and significant reductions in rotation and translation errors.

0 favorites 0 likes
#diffusion-model

APCyc: Property-Informed Design of Cyclic Peptides via Automated Cyclization

arXiv cs.AI · 2026-06-12 Cached

APCyc is a target-aware generative framework that designs cyclic peptides with controlled physicochemical properties by explicitly modeling cyclization patterns and using Bayesian posterior guidance.

0 favorites 0 likes
#diffusion-model

Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

arXiv cs.AI · 2026-06-12 Cached

Pythagoras-Prover is a compute-efficient family of Lean theorem provers that achieves strong performance using curriculum supervised fine-tuning and a novel Augmented Lean Formalisation technique. The 4B model surpasses DeepSeek-Prover-V2-671B at pass@32 on MiniF2F-Test, and the 32B model sets a new state-of-the-art among open-source provers.

0 favorites 0 likes
#diffusion-model

[Talk] Text Diffusion — Google DeepMind's Brendan O’Donoghue

Reddit r/LocalLLaMA · 2026-06-11 Cached

DeepMind researcher Brendan O'Donoghue provides an in-depth introduction to text diffusion models, which generate text through iterative denoising. Compared to autoregressive models, they offer lower latency but limited throughput, and demonstrate unique advantages such as self-correction and dynamic computation.

0 favorites 0 likes
#diffusion-model

Google's latest DiffusionGemma open AI model comes with a 4x speed boost

Ars Technica · 2026-06-10 Cached

Google released DiffusionGemma, an experimental open-source diffusion model for text generation that achieves 4x speed boost over autoregressive models, optimized for local processing.

0 favorites 0 likes
#diffusion-model

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

Hugging Face Daily Papers · 2026-06-09 Cached

This paper introduces Lip Forcing, the first autoregressive diffusion method for real-time video-to-video lip synchronization. By distilling a 14B teacher into causal students and using only two denoising steps, it achieves 31 FPS streaming at 1.3B scale, 17.6x faster than same-scale bidirectional models.

0 favorites 0 likes
#diffusion-model

@XAMTO_AI: ControlNet author Min Shen has come up with something new! The newly open-sourced FramePack directly lowers the barrier for video generation — runs on just 6GB VRAM, generates a 1-minute 30fps video with a 13B model, and on an RTX 4090 it takes only 1.5 seconds per frame. Such configuration requirements were unimaginable before. The core idea is frame-by-frame…

X AI KOLs Timeline · 2026-06-08 Cached

ControlNet author Min Shen has open-sourced the FramePack video generation model, which requires only 6GB of VRAM to run a 13B model, generates a 1-minute 30fps video, takes 1.5 seconds per frame on an RTX 4090, and comes with a one-click Windows package.

0 favorites 0 likes
#diffusion-model

MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

Hugging Face Daily Papers · 2026-06-08 Cached

The paper introduces MilliVid, a method for improving long-range consistency in video generation by using a multi-scale autoencoder to compress frames into hierarchical tokens and then generating them with a coarse-to-fine diffusion model, outperforming baselines on Minecraft videos.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback