Tag
Celeris-1 Magnus is a hybrid diffusion model optimized for agentic work, achieving a 41.2% solve rate on the τ³-bench banking benchmark at a 55-second median time, outperforming models like GPT-5.6-sol.
LayerRecall improves long-video consistency in diffusion models by selectively routing historical memory to specific layers, supervised by cross-horizon prediction matching.
I trained a 1.2B DiT model for instrumental game music generation, using Stable Audio's VAE and aiming to cover diverse styles. The project is open-source with a WebUI and samples available on HuggingFace.
ADAPT is a physics-aware conditional diffusion model for HVAC control that reduces energy consumption and occupant discomfort, with robust transferability to unseen environments.
Block3D accelerates text-to-3D generation by using block-wise diffusion with confidence-guided correction to reduce inference time while preserving geometric fidelity, achieving a 5.15x speedup.
An individual trained a diffusion model to generate 32x32 pixel images on a Shrike lite microcontroller with only 264KB RAM, experimenting with FPGA acceleration that hit memory bottlenecks, resulting in noisy but sometimes interesting outputs.
DFlash 2 is a block-diffusion draft model for speculative decoding that improves inference speed for the Qwen3.8-27B language model, offering higher acceptance length and throughput in benchmarks.
Molei Tao introduces FLARE, a diffusion language model that achieves near GPT5 performance with significantly faster inference speed.
UniSwap is a new framework for joint audio-visual identity swapping in talking videos, using a unified streaming audio-visual diffusion transformer to replace appearance and vocal timbre while preserving source content and dynamics.
Surg-UniWorld is a unified surgical world model with multimodal control experts, enabling controllable generation of coherent instrument-tissue interaction videos using edge, depth, and optical-flow inputs. It introduces a new benchmark (Cholec80-SurgWAM) and outperforms existing controllable video generation methods.
This article provides instructions for placing repackaged MiniMax-Music-3 model files into ComfyUI directories for music generation.
This paper introduces a probabilistic ensemble model based on a conditional diffusion model for real-time tsunami inundation forecasting, offering uncertainty quantification in contrast to deterministic warnings. Validated with 2011 Tohoku-oki data, it demonstrates that generative AI can shift tsunami forecasting from deterministic to probabilistic approaches.
Scenema Audio, an expressive text-to-speech model with zero-shot voice cloning, is now available as a native ComfyUI custom node, quantized to run on 8GB VRAM. The release adds inline stage direction cues, 12 preset voices, and simplifies the prompt format for ComfyUI.
This paper introduces SynEnergy, a two-stage diffusion-based framework for generating synthetic energy consumption data while preserving rare anomalous events using heterogeneous graph-based anomaly semantic learning and anomaly semantic-guided diffusion.
UniNav is a unified world-action diffusion model for image-goal visual navigation that jointly predicts future visual observations and waypoint trajectories in a single diffusion process, achieving strong benchmark results with efficient inference.
Presents MBDiff, a multi-view behavior-aware diffusion model for probabilistic utility data imputation that learns user behavior from global, local, and instance-level views and uses a conditional attentional denoising network. Evaluated on real utility data from Florida, it outperforms state-of-the-art baselines.
Proposes the existence-field diffusion model (EFDM) that jointly models spatial locations and cardinality of point sets via a unified diffusion process with existence variables, eliminating the need for explicit discrete transitions.
This paper introduces DiffTilt, a distributional framework that exponentially tilts a diffusion model-induced joint distribution over environments and executions to efficiently discover rare safety-critical failures in autonomous and cyber-physical systems, outperforming conditional sampling strategies on ARCH-COMP benchmarks and a new tractor-trailer benchmark.
This paper introduces TriLayer, a large-scale video dataset with foreground-background-composite triplets, and DBL-Diffusion, a dual-branch diffusion framework for explicit layered video representation, enabling high-fidelity object insertion and layer decomposition.
The lab behind LLaDA2.2 released a diffusion model benchmarked against its own autoregressive model, showing diffusion lags on general knowledge and coding but wins on speed and agent tasks, providing a clean tradeoff data point.