efficient-generation

Tag

Cards List
#efficient-generation

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Hugging Face Daily Papers · 2026-07-05 Cached

Flash-BoN improves text-to-image generation efficiency by generating cheap draft candidates via timestep truncation, layer skipping, and activation proxies, then using multi-stage verification to select the best draft for full refinement, outperforming baselines under fixed wall-clock budgets.

0 favorites 0 likes
#efficient-generation

NVIDIA has released Nemotron-TwoTower-30B-A3B-Base-BF16, an unusual diffusion-based language model built from the Nemotron 3 Nano 30B-A3B backbone.

Reddit r/LocalLLaMA · 2026-06-25 Cached

NVIDIA released Nemotron-TwoTower-30B-A3B-Base-BF16, a diffusion-based language model that uses block-wise autoregressive diffusion to generate text by iterative denoising of token blocks, achieving 2.42× the generation throughput of the autoregressive baseline while retaining 98.7% of benchmark quality.

0 favorites 0 likes
#efficient-generation

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

arXiv cs.CL · 2026-06-25 Cached

Dustin introduces a sparse verification framework for speculative decoding that leverages draft model signals and sparse attention head scoring to overcome the KV cache verification bottleneck, achieving up to 27.85x speedup in self-attention and 9.17x end-to-end decoding speedup on long-context tasks with negligible accuracy loss.

0 favorites 0 likes
#efficient-generation

Latent Spatial Memory for Video World Models

Hugging Face Daily Papers · 2026-06-08 Cached

This paper introduces latent spatial memory for video world models, storing 3D scene information directly in diffusion latent space to avoid costly pixel-space reconstruction. The proposed Mirage framework achieves up to 10.57x faster generation and 55x memory reduction while achieving state-of-the-art performance on WorldScore and RealEstate10K.

0 favorites 0 likes
#efficient-generation

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

Hugging Face Daily Papers · 2026-05-07 Cached

SwiftI2V is a new efficient framework for high-resolution image-to-video generation that uses conditional segment-wise generation to achieve 2K synthesis with significantly reduced computational costs. It enables practical generation on single consumer or datacenter GPUs while maintaining input fidelity.

0 favorites 0 likes
#efficient-generation

zhen-nan/L2P

Hugging Face Models Trending · 2026-05-03 Cached

L2P proposes an efficient transfer paradigm that leverages pre-trained latent diffusion models to build pixel-space diffusion models, enabling high-quality generation with minimal computational overhead and data requirements, and supporting native 4K resolution.

0 favorites 0 likes
#efficient-generation

NVlabs/Sana

GitHub Trending (daily) · 2026-05-18 Cached

NVlabs/Sana is an efficiency-oriented open-source codebase for high-resolution image and video generation, including multiple model variants and training/inference pipelines.

0 favorites 0 likes
← Back to home

Submit Feedback