SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Summary
SANA-Streaming enables real-time high-resolution video-to-video editing on consumer GPUs using a hybrid diffusion transformer architecture, cycle-reverse regularization, and efficient system co-design, achieving 24 FPS at 1280x704 resolution on a single RTX 5090.
View Cached Full Text
Cached at: 06/01/26, 03:17 AM
Paper page - SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Source: https://huggingface.co/papers/2605.30409 Published on May 28
·
Submitted byhttps://huggingface.co/Yuyang-z
Yuyangon Jun 1
Abstract
SANA-Streaming enables real-time high-resolution video-to-video editing through a hybrid diffusion transformer architecture, cycle-reverse regularization, and efficient system co-design optimized for consumer GPUs.
Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements fortemporal consistencyand inference throughput. In this paper, we present SANA-Streaming, asystem-algorithm co-designed framework for high-resolution, real-time streaming video editing on consumer GPUs, with the following three core designs: (1) HybridDiffusion Transformerarchitecture introducessoftmax attentionin part of the blocks to improve local modeling capabilities while preserving the efficiency of linear layers. (2) Cycle-Reverse Regularization is a novel training strategy that enforces semantic consistency by predicting source frames from generated content viaflow matching, improvingtemporal consistencywithout requiring paired long edited videos. (3) Efficient System Co-design combines fused GDN kernels andMixed-Precision Quantization(MPQ) optimized for the NVIDIA Blackwell (RTX 5090) architecture. By profiling real-world throughput, our MPQ maximizes Tensor Core utilization while maintaining generation quality. The resulting system achieves real-time 1280 x 704 resolution editing at 24 end-to-end FPS on a single RTX 5090 GPU, with the DiT core running at 58 FPS. Experimental results demonstrate that our co-design approach significantly outperforms existing SOTA methods in both temporal coherence and system throughput.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.30409
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.30409 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.30409 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.30409 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Google’s free streaming service now lets you pick shows and movies to watch
Google TV Freeplay, Google's free ad-supported streaming service, now offers on-demand access to over 10,000 shows and movies, in addition to expanding its live TV lineup to more than 300 channels.
@Ryrenz: Damn, you can find the clip you want from hours of footage with just one sentence — MIT License, Whisper transcription plus LLM analysis. The most painful part of editing long videos is finding footage. With hours of recordings, to find every segment about a topic, you can only drag the timeline and listen bit by bit, burning an entire afternoon. PreenCut ...
PreenCut is an open-source tool (MIT License) based on Whisper transcription and LLM analysis, allowing users to search for segments in long videos using natural language, with support for batch processing, export/merging, and a REST API.
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Ex-Omni-2D is an omni-modal dialogue framework that generates coordinated text, speech, and reference-conditioned video responses via a visual thought plan and a distilled streaming video generator, achieving a practical quality-efficiency trade-off.
YouTube is making it harder to earn money on YouTube
YouTube is raising the requirements for creators to earn money through its Partner Program, increasing watch time and Shorts view thresholds starting February 1st, 2027. Existing creators must accept the new terms by January 31st, 2027, as YouTube pushes toward a premium TV-style service.
@addyosmani: You know how refreshing the page kills your AI chat mid-response? @triggerdotdev's new chat agent fixes that so it surv…
Trigger.dev's new chat agent lets developers build durable, stateful AI chat experiences that survive refreshes and crashes, with no timeouts and the ability to pause for permission before risky tool calls.