SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

Hugging Face Daily Papers Papers

Summary

SoL-Refiner is a one-step video refiner that transforms low-resolution video outputs into 4K resolution with a single denoising step, achieving significant speed improvements and outperforming existing refineries in quality metrics.

High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at 3840!times!2176 it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an 8.91times speedup in refinement latency over the same baseline in our 2K latency setting.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:19 AM

Paper page - SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

Source: https://huggingface.co/papers/2609.37969 Published on Sep 29

·

Submitted byhttps://huggingface.co/Owen777

Owenon Sep 30

Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

High-resolutionvideogenerationisexpensive,asitscostgrowsrapidlywiththenumberofspatiotemporaltokens.Apracticalalternativefirstgeneratesalower-resolutionvideoandthenappliesarefiner,butconventionalmulti-steprefinementintroducesasecondsamplingbottleneck.WepresentSoL-Refiner,aone-stepvideorefinerthattransformslow-resolutionmodeloutputsinto4Kvideoswithasingledenoisingstep.Ourthree-stagerecipecombineshigh-resolutioncontinualtraining,reinforcementlearning(RL)post-training,andafinalone-stepdistillation.WeintroduceRefiner-Bench,avideorefinementbenchmarkconstructedfromtheoutputsofdifferentvideogenerators,anduseashared-inputprotocoltocomparerefinersatapproximately2Koutputresolution.At2K,theone-stepSoL-RefineroutperformsallexternalrefinersontheVBenchandUniPerceptaverages,whileat3840!times!2176itimprovesbothmetricsoverthethree-stepLTX-2.3Refiner.Withthecompleteaccelerationstack,SoL-Refinerachievesan8.91timesspeedupinrefinementlatencyoverthesamebaselineinour2Klatencysetting.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.37969

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37969 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37969 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37969 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

LiteFrame Scales Video LLM Efficiency (6 minute read)

TLDR AI

LiteFrame introduces a highly efficient video encoder for Video LLMs that uses Compressed Token Distillation to enable up to 8x more frames and 35% latency reduction while maintaining accuracy, setting a new Pareto frontier for long-form video understanding.