Tag
EditaLive is a novel framework for real-time human-centric live-stream video editing that adapts image animation models using causal generation and distillation for efficient streaming inference.
Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight preview of the Qwen4 architecture that activates only 6B parameters out of 125B to reduce inference costs and address hardware limitations.
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active, utilizing a hybrid sparse-linear attention architecture to reduce costs while outperforming previous versions and approaching Claude Opus 4.8 on benchmarks.
Alibaba will open-source the Qwen3.8-Flash-Next model on ModelScope at 23:00 tonight, featuring Qwen4's GDN hybrid layers and sparse attention architecture, but lacking benchmark data.
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters, outperforming previous versions and approaching Claude Opus 4.8 through a redesigned hybrid architecture for improved efficiency.
The paper introduces BF1, a causal dyadic sparse-attention retrofit designed to improve the efficiency of transformers for long-context processing.
This paper evaluates adaptive computation in transformers using SEWN, a two-stream model with a learned gate for token routing, demonstrating that SEWN-sparse achieves significant throughput improvements over BERT-base and DistilBERT with modest accuracy trade-offs while providing interpretable token-importance signals.
This paper presents a novel method for fine-tuning transformer language models with sparse attention to enable efficient long-context inference, often outperforming models trained with exact attention, and introduces an efficient implementation and a new open-source library.
FlashPrefill V2 improves long-context LLM serving through mean-corrected sparse attention and optimized GPU operators, delivering substantial speedups over FlashAttention-2 and dense baselines.
The author describes building an experimental Loop Language Model with exit gates, sparse/compressed attention, and diffusion-based optimization, trained on limited hardware with preliminary results showing potential efficiency gains.
This paper introduces Daedalus-150M, a hybrid language model combining convolution and attention mechanisms optimized for CPU inference, achieving better benchmark performance than larger models with significantly less training data.
SparsePR is a training-free method that accelerates video transformers by using response-coupled partitioning and probe-fitted residual reconstruction to reduce attention error while maintaining generation quality.
EditBridge is a diffusion bridge framework that enables efficient ultra-high-resolution image editing up to 4K by translating low-resolution edits to high-resolution outputs with sparse attention, achieving significant speed improvements and preserving source details.
The article critiques common practices in AI research to make sparse attention and KV cache compression techniques appear more effective than they might be, by exploiting benchmark flaws and implementation details.
AoH is a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections, enabling sparse attention without runtime attention scores. At 50% sparsity it retains 96.5% of full-attention performance while reducing prefill/decode latency and KV-cache memory.
OasisKV is a memory-centric LLM inference system that decouples full KV-cache storage from HBM by prefetching sparse, important KV blocks using lookahead tokens from speculative decoding, achieving up to 2.1x throughput gains over dense vLLM with minimal accuracy loss.
Introduces BinaryPC, a training-free hashing-based sparse attention method for long-context LLMs that uses binary principal components to construct hash codes, preserving accuracy while improving decoding throughput by 3.56x over FlashAttention.
ATFlash introduces a per-RoPE-wavelength distance window that prunes query-key inner-product terms proportional to each frequency pair's wavelength, cutting 37-48% of attention compute with minimal quality loss and up to 1.31x speedups on long-context LLM inference.
This paper proposes LTGA, a graph attention layer that learns a per-edge Tsallis entropic index to interpolate between heavy-tailed, softmax, and compact-support attention, offering interpretable sparse attention and competitive performance on graph benchmarks.
Meituan released LongCat-Flash-Lite-Sparse, a sparse-attention model supporting 1M-token context with only ~3B active parameters per token, achieving strong SWE-Bench scores under an MIT license.