sparse-attention

Tag

Cards List
#sparse-attention

EditaLive! Unified Character Video Editing for Live Streaming

Hugging Face Daily Papers ↗ · 2026-08-27 Cached

EditaLive is a novel framework for real-time human-centric live-stream video editing that adapts image animation models using causal generation and distillation for efficient streaming inference.

0 favorites 0 likes
#sparse-attention

Qwen4's architecture is here early, firing 6B parameters out of 125B (3 minute read)

TLDR AI ↗ · 2026-08-27 Cached

Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight preview of the Qwen4 architecture that activates only 6B parameters out of 125B to reduce inference costs and address hardware limitations.

0 favorites 0 likes
#sparse-attention

unsloth/GLM-5.3-Flash-GGUF

Hugging Face Models Trending ↗ · 2026-08-26 Cached

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active, utilizing a hybrid sparse-linear attention architecture to reduce costs while outperforming previous versions and approaching Claude Opus 4.8 on benchmarks.

0 favorites 0 likes
#sparse-attention

@threerouter_com: Tonight at 23:00, Alibaba is going to open-source Qwen3.8-Flash-Next on ModelScope, and the news is accurate. But honestly, releasing it without benchmarks, I'm skeptical. The official reason is to adapt for the complete Qwen4 family, meaning 'you try it first, and we'll add the report card later.' The architecture looks impressive—Qwen…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

Alibaba will open-source the Qwen3.8-Flash-Next model on ModelScope at 23:00 tonight, featuring Qwen4's GDN hybrid layers and sparse attention architecture, but lacking benchmark data.

0 favorites 0 likes
#sparse-attention

zai-org/GLM-5.3-Flash

Hugging Face Models Trending ↗ · 2026-08-25 Cached

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters, outperforming previous versions and approaching Claude Opus 4.8 through a redesigned hybrid architecture for improved efficiency.

0 favorites 0 likes
#sparse-attention

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

arXiv cs.LG ↗ · 2026-08-24 Cached

The paper introduces BF1, a causal dyadic sparse-attention retrofit designed to improve the efficiency of transformers for long-context processing.

0 favorites 0 likes
#sparse-attention

Sparse Token Routing in Efficient Transformers

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper evaluates adaptive computation in transformers using SEWN, a two-stream model with a learned gate for token routing, demonstrating that SEWN-sparse achieves significant throughput improvements over BERT-base and DistilBERT with modest accuracy trade-offs while providing interpretable token-importance signals.

0 favorites 0 likes
#sparse-attention

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper presents a novel method for fine-tuning transformer language models with sparse attention to enable efficient long-context inference, often outperforming models trained with exact attention, and introduces an efficient implementation and a new open-source library.

0 favorites 0 likes
#sparse-attention

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

FlashPrefill V2 improves long-context LLM serving through mean-corrected sparse attention and optimized GPU operators, delivering substantial speedups over FlashAttention-2 and dense baselines.

0 favorites 0 likes
#sparse-attention

Did anyone Tried making a loop LM with exit gate, sparced, compressed and highly compressed attention and layer attention with diffusion optimize?

Reddit r/ArtificialInteligence ↗ · 2026-08-20

The author describes building an experimental Loop Language Model with exit gates, sparse/compressed attention, and diffusion-based optimization, trained on limited hardware with preliminary results showing potential efficiency gains.

0 favorites 0 likes
#sparse-attention

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

This paper introduces Daedalus-150M, a hybrid language model combining convolution and attention mechanisms optimized for CPU inference, achieving better benchmark performance than larger models with significantly less training data.

0 favorites 0 likes
#sparse-attention

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Hugging Face Daily Papers ↗ · 2026-08-19 Cached

SparsePR is a training-free method that accelerates video transformers by using response-coupled partitioning and probe-fitted residual reconstruction to reduce attention error while maintaining generation quality.

0 favorites 0 likes
#sparse-attention

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

EditBridge is a diffusion bridge framework that enables efficient ultra-high-resolution image editing up to 4K by translating low-resolution edits to high-resolution outputs with sparse attention, achieving significant speed improvements and preserving source details.

0 favorites 0 likes
#sparse-attention

🤡 How to make any Sparse Attention / KV Compression look good? 🤡

Reddit r/ArtificialInteligence ↗ · 2026-08-17

The article critiques common practices in AI research to make sparse attention and KV cache compression techniques appear more effective than they might be, by exploiting benchmark flaws and implementation details.

0 favorites 0 likes
#sparse-attention

Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry

arXiv cs.CL ↗ · 2026-08-10 Cached

AoH is a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections, enabling sparse attention without runtime attention scores. At 50% sparsity it retains 96.5% of full-attention performance while reducing prefill/decode latency and KV-cache memory.

0 favorites 0 likes
#sparse-attention

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

Hugging Face Daily Papers ↗ · 2026-08-08 Cached

OasisKV is a memory-centric LLM inference system that decouples full KV-cache storage from HBM by prefetching sparse, important KV blocks using lookahead tokens from speculative decoding, achieving up to 2.1x throughput gains over dense vLLM with minimal accuracy loss.

0 favorites 0 likes
#sparse-attention

Training-Free Hashing-Based Attention via Binary Principal Components

arXiv cs.LG ↗ · 2026-08-06 Cached

Introduces BinaryPC, a training-free hashing-based sparse attention method for long-context LLMs that uses binary principal components to construct hash codes, preserving accuracy while improving decoding throughput by 3.56x over FlashAttention.

0 favorites 0 likes
#sparse-attention

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

arXiv cs.LG ↗ · 2026-08-05 Cached

ATFlash introduces a per-RoPE-wavelength distance window that prunes query-key inner-product terms proportional to each frequency pair's wavelength, cutting 37-48% of attention compute with minimal quality loss and up to 1.31x speedups on long-context LLM inference.

0 favorites 0 likes
#sparse-attention

When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper proposes LTGA, a graph attention layer that learns a per-edge Tsallis entropic index to interpolate between heavy-tailed, softmax, and compact-support attention, offering interpretable sparse attention and competitive performance on graph benchmarks.

0 favorites 0 likes
#sparse-attention

@ModelScope2022: 1M-token context with only ~3B parameters active per token. Meituan’s LongCat-Flash-Lite-Sparse brings sparse attention…

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Meituan released LongCat-Flash-Lite-Sparse, a sparse-attention model supporting 1M-token context with only ~3B active parameters per token, achieving strong SWE-Bench scores under an MIT license.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback