hybrid-attention

Tag

Cards List
#hybrid-attention

OpenVDN/vdn-minimax-h3

Hugging Face Models Trending · 2026-09-02 Cached

VDN-Minimax-H3 is an open-source hybrid-attention model that speeds up video generation with near-lossless quality, featuring fast inference and plug-and-play adapters powered by MiniMax H3.

0 favorites 0 likes
#hybrid-attention

Qwen/Qwen3.8-Flash-Next-FP8

Hugging Face Models Trending · 2026-08-24 Cached

The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.

0 favorites 0 likes
#hybrid-attention

Qwen/Qwen3.8-Flash-Next

Hugging Face Models Trending · 2026-08-24 Cached

Release of Qwen3.8-Flash-Next, an open-weight AI model introducing architectural innovations like Hybrid Attention with QSA and Gated Residual for improved efficiency and scalability, previewing the future Qwen4 architecture.

0 favorites 0 likes
#hybrid-attention

Retrofitting Linear Attention into Diffusion Language Models

arXiv cs.LG · 2026-08-10 Cached

This paper introduces block-hybrid attention, which combines exact softmax attention within active denoising blocks and linear attention over previous blocks, to accelerate inference in pretrained diffusion language models. The authors retrofit this hybrid attention into LLaDA 2.1, achieving up to 1.7x higher decoding throughput with minimal post-training.

0 favorites 0 likes
#hybrid-attention

inclusionAI/Ling-3.0-flash · Hugging Face

Reddit r/LocalLLaMA · 2026-08-04 Cached

inclusionAI released Ling-3.0-flash, a native hybrid reasoning model with 124B total/5.1B active parameters using a hybrid linear attention architecture (KDA+MLA) and sparse MoE. It matches or outperforms its 1T-class predecessor Ring-2.6-1T while being far more compute-efficient, with built-in agentic and long-context optimizations.

0 favorites 0 likes
#hybrid-attention

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

arXiv cs.AI · 2026-07-29 Cached

PreDiff-LM proposes a hybrid attention mechanism that preserves causal attention for prompt tokens and bidirectional attention for masked target tokens, enabling adaptation of pretrained autoregressive models for discrete masked diffusion language modeling, achieving improvements in perplexity and downstream tasks over prior diffusion baselines.

0 favorites 0 likes
#hybrid-attention

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

arXiv cs.AI · 2026-07-29 Cached

GLIDE introduces a layer-wise adaptive mechanism that strategically integrates sliding-window softmax attention with linear recurrent aggregation for efficient LLM inference, reducing KV cache I/O and latency for long contexts without compromising quality.

0 favorites 0 likes
#hybrid-attention

The Inference Engine Guide for K3 Deployment (10 minute read)

TLDR AI · 2026-07-29 Cached

Kimi K3, a 2.8-trillion-parameter multimodal MoE model with 1M context window, now has day-0 vLLM support. This guide details how vLLM serves K3's novel architecture including Kimi Delta Attention, Attention Residuals, and speculative decoding achieving up to 370 tok/s.

0 favorites 0 likes
#hybrid-attention

@jerryjliu0: If I only went off X posts, I'd think Ramp was an AI lab

X AI KOLs Following · 2026-07-28 Cached

Ramp Labs open-sourced PorTAL, a framework for shared task representations and cross-model LoRA adaptation, supporting hybrid attention models and multimodal systems including Gemma 4, Mistral 7B, and Inkling.

0 favorites 0 likes
#hybrid-attention

Nvidia's New Long-Form Video Generation (12 minute read)

TLDR AI · 2026-07-27 Cached

NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.

0 favorites 0 likes
#hybrid-attention

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Hugging Face Daily Papers · 2026-07-23 Cached

SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.

0 favorites 0 likes
#hybrid-attention

@_akhaliq: Solar Open2 250B just dropped on Hugging Face https://huggingface.co/upstage/Solar-Open2-250B…

X AI KOLs Following · 2026-07-22 Cached

Upstage released Solar Open2 250B, a 250B-parameter mixture-of-experts model with only 15B active parameters, featuring hybrid attention for efficient long-context inference up to 1M tokens, designed for agentic use cases such as tool calling and multi-step reasoning.

0 favorites 0 likes
#hybrid-attention

Upstage 'Solar open2' release. performance on par with DeepSeek V4 Flash.

Reddit r/LocalLLaMA · 2026-07-22 Cached

Upstage released Solar Open 2, a 250B-parameter open-weight MoE model with hybrid attention, delivering performance comparable to DeepSeek V4 Flash while being efficient for agentic workflows.

0 favorites 0 likes
#hybrid-attention

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation

arXiv cs.CL · 2026-06-29 Cached

Proposes a training-free NLL-guided method for selecting which layers to retain full attention in hybrid attention models, achieving comparable accuracy with 1/4 full-attention layers against a 1/2 periodic baseline on long-context tasks.

0 favorites 0 likes
#hybrid-attention

Morphing into Hybrid Attention Models

Hugging Face Daily Papers · 2026-06-29 Cached

FlashMorph is an efficient layer selection method that formulates hybrid layer selection as a budget-constrained optimization problem, using morphable models and linearization regularization to improve long-context efficiency in Transformers.

0 favorites 0 likes
#hybrid-attention

deepseek-ai/DeepSeek-V4-Flash-DSpark

Hugging Face Models Trending · 2026-06-27 Cached

DeepSeek releases V4 series of Mixture-of-Experts language models (Pro 1.6T/49B activated, Flash 284B/13B activated) supporting one-million-token context with hybrid attention and speculative decoding, claiming best open-source model performance.

0 favorites 0 likes
#hybrid-attention

HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization

Hugging Face Daily Papers · 2026-06-18 Cached

HydraHead is a novel attention hybridization architecture that combines Full and Linear Attention at the head level, achieving superior long-context performance with reduced training overhead via interpretability-driven selection and scale-normalized fusion.

0 favorites 0 likes
#hybrid-attention

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation

arXiv cs.CL · 2026-06-17 Cached

ConSA is a framework that learns optimal assignment between full attention and sliding-window attention under a user-specified sparsity target, using L0 regularization and augmented Lagrangian constraint. It demonstrates consistent gains over rule-based baselines on LLMs at 0.6B and 1.7B scales.

0 favorites 0 likes
#hybrid-attention

Chiaroscuro Attention: Spending Compute in the Dark

Hugging Face Daily Papers · 2026-06-06 Cached

CHIAR-Former uses spectral entropy-based routing to dynamically select between DCT, RBF, and self-attention operators, achieving improved efficiency on large text datasets while maintaining performance through hybrid attention mechanisms.

0 favorites 0 likes
#hybrid-attention

InstructSAM: Segment Any Instance with Any Instructions

Hugging Face Daily Papers · 2026-05-25 Cached

InstructSAM presents a unified framework for multi-instance segmentation using instruction-driven queries that bridge vision-language models and SAM3, achieving strong results across complex benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback