bidirectional-attention

Tag

Cards List
#bidirectional-attention

@nathanrs: New post! A drawback of diffusion LLMs is that bidirectional attention causes keys and values to drift across steps, br…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

A new post highlights a drawback of diffusion LLMs: bidirectional attention causes keys and values to drift across steps, breaking KV caching. However, generation quality is robust to slight KV drift, and research has focused on maximizing stale KV reuse without quality degradation.

0 favorites 0 likes
#bidirectional-attention

Improved Large Language Diffusion Models

arXiv cs.CL ↗ · 2026-06-25 Cached

iLLaDA is an 8B parameter masked diffusion language model with fully bidirectional attention, trained from scratch on 12T tokens. It shows broad improvements over LLaDA and remains competitive with Qwen2.5 7B on several benchmarks. The model and code are open-sourced.

0 favorites 0 likes
#bidirectional-attention

Why might DiffusionGemma be better at tool calls than its benchmark quality suggests

Reddit r/LocalLLaMA ↗ · 2026-06-16

Analyzes how DiffusionGemma's bidirectional attention and parallel block generation could potentially yield higher valid tool call rates due to its ability to revise tokens, even though its base quality is lower than Gemma 4.

0 favorites 0 likes
#bidirectional-attention

[Talk] Text Diffusion — Google DeepMind's Brendan O’Donoghue

Reddit r/LocalLLaMA ↗ · 2026-06-11 Cached

DeepMind researcher Brendan O'Donoghue provides an in-depth introduction to text diffusion models, which generate text through iterative denoising. Compared to autoregressive models, they offer lower latency but limited throughput, and demonstrate unique advantages such as self-correction and dynamic computation.

0 favorites 0 likes
#bidirectional-attention

Enabling KV Caching of Shared Prefix for Diffusion Language Models

arXiv cs.LG ↗ · 2026-06-09 Cached

This paper proposes BiCache, a novel KV caching technique for shared prefixes in diffusion language models, which avoids accuracy collapse by dynamically reusing cached keys and values in shallow layers and achieves 36.3%–98.3% throughput improvement.

0 favorites 0 likes
← Back to home

Submit Feedback