layerwise-adaptation

Tag

Cards List
#layerwise-adaptation

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

arXiv cs.AI · 4d ago Cached

GLIDE introduces a layer-wise adaptive mechanism that strategically integrates sliding-window softmax attention with linear recurrent aggregation for efficient LLM inference, reducing KV cache I/O and latency for long contexts without compromising quality.

0 favorites 0 likes
← Back to home

Submit Feedback