kvcache

Tag

Cards List
#kvcache

Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes

Reddit r/LocalLLaMA · 22h ago

The article expresses hope for future AI models with optimizations like KVCache and Engram to reduce memory usage, enabling larger models to run on consumer GPUs with limited VRAM.

0 favorites 0 likes
#kvcache

@akshay_pachaar: RadixAttention, clearly explained. (how SGLang makes prefix caching highly efficient) Prefix caching sounds simple unti…

X AI KOLs Timeline · 2d ago Cached

The article explains RadixAttention, a technique in SGLang that uses a radix tree to efficiently cache and reuse overlapping KV caches across branching requests in AI serving.

0 favorites 0 likes
#kvcache

@davideciffa: Very proud to share that we just release Luce KVFlash. Run your preferred model inside Lucebox at 256k context, without…

X AI KOLs Timeline · 2026-06-14 Cached

Announced release of Luce KVFlash, a tool to run models inside Lucebox at 256k context without KVCache OOM, achieving up to 2.9x faster decoding at long context using speculative prefill and dynamic offloading.

0 favorites 0 likes
#kvcache

@FGuzmanAI: 56,000+ tokens/sec at just 80 MHz. I burned a full Transformer with KV cache into a custom chip. Designed gate by gate …

X AI KOLs Timeline · 2026-06-13 Cached

A custom digital chip designed gate-by-gate achieves over 56,000 tokens/sec running a Transformer with KV cache at just 80 MHz, prototyped on an FPGA.

0 favorites 0 likes
#kvcache

Qwen3.6-35B-A3B tool calling benchmark: ByteShape vs. Unsloth GGUFs, KV cache quants & long context performance

Reddit r/LocalLLaMA · 2026-06-08

A detailed benchmark comparing ByteShape and Unsloth quantizations of Qwen3.6-35B-A3B on tool calling performance, KV cache quantization effects, and long context degradation using llama.cpp and tool-eval-bench.

0 favorites 0 likes
#kvcache

@MaxForAI: http://Z.ai and this ZCube paper from Tsinghua—worth a read for anyone in Infra. Many people's first reaction when talking about AI infra is still GPU, memory, quantization, and inference frameworks. But once you get into long context and Prefill-Decode separation, the network is no longer just a 'supporting role' in the data center. Every...

X AI KOLs Timeline · 2026-05-21

ZCube is a new network architecture that flattens the topology and mixes single/multi-rail access to optimize KV Cache transmission in long-context and PD separation scenarios. In the GLM-5.1 production cluster, it achieved a 33% reduction in switch/optical module costs, a 15% increase in GPU inference throughput, and a 40.6% decrease in TTFT P99.

0 favorites 0 likes
← Back to home

Submit Feedback