transformer-inference

Tag

Cards List
#transformer-inference

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv cs.AI · 2026-07-08 Cached

DepthWeave-KV is a token-adaptive cross-layer residual factorization method for compressing KV cache in long-context transformer inference, achieving 8.3x memory reduction and 72.8 tokens/s at 64K context while preserving near-full-cache task quality across benchmarks.

0 favorites 0 likes
#transformer-inference

@bojie_li: If you run LLM agents in production you know the tax. Every turn re-reads the same long context — a system policy, tool…

X AI KOLs Timeline · 2026-07-06 Cached

Researchers introduce Programmable KV Cache, a method for editing and composing KV caches to avoid re-prefilling long contexts during LLM agent inference, achieving 53–398× reduction in p90 time-to-first-token while maintaining decision identity.

0 favorites 0 likes
#transformer-inference

@Michaelzsguo: Today Etched made a high-profile entry into the public eye. What this company wants to do is not yet another GPU replacement, but an AI chip designed specifically for Transformer inference, along with a complete inference cluster. The investors and backers behind it are practically a Who’s Who of the AI industry…

X AI KOLs Timeline · 2026-06-30 Cached

Etched made a high-profile debut, announcing an AI chip and full inference cluster built specifically for Transformer inference. It has secured over $1 billion in customer contracts and $800 million in funding, with the first cabinet set to ship this summer.

0 favorites 0 likes
#transformer-inference

@FGuzmanAI: 56,000+ tokens/sec at just 80 MHz. I burned a full Transformer with KV cache into a custom chip. Designed gate by gate …

X AI KOLs Timeline · 2026-06-13 Cached

A custom digital chip designed gate-by-gate achieves over 56,000 tokens/sec running a Transformer with KV cache at just 80 MHz, prototyped on an FPGA.

0 favorites 0 likes
#transformer-inference

@charles_irl: Tried to squeeze the most important bits about the entire stack for cloud deployment of transformer inference, from app…

X AI KOLs Following · 2026-06-10 Cached

This article provides a comprehensive overview of the complete technology stack for cloud deployment of Transformer inference, covering application scenarios, workload definition, models, inference engines, hardware, observability, and performance optimization, along with future trends.

0 favorites 0 likes
← Back to home

Submit Feedback