inference-throughput

Tag

Cards List
#inference-throughput

DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference

Reddit r/singularity · 2026-06-24 Cached

DualPath is a system that breaks the storage bandwidth bottleneck in agentic LLM inference by introducing a dual-path KV-cache loading mechanism, improving throughput by up to 1.87x offline and 1.96x online.

0 favorites 0 likes
#inference-throughput

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

arXiv cs.CL · 2026-05-12 Cached

This paper introduces the Structured Recurrent Mixer (SRM), an architecture enabling algebraic conversion between parallel training and recurrent inference without specialized kernels. Experiments show SRMs achieve significantly higher throughput and concurrency compared to Transformers, with effective performance in reinforcement learning tasks.

0 favorites 0 likes
← Back to home

Submit Feedback