Tag
The article deconstructs the core system architecture of TigerBeetle, a financial ledger database, focusing on performance engineering techniques like static memory allocation and custom zero-copy interfaces to achieve high throughput and predictable latency.
A deep dive into vLLM's architecture and components for high-throughput LLM inference, covering scheduling, paged attention, continuous batching, advanced features, scaling, serving, and benchmarking.
Olio Labs is introducing its in vivo platform, designed to predict diverse human clinical outcomes from a single high-throughput experiment, bringing predictive validity to drug discovery.
Tsinghua University and AIRTHU introduce GalaxyVS, a method that transforms protein–ligand docking into ultra-large-scale parallel vector retrieval, setting a new global benchmark for high-throughput drug discovery.
Laguna S 2.1 118B-A8B model runs on DGX Station with NVFP4 and FP8 KV Cache, achieving ~1k tokens/second for 10 parallel agents with 256k context each.
misa77 is a new LZ-based codec that achieves decompression throughput up to 2x faster than LZ4 while also offering better compression ratios. It targets write-once read-many workloads and has constant memory usage.
BlockServe introduces block-grained continuous batching to address convergence heterogeneity in diffusion LLMs, enabling 1.9–10.6x throughput improvement over Fast-dLLM while maintaining generation quality.
LFM2.5 230M model achieves 1,400 tokens per second in-browser using custom WebGPU kernels, demonstrating efficient local inference.
A custom FPGA implementation of a Transformer with KV cache achieves 56,000 tokens per second at 80 MHz, running microGPT on a tiny LCD.
The article shares practical lessons for building low-latency, high-throughput AI agents, including workload estimation, token reduction, parallelism, microservices, and handling LLM failures.
Kog announces real-time LLM inference achieving 3000+ output tokens per second per request on standard datacenter GPUs, bringing high-speed inference previously limited to custom silicon to production hardware.
Arc Institute's PerturbSpace enables high-throughput single-cell profiling of transcriptome, location, CRISPR guides, clonal relationships, and surface proteins from many samples in one day, using standard single-cell sequencing.
H Company releases Holotron-12B, a multimodal computer-use agent optimized for high-throughput inference using a hybrid SSM architecture. The model, post-trained on NVIDIA Nemotron, demonstrates superior efficiency and scalability for interactive agentic workloads.
Google introduces Gemini 3.1 Flash-Lite, a high-speed, cost-efficient AI model available in preview via Google AI Studio and Vertex API, designed for high-volume developer workloads.