high-throughput

Tag

Cards List
#high-throughput

TigerBeetle Core System Architecture: Deconstructing Performance Engineering

Hacker News Top · 2026-08-21 Cached

The article deconstructs the core system architecture of TigerBeetle, a financial ledger database, focusing on performance engineering techniques like static memory allocation and custom zero-copy interfaces to achieve high throughput and predictable latency.

0 favorites 0 likes
#high-throughput

Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)

Hacker News Top · 2026-08-06 Cached

A deep dive into vLLM's architecture and components for high-throughput LLM inference, covering scheduling, paged attention, continuous batching, advanced features, scaling, serving, and benchmarking.

0 favorites 0 likes
#high-throughput

@OlioLabsInc: We’re bringing predictive validity to drug discovery’s most unpredictable stage. Today we are introducing Olio Labs’ in…

X AI KOLs Following · 2026-08-05 Cached

Olio Labs is introducing its in vivo platform, designed to predict diverse human clinical outcomes from a single high-throughput experiment, bringing predictive validity to drug discovery.

0 favorites 0 likes
#high-throughput

@Tsinghua_Uni: Breaking records across all four key benchmarks! Innovative Tsinghua's @AIRTHU1201 and collaborators have unveiled Gala…

X AI KOLs Timeline · 2026-07-27 Cached

Tsinghua University and AIRTHU introduce GalaxyVS, a method that transforms protein–ligand docking into ultra-large-scale parallel vector retrieval, setting a new global benchmark for high-throughput drug discovery.

0 favorites 0 likes
#high-throughput

@TheAhmadOsman: Laguna S 2.1 118B-A8B on DGX Station using - NVFP4 - FP8 KV Cache Can run 10 parallel agents > with 256k context each >…

X AI KOLs Timeline · 2026-07-25 Cached

Laguna S 2.1 118B-A8B model runs on DGX Station with NVFP4 and FP8 KV Cache, achieving ~1k tokens/second for 10 parallel agents with 256k context each.

0 favorites 0 likes
#high-throughput

Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios)

Hacker News Top · 2026-07-15 Cached

misa77 is a new LZ-based codec that achieves decompression throughput up to 2x faster than LZ4 while also offering better compression ratios. It targets write-once read-many workloads and has constant memory usage.

0 favorites 0 likes
#high-throughput

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

arXiv cs.LG · 2026-07-13 Cached

BlockServe introduces block-grained continuous batching to address convergence heterogeneity in diffusion LLMs, enabling 1.9–10.6x throughput improvement over Fast-dLLM while maintaining generation quality.

0 favorites 0 likes
#high-throughput

LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels

Reddit r/LocalLLaMA · 2026-06-25

LFM2.5 230M model achieves 1,400 tokens per second in-browser using custom WebGPU kernels, demonstrating efficient local inference.

0 favorites 0 likes
#high-throughput

GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz

Hacker News Top · 2026-06-16 Cached

A custom FPGA implementation of a Transformer with KV cache achieves 56,000 tokens per second at 80 MHz, running microGPT on a tiny LCD.

0 favorites 0 likes
#high-throughput

What I learned building low latency and high throughput AI agents

Reddit r/AI_Agents · 2026-06-05

The article shares practical lessons for building low-latency, high-throughput AI agents, including workload estimation, token reduction, parallelism, microservices, and handling LLM failures.

0 favorites 0 likes
#high-throughput

@HotAisle: This is awesome. I wonder who's MI300x they used... ;-)

X AI KOLs Following · 2026-05-29 Cached

Kog announces real-time LLM inference achieving 3000+ output tokens per second per request on standard datacenter GPUs, bringing high-speed inference previously limited to custom silicon to production hardware.

0 favorites 0 likes
#high-throughput

@arcinstitute: Because PerturbSpace uses standard single-cell sequencing, it's compatible with any single-cell readout. In one day, th…

X AI KOLs Timeline · 2026-05-26 Cached

Arc Institute's PerturbSpace enables high-throughput single-cell profiling of transcriptome, location, CRISPR guides, clonal relationships, and surface proteins from many samples in one day, using standard single-cell sequencing.

0 favorites 0 likes
#high-throughput

Holotron-12B - High Throughput Computer Use Agent

Hugging Face Blog · 2026-03-17 Cached

H Company releases Holotron-12B, a multimodal computer-use agent optimized for high-throughput inference using a hybrid SSM architecture. The model, post-trained on NVIDIA Nemotron, demonstrates superior efficiency and scalability for interactive agentic workloads.

0 favorites 0 likes
#high-throughput

Gemini 3.1 Flash-Lite: Built for intelligence at scale

Google DeepMind Blog · 2026-03-03 Cached

Google introduces Gemini 3.1 Flash-Lite, a high-speed, cost-efficient AI model available in preview via Google AI Studio and Vertex API, designed for high-volume developer workloads.

0 favorites 0 likes
← Back to home

Submit Feedback