@h100envy: Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs. She also built Fle…

X AI KOLs Timeline Tools

Summary

Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs, achieving 5x cost cuts over DeepSeek's API; she also built FlexGen and helped build Chatbot Arena.

Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs. She also built FlexGen, which made a 175-billion model run on a single consumer GPU, and helped build Chatbot Arena. Three artifacts the whole field uses, one researcher. SGLang hit a 5x cost cut over DeepSeek's own API, and a dozen teams reproduced it. Everyone argues about models. She builds the engines that actually serve them cheaply enough to survive.
Original Article
View Cached Full Text

Cached at: 06/20/26, 02:38 PM

Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs.

She also built FlexGen, which made a 175-billion model run on a single consumer GPU, and helped build Chatbot Arena.

Three artifacts the whole field uses, one researcher. SGLang hit a 5x cost cut over DeepSeek’s own API, and a dozen teams reproduced it.

Everyone argues about models. She builds the engines that actually serve them cheaply enough to survive.

Similar Articles

Agent-Assisted SGLang Development (18 minute read)

TLDR AI

The article explores how agent tools are being used to encode development workflows for SGLang, turning debugging, benchmarking, and profiling into executable skills and reproducible experiments, with efforts like KDA-Pilot already producing merged PRs.