@PyTorch: How do we get more useful work—not just more tokens—from every AI dollar? This Wednesday at 11:50 AM at @AMD #Advancing…
Summary
PyTorch Foundation CTO Matt White will speak at AMD's AdvancingAI event about optimizing AI inference economics using open-source tools like vLLM and SGLang, advocating for right-sized models and intelligent routing to improve dollar per intelligence.
View Cached Full Text
Cached at: 07/21/26, 06:49 PM
How do we get more useful work—not just more tokens—from every AI dollar? This Wednesday at 11:50 AM at @AMD #AdvancingAI, PyTorch Foundation CTO @matthew_d_white will explore how open source, @vllm_project and @radixark’s @sgl_project are reshaping inference economics.
In “Accelerating Open Source AI: Just Enough Intelligence,” Matt will outline a practical alternative to using frontier models for every task: → Decompose workflows → Select right-sized models → Route and cache intelligently → Escalate only when necessary → Measure cost per successfully completed task
The goal is better Dollar per Intelligence: more correct, business-relevant outcomes from every dollar spent, with failures, retries, and review included in the calculation.
Matt opens a series of talks in the AI Training & Inference track ahead of @simon_mo_, co-founder and CEO of @inferact and lead contributor to vLLM, and @ying11231, co-founder and CEO of RadixArk and co-creator of SGLang.
Moscone West, San Francisco Wednesday, July 22 11:50 AM PDT
Explore the track:
Similar Articles
@AMD: From bring-up to tuning, AMD and @OpenAI engineers are sharing insights to push performance further. Go behind the coll…
AMD and OpenAI engineers collaborate to share insights on performance optimization, featuring a behind-the-scenes look with OpenAI's VP of Compute Strategy, Sachin Katti.
@PyTorch: Interested in optimizing Mixture of Experts LLM inference on ARM CPUs? In his talk at PyTorch Conference North America,…
At PyTorch Conference North America, Maajid Khan from Fujitsu Research India will present a talk on optimizing Mixture of Experts LLM inference on ARM CPUs using vLLM and OpenVINO.
@twimlai: As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens ac…
This podcast episode discusses AI tokenomics and the emerging issue of 'tokenflation', focusing on measuring the value of AI tokens, the limitations of benchmarks, and future innovations in AI efficiency.
@PyTorch: At PyTorch Conference North America 2026, hear directly from and connect with those working on PyTorch, @vllm_project, …
The PyTorch Conference North America 2026 will feature discussions with experts on PyTorch and related projects, focusing on advancements in AI frameworks, inference engines, and hardware integration.
@PyTorch: Discover how open source agentic search and hardware-guided workflows are unlocking massive speedups across GPUs and TP…
The PyTorch Conference North America will be held in San Jose, featuring sessions on agentic search, hardware-guided workflows, and AI performance optimization with speakers from major tech companies.