Tag
Google is reportedly collaborating with AMD to design a next-generation TPU that could integrate CPU cores on-package for reinforcement learning workloads, potentially advancing AI hardware architecture.
Google's aggressive monetization of TPU capacity to outside customers like Anthropic is fueling internal frustration and driving key AI researchers to depart for competitors, intensifying talent and competitive tensions.
At Startup School 2026, Google's Chief Scientist Jeff Dean recounts the napkin math that led to Google's search index fitting in RAM and the TPU development, discussing inference hardware specialization and how startups can still compete.
Google Cloud and RadixArk are partnering to bring the SGLang open-source inference framework to Google Cloud TPUs, initially via SGL-JAX and later with SGL-torchtpu for PyTorch-native support, enabling developers to run production workloads seamlessly across GPUs and TPUs.
JAXBench is a new benchmark suite of 50 JAX workloads for evaluating AI-generated kernel optimization on Google Cloud TPUs, with hand-tuned baselines and an agent evaluation harness. The paper finds that conditioning on curated TPU documentation significantly improves correctness and speedup, with Autocomp beam-search achieving up to 1.6x geomean speedup over XLA on hand-tuned kernels.
Google announces a major update to Tunix, its post-training library, with an asynchronous decoupled rollout engine to scale agentic reinforcement learning on JAX/TPU, eliminating idle time and improving throughput.
This paper presents an MLIR-based compilation method for large language models, using two custom dialects (TopOp and TpuOp) to lower models from framework-agnostic semantics to hardware-specific instructions. It also introduces a three-stage static compilation for autoregressive inference stages: prefill, prefill_kv, and decode.
Google Cloud details how they optimized Qwen 3.5-397B MoE on Ironwood TPUs using a modular, model-agnostic engineering playbook, achieving 3.1× decode and 4.7× prefill performance gains.
Google launched the TPU Developer Hub, a centralized resource with documentation and framework recipes for building, training, and serving AI on Google Cloud TPUs, supporting JAX, PyTorch, and vLLM.
Google announced Gemma 4 E2B optimized for the Pixel 10's TPU, enabling on-device multimodal AI capabilities like offline chat, image recognition, and audio transcription.
An interview with Turing Award winner David Patterson covering RISC vs CISC history, GPU/TPU comparisons, Moore's Law, and career advice.
A deep dive into Google's TPU architecture, explaining the design philosophy of systolic arrays, pipelining, and ahead-of-time compilation that enables high throughput and energy efficiency.
Google's TPU uses the systolic array architecture from 1978 to accelerate matrix multiplication with less memory movement. The post shares links to the original paper and TPU design, and suggests building a small-scale version on an FPGA.
Google is adopting Nvidia's strategy to build a competitive AI chip business, renting TPU computing power to Anthropic and boosting inference performance to rival Nvidia's dominance.
A blog post from LMSYS Org details optimizing Ling-2.6-1T, a 1 trillion parameter hybrid MoE model, on TPU v7x using SGLang-JAX, achieving efficient inference by hiding MoE data movement behind computation with a single Pallas kernel.
Google is in talks with Samsung to manufacture a component of its next-generation AI chip, codenamed Icefish, using 2-nanometer technology, while the main part will be made by TSMC. The chip aims to offer an alternative to Nvidia's GPUs and is expected to enter mass production as soon as 2028.
A new CLI tool for Google Colab enables GPU/TPU provisioning, remote script execution, and interactive REPL access from the terminal, with built-in agent skills for automated tasks like fine-tuning models.
This article is the first part of the AI Engineering Panorama series. From a historical perspective, it reviews the evolution of GPUs from gaming graphics cards to AI accelerators, the bold bet of CUDA, the independent path of Google's TPU, and why NVIDIA ultimately prevailed. It also provides a detailed analysis of the underlying logic of AI infrastructure such as chips, supply chain, networking, and power.
Midjourney stated that using Google TPUs set their research back by a year, expressing regret for not sticking exclusively with Nvidia hardware.
Google demonstrated Gemini Flash model achieving 600-1400 tokens per second on TPU 8i, rivaling Groq's inference speeds.