tpu

Tag

Cards List
#tpu

Google reportedly taps AMD to design next-generation TPU — hybrid AI ASIC could integrate on-package CPU cores for reinforcement learning (3 minute read)

TLDR AI ↗ · 2026-08-17 Cached

Google is reportedly collaborating with AMD to design a next-generation TPU that could integrate CPU cores on-package for reinforcement learning workloads, potentially advancing AI hardware architecture.

0 favorites 0 likes
#tpu

Google's TPU Sales to Anthropic Squeeze Its Own Researchers

Reddit r/ArtificialInteligence ↗ · 2026-08-06

Google's aggressive monetization of TPU capacity to outside customers like Anthropic is fueling internal frustration and driving key AI researchers to depart for competitors, intensifying talent and competitive tensions.

0 favorites 0 likes
#tpu

@ycombinator: In 2001, @JeffDean and Sanjay Ghemawat did the math and realized Google’s entire search index would fit in RAM — then s…

X AI KOLs Timeline ↗ · 2026-07-30 Cached

At Startup School 2026, Google's Chief Scientist Jeff Dean recounts the napkin math that led to Google's search index fitting in RAM and the TPU development, discussing inference hardware specialization and how startups can still compete.

0 favorites 0 likes
#tpu

@googledevs: Big news: @Google and @RadixArk are partnering to bring @sgl_project to Google Cloud TPUs! Run SGLang on TPU today via …

X AI KOLs Timeline ↗ · 2026-07-30 Cached

Google Cloud and RadixArk are partnering to bring the SGLang open-source inference framework to Google Cloud TPUs, initially via SGL-JAX and later with SGL-torchtpu for PyTorch-native support, enabling developers to run production workloads seamlessly across GPUs and TPUs.

0 favorites 0 likes
#tpu

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

arXiv cs.AI ↗ · 2026-07-24 Cached

JAXBench is a new benchmark suite of 50 JAX workloads for evaluating AI-generated kernel optimization on Google Cloud TPUs, with hand-tuned baselines and an agent evaluation harness. The paper finds that conditioning on curated TPU documentation significantly improves correctness and speedup, with Autocomp beam-search achieving up to 1.6x geomean speedup over XLA on hand-tuned kernels.

0 favorites 0 likes
#tpu

@googledevs: A major update to Tunix for scaling Agentic RL is here The new asynchronous, decoupled rollout engine solves multi-turn…

X AI KOLs Following ↗ · 2026-07-21 Cached

Google announces a major update to Tunix, its post-training library, with an asynchronous decoupled rollout engine to scale agentic reinforcement learning on JAX/TPU, eliminating idle time and improving throughput.

0 favorites 0 likes
#tpu

An MLIR-Based Compilation Method for Large Language Models

arXiv cs.CL ↗ · 2026-07-20 Cached

This paper presents an MLIR-based compilation method for large language models, using two custom dialects (TopOp and TpuOp) to lower models from framework-agnostic semantics to hardware-specific instructions. It also introduces a three-stage static compilation for autoregressive inference stages: prefill, prefill_kv, and decode.

0 favorites 0 likes
#tpu

@googledevs: Scaling frontier Mixture-of-Experts (MoE) models takes more than trial-and-error tuning Explore how Qwen 3.5-397B was o…

X AI KOLs Following ↗ · 2026-07-15 Cached

Google Cloud details how they optimized Qwen 3.5-397B MoE on Ironwood TPUs using a modular, model-agnostic engineering playbook, achieving 3.1× decode and 4.7× prefill performance gains.

0 favorites 0 likes
#tpu

@googledevs: Build, train, serve. The new TPU Developer Hub is live. Access documentation and framework recipes in one place to buil…

X AI KOLs Following ↗ · 2026-07-15 Cached

Google launched the TPU Developer Hub, a centralized resource with documentation and framework recipes for building, training, and serving AI on Google Cloud TPUs, supporting JAX, PyTorch, and vLLM.

0 favorites 0 likes
#tpu

Google announces Gemma 4 optimized for the Pixel 10's TPU (2 minute read)

TLDR AI ↗ · 2026-07-15 Cached

Google announced Gemma 4 E2B optimized for the Pixel 10's TPU, enabling on-device multimodal AI capabilities like offline chat, image recognition, and audio transcription.

0 favorites 0 likes
#tpu

@ryanlpeterman: David Patterson is a Turing Award winner famous for his contributions to computer architecture. I interviewed him about…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

An interview with Turing Award winner David Patterson covering RISC vs CISC history, GPU/TPU comparisons, Moore's Law, and career advice.

0 favorites 0 likes
#tpu

@vivekgalatage: It's super interesting to know the system architecture of the TPUs. https://henryhmko.github.io/posts/tpu/tpu.html…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

A deep dive into Google's TPU architecture, explaining the design philosophy of systolic arrays, pipelining, and ahead-of-time compilation that enables high throughput and energy efficiency.

0 favorites 0 likes
#tpu

@mkvenkit: Google’s Tensor Processing Unit (TPU) uses the systolic array architecture - an idea from 1978 - to accelerate matrix m…

X AI KOLs Timeline ↗ · 2026-06-27 Cached

Google's TPU uses the systolic array architecture from 1978 to accelerate matrix multiplication with less memory movement. The post shares links to the original paper and TPU design, and suggests building a small-scale version on an FPGA.

0 favorites 0 likes
#tpu

Google Is Using Nvidia's Playbook to Build a Rival AI Chip Business (11 minute read)

TLDR AI ↗ · 2026-06-19

Google is adopting Nvidia's strategy to build a competitive AI chip business, renting TPU computing power to Anthropic and boosting inference performance to rival Nvidia's dominance.

0 favorites 0 likes
#tpu

@ying11231: Impressive performance on TPU.

X AI KOLs Timeline ↗ · 2026-06-17 Cached

A blog post from LMSYS Org details optimizing Ling-2.6-1T, a 1 trillion parameter hybrid MoE model, on TPU v7x using SGLang-JAX, achieving efficient inference by hiding MoE data movement behind computation with a single Pallas kernel.

0 favorites 0 likes
#tpu

Google in talks with Samsung to make part of next-gen chip

Reddit r/singularity ↗ · 2026-06-11

Google is in talks with Samsung to manufacture a component of its next-generation AI chip, codenamed Icefish, using 2-nanometer technology, while the main part will be made by TSMC. The chip aims to offer an alternative to Nvidia's GPUs and is expected to enter mass production as soon as 2028.

0 favorites 0 likes
#tpu

@_philschmid: Google Colab CLI and Skills are out. Full Colab runtimes from your terminal. - GPU/TPU provisioning (colab --gpu A100) …

X AI KOLs Following ↗ · 2026-06-09 Cached

A new CLI tool for Google Colab enables GPU/TPU provisioning, remote script execution, and interactive REPL access from the terminal, with built-in agent skills for automated tasks like fine-tuning models.

0 favorites 0 likes
#tpu

@snowboat84: https://x.com/snowboat84/status/2061962883651731602

X AI KOLs Timeline ↗ · 2026-06-03 Cached

This article is the first part of the AI Engineering Panorama series. From a historical perspective, it reviews the evolution of GPUs from gaming graphics cards to AI accelerators, the bold bet of CUDA, the independent path of Google's TPU, and why NVIDIA ultimately prevailed. It also provides a detailed analysis of the underlying logic of AI infrastructure such as chips, supply chain, networking, and power.

0 favorites 0 likes
#tpu

Midjourney says their research was set back by a year by using TPU, regrets not sticking purely with nvidia

Reddit r/singularity ↗ · 2026-05-20

Midjourney stated that using Google TPUs set their research back by a year, expressing regret for not sticking exclusively with Nvidia hardware.

0 favorites 0 likes
#tpu

@mweinbach: Who said TPUs can't be fast! This is roughly Groq speeds out of TPU 8i, but with Gemini Flash model so far more intelli…

X AI KOLs Timeline ↗ · 2026-05-19 Cached

Google demonstrated Gemini Flash model achieving 600-1400 tokens per second on TPU 8i, rivaling Groq's inference speeds.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback