Google unveils eighth-generation TPUs, split into TPU 8t (training) and TPU 8i (inference)
At Google Cloud Next 2026 (April) Google announced its first split TPU generation: TPU 8t for training (pods of 9,600 chips, 2 PB shared memory, 121 exaFLOPS) and TPU 8i for inference (288 GB HBM, 80% better perf/$), both up to 2x better performance-per-watt than Ironwood, which became generally available at the same event.
Key facts
- TPU 8t: ~3x compute per pod vs previous generation; scales to 9,600 chips with 2 PB shared memory; 121 ExaFLOPS; >97% goodput target
- TPU 8i: 80% better performance-per-dollar; 288 GB HBM + 384 MB on-chip SRAM; 19.2 Tb/s interconnect for MoE; up to 5x lower on-chip latency
- Both: up to 2x performance-per-watt vs Ironwood (TPU v7)
- Ironwood (v7) GA: 4.6 PFLOPS per chip, 42.5 EFLOPS per 9,216-chip superpod (press figures)
- Press reports: TPU 8t designed with Broadcom and TPU 8i with MediaTek on TSMC 2nm (not confirmed in Google's post)
What happened
Google introduced two purpose-built eighth-generation TPUs at Cloud Next 2026 in Las Vegas, with general availability promised later in 2026 as part of AI Hypercomputer.
Why it matters
Separate training and inference silicon reflects how agentic, long-running inference now dominates compute demand, and strengthens Google's position as the main non-NVIDIA accelerator supplier (Anthropic is reported as an anchor customer).
Changelog
- 2026-09-29: created (exact announcement day inferred from press dated 2026-04-22; confidence medium)
Related events
Sources (3)
- officialGoogle: Our eighth generation TPUs — two chips for the agentic era
- officialGoogle Cloud: TPU 8t and TPU 8i technical deep dive
- pressThe Next Web: Ironwood launches, eighth-gen split previewed
id: 2026-04-22-google-tpu-8t-8i · updated 2026-09-29 · open in the interactive timeline