← Frontier Inference Margins · all research reports
RECOVERED from ChatGPT history — dispatcher timer died before filing (commissioned 2026-07-15, Targeted 4). Recovered from the ChatGPT conversation history after the dispatching session failed. No new prompt was sent; this is the verbatim assistant report as it already existed in-account. The dispatching (pro-dive-fleet) session's in-session MCP request registry expired unread, but the ChatGPT conversation itself survived.
| Field | Value |
|---|---|
| Model | gpt-5-6-pro (from message metadata.model_slug) |
| Date | 2026-07-15 |
| Dive title | TPU Inference Economics Analysis |
| Created / final answer (UTC) | 2026-07-15T20:33:52.159Z / 2026-07-15T21:31:11.150Z |
Citation anchors removed: citation anchors from the original interface were removed for publication; original preserved in the archive copy.
Provenance routing: any site change this dive justifies must enter through the update queue
(research/update-queue.md), not via direct edits.
You are conducting a rigorous source-hunting deep dive for a quantitative model of frontier-LLM inference serving economics (cost per token, and margin at list prices). I need the strongest available PUBLIC anchors for Google TPU inference/serving economics — or a rigorous, well-documented finding that no such public anchor exists.
SCOPE:
- TPU v5e, v6e (Trillium), and v7 (Ironwood) used for LLM *inference/serving*, not training.
- Published inference throughput on frontier-class models: MLPerf Inference submissions (v4.x/v5.x rounds) on TPU; Google's own published serving claims (Gemini serving on TPU, JetStream / MaxText / vLLM-on-TPU benchmarks); and any third-party TPU inference benchmarks (via GKE, JetStream, SGLang, etc.). Prefer figures expressed as tokens/sec/chip (or per host) on a named model at a named precision.
- GCP TPU pricing: on-demand $/chip-hour and $/host-hour, 1-yr and 3-yr committed-use discounts, spot/preemptible rates, for v5e / v6e / v7 where published.
- Power and TCO comparisons vs NVIDIA H100/H200/GB200 (perf-per-watt, perf-per-dollar), from Google or credible third parties.
- Any disclosures about Gemini serving cost or TPU fleet economics.
GOAL: a candidate quantitative anchor for cost-per-token on TPU — ideally (tokens/sec/chip on a named model at a named precision) paired with ($/chip-hour) so a $/token can be derived. If no such public serving anchor exists (e.g. Google publishes peak FLOPS and rental pricing but no real-model serving throughput), say so explicitly and precisely — a rigorous negative result is exactly as valuable here. TPU is currently the single weakest-anchored platform in the model, so precision about what is and is not publicly disclosed matters more than breadth.
Structure your findings as dated claims. For each: the claim, primary source (URL + date), your assessment of source quality (vendor claim / third-party benchmark / production disclosure / analyst estimate), and whether it could serve as a quantitative anchor (tokens/sec/chip, $/hr, $/token). If no public anchor exists for a platform, say so explicitly as a finding — a rigorous negative result is wanted. End with a numbered list of your 5 most load-bearing sources.
A public cost-per-output-token anchor now exists for TPU v5e, v6e, and v7/Ironwood. The evidence is materially better than a model built only from peak FLOPS:
| Platform | Public named-model throughput? | Public rental price? | Defensible accelerator-rental $/token? | Main limitation |
|---|---|---|---|---|
| TPU v5e | Yes | Yes | Yes | Strongest large-model result is Llama 2 70B; several results are vendor saturation benchmarks. |
| TPU v6e / Trillium | Yes | Yes | Yes | Best vendor numbers sometimes omit precision; the best independent comparison is a 2026 preprint rather than an audited benchmark. |
| TPU v7 / Ironwood | Yes, as of 2026 | Yes | Yes | The useful results are new Google reference recipes, not MLPerf or independent production disclosures; no latency-qualified result accompanies the headline throughput. |
| Gemini production serving | No model-level rate | Internal fleet cost not public | No | Google discloses aggregate energy per Gemini prompt, but not model, TPU generation, token lengths, throughput, utilization, or dollars. |
The strongest frontier-scale public anchor I found is Qwen3-Coder-480B-A35B-Instruct-FP8 on four Ironwood chips:
That is a real named-model, named-precision, named-topology serving benchmark. It is nevertheless a maximum-throughput vendor reference run at infinite offered load, not a production cost disclosure or a latency-constrained serving result.
The rigorous negative finding is narrower but important:
There is no public basis for calculating Gemini gross margin from TPU economics. Public TPU rental prices are external customer list prices, not Google’s internal TPU fleet cost, and Google has not published a mapping from a named Gemini API SKU to TPU generation, topology, precision, output throughput, traffic mix, utilization, or supporting-model overhead.
For a benchmark reporting aggregate output-token throughput, I use:
\frac{ N_{\text{chips}}\times $/{\text{chip-hour}} }{ \text{output tokens/sec}\times 3600 } \times 10^6 ]
This is an accelerator-rental floor, not fully loaded serving COGS. It excludes:
Where a benchmark has a specified input length, its reported output-token rate implicitly includes the cost of processing those inputs. It should therefore be read as cost per output token for that particular input/output workload, not as a universal decode-only cost.
All calculated prices below use the current lowest listed US price for each generation as of July 15, 2026. Pairing a current price with an older benchmark is a current repricing of the benchmark, not necessarily the economics prevailing when it was published.
Primary source: https://cloud.google.com/tpu/pricing — accessed 2026-07-15
Prices are USD per chip-hour. Google explicitly says the TPU price table is expressed per chip-hour even where billing interfaces may display VM-hours.
| Generation | On-demand | DWS Flex | DWS Calendar | 1-year commitment | 3-year commitment | Useful host equivalents |
|---|---|---|---|---|---|---|
| v5e | $1.20 | $0.60 | $0.84 | $0.84 | $0.54 | v5e-8: $9.60 on-demand; $4.32 at 3 years |
| v6e | $2.70 | $1.35 | $1.89 | $1.89 | $1.22 | v6e-4: $10.80 on-demand; v6e-8: $21.60 |
| v7 / Ironwood | $12.00 | $6.00 | $8.40 | $8.40 | $5.40 | Current standard four-chip VM: $48 on-demand; $21.60 at 3 years |
Spot/preemptible: Google’s current public TPU pricing page does not provide a stable numeric spot rate. It says Spot prices are dynamic and can change as frequently as once every 30 days. DWS Flex is a separately quoted queue-based rate and should not be relabeled as Spot. Therefore, a timeless numeric “TPU spot price” is not available from the public static price card.
Current topology documentation says:
The table uses output-token throughput, not total input-plus-output throughput.
| Platform and model | Configuration and workload | Published output throughput | Derived $/1M output tokens: on-demand / DWS Flex / 3-year | Anchor assessment |
|---|---|---|---|---|
| v5e — Llama 2 13B | v5e-8; INT8 weights, activations, and KV cache; maximum input/output 1,024 | 3,938 tok/s | $0.677 / $0.339 / $0.305 | Clean, precisely quantized vendor anchor; model is small by current frontier standards. |
| v5e — Llama 2 70B | Eight chips; INT8 weight quantization; ShareGPT variable-length workload | 1,510 tok/s | $1.766 / $0.883 / $0.795 | Best straightforward public 70B v5e anchor; activation/KV precision is not fully specified. |
| v6e — Llama 2 70B | v6e-8; MaxText/JetStream; maximum input/output 1,024 | 6,389 tok/s | $0.939 / $0.470 / $0.424 | Very strong throughput figure, but precision and latency/load point are omitted. Treat as optimistic saturation floor. |
| v6e — Qwen3-32B | v6e-4; W8A8 INT8; 1,800 input / 128 output; 1,000 requests | 1,396.35 tok/s | $2.148 / $1.074 / $0.971 | Exact and reproducible, but mean TTFT was 44.4 seconds; relevant to batch/high-throughput economics, not interactive serving. |
| v6e — Gemma 4 31B | v6e-8; 4,096 input / 512 output; TPU KV cache FP8 | 1,206 tok/s | $4.975 / $2.488 / $2.248 | Best independent TPU anchor found; small sample and not perfectly precision-matched to the H100 comparison. |
| v7 — Qwen3-32B | One-chip-normalized recipe; FP8 weights, activations, and KV; 1K/1K; concurrency 320 | 6,556.72 tok/s/chip | $0.508 / $0.254 / $0.229 | Excellent normalized vendor figure, but the recipe’s one-chip machine type conflicts with current four-chip-minimum topology documentation. |
| v7 — Qwen3-Coder-480B-A35B | Four chips; FP8 model and KV; 1K/8K; concurrency 64 | 2,075.47 total; 518.86/chip | $6.424 / $3.212 / $2.891 | Best frontier-scale public TPU anchor. Exact model, topology, precision, and workload; vendor saturation benchmark with no published achieved latency. |
The v5e figures come from Google’s 2024 JetStream and Hex-LLM publications; v6e figures come from Google’s 2025–2026 JetStream/vLLM work and an independent Gemma 4 paper; Ironwood figures come from Google’s 2026 AI Hypercomputer reference recipes.
Claim. GCP publishes stable on-demand, DWS Flex, DWS Calendar, one-year, and three-year per-chip-hour rates for v5e, v6e, and Ironwood. Ironwood is currently $12/chip-hour on demand; v6e is $2.70; v5e is $1.20. Three-year rates are $5.40, $1.22, and $0.54 respectively.
Primary source and date. https://cloud.google.com/tpu/pricing — accessed 2026-07-15
Source quality. Official commercial price card; strongest possible source for external list-price economics.
Quantitative-anchor status. Yes, for the $/hour leg. It does not disclose Google’s internal fleet cost. Spot remains a dynamic price rather than a stable published number.
Claim. Google reported JetStream on a v5e-8 slice with continuous batching and INT8 weights, activations, and KV cache. With maximum input and output lengths of 1,024 tokens, it reported:
The original publication also gave an explicit three-year-CUD estimate of approximately $0.30 per million output tokens for Gemma 7B and Llama 2 13B and $0.25 for Llama 2 7B.
Primary source and date. https://cloud.google.com/blog/products/compute/accelerating-ai-inference-with-google-cloud-tpus-and-gpus — 2024-04-10
Source quality. Vendor benchmark. The workload, slice size, continuous-batching mode, and quantization are unusually well specified. It is not independently reproduced and represents a maximum-throughput operating point.
Quantitative-anchor status. Yes. Repricing Llama 2 13B at today’s rates gives $0.677/M output tokens on demand and $0.305/M under the current three-year rate. This is probably the cleanest v5e anchor, although not a frontier-size model.
Claim. Google reported Hex-LLM on eight v5e chips using a ShareGPT variable-length workload. At its highest-throughput point:
Google separately reported low-load per-output-token latencies, but the peak-throughput and minimum-latency figures are different points on the load curve and must not be combined into one serving configuration.
Primary source and date. https://cloud.google.com/blog/products/ai-machine-learning/hex-llm-on-tpus-in-vertex-ai-model-garden/ — 2024-07-26
Source quality. Vendor benchmark with a realistic variable-length dataset and a named 70B model. The source specifies INT8 weight quantization but does not fully specify activation and KV-cache precision.
Quantitative-anchor status. Yes. At current list pricing, 1,510 tok/s on a $9.60/hour v5e-8 slice implies:
This is the best simple public v5e anchor for a large dense model.
Claim. Google did submit a TPU v5e LLM result to MLPerf Inference v4.0, but it is not a strong cost-per-token anchor:
The variable generated length means samples/s cannot be converted exactly to output tokens/s from the summary result. GPT-J 6B also does not meet the requested frontier-class criterion.
Primary sources.
Source quality. Auditable MLCommons submission; very high benchmark integrity, but the reported unit and workload are unsuitable for exact $/output-token conversion.
Quantitative-anchor status. No exact $/token anchor. It can support a cost-per-sample calculation—about $0.000131 per offline sample at current on-demand v5e pricing—but not exact cost per generated token.
In the Google closed-submission directories inspected:
Therefore, MLPerf Inference v4.0 through v5.1 does not provide the desired frontier TPU LLM tokens/sec/chip anchor.
Claim. At the Trillium announcement, Google claimed 4.7 times the peak compute per chip of v5e and “over 67%” greater energy efficiency. It also said Gemini 1.5 Flash, Imagen 3, and Gemma 2 were trained and served using TPUs, but did not attribute a named model’s production serving rate to a specified Trillium topology.
Primary source and date. https://cloud.google.com/blog/products/compute/introducing-trillium-6th-gen-tpus — 2024-05-14
Source quality. Vendor generation-level specification and production-use statement.
Quantitative-anchor status. No. It does not state the LLM, precision, tokens/s, watts, utilization, or latency corresponding to the energy-efficiency claim. It is a useful architectural prior, not a serving tokens/J measurement.
Claim. Google reported the following MaxText/JetStream output throughput, with maximum input and output lengths of 1,024:
The article states that JetStream uses the same inference stack Google uses to serve Gemini.
Primary source and date. https://cloud.google.com/blog/products/compute/ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu — 2025-05-09
Source quality. Vendor benchmark with named models, chip counts, and workloads. The notable omission is precision/quantization. Load level, TTFT, and output latency are also not reported.
Quantitative-anchor status. Yes, with a precision caveat.
For Llama 2 70B:
This is the lowest public v6e 70B accelerator-floor estimate I found, but it should be modeled as an optimistic maximum-throughput case because the precision and latency point are undisclosed.
The same article reported 1,703 tok/s for Llama 3.1 405B and a threefold inference-per-dollar improvement over v5e. However, it did not disclose the Trillium chip count/topology or precision for that figure. Consequently, the 405B result cannot be converted into $/token and is not a valid frontier anchor despite being the most frontier-relevant model in the article.
Claim. A Google-authored paper on Ragged Paged Attention reports up to 86% memory-bandwidth utilization in decode and 73% model-FLOP utilization in prefill on Ironwood. A fixed-workload software-progress figure also reports, for 1,800 input and 128 output tokens in BF16:
Because every request has 128 output tokens, these imply:
Primary source and date. https://arxiv.org/abs/2604.15464 — submitted 2026-04-16
Source quality. Vendor-authored research preprint. Precision, model, topology, and fixed output length are specified; the chart is primarily a history of software improvements rather than a complete latency-qualified serving benchmark.
Quantitative-anchor status. Yes, as a conservative technical anchor. Current on-demand accelerator floors are approximately $0.505/M for the 8B run and $6.13/M for the 70B run. The large difference from the 2025 JetStream 70B figure illustrates how strongly TPU $/token depends on serving engine, quantization, batching, and benchmark workload.
The 86%/73% utilization results are not themselves cost anchors because they do not disclose end-to-end tokens/s for the corresponding Ironwood runs.
Claim. Google’s vLLM-on-TPU tutorial used:
It reported:
Primary source and date. https://cloud.google.com/blog/ja/products/infrastructure/lets-try-llm-inference-with-cloud-tpu-and-vllm — 2026-05-01
Source quality. Official, reproducible vendor tutorial with a named model, exact quantization, topology, workload, throughput, and latency metrics.
Quantitative-anchor status. Yes—but specifically for high-throughput/batch economics.
Using output tokens:
Using the published total-token throughput instead would produce $0.143/M total processed tokens on demand, but that is a different denominator and should not be compared directly with API output-token prices.
The 44-second mean TTFT makes this unsuitable as a low-latency interactive serving anchor.
Claim. An independent preprint compared Gemma 4 31B on a v6e-8 system with the same model on two H100 GPUs. At comparable hourly rental prices, reported output throughput included:
| Input/output workload | v6e-8 | 2×H100 |
|---|---|---|
| 512 / 256 | 1,403 tok/s | 1,490 tok/s |
| 1,024 / 512 | 1,404 tok/s | 1,387 tok/s |
| 4,096 / 512 | 1,206 tok/s | 728 tok/s |
| 8,192 / 512 | 482 tok/s | 449 tok/s |
| 16,384 / 512 | 474 tok/s | 326 tok/s |
The paper’s own cost table put the TPU at approximately $4.27/M output tokens for the shortest workload and $4.95/M at 4,096/512, versus $4.13/M and $8.44/M respectively for the H100 pair.
Primary source and date. https://arxiv.org/abs/2605.25645 — submitted 2026-05-25; version 3 dated 2026-06-05
Source quality. Independent third-party benchmark/preprint with scripts and multiple sequence lengths. It is not peer-reviewed or audited, and the runs used only 30 prompts per QPS point.
A significant comparability caveat is that the TPU used FP8 KV cache while the GPU path used BF16 KV cache, and the two paths used different software versions. Thus it is not a perfect hardware-only comparison.
Quantitative-anchor status. Yes. At the current official $21.60/hour v6e-8 rate:
This is the strongest independent evidence I found for TPU serving performance per dollar versus NVIDIA. It compares against H100, not H200 or GB200.
Claim. Google announced Ironwood as an inference-oriented TPU, with 2,307 BF16 TFLOPS/chip, 4,614 FP8 TFLOPS/chip, 192 GB HBM, and 7.38 TB/s HBM bandwidth. It claimed twice Trillium’s performance per watt.
Primary sources and dates.
Source quality. Official vendor hardware specifications and peak-efficiency claim.
Quantitative-anchor status. No, by themselves. The performance-per-watt footnote is based on peak FP8 FLOPS divided by TDP, not measured end-to-end LLM output tokens per joule. Peak FLOPS, bandwidth, and rental price are insufficient to infer serving $/token without a real-model utilization and serving benchmark.
This was the correct negative conclusion at Ironwood’s launch. The result changed only when Google began publishing the 2026 reference-recipe benchmarks below.
Claim. Google’s AI Hypercomputer recipe reports Qwen3-32B on Ironwood with:
Primary source and date. https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/ironwood/vLLM/Qwen3-32B — throughput update merged 2026-04-24
The relevant commit explicitly updated the per-chip values to 6,556.72, 3,959.56, and 1,382.18 tok/s.
Source quality. Google-controlled vendor reference recipe using the open InferenceX client pinned to a specific commit. Highly reproducible, but not independent and clearly optimized for saturation throughput.
Quantitative-anchor status. Yes, on a normalized per-chip basis.
For the 1K/1K result:
For 1K/8K, the corresponding values are approximately $0.842/M, $0.421/M, and $0.379/M.
Material topology caveat. The recipe uses a tpu7x-standard-1t one-chip configuration, while current TPU7x documentation describes standard four-chip full-host VMs. This may reflect a preview or superseded machine type. The reported per-chip throughput remains a useful normalized engineering result, but a current customer may face a four-chip minimum bill. Unless four independent model replicas can fill that VM, the actual deployment cost can exceed the simple per-chip estimate by as much as fourfold.
Claim. Google’s Ironwood vLLM recipe reports:
Published rates:
| Workload | Aggregate output throughput | Per chip |
|---|---|---|
| 1K/1K | 1,996.00 tok/s | 499.00 tok/s |
| 1K/8K | 2,075.47 tok/s | 518.86 tok/s |
| 8K/1K | 1,053.10 tok/s | 263.27 tok/s |
Primary source and date. https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/ironwood/vLLM/Qwen3-Coder-480B-A35B — values updated 2026-05-04
The dated commit records both the aggregate and per-chip figures.
Source quality. Vendor reference benchmark using an open benchmark client. The model, checkpoint precision, topology, tensor/expert parallelism, input/output lengths, and concurrency are disclosed. No achieved TTFT, TPOT, or latency percentiles are supplied with the final table.
Quantitative-anchor status. Yes; this is the strongest public frontier-scale TPU cost anchor found.
For 1K input / 8K output:
The fact that the 1K/8K workload slightly outperforms 1K/1K is plausible for a saturation benchmark because the longer decode phase amortizes prefill and scheduling overhead. It should not be interpreted as an interactive single-user rate.
Claim. A second frontier-scale recipe reports Qwen3.5-397B-A17B-FP8 on four Ironwood chips:
Primary source and date. https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/ironwood/vLLM/Qwen3.5-397B — current recipe accessed 2026-07-15; benchmark image/build dated 2026-06-26
Source quality. Vendor reference recipe; reproducible but not independent or latency-qualified.
Quantitative-anchor status. Yes. The 1K/8K result implies approximately:
This corroborates that Ironwood’s public serving evidence is no longer limited to small dense models. However, the difference between this result and Qwen3-Coder also shows that active parameters, expert routing, architecture, kernels, and software maturity dominate simple parameter-count comparisons.
Claim. Google’s production environmental-impact paper measured the median Gemini Apps text prompt in May 2025 at a total of 0.24 Wh, broken down as:
The methodology covers the full production serving stack, including supporting models and machines. Google also reported a 33-fold reduction in energy per median prompt over the preceding year.
Primary source and date. https://arxiv.org/abs/2508.15734 — submitted 2025-08-21
Source quality. Primary production disclosure based on Google’s actual fleet, rather than a laboratory benchmark. This is the strongest public evidence about full-stack Gemini serving resource use.
Quantitative-anchor status. No $/token anchor. The paper does not disclose:
The 0.24 Wh figure can support an energy-per-request model. It cannot be converted rigorously into $/token or gross margin without inventing token counts, hardware assignment, and internal cost assumptions.
Google publishes generation-to-generation peak/specification claims:
Google’s v5e and v6e spec pages also expose a relative “expected performance/TCO” index, with v5e shown below v6e, but the index is not tied to a named model, serving workload, price basis, latency target, or disclosed methodology. It is not usable as a quantitative serving-TCO anchor.
The independent Gemma 4 study provides a useful performance-per-rental-dollar comparison against two H100s, especially at long context, but it does not measure wall power.
I found no public, credible, apples-to-apples benchmark that simultaneously provides:
Therefore:
A valid model-level API margin calculation would require:
1- \frac{\text{fully loaded serving COGS per token}} {\text{API list revenue per token}} ]
Even the narrower accelerator-only approximation,
[ 1- \frac{\text{TPU rental floor per output token}} {\text{Gemini API output list price}}, ]
is not defensible because the numerator and denominator would refer to different things:
request-rate=inf saturation tests.Accordingly:
No rigorous public Gemini cost-per-token or list-price gross-margin anchor exists.
The production 0.24 Wh/prompt disclosure is valuable, and JetStream’s use in Gemini serving establishes technological relevance, but neither closes the missing token, model, topology, utilization, and dollar-cost fields.
For a quantitative serving-economics model, I would represent TPU with three separate cases rather than one point estimate:
Use the published output throughput and current per-chip rental rate directly. This captures the best case at effectively full utilization.
Best examples:
Use DWS Flex or a three-year commitment according to workload interruptibility and capacity certainty. Do not describe DWS as Spot.
Examples:
Apply explicit, separately modeled multipliers for:
For example, if a service realizes only 40% of saturation throughput after SLO headroom and utilization, the accelerator component is 2.5 times the benchmark floor before adding non-accelerator costs. This should be an explicit scenario assumption, not silently embedded in the TPU benchmark.
The public evidence supports the accelerator-floor cases. It does not identify a single empirically justified production-utilization multiplier for Gemini.
Google Cloud TPU pricing — current per-chip on-demand, DWS, and committed-use price legs for v5e, v6e, and Ironwood. https://cloud.google.com/tpu/pricing — accessed 2026-07-15.
Ironwood Qwen3-Coder-480B-A35B vLLM recipe — best public frontier-scale TPU benchmark: named FP8 model, four-chip topology, explicit input/output workloads, and per-chip throughput. https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/ironwood/vLLM/Qwen3-Coder-480B-A35B — values updated 2026-05-04.
AI Hypercomputer inference updates for Trillium — Llama 2 70B and Mixtral throughput with explicit chip counts; also establishes JetStream’s relationship to Gemini serving. https://cloud.google.com/blog/products/compute/ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu — 2025-05-09.
Original JetStream v5e publication — unusually clean specification of v5e-8, INT8 weights/activations/KV cache, model, context limits, output throughput, and contemporary cost per million tokens. https://cloud.google.com/blog/products/compute/accelerating-ai-inference-with-google-cloud-tpus-and-gpus — 2024-04-10.
Independent Gemma 4 31B v6e-versus-H100 study — strongest non-Google throughput-per-dollar check, including sequence-length sensitivity and a direct H100 comparison. https://arxiv.org/abs/2605.25645 — submitted 2026-05-25, v3 2026-06-05.