| Platform | Dense FP8 (PF) | Dense FP4 (PF) | HBM | BW | TDP | Rental (Jul 2026) |
| H100 SXM (2022) | 1.98 | — | 80 GB | 3.35 TB/s | 700 W | $1.99–3.90/hr, spot ≈ $2.40 |
| H200 (2024) | 1.98 | — | 141 GB | 4.8 TB/s | 700 W | $2.45–4.50/hr |
| B200 HGX, per GPU (2024) | 4.5 | 9 | 180 GB | 8 TB/s | 1.0 kW | $6.69/hr on-demand (Lambda) |
| GB200 NVL72, per GPU (2024-25) | 5.0 | 10 | 186 GB | 8 TB/s | 1.2 kW | $3.50–6.00/hr; rack ≈ $3–3.5M; AWS Capacity Block rack rate $761.904/hr ($10.582/GPU-hr) — cleanest full-rack market datum |
| GB300 NVL72 "Blackwell Ultra" (2025-26) | 5.0 | 15 | 288 GB | 8 TB/s | 1.4 kW | No numeric rack rate public on any major provider (Jul 2026); B300 node anchor: Nebius $7.85/GPU-hr ⇒ $0.289/M output (§4) |
| TPU v7 Ironwood (GA Mar 2026) | 4.61 | — | 192 GB | 7.37 TB/s | ~1 kW | $12/chip-hr on-demand, $5.40 3-yr (Jul 2026); Anthropic deal: up to 1M chips; rental-floor anchor: $6.42/M output tok on-demand (§10) — a vendor saturation run at infinite offered load; no achieved TTFT/TPOT published |
| Trainium2 (GA Dec 2024) / Trainium3 (GA Dec 2025) | 1.3 / 2.51 | — | 96 / 144 GB | 2.9 / 4.9 TB/s | ~0.5–0.8 kW | Trn2 Capacity Block ≈$2.235/chip-hr (Jul 2026), engineering anchor $68.90–98.65/M output tok; Rainier launched with ~500k Trainium2 confirmed running Claude inference (Nov 2025), Anthropic reported >1M in use by Apr 2026, zero Trn3/Rainier economics disclosed |
| H800 — China export SKU (2023) | 1.98 | — | 80 GB | 3.35 TB/s | 700 W | IDC annual-commit $1.47–2.06/hr (mid-2026, +30% post-Spring-Festival); DeepSeek's disclosure assumed $2 |
| H20 — China-legal SKU (2024) | 0.296 | — | 96 GB | 4.0 TB/s | 400 W | ~$10–12k chip / ~$20k installed; rental class spans ~$0.76 (IDC annual) to $7+ (cloud on-demand) |
| Huawei Ascend 910C (2024–25) | 1.50 (INT8; no FP8) | — | 128 GB | 3.2 TB/s | ~0.6 kW | ~$23k/chip installed; Huatai procurement $1.71–2.25/hr; CloudMatrix 384 ≈ RMB 60M ($8.2M); CM384 throughput anchored (§4), cost unanchored; 910B (older chip) cost proxy ≈$2.09/M output |
| Vera Rubin NVL72, per GPU (preliminary 2026 specs) | 17.5 (dense FP8/FP6) | 50 sparse NVFP4 inference (35 dense-class training) | 288 GB HBM4 | 22 TB/s | ~1.8 kW | production power, throughput and cost TBD — 2026-07-15 dive: no MLPerf, no InferenceX result, no public price; only a relative Kimi-K2-Thinking claim vs GB200 (10× tok/s/MW, ~1/10 cost) |
Measured, not marketing: SemiAnalysis InferenceX (the successor to InferenceMAX, the field's independent benchmark) finds the most-optimized GB300 NVL72 delivers ~17× the best H100 config in FP8 and ~32× in FP4 on DeepSeek R1 (Jun 27, 2026) — and, critically, that software alone was a 14× gain on the same silicon (baseline FP8 ~1k → wideEP+disagg ~8k → +MTP ~14k tok/s/GPU). SGLang's reproducible benchmark runs agree: ~11–12k tok/s/GPU on V4 Pro 1.6T (an InferenceX/SGLang benchmark, not production telemetry), 6.5× over B200 with Dynamo disaggregation. At rack level: an H100 rack two years ago did ~8.8k tok/s; a GB300 NVL72 does ~370k — 42× in two years (8× of it HBM growth).
2026-07-15 — GB300 is now audited-throughput-anchored, but its rack price is still not public. A round-2 dive found MLPerf Inference v6.0 (Apr 1, 2026) contains valid, reproducible single-rack GB300 NVL72 results on DeepSeek-R1 FP4 — refuting any "GB300 has no public serving benchmark" framing: NVIDIA Interactive 250,634 / Server 400,437 / Offline 647,076 generated tok/s; Nebius submitted the strongest one-rack results, Server 575,580 / Offline 673,936. These are generated-output tokens (LoadGen sums completed-output-token counts), not input+output totals, and NVIDIA's headline 2.49M tok/s Offline figure is four racks (288 GPUs), not one. What remains missing: no numeric GB300 rack rental rate is public on AWS, CoreWeave, Nebius, GCP, Azure, OCI or Crusoe's pricing pages as of Jul 15, 2026 — so the honest model is GB300 $/M-output as a function of rack-hour price, not a point estimate (SemiAnalysis's own $190.94/rack-hr GB300 figure is explicitly labeled a temporary 1.2×-GB200 placeholder in its source code, not an observed rate). Note on the live calculator: that "function of price" framing describes the honest representation of the newly-documented bridge anchor — the deployed calculator still prices GB300 from an analyst-estimated $6/GPU-hr scenario (no observed market rate), which is 12% of the activated NA-blend default fleet; treat that $6 as a labeled scenario input, not a market anchor. The cleanest present Blackwell-Ultra anchor instead comes from the same-provider B300 node: Nebius's 8-GPU MLPerf Server result (60,413 gen tok/s) paired with its own public $7.85/GPU-hr rate gives $0.289/M generated output tokens — an accelerator-rental floor, before non-GPU serving overhead and utilization reserve (B200 comparably: $0.307/M — B300 is only ~6% cheaper despite ~17% more throughput, because its hourly price is higher). A GB200 rack bridge exists via AWS's Capacity Block rate ($761.904/rack-hr, the cleanest explicit full-rack market datum — though a reserved/upfront Capacity Block price, not ordinary on-demand): $0.881/$0.630/$0.435 per M output at Interactive/Server/Offline. Rubin remains absolute-economics-unanchored — no MLPerf submission, no InferenceX result (listed "Coming Soon"), no public rental or purchase price for Rubin or Rubin Ultra; NVIDIA's only public claim is relative (Kimi-K2-Thinking, 32K-in/8K-out): up to 10× tok/s/MW and ~1/10 cost per M tokens vs GB200 NVL72 — not an absolute anchor, and not to be confused with the different NVL144/Rubin CPX/R100/VR200 configurations. None of this changes the deployed GB200/GB300/Rubin roofline parameters above — it upgrades the evidence quality behind GB300 and documents, rather than fits, the new bridge anchors. (Full NVIDIA forward inference-economics dive →)
The China stack is its own cost universe. Under export controls the Chinese labs serve on three tiers: hoarded pre-ban H800s (H100 compute, capped NVLink — where DeepSeek's disclosure happened), the deliberately compute-starved but bandwidth-rich H20 (the main legal SKU since mid-2025; decode-friendly because decode is bandwidth-bound — Ant Group's production SGLang deployment is the best public anchor), and Huawei's Ascend 910C (Huawei's CloudMatrix-Infer paper reports 1,943 tok/s/NPU decode on R1, while DeepSeek's internal evaluation put the chip at ~60% of H100 — vendor and customer numbers disagree, so the calculator's Ascend row carries wide error bars). H200s were license-cleared for ten named Chinese firms in early 2026 but essentially none had shipped as of May; Beijing reportedly moved to allow those purchases this week (Jul 7–8). Chip scarcity also cuts the other way on price: Chinese H-series lease rates rose 20–30% over early 2026 on a >1,000× surge in national token volume (TrendForce) — Chinese margins are being squeezed from the cost side at exactly the moment Western $/token falls. One more trap: "the China rental rate" is a category error — the same H20 spans roughly ¥5–72 per card-hour (>10×) between annual IDC bare-metal leases and hyperscaler on-demand instances. The calculator's China rows use annual-commit rates; the GPU-hour cost multiplier is the dial for other rental classes.
2026-07-15 — CloudMatrix 384 gets a real generated-throughput anchor; production cost stays unmatched. A round-2 dive found the strongest public CM384 result to date: FlexNPU (arXiv:2606.04415, Jun 3, 2026) served DeepSeek-R1 W8A8 on a full 384-card 910C system at ≈633,000 generated tok/s system-wide (≈1,646/card) under stated TTFT ≤1s / TPOT ≤50ms constraints — end-to-end, SLO-constrained, and a different (full-system) measurement from the isolated-decode figure already in the row above. No public CM384 hourly rental price accompanies it, so throughput is anchored but cost is not. This dive also nails down a total-vs-generated conflation in the widely-repeated CloudMatrix numbers: the "6,688 tok/s/NPU" figure is PREFILL/INPUT throughput and an idealized perfect-expert-balancing projection (the directly measured default is 5,655 prefill tok/s/NPU), while the "1,943 tok/s/NPU" figure already cited above is DECODE at batch 96 (49.4 ms TPOT) — only ≈20 generated tok/s per active sequence, not an interactive per-user rate; these two figures must never be blended as simultaneous end-to-end throughput. A cross-source proxy exists for the older 910B chip: JD's xLLM paper (709 generated tok/s/card on 16 cards) paired with China Telecom CTyun's public RMB 38.45/hr 910B instance rate gives ≈$2.09/M generated output tokens (public-evidence range $1.61–2.80/M) — a cross-source reconstruction, not audited provider COGS, and not the 910C row above. Ascend 920 has no official product page, benchmark, deployment, cloud SKU or price — Huawei's own Sept-2025 roadmap goes 910C → 950PR (Q1'26) → 950DT (Q4'26) → 960 (Q4'27) → 970 (Q4'28), with no 920 — so any "920" cost figure would be fabricated from rumor and stays excluded from this model. None of this changes the deployed Ascend 910C roofline parameters above: the 1,943 and 538 isolated-decode observations remain historical inputs to the frozen joint calibration, while the live row is the separate source-informed neutral η=0.299324 identity and does not reproduce either observation. The FlexNPU full-system result remains an evidence annotation, not a pending automatic refit. (Full Ascend inference-economics dive →)
But cost per token falls slower than throughput rises, because rental prices track capability: $2.40/hr H100 → $6/hr GB300 eats ~2.5× of the 17×. Net hardware-economics gain per generation in $/Mtok: H100→H200 ~1.2×, H200→GB200 ~2–3×, GB200→GB300 ~1.1–1.5× (GB300's edge is HBM capacity for reasoning/long-context, not FLOPS), Rubin ~2–3× again. Cumulative 2024→2027: roughly 5–16× cheaper per token at the hardware level (the product of the per-generation ranges just listed; the commissioned fixed-model projection lands in the same band — ~4.5–10× by the 2027 Rubin ramp, ~7–17× by mature Rubin), before model-side efficiency (sparser MoEs, MTP, quantization) which historically contributed as much again. This is why margins at constant list prices would, on these hardware-cost assumptions alone, trend toward 95%+ — a mechanical consequence of the cost model under a fixed-price counterfactual, not a prediction, and one the market visibly pre-empts, which is why in practice prices fall instead (Opus's 2025-11 cut to $5/$25 was 3×; fast mode — $10/$50 on Opus 4.8, down from 4.7's $30/$150 tier retiring July 24 — shows the latency premium being monetized separately, and itself falling).