← Frontier Inference Margins · all research reports

Research artifact — each page states its own provenance class in its header: verbatim originals carry SHA-256 stamps; adopted-findings summaries and reconstructions say so explicitly. Conclusions are synthesized (and where needed corrected) in the main report; each page carries its own run date.

GPT-5.6 Pro deep-dive: AMD Instinct LLM serving economics

Field Value
Model gpt-5-6-pro (confirmed via check() model_slug, is_pro: true)
Date 2026-07-15
Dive title AMD serving economics (Targeted 4 — dive 3 of 4)
Request ID req_1784147717347_pvneg3
Commissioned 2026-07-15 ("Targeted 4" package)

Provenance — READ BEFORE USING (important)

Goal (as dispatched)

Strongest available PUBLIC anchors for AMD Instinct (MI300X / MI325X / MI350X / MI355X / MI400) LLM inference/serving economics — or a rigorous negative result. Special task: corroborate or challenge the DigitalOcean/RadixArk claim of DeepSeek V4 at 3,500+ tok/s/GPU on MI350X; pin down benchmark conditions; assess whether it is a defensible quantitative anchor. Output contract: dated claims (claim / primary source URL+date / source-quality class / anchor-usability), explicit negative results where no anchor exists, and a numbered list of the 5 most load-bearing sources.


Verbatim captured text (reasoning-summary; 120-min hard-timeout truncation)

I'll trace the MI350X claim to its earliest source and reconstruct omitted conditions, while inventorying deployments, pricing, MLPerf, and ROCm evidence to separate defensible anchors from marketing.

The source resolves to MI355X, not MI350X: synthetic 8,192/1,024-token FP4/FP8 sweeps at concurrency 1,024+, with 3.5K on a low-interactivity frontier. HIP graphs were one of several optimizations.

The exact artifact shows 3.56–3.67K total tokens/s/GPU but only 396–408 generated, with 64–178s TTFT. At TensorWave's $2.95/hour, that is roughly $2/M output tokens.

MLPerf offers audited MI355X latency-qualified throughput, but its 405B submission prunes layers, limiting economics use. Azure lacks public chip-tied rates; Oracle's audited result lacks cloud pricing.

[NOTE: Hard timeout at 120 minutes; returned latest captured text — may be incomplete.]


Interpretation for the model (from the captured synopsis; sources NOT yet URL-pinned — pending recovery from ChatGPT conversation history)

The synopsis makes four load-bearing points, all of which sharpen (rather than overturn) the site's low-confidence stance on non-NVIDIA anchors:

  1. The viral "3.5K tok/s/GPU MI350X" headline is mis-attributed and mis-framed. The primary artifact is MI355X (not MI350X), a synthetic 8,192-input / 1,024-output sweep at concurrency 1,024+, FP4/FP8. The "3.5K" is total throughput (prefill-dominated); generated (decode) throughput is only ~396–408 tok/s/GPU at 64–178 s TTFT — a low-interactivity operating point unusable for interactive serving. HIP graphs were one of several optimizations, not a standalone 10× lever.
  2. Derived cost anchor (heavily caveated): at TensorWave ~$2.95/GPU-hour, the synthetic generated-throughput figure implies ≈ $2 / M output tokens — a candidate AMD rental-path $/token anchor, valid only at that low-interactivity point.
  3. MLPerf has an audited MI355X latency-qualified result, but the 405B submission prunes layers, limiting its economics usefulness.
  4. Anchor-quality negatives: Azure (ND MI300X v5) exposes no public chip-tied rate; Oracle's audited result lacks cloud pricing. Net: no clean production interactive AMD $/token anchor exists publicly as of 2026-07-15; the only derivable number is the synthetic-sweep figure in (2).

These items are queued in research/update-queue.md as Q-AUTO-2026-07-15 entries, marked QUEUED-AUTO and flagged as reasoning-summary-derived pending recovery from ChatGPT conversation history for exact source URLs/dates.