In today's words — This is an archived research artifact. Where it says
marginal serving gross margin,
unit CM or
serving contribution margin, the site now says
serving margin; where it states a confidence interval, the site treats it as a judgment range, not a calibrated one. What the site adopted from any artifact is stated in the
main report, which is authoritative where the two differ.
Research artifact — each page states its own provenance class in its header: verbatim originals carry SHA-256 stamps; adopted-findings summaries and reconstructions say so explicitly. Conclusions are synthesized (and where needed corrected) in the
main report; each page carries its own run date.
GPT-5.6 Pro deep-dive: AMD Instinct LLM serving economics
| Field |
Value |
| Model |
gpt-5-6-pro (confirmed via check() model_slug, is_pro: true) |
| Date |
2026-07-15 |
| Dive title |
AMD serving economics (Targeted 4 — dive 3 of 4) |
| Request ID |
req_1784147717347_pvneg3 |
| Commissioned |
2026-07-15 ("Targeted 4" package) |
Provenance — READ BEFORE USING (important)
- This dive hit the ChatGPT-Pro MCP's 120-minute HARD TIMEOUT.
completion_path: hard_timeout_api, elapsed 2h0m13s. The MCP returned its latest captured text and flagged it "may be incomplete."
- What this means: GPT-5.6 Pro was still in its extended reasoning/tool-use (search) phase at the 2-hour cap and had not yet emitted a full structured report. Across ~10 polls over the final ~15 minutes the captured
partial never grew beyond the same 770-character reasoning summary reproduced verbatim below — so this is the model's own dense reasoning-trace synopsis of its findings, NOT the requested full "dated claims + 5 sources" report.
- The reasoning synopsis is still substantively useful — it contains concrete, specific findings (MI355X-not-MI350X source correction; total-vs-generated throughput split; a derived $/M-output-token anchor; MLPerf/Azure/Oracle anchor-quality negatives). Those are captured in the update queue.
- Recovery: the conversation persists in the ChatGPT conversation history. Checking it directly (same approach as
recovered-deepseek-dive-2026-07-15.md) may yield more if the model wrote any additional text server-side after the MCP stopped polling — worth attempting, but the flat 770-char partial suggests little-to-no final-answer body was produced within 120 min.
- Sibling dives: TPU, Trainium, and the blinded replication completed during a poller dormancy gap and expired from the 5-min MCP cache (
check() = not_found) — recovered separately from the ChatGPT conversation history.
- Routing: any site change this justifies enters through
research/update-queue.md (Q-AUTO items), never via direct edits to the site/engine/tracked numbers.
- Citation anchors: no raw citation markers were present in this dive's captured text (hard-timeout truncation); as with the sibling dives, the original is preserved unmodified in the archive copy.
Goal (as dispatched)
Strongest available PUBLIC anchors for AMD Instinct (MI300X / MI325X / MI350X / MI355X / MI400) LLM inference/serving economics — or a rigorous negative result. Special task: corroborate or challenge the DigitalOcean/RadixArk claim of DeepSeek V4 at 3,500+ tok/s/GPU on MI350X; pin down benchmark conditions; assess whether it is a defensible quantitative anchor. Output contract: dated claims (claim / primary source URL+date / source-quality class / anchor-usability), explicit negative results where no anchor exists, and a numbered list of the 5 most load-bearing sources.
Verbatim captured text (reasoning-summary; 120-min hard-timeout truncation)
I'll trace the MI350X claim to its earliest source and reconstruct omitted conditions, while inventorying deployments, pricing, MLPerf, and ROCm evidence to separate defensible anchors from marketing.
The source resolves to MI355X, not MI350X: synthetic 8,192/1,024-token FP4/FP8 sweeps at concurrency 1,024+, with 3.5K on a low-interactivity frontier. HIP graphs were one of several optimizations.
The exact artifact shows 3.56–3.67K total tokens/s/GPU but only 396–408 generated, with 64–178s TTFT. At TensorWave's $2.95/hour, that is roughly $2/M output tokens.
MLPerf offers audited MI355X latency-qualified throughput, but its 405B submission prunes layers, limiting economics use. Azure lacks public chip-tied rates; Oracle's audited result lacks cloud pricing.
[NOTE: Hard timeout at 120 minutes; returned latest captured text — may be incomplete.]
Interpretation for the model (from the captured synopsis; sources NOT yet URL-pinned — pending recovery from ChatGPT conversation history)
The synopsis makes four load-bearing points, all of which sharpen (rather than overturn) the site's low-confidence stance on non-NVIDIA anchors:
- The viral "3.5K tok/s/GPU MI350X" headline is mis-attributed and mis-framed. The primary artifact is MI355X (not MI350X), a synthetic 8,192-input / 1,024-output sweep at concurrency 1,024+, FP4/FP8. The "3.5K" is total throughput (prefill-dominated); generated (decode) throughput is only ~396–408 tok/s/GPU at 64–178 s TTFT — a low-interactivity operating point unusable for interactive serving. HIP graphs were one of several optimizations, not a standalone 10× lever.
- Derived cost anchor (heavily caveated): at TensorWave ~$2.95/GPU-hour, the synthetic generated-throughput figure implies ≈ $2 / M output tokens — a candidate AMD rental-path $/token anchor, valid only at that low-interactivity point.
- MLPerf has an audited MI355X latency-qualified result, but the 405B submission prunes layers, limiting its economics usefulness.
- Anchor-quality negatives: Azure (ND MI300X v5) exposes no public chip-tied rate; Oracle's audited result lacks cloud pricing. Net: no clean production interactive AMD $/token anchor exists publicly as of 2026-07-15; the only derivable number is the synthetic-sweep figure in (2).
These items are queued in research/update-queue.md as Q-AUTO-2026-07-15 entries, marked QUEUED-AUTO and flagged as reasoning-summary-derived pending recovery from ChatGPT conversation history for exact source URLs/dates.