← Frontier Inference Margins · all research reports
| Field | Value |
|---|---|
| Model | gpt-5-6-pro (confirmed via check() model_slug, is_pro: true) |
| Date | 2026-07-15 |
| Dive title | AMD serving economics (Targeted 4 — dive 3 of 4) |
| Request ID | req_1784147717347_pvneg3 |
| Commissioned | 2026-07-15 ("Targeted 4" package) |
completion_path: hard_timeout_api, elapsed 2h0m13s. The MCP returned its latest captured text and flagged it "may be incomplete."partial never grew beyond the same 770-character reasoning summary reproduced verbatim below — so this is the model's own dense reasoning-trace synopsis of its findings, NOT the requested full "dated claims + 5 sources" report.recovered-deepseek-dive-2026-07-15.md) may yield more if the model wrote any additional text server-side after the MCP stopped polling — worth attempting, but the flat 770-char partial suggests little-to-no final-answer body was produced within 120 min.check() = not_found) — recovered separately from the ChatGPT conversation history.research/update-queue.md (Q-AUTO items), never via direct edits to the site/engine/tracked numbers.Strongest available PUBLIC anchors for AMD Instinct (MI300X / MI325X / MI350X / MI355X / MI400) LLM inference/serving economics — or a rigorous negative result. Special task: corroborate or challenge the DigitalOcean/RadixArk claim of DeepSeek V4 at 3,500+ tok/s/GPU on MI350X; pin down benchmark conditions; assess whether it is a defensible quantitative anchor. Output contract: dated claims (claim / primary source URL+date / source-quality class / anchor-usability), explicit negative results where no anchor exists, and a numbered list of the 5 most load-bearing sources.
I'll trace the MI350X claim to its earliest source and reconstruct omitted conditions, while inventorying deployments, pricing, MLPerf, and ROCm evidence to separate defensible anchors from marketing.
The source resolves to MI355X, not MI350X: synthetic 8,192/1,024-token FP4/FP8 sweeps at concurrency 1,024+, with 3.5K on a low-interactivity frontier. HIP graphs were one of several optimizations.
The exact artifact shows 3.56–3.67K total tokens/s/GPU but only 396–408 generated, with 64–178s TTFT. At TensorWave's $2.95/hour, that is roughly $2/M output tokens.
MLPerf offers audited MI355X latency-qualified throughput, but its 405B submission prunes layers, limiting economics use. Azure lacks public chip-tied rates; Oracle's audited result lacks cloud pricing.
[NOTE: Hard timeout at 120 minutes; returned latest captured text — may be incomplete.]
The synopsis makes four load-bearing points, all of which sharpen (rather than overturn) the site's low-confidence stance on non-NVIDIA anchors:
These items are queued in research/update-queue.md as Q-AUTO-2026-07-15 entries, marked QUEUED-AUTO and flagged as reasoning-summary-derived pending recovery from ChatGPT conversation history for exact source URLs/dates.