← Frontier Inference Margins · all research reports

Research artifact — each page states its own provenance class in its header: verbatim originals carry SHA-256 stamps; adopted-findings summaries and reconstructions say so explicitly. Conclusions are synthesized (and where needed corrected) in the main report; each page carries its own run date.

The final answer — rationale and evidence chain

What this document is. The owner-level requirement behind this page's FINAL-ANSWER block (2026-07-22, verbatim in the project record): "we do need an obvious final answer … it needs to have the rationale clearly explained with links back to the evidence and therefore be defensible in the face of expert scrutiny." This annex is that rationale. The numbers themselves render LIVE on the landing block and through the MCP get_report id final-answer — this document explains where each one comes from, what it assumes, and what disclosure would move it. Values quoted here are execution outputs re-derived during the external review (2026-07-27) and are marked ≈; the live surfaces are authoritative.

What the answer is NOT. Every calculator figure is a policy-labeled scenario estimate, and the above-80 reading is a separately labeled adopted analyst judgment — this page's source-reliability adjudication, not a calculator output and not provider disclosure. No placement map for Claude Opus 4.x is public to this page, so this page does not label any calculator result verified or central, and the comparison slot renders honestly empty ("no verified central comparator exists" — an owner-ratified design decision, 2026-07-22). The spans below are spans across declared alternatives, never statistical statements about a distribution — this page does not have the evidence to propagate input distributions, and pretending otherwise would manufacture precision.

1. The estimand

Unit direct-serving contribution margin for Claude Opus 4.x at the published tariff schedule under the reference cache/batch/discount mix: 1 − (modeled direct serving cost ÷ modeled effective billings) at the Reference traffic mix (15:1 input:output, 60% cache hit). NOT a company gross margin — the report's §7 bridges the two. The fleet is the serve-feasibility-filtered evidence-informed default (owner rulings 2026-07-21 + 2026-07-23): the NA-blend family shares, with any leg that cannot serve the model under the loaded-bytes policy excluded at weight 0 (disclosed inline) and the remaining declared weights renormalized.

2. The conservative low/committed-rate planning case (≈59% at the public-evidence reference, policy-labeled scenario)

Basis note (b9 M5). Every calculator figure in this annex is the PUBLIC-EVIDENCE REFERENCE reading unless that figure explicitly says otherwise — algorithmic lead 0 months, family multipliers 1.0×. (Exactly one figure below says otherwise, and labels itself where it stands: the ratified-prior default named in the next sentence.) The calculator's own default state additionally carries the owner-ratified +3-month algorithmic-lead prior (a labeled scenario prior, not a measurement) and reads ≈69%. The FINAL-ANSWER surface this annex explains RETAINS that reference computation unchanged and, since the b9 M6 rework, additionally shows the ratified-prior reading beside it — this annex remains reference-basis throughout.

The chain, innermost first — each link names its evidence class:

  1. Model size. Opus total = the page-adopted 2–3 T planning band, scalar 2.5 T (adjudicated 2026-07-24; size case revised-band-central-2.5) — informed by newer community estimates; Musk's public "Grok is 0.5T = 1/10 Opus" (April 2026) is read as likely referring to the prior flagship, Opus 4.6 (owner-adjudicated planning interpretation; the referent is unverified), and its 5 T deduction is preserved as the labeled community-central-5.0 size case. The headline is size-dependent and decreases monotonically across the three sampled totals: 2.0/2.5/3.0 T compute ≈59.87% / ≈59.18% / ≈58.49% (executed). Active ≈300B is this page's working estimate (the open question is flagged on the model note). COMMUNITY ESTIMATE class, page-adopted.
  2. Fleet membership + weights. NA-blend: NVIDIA 50% (a SECONDHAND Morgan Stanley summary of an NVIDIA NDR comment about an unnamed ASIC-heavy lab, attribution inference by a named relayer — the full chain and its limits are in research/primary-sources/nvidia-anthropic-share-x-2026-07/), TPU 25% / Trainium 25% (equal residual split across the two ~1M-chip-scale public commitments — this page's declared choice). Within-family splits carried from the v2.1 declared topology (no public basis exists). At the revised flagship size EVERY declared leg serves its declared operating point (H100 holds b=96 at solver width 112), so the serve-feasibility rule removes nothing and all seven declared weights bind unrenormalized. The exclusion story is preserved, reproducible, and disclosed under the 5 T size case: there H100's declared operating point (b=96) is not satisfiable within its registered 144-GPU domain under the policy (best feasible batch 75 — it would render capped, never policy-clean), so it enters at weight 0 with the remaining six legs renormalized, and the include-H100-at-its-capped-77 counterfactual is computed and disclosed there. Capacity uses peak live KV (ISL + OSL), while decode performance uses the representative position (ISL + OSL/2). Membership derives at the native Reference traffic anchor — it never re-derives under a user's traffic selection (the KV term makes serve-feasibility traffic-dependent; the anchor keeps the default ONE thing).
  3. Per-leg throughput. Calibrated decode rooflines with typed evidence classes: gb200 is FITTED to the vLLM R1 decode observation (evidence-instances-v22.json# gb200-vllm-r1, frozen d2 §3.1 F4); gb300 is ANALYST-SET at an assumed operating point with no measured value and is excluded from the fit set; h100/h200 are FAMILY-TRANSFER rows off the DeepSeek H800 production fit (F1); TPU v7 is an ANALYST-SET same-platform aggregate-form bridge; Trainium2/3 carry the out-of-family joint coefficient and have no matched serving anchor. Prefill: the single frozen F7 identity fit, universal-transfer rule — prefill never carries a measured basis. Widths are CAPACITY-SOLVED per leg over registered hardware domains (solver receipts on every leg), never fixed constants; solved widths at the revised default: h100 112, h200 32, gb200 24, gb300 12, tpu7 16, trn2 48, trn3 32.
  4. Prices/rents. gb200, TPU v7 and Trainium2 are observed-source-named; h100/h200/gb300/Trainium3 are analyst-set — four of the seven member rents (each row cites its basis). List prices $5/$25 (published tariff). The five-status vector's economics slot reports the WORST input class present — at the default: "evidence-quality: joint-fit throughput · analyst-set price".
  5. Loaded-bytes policy. fp8 scenario → 1.0 B/param loaded (capacity-only planning default, dive §B; firewalled from the calibrated throughput tuple — a release-gated invariant). This is the planning point's policy identity and it is NOT one of the three sampled band points.

Composition: weighted blend of all seven per-leg costs at the declared operating points → margin ≈59% (59.1806%) at the current engine data, at the public-evidence reference.

Open form-correction debt — not a repaired estimate. Holding each legacy calibration coefficient fixed while re-expressing its traffic form produces 53.24% to 64.17% at the flagship baseline computed at the public-evidence reference: a 10.93-point open calibration debt from the declared replica width alone. More seriously, the Trainium2/3 operating point registry labels batch as replica-global while the shipped engine consumes it per chip; the alternate reading makes those affected legs about 15.2× lower-throughput. Public evidence does not identify which form is right. The live calculator and MCP now disclose this beside every result; none of these counterfactual values is presented as a corrected margin.

3. The three sampled policy points

Loaded-bytes at {0.55, 0.65, 1.05} B/param on the SAME fixed membership: ≈59% / ≈59% / ≈59% at the current engine data, at the public-evidence reference (no continuity is implied between samples). Membership is held FIXED at the central-policy derivation — and at the revised flagship size it is policy-STABLE: no leg would enter or leave at any sampled point. The selected operating batches also remain unchanged at all three points, so this sampled band is flat rather than a 12-point margin span (the typed membership-sensitivity record is empty everywhere; the old H100 would-re-enter counterfactual belongs to the 5 T size case).

4. The spans across declared alternatives

Scenario-only ledger — fleet utilization (r4 §C3)

The 50% paid-capacity occupancy this page holds by default is a declared planning convention, not a measurement: no provider publishes occupancy telemetry by phase, so there is nothing to calibrate it against. It is enumerated in the r4 adversarial review's scenario-only ledger (§C3) for exactly that reason. Moving it is a legitimate scenario question and the calculator exposes it as a live control — which is why the analyst-gap summary's first two rows move this one control and show what the engine then computes.

Scenario-only ledger — speculative decode / MTP (r4 §B10, §C3)

Acceptance rates for speculative decoding are unpublished for the fleet this page models. Its cited evidence set carries four non-fleet acceptance figures and no fleet-specific one: two non-flagship anchors — one reporting an average acceptance length of about 1.8–1.9 (SGLang's acceptance-length metric counts accepted draft tokens plus the bonus token produced per verification step), one assuming 70% acceptance for a single speculative token — and, in the open-stack post cited below, average acceptance lengths of 2.18 and 2.44 at two draft-window settings. These are the figures this page has found, not a claim about every figure that exists. The r4 review's verdict on applying a fleet-wide credit is explicit: "do not apply one universal multiplier" (§B10). Published gains are workload-dependent — about 14% at production-like batch against about 60% at modest concurrency — so a single multiplier would be a workload assumption wearing a mechanism's clothes. Those two figures are the SAME model on the SAME stack (DeepSeek V3 under SGLang), differing in cluster scale, concurrency, sequence lengths and draft window. The larger figure is the MTP-versus-no-MTP delta with overlap scheduling absent from both arms: 82.0 versus 51.0 tokens/s/rank (+60.8%). The post separately reports 60.4 tokens/s/rank for overlap scheduling without MTP; because that SGLang version did not support MTP together with overlap scheduling, it does not report MTP's incremental gain on top of overlap. Both are cited from that post and are not registered evidence rows of this page. Since 2026-07-29 the MECHANISM is vendor-officially on the record at one frontier lab (OpenAI's engineering post credits an improved draft/speculator model with more than 15% additional token-generation efficiency, and its 2026-07-30 pricing post says it is passing those gains on); that lab is not this page's flagship, and the vendor claim and the SGLang-reported open-stack measurements are kept in separate classes and never summed. This page therefore holds the credit at zero in every reading it authors and says so. A reader may price the lever themselves with the scenario control, from the "no MTP/disagg" stack setting only; the calculator then computes and labels that reading as the reader's. It is off by default, never applied to a leg whose deployed efficiency already absorbs speculation or whose status this page cannot establish, and never used to select a reading of this page's own.

4-bis. The most plausible reading, and why not the higher numbers

The FINAL-ANSWER surface carries a calculator reading beside the adopted analyst reading (ruling of 2026-07-24, recorded in the project ledger): the conservative low/committed-rate planning case (≈59% at the public-evidence reference — the reproducible scenario output above) and the most plausible reading of the actual figure — above 80%. Since the b9 M6 rework the surface carries three labeled readings in total, because the calculator's own ratified-prior default (≈69%) is now shown beside the reference one; "two labeled readings" below always means calculator vs analyst, and each reading names its own basis where it stands. The above-80 reading is an adopted analyst reading, not a disclosure: it rests on Dylan Patel's first-party Sequoia-transcript statement ("north of 80 percent for the API price" on an Opus token) and SemiAnalysis's 3Q26 estimate of an above-80% API-business gross margin, and this page adopts SemiAnalysis as the most reliable analyst tier for such figures while stating plainly that its private calculations are unpublished. Where that claim targets the same estimand this page models, the two readings genuinely disagree — the surface says so rather than blending them.

The "Why not the higher numbers?" block renders one entry per higher justification (the 90–95 conditional cluster, the 80+ tier, the model-generated 92–94 consult, and every remaining ≥80 registry row — a superset by enumeration, fixture-enforced). Each entry states, in plain language: what the claim says (verbatim from the claims registry), what it does not say, and what evidence would flip it; the fully-bridged groups additionally identify the calculator changes that move toward each higher claim and quantify any remaining unreproduced gap (the executed ladder from ≈59 runs ≈71 / ≈71 / ≈80 / ≈83 / ≈85, while the separately composed strategic-partner lens lands at ≈83.3 — only where the calculator actually reaches a claim's neighborhood does the entry say so), and the compact entries say honestly where no bridge is constructed. The current decomposition (≈60.18 → ≈59.18) carries its own line: the page-adjudicated evidence-informed weight rebind moves the result by −1.0 point; at the revised size the serve-feasibility rule removes nothing (under the 5 T case it removes H100 and Trainium2, landing ≈57.22). Full drafting history and review chain: research/im4-fa-justifications-memo.md (design gate, seven rounds + the dual GPT Pro review), raws under research/reviews/, and the cumulative concern ledger research/gptpro-concern-ledger.md.

5. What would move or verify this answer

Only a direct same-scope margin disclosure, or matched disclosures of serving cost and realized billings, could verify actual margin. Short of that, the disclosures below would materially move or narrow the calculator scenario — in descending order of impact: (1) any provider disclosure of actual fleet composition for frontier serving; (2) a measured loaded-bytes/checkpoint disclosure or placement map for a closed model (would remove the placement-specific blocker and materially narrow the calculator scenario — not, by itself, verify actual margin without matched cost and realized-billing evidence); (3) strategic-vs-market rate confirmation for TPU/ Trainium at scale; (4) billable-cache share telemetry (the cacheHit dual-use assumption is labeled and material); (5) any first-party statement of Opus parameters. Expert scrutiny is invited precisely to surface such disclosures — that is this project's stated north star: the best possible answer from public evidence plus explicitly labeled page assumptions, revised the moment better data exists.

6. Provenance pointers