← Frontier Inference Margins · all research reports
What this document is. The owner-level requirement behind this page's FINAL-ANSWER
block (2026-07-22, verbatim in the project record): "we do need an obvious final
answer … it needs to have the rationale clearly explained with links back to the
evidence and therefore be defensible in the face of expert scrutiny." This annex is
that rationale. The numbers themselves render LIVE on the landing block and through
the MCP get_report id final-answer — this document explains where each one comes
from, what it assumes, and what disclosure would move it. Values quoted here are
execution outputs re-derived during the external review (2026-07-27) and are marked ≈; the live
surfaces are authoritative.
What the answer is NOT. Every calculator figure is a policy-labeled scenario estimate, and the above-80 reading is a separately labeled adopted analyst judgment — this page's source-reliability adjudication, not a calculator output and not provider disclosure. No placement map for Claude Opus 4.x is public to this page, so this page does not label any calculator result verified or central, and the comparison slot renders honestly empty ("no verified central comparator exists" — an owner-ratified design decision, 2026-07-22). The spans below are spans across declared alternatives, never statistical statements about a distribution — this page does not have the evidence to propagate input distributions, and pretending otherwise would manufacture precision.
Unit direct-serving contribution margin for Claude Opus 4.x at the published tariff schedule under the reference cache/batch/discount mix: 1 − (modeled direct serving cost ÷ modeled effective billings) at the Reference traffic mix (15:1 input:output, 60% cache hit). NOT a company gross margin — the report's §7 bridges the two. The fleet is the serve-feasibility-filtered evidence-informed default (owner rulings 2026-07-21 + 2026-07-23): the NA-blend family shares, with any leg that cannot serve the model under the loaded-bytes policy excluded at weight 0 (disclosed inline) and the remaining declared weights renormalized.
Basis note (b9 M5). Every calculator figure in this annex is the PUBLIC-EVIDENCE REFERENCE reading unless that figure explicitly says otherwise — algorithmic lead 0 months, family multipliers 1.0×. (Exactly one figure below says otherwise, and labels itself where it stands: the ratified-prior default named in the next sentence.) The calculator's own default state additionally carries the owner-ratified +3-month algorithmic-lead prior (a labeled scenario prior, not a measurement) and reads ≈69%. The FINAL-ANSWER surface this annex explains RETAINS that reference computation unchanged and, since the b9 M6 rework, additionally shows the ratified-prior reading beside it — this annex remains reference-basis throughout.
The chain, innermost first — each link names its evidence class:
revised-band-central-2.5) — informed
by newer community estimates; Musk's public "Grok is 0.5T = 1/10 Opus"
(April 2026) is read as likely referring to the prior flagship, Opus 4.6
(owner-adjudicated planning interpretation; the referent is unverified), and
its 5 T deduction is preserved as the labeled community-central-5.0 size
case. The headline is size-dependent and decreases monotonically across the
three sampled totals: 2.0/2.5/3.0 T compute ≈59.87% / ≈59.18% / ≈58.49%
(executed). Active ≈300B is this page's working
estimate (the open question is flagged on the model note). COMMUNITY ESTIMATE
class, page-adopted.research/primary-sources/nvidia-anthropic-share-x-2026-07/), TPU 25% /
Trainium 25% (equal residual split across the two ~1M-chip-scale public
commitments — this page's declared choice). Within-family splits carried from
the v2.1 declared topology (no public basis exists). At the revised flagship
size EVERY declared leg serves its declared operating point (H100 holds b=96 at
solver width 112), so the serve-feasibility rule removes nothing and all seven
declared weights bind unrenormalized. The exclusion story is preserved,
reproducible, and disclosed under the 5 T size case: there H100's declared
operating point (b=96) is not satisfiable within its registered 144-GPU domain
under the policy (best feasible batch 75 — it would render capped, never
policy-clean), so it enters at weight 0 with the remaining six legs
renormalized, and the include-H100-at-its-capped-77 counterfactual is computed
and disclosed there. Capacity uses peak live KV (ISL + OSL), while decode
performance uses the representative position (ISL + OSL/2). Membership derives at the native Reference traffic
anchor — it never re-derives under a user's traffic selection (the KV term
makes serve-feasibility traffic-dependent; the anchor keeps the default ONE
thing).evidence-instances-v22.json# gb200-vllm-r1, frozen d2 §3.1 F4); gb300 is ANALYST-SET at an assumed operating
point with no measured value and is excluded from the fit set; h100/h200 are
FAMILY-TRANSFER rows off the DeepSeek H800 production fit (F1); TPU v7 is an
ANALYST-SET same-platform aggregate-form bridge; Trainium2/3 carry the out-of-family
joint coefficient and have no matched serving anchor. Prefill: the single frozen
F7 identity fit, universal-transfer rule —
prefill never carries a measured basis. Widths are CAPACITY-SOLVED per leg over
registered hardware domains (solver receipts on every leg), never fixed
constants; solved widths at the revised default: h100 112, h200 32, gb200 24,
gb300 12, tpu7 16, trn2 48, trn3 32.Composition: weighted blend of all seven per-leg costs at the declared operating points → margin ≈59% (59.1806%) at the current engine data, at the public-evidence reference.
Open form-correction debt — not a repaired estimate. Holding each legacy calibration coefficient fixed while re-expressing its traffic form produces 53.24% to 64.17% at the flagship baseline computed at the public-evidence reference: a 10.93-point open calibration debt from the declared replica width alone. More seriously, the Trainium2/3 operating point registry labels batch as replica-global while the shipped engine consumes it per chip; the alternate reading makes those affected legs about 15.2× lower-throughput. Public evidence does not identify which form is right. The live calculator and MCP now disclose this beside every result; none of these counterfactual values is presented as a corrected margin.
Loaded-bytes at {0.55, 0.65, 1.05} B/param on the SAME fixed membership: ≈59% / ≈59% / ≈59% at the current engine data, at the public-evidence reference (no continuity is implied between samples). Membership is held FIXED at the central-policy derivation — and at the revised flagship size it is policy-STABLE: no leg would enter or leave at any sampled point. The selected operating batches also remain unchanged at all three points, so this sampled band is flat rather than a 12-point margin span (the typed membership-sensitivity record is empty everywhere; the old H100 would-re-enter counterfactual belongs to the 5 T size case).
The 50% paid-capacity occupancy this page holds by default is a declared planning convention, not a measurement: no provider publishes occupancy telemetry by phase, so there is nothing to calibrate it against. It is enumerated in the r4 adversarial review's scenario-only ledger (§C3) for exactly that reason. Moving it is a legitimate scenario question and the calculator exposes it as a live control — which is why the analyst-gap summary's first two rows move this one control and show what the engine then computes.
Acceptance rates for speculative decoding are unpublished for the fleet this page models. Its cited evidence set carries four non-fleet acceptance figures and no fleet-specific one: two non-flagship anchors — one reporting an average acceptance length of about 1.8–1.9 (SGLang's acceptance-length metric counts accepted draft tokens plus the bonus token produced per verification step), one assuming 70% acceptance for a single speculative token — and, in the open-stack post cited below, average acceptance lengths of 2.18 and 2.44 at two draft-window settings. These are the figures this page has found, not a claim about every figure that exists. The r4 review's verdict on applying a fleet-wide credit is explicit: "do not apply one universal multiplier" (§B10). Published gains are workload-dependent — about 14% at production-like batch against about 60% at modest concurrency — so a single multiplier would be a workload assumption wearing a mechanism's clothes. Those two figures are the SAME model on the SAME stack (DeepSeek V3 under SGLang), differing in cluster scale, concurrency, sequence lengths and draft window. The larger figure is the MTP-versus-no-MTP delta with overlap scheduling absent from both arms: 82.0 versus 51.0 tokens/s/rank (+60.8%). The post separately reports 60.4 tokens/s/rank for overlap scheduling without MTP; because that SGLang version did not support MTP together with overlap scheduling, it does not report MTP's incremental gain on top of overlap. Both are cited from that post and are not registered evidence rows of this page. Since 2026-07-29 the MECHANISM is vendor-officially on the record at one frontier lab (OpenAI's engineering post credits an improved draft/speculator model with more than 15% additional token-generation efficiency, and its 2026-07-30 pricing post says it is passing those gains on); that lab is not this page's flagship, and the vendor claim and the SGLang-reported open-stack measurements are kept in separate classes and never summed. This page therefore holds the credit at zero in every reading it authors and says so. A reader may price the lever themselves with the scenario control, from the "no MTP/disagg" stack setting only; the calculator then computes and labels that reading as the reader's. It is off by default, never applied to a leg whose deployed efficiency already absorbs speculation or whose status this page cannot establish, and never used to select a reading of this page's own.
The FINAL-ANSWER surface carries a calculator reading beside the adopted analyst reading (ruling of 2026-07-24, recorded in the project ledger): the conservative low/committed-rate planning case (≈59% at the public-evidence reference — the reproducible scenario output above) and the most plausible reading of the actual figure — above 80%. Since the b9 M6 rework the surface carries three labeled readings in total, because the calculator's own ratified-prior default (≈69%) is now shown beside the reference one; "two labeled readings" below always means calculator vs analyst, and each reading names its own basis where it stands. The above-80 reading is an adopted analyst reading, not a disclosure: it rests on Dylan Patel's first-party Sequoia-transcript statement ("north of 80 percent for the API price" on an Opus token) and SemiAnalysis's 3Q26 estimate of an above-80% API-business gross margin, and this page adopts SemiAnalysis as the most reliable analyst tier for such figures while stating plainly that its private calculations are unpublished. Where that claim targets the same estimand this page models, the two readings genuinely disagree — the surface says so rather than blending them.
The "Why not the higher numbers?" block renders one entry per higher
justification (the 90–95 conditional cluster, the 80+ tier, the model-generated
92–94 consult, and every remaining ≥80 registry row — a superset by enumeration,
fixture-enforced). Each entry states, in plain language: what the claim says
(verbatim from the claims registry), what it does not say, and what evidence
would flip it; the fully-bridged groups additionally identify the calculator
changes that move toward each higher claim and quantify any remaining
unreproduced gap (the executed ladder from ≈59 runs ≈71 / ≈71 / ≈80 / ≈83 /
≈85, while the separately composed strategic-partner lens lands at ≈83.3 — only where the calculator actually
reaches a claim's neighborhood does the entry say so), and the compact entries
say honestly where no bridge is constructed.
The current decomposition (≈60.18 → ≈59.18) carries its own line: the
page-adjudicated evidence-informed weight rebind moves the result by −1.0 point;
at the revised size the serve-feasibility rule removes nothing (under the 5 T case
it removes H100 and Trainium2, landing ≈57.22). Full drafting
history and review chain: research/im4-fa-justifications-memo.md (design gate,
seven rounds + the dual GPT Pro review), raws under research/reviews/, and the
cumulative concern ledger research/gptpro-concern-ledger.md.
Only a direct same-scope margin disclosure, or matched disclosures of serving cost and realized billings, could verify actual margin. Short of that, the disclosures below would materially move or narrow the calculator scenario — in descending order of impact: (1) any provider disclosure of actual fleet composition for frontier serving; (2) a measured loaded-bytes/checkpoint disclosure or placement map for a closed model (would remove the placement-specific blocker and materially narrow the calculator scenario — not, by itself, verify actual margin without matched cost and realized-billing evidence); (3) strategic-vs-market rate confirmation for TPU/ Trainium at scale; (4) billable-cache share telemetry (the cacheHit dual-use assumption is labeled and material); (5) any first-party statement of Opus parameters. Expert scrutiny is invited precisely to surface such disclosures — that is this project's stated north star: the best possible answer from public evidence plus explicitly labeled page assumptions, revised the moment better data exists.
research/evidence-instances-v22.json (WS-B schema, three-status)research/d2-receipt-pack.mdWEIGHT_PLACEMENT in engine data +
sha256-pinned config captures (research/primary-sources/hf-configs-placement-2026-07-22/)research/primary-sources/nvidia-anthropic-share-x-2026-07/MARGIN_CLAIMSresearch/im-gates-ledger.md