Fleet
The same GPU, at seven different prices, all week.
Serving cost is dominated by two things a model author never touches: what you pay for a GPU-hour, and how much of that hour is doing useful work. Both vary far more across the market than the hardware does. This page shows the dispersion, names both ends of it, and works the arithmetic from a rental rate to a cost per million tokens so it can be checked.
$7.24 against $2.70 per GPU-hour on 2026-07-15. Both legs like-for-like normalized by the same index provider on the same day, which is what makes this the defensible number rather than the dramatic one.
01The accelerator lineup
Decode is a memory bandwidth problem, not a FLOPS problem
Autoregressive decode reads the whole weight set from HBM to produce one token. Arithmetic intensity is roughly one multiply-accumulate per parameter loaded, so the tensor cores idle while the memory system works, and output tokens per second per GPU tracks bandwidth rather than peak FLOPS. Prefill is the opposite: it is compute bound and batches well. A fleet tuned for cost per output token is therefore bought on TB/s, not on PFLOPS.
H100 SXM5 and H200 SXM share a die and an FP8 rating. The H200 carries 76% more memory and moves 43% more bytes per second, at identical compute, and that is the whole of the decode gain.
The inverse case is in the same table. GB300 NVL72 adds HBM capacity and 1.5x the dense FP4 rate over GB200 NVL72 while sitting at the same 8 TB/s per GPU. It is a capacity and compute upgrade that moves no additional bytes per second, which is why it changes what you can fit far more than it changes what you can decode.
| Accelerator | HBM | Bandwidth | TB/s per $/hr | On-demand / hr | Floor / hr |
|---|---|---|---|---|---|
H100 SXM5NVIDIA | 80 GB | 3.35 TB/s | 0.88reported | $3.82reported | $2.15 |
H200 SXMNVIDIA | 141 GB | 4.8 TB/s | 1.11reported | $4.32reported | $2.45 |
B200 (HGX, 1000 W)NVIDIAreported | 180 GB | 8 TB/s | 1.11reported | $7.22reported | $3.60reported |
B300 (Blackwell Ultra, HGX)NVIDIA | 288 GB | 8 TB/s | 0.88reported | $9.09reported | $4.30 |
GB200 NVL72 (per GPU)NVIDIA | 192 GB | 8 TB/s | 0.37reported | $21.53reported | $10.50 |
GB300 NVL72 (per GPU)NVIDIAest. | 288 GB | 8 TB/s | 0.88est. | $9.09est. | $4.48est. |
Instinct MI300XAMDreported | 192 GB | 5.3 TB/s | 1.71reported | $3.10reported | $1.65reported |
Instinct MI325XAMDest. | 256 GB | 6 TB/s | 1.03est. | $5.85est. | $2.59est. |
Instinct MI355XAMD | 288 GB | 8 TB/s | 0.93reported | $8.60reported | $2.59reported |
TPU v7 (Ironwood)Googlereported | 192 GB | 7.38 TB/s | 4.61est. | $1.60est. | $1.60est. |
All FLOPS figures are DENSE, in PFLOPS. Vendors headline structured 2:4 sparsity numbers that are essentially never used in production LLM serving, and where a datasheet published only a sparse figure the halved dense figure is marked as estimated. The on-demand column is a market average across many providers rather than any single rate card; named provider rates are in the next section, because the dispersion across providers is larger than the differences between these chips. Floor is the best observed spot, preemptible, or reserved rate. Bandwidth per dollar is derived from the two columns beside it and inherits the weaker confidence of the pair. Snapshot 2026-07-20.
Read the price columns carefully
- GB300 NVL72 (per GPU): NO PUBLIC GB300 NVL72 RATE EXISTS. CoreWeave, Nebius, Lambda, AWS, Azure and GCP all say contact sales as of 2026-07-20. The HGX B300 street average is shown only so the row renders; the rack-scale SKU should be assumed to price above it, since GB200 NVL72 prices above HGX B200. Do not quote this as a GB300 price.
- Instinct MI325X: NO PUBLIC $/GPU-HR RATE WAS FOUND for MI325X at any provider. It is listed at Vultr, TensorWave and DigitalOcean without a published rate. This is a bracket midpoint, not a price. Do not quote it.
- TPU v7 (Ironwood): Ironwood is NOT on Google's public rate card. This is the only figure available and it is a negotiated large-account rate, not a price anyone can buy at self-serve. Published TPU rates for reference: v5e $1.20/chip-hr, v6e $2.70/chip-hr.
Runs on
Accelerators the serving stack is built and tuned for. Kernel work targets each generation directly rather than relying on a portability layer.
Vendor names in the table identify silicon we run on and tune kernels for. They are not partnerships, endorsements, or commercial relationships, and we do not describe them as any of those.
02Capacity arbitrage
An identical H100, priced across an order of magnitude
The cheapest posted on-demand rate for an H100 in this dataset is $1.79 at Vast.ai. The dearest is $12.29 at Azure. Same silicon, same rate tier, 6.9x apart. A spread that wide on a fixed, fungible input is the definition of an inefficient market, and sourcing across it is the largest single lever on cost per token that does not require touching the model.
Silicon Data SDH100RT indices, distributed on Bloomberg, 2026-07-15. The only pair in this dataset where both legs were normalized by the same methodology on the same day.
Both legs non-interruptible and posted. This is the number that survives a like-for-like challenge.
Mixes a single marketplace spot listing against a hyperscaler list price. Shown for context only. Quoting it on its own would be dishonest, so we do not.
NVIDIA H100 80GB, USD per GPU-hour by provider and rate tier
Every quote names a provider, a rate tier, and a source. The two highlighted rows are the Silicon Data neocloud and hyperscaler indices: the only pair here normalized by one methodology on one day, and therefore the only pair whose ratio is a clean like-for-like reading.
| Item | Detail | Value ($) |
|---|---|---|
| Vast.ai | spot / marketplace | $0.45 |
| Compute Exchange (marketplace median) | reserved / marketplace | $1.70 |
| Vast.ai | on-demand / marketplace | $1.79 |
| RunPod Community | on-demand / marketplace | $1.99 |
| Nebius | preemptible / neocloud | $2.15 |
| SF Compute | clearing / marketplace | $2.16 |
| Azure | spot / hyperscaler | $2.27 |
| CoreWeave | spot / neocloud | $2.46 |
| AWS | spot / hyperscaler | $2.62 |
| Silicon Data neocloud index | on-demand / index | $2.70 |
| AWS | reserved / hyperscaler | $2.97 |
| Nebius | on-demand / neocloud | $3.85 |
| Crusoe | on-demand / neocloud | $3.90 |
| Lambda | on-demand / neocloud | $3.99 |
| Together | on-demand / neocloud | $3.99 |
| GCP | reserved / hyperscaler | $4.86 |
| AWS | capacity-block / hyperscaler | $5.19 |
| GCP | spot / hyperscaler | $6.02 |
| CoreWeave | on-demand / neocloud | $6.16 |
| AWS | on-demand / hyperscaler | $6.88 |
| Azure | on-demand / hyperscaler | $6.98 |
| Silicon Data hyperscaler index | on-demand / index | $7.24 |
| GCP | reserved / hyperscaler | $7.67 |
| GCP | on-demand / hyperscaler | $11.06 |
| Azure | on-demand / hyperscaler | $12.29 |
Three failure modes recur in public reporting of these figures: mixing rate tiers, mixing SKUs (H100 PCIe, SXM and NVL are three different products), and treating a posted price as a transacted one, which for the contract market it is not. Read the tier under each label before comparing any two bars.
Live provider pricing pages, published indices, and marketplace APIs. Snapshot 2026-07-20.
| Silicon | Cheapest observed | Cheapest on-demand | Dearest observed | On-demand spread |
|---|---|---|---|---|
| NVIDIA H100 80GB | $0.45reportedVast.ai, spot | $1.79reportedVast.ai, on-demand | $12.29Azure, on-demand | 6.9x |
| NVIDIA H200 141GB | $2.45Nebius, preemptible | $3.59RunPod Community, on-demand | $13.78reportedAzure, on-demand | 3.8x |
| NVIDIA B200 180GB | $3.67AWS, spot | $3.75reportedPacket.ai, on-demand | $14.24AWS, on-demand | 3.8x |
On-demand spread compares the dearest observation against the cheapest posted on-demand rate, so both legs are non-interruptible and directly comparable. Raw spread compares the cheapest observation of any kind, usually a single marketplace spot listing, against the dearest, and therefore mixes rate tiers. It is shown greyed because it should never be quoted on its own, and it is the number most often reproduced in press coverage of this market.
Why the spread does not close
Contract structure. A hyperscaler list price is not a clearing price. It is an administratively set queue price sitting above a large book of capacity already sold on multi-year commitments, and it moves when a committee decides it moves rather than when supply does. A marketplace clearing price moves hourly. Two different price formation processes over identical silicon is exactly the condition under which a large spread survives.
Utilization risk. The cheap end of the ladder is interruptible. That discount is payment for absorbing eviction, and most buyers cannot absorb it, because a serving stack that loses a replica mid-decode drops requests. Checkpointing KV state and draining a node inside a preemption notice is an engineering problem, and a solved one, but it is capital expenditure in software that most teams will not make for a workload they run occasionally.
Switching costs. Egress fees, committed spend agreements, security review of a new operator, and the engineering time to re-qualify a serving image on unfamiliar drivers all raise the threshold a discount has to clear before a team will act. A buyer already inside a hyperscaler commitment often cannot move at all without forfeiting the commitment. The spread persists because moving is expensive, not because anyone thinks the dear end is a fair price.
What a GPU-hour costs to produceest.
- H100, all in
- $1.45 / hr
- H200, all in
- $1.85 / hr
- B200, all in
- $2.07 / hr
- Assumes 70% fleet utilization. Across a 65-75% band the H100 hardware bucket moves only $1.04 to $1.20, so the conclusion is robust to any utilization assumption a competent operator would admit to.
- Electricity is a rounding error. An H100 at 700W with PUE 1.12 at Northern Virginia rates burns $0.045/hr of actual electricity, about 1% of the rental price. The facility slice is 4-6x larger than the electrons.
- Scarcity premium as a share of price: 52-64% at neoclouds, 77-88% at hyperscalers. There is no tier in this market, at any provider, where GPU-hours rent near cost.
- Two of the source's reference prices look stale or tier-ambiguous as of July 2026 (Lambda B200 at $5.50 against a live $6.69, CoreWeave H200 at $3.89 against a live $6.31).
Silicon Analysts, GPU Rental Price Decomposition, 2026-07-10
A posted price is not a clearing price
Between 1 Mar and 20 Apr 2026 this series was unchanged on 39 of 50 daily transitions, coefficient of variation 0.47%. It is not a price; it is a posted ceiling revised administratively above a large book of already-priced reserved capacity.
Silicon Data hyperscaler index, on-demand
The discount is payment for eviction risk
Same silicon, same provider, same region as the $12.29 on-demand row below. A 5.4x spread without changing vendors. The constraint is a 30-second eviction notice.
Azure, spot
Committing can cost more than not committing
Lambda's own 1-Click Cluster reservation product prices the same GPU at $5.54 to $6.16, i.e. committing costs 39% to 54% more than not committing.
Lambda, on-demand
03Unit economics
Six inputs, one output, and no step you cannot check
A cost per token figure with no stated throughput, no stated interactivity target, and no stated utilization is uninterpretable. Here is the whole model, the provenance of every input, and three cases worked end to end. The third is a back-test against a published production bill.
CPMT = 277.7778 x hourly x [ r * (1 - h) / P + 1 / D ] / U # prefill term: r * (1 - h) / P# decode term: 1 / D# result: USD per 1,000,000 output tokensThe first bracketed term is prefill, the second is decode. Every other technique in inference optimization is a modifier on one of those two denominators: quantization and speculative decoding raise D, prefix caching raises h, disaggregation lets P and D be served by different hardware. The constant is 1,000,000 / 3600, which converts a per second rate and a per hour price into a price per million tokens.
- hourlyRate paid for the accelerator
- USD / GPU-hr
- DDecode throughput at a stated interactivity target
- out tok/s/GPU
- PPrefill throughput
- in tok/s/GPU
- rInput to output token ratio
- ratio
- hShare of input served from prefix cache
- 0 to 1
- UFleet utilization, paid hours actually serving
- 0 to 1
Llama 3.3 70B
1x H200 SXM 141GB
Chat, 1K input / 1K output, no prefix cache, 100% utilization
- Geometry
- 70B dense, FP8 weights
- Prefill share
- 31%
- Stated result
- $0.704
Inputs
- hourly, rate paid
- $4.39
- D, decode throughput
- 2,500 tok/s
- P, prefill throughput
- 5,654 tok/sest.
- r, input to output ratio
- 1.00
- h, prefix cache hit rate
- 0.0%
- U, fleet utilization
- 100%
The arithmetic
277.7778 x 4.39 = 1219.44prefill 1219.44 x (1/5654) = 0.2157decode 1219.44 x (1/2500) = 0.4878total 0.7035 per 1M output tokens.Where each input comes from
hourly: RunPod H200 Secure, live page updated 2026-07-17
D: SemiAnalysis InferenceMAX, H200 at a 50 tok/s/user interactivity SLA
Interactivity is load bearing. The same model on the same GPU produces a different number at a different SLA, and a cost-per-token figure without an attached tok/s/user is uninterpretable.
P: Weights-only roofline at an assumed 40% MFUest.
The weakest input in this case. It ignores the quadratic attention term, which becomes material above roughly 8K context, so it is optimistic. Measure P directly before relying on it.
gpt-oss-120b
1x B200 SXM 180GB
Chat, 1K input / 1K output, no prefix cache, 100 tok/s/user SLA, 100% utilization
- Geometry
- 117B total / 5.1B active mixture of experts, MXFP4 weights
- Prefill share
- 10%
- Stated result
- $0.061
Inputs
- hourly, rate paid
- $5.89
- D, decode throughput
- 30,000 tok/s
- P, prefill throughput
- 264,700 tok/sest.
- r, input to output ratio
- 1.00
- h, prefix cache hit rate
- 0.0%
- U, fleet utilization
- 100%
The arithmetic
277.7778 x 5.89 = 1636.11prefill 1636.11 x (1/264700) = 0.00618decode 1636.11 x (1/30000) = 0.05454total 0.06072 per 1M output tokens.Where each input comes from
hourly: RunPod B200 Secure, live page updated 2026-07-17
D: NVIDIA / InferenceMAX, TensorRT-LLM with EAGLE-3 speculative decoding, at 100 tok/s/user
The same model on the same silicon went from 6,000 to 30,000 tok/s/GPU at fixed interactivity in two months of pure software work, Aug to Oct 2025. Any cost model with a hardcoded throughput constant has a shelf life measured in weeks.
P: Weights-only roofline at an assumed 30% MFUest.
Prefill scales with ACTIVE parameters, so a sparse 5.1B-active model prefills roughly 47x faster per GPU than a dense 70B. This is why long-context products have consolidated on mixture-of-experts architectures.
DeepSeek V3 / R1
H800 x8 nodes, 226.75 nodes average over 24 hours
Real production traffic, r = 3.62, 56.3% prefix cache hit rate, 20-22 tok/s/user
- Geometry
- 671B total / 37B active mixture of experts, FP8 weights, MLA attention
- Prefill share
- 42%
- Stated result
- $0.519
Inputs
- hourly, rate paid
- $2.00
- D, decode throughput
- 1,850 tok/s
- P, prefill throughput
- 4,026 tok/s
- r, input to output ratio
- 3.62
- h, prefix cache hit rate
- 56.3%
- U, fleet utilization
- 100%
The arithmetic
277.7778 x 2.00 = 555.556prefill 555.556 x (3.62 x 0.437 / 4026) = 0.21830decode 555.556 x (1/1850) = 0.30030total 0.51859 per 1M output tokens.Where each input comes from
hourly: DeepSeek's own stated H800 lease assumption
D: Published 14,800 output tok/s per 8-GPU node, divided by 8
P: Derived: published 73,700 input tok/s per node INCLUDES cache hits, so the true uncached rate is 73,700 x (1 - 0.563) / 8
The model forces this interpretation. Reading the 73.7k figure as excluding cache hits does not reconcile with the published bill.
Modelled $0.519 against an actual $0.518 per 1M output tokens.
The back-test, and its honest limit
Case three reproduces a real published bill to within 0.06%. That is not an independent oracle: the published cost figure is itself a pure GPU-hours calculation, so the match demonstrates that their published per node throughput, applied to their published traffic mix, exactly accounts for their published node-hours. What it does establish is that the prefill plus decode decomposition is complete, with no missing term large enough to show at that tolerance.
It also prices the cache. With the 56.3% prefix cache hit rate in place, prefill is 42% of cost against a 58% decode majority. Without it, prefill would have been 63%, because prefill work scales with uncached input and decode does not. Total cost per million output tokens moves from $0.519 with the cache to roughly $0.80 without it, so the cache is worth about 36% of their total serving cost.
DeepSeek open-infra-index, DeepSeek-V3/R1 Inference System Overview, 24 hours to 2025-02-28
Utilization is a division, and it is the term most often omitted
Every figure above assumes the GPU is serving. It usually is not. Cost per token divides by fleet utilization, so a serving stack twice as efficient as a competitor’s, running at 35 percent against their 80, is more expensive per token. This is the operational variable, and it is larger than the entire sourcing spread on the previous section.
| Fleet utilization | Cost multiplier | Cost / 1M output tokens |
|---|---|---|
| 100% | 1.00x | $0.704 |
| 80% | 1.25x | $0.879 |
| 65% | 1.54x | $1.08 |
| 60% | 1.67x | $1.17 |
| 40% | 2.50x | $1.76 |
| 20% | 5.00x | $3.52 |
Applied to the first worked case above, a Llama 3.3 70B on 1x H200 SXM 141GB. The top row is the roofline: every paid GPU-hour serving, which no fleet achieves. Measured enterprise Kubernetes GPU fleets sit an order of magnitude below it, and the difference between a well run fleet and a typical one is larger than the entire provider price spread in the section above. Cost per token divides by utilization, so this is a multiplier on whatever sourcing achieves, not an addition to it.
04What we do not claim
Silicon that is not in the fleet, and why
Absence here is a finding rather than an oversight. Each of these is either unshipped, unpriced in any public venue, or structurally impossible to place in a dollars per GPU-hour table without inventing a denominator. We would rather print the gap than fill it with a number nobody can check.
NVIDIA Vera Rubin VR200 / NVL72
Not in the fleetHas not shipped. NVIDIA's own Vera Rubin NVL72 page is marked 'Preliminary information, subject to change'; TDP is unpublished; no provider anywhere has a rate card. Lambda merchandises VR200 on a public Superclusters page, but that is a restatement of NVIDIA platform specs and is not evidence of racked hardware. A separate report that NVIDIA lowered the HBM4 target could not be read (HTTP 403).
Published specifications
288 GB HBM4, ~22 TB/s per GPU, 35 PFLOPS dense NVFP4, 3.6 TB/s NVLink 6, 1,580 TB/s aggregate across an NVL72 rack.
Vendor figures, restated. We have not measured any of them and do not present them as our own results.
AMD Instinct MI455X / Helios
Not in the fleetAnnounced for 2H 2026, not shipping. No rate card exists.
Published specifications
432 GB HBM4, 19.6 TB/s, 40 PFLOPS dense FP4, 72-GPU Helios rack.
Vendor figures, restated. We have not measured any of them and do not present them as our own results.
AWS Trainium2 / Trainium3
Not in the fleetNo defensible on-demand list price could be found for any Trainium SKU. AWS does not publish trn2/trn3 on the instance product pages, one aggregator's price field is visibly corrupt, and another paywalls it. The only observed figure is $8.5964/hr spot for trn2.48xlarge in us-east-2, which is $0.537 per chip-hour. Anyone quoting a Trainium $/hr should be asked for their source.
Published specifications
Trn2: 96 GB HBM3, 2.9 TB/s, 1.3 PFLOPS FP8. Trn3: 144 GB HBM3e, 4.9 TB/s, 2.517 PFLOPS MXFP8.
Vendor figures, restated. We have not measured any of them and do not present them as our own results.
Groq LPU, Cerebras WSE-3
Not in the fleetSold per token, not per hour, because per-chip capacity is tiny (230 MB SRAM per LPU, 44 GB per wafer) and the value is latency. Neither can be placed in a $/GPU-hr table without inventing a denominator. NVIDIA licensed Groq's LPU technology for roughly $20B in Dec 2025, which is the clearest signal that decode wants different silicon than prefill.
Published specifications
Groq LPU: 230 MB on-die SRAM, 80 TB/s on-die bandwidth, ridge point near zero. Cerebras WSE-3: 44 GB on-wafer SRAM, 21,000 TB/s on-wafer.
Vendor figures, restated. We have not measured any of them and do not present them as our own results.
Three further things we do not claim
- We do not claim to own any of this hardware. We source capacity across the market described above, and where a rate is a spot or preemptible rate we say so, because the interruption behaviour is the product difference.
- We do not claim a partnership, membership, or endorsement with any silicon vendor, cloud, or marketplace named on this page. They are hardware we run on and venues we buy from, which is a technical statement and not a commercial one.
- We do not claim these prices are stable. Every figure here has a shelf life measured in weeks, and the current memory shortage has moved individual SKUs by more than 100 percent inside a year. The snapshot date is printed on the table for that reason.
Price your own traffic against this model.
Send the shape rather than the data: input and output token mix, prefix cache hit rate, target tokens per second per user, and peak to average ratio. We will run it through the same six inputs used above, show the arithmetic, and tell you where our number is weakest before you find it yourself.