Skip to content
Seldon

Fleet

The same GPU, at seven different prices, all week.

Serving cost is dominated by two things a model author never touches: what you pay for a GPU-hour, and how much of that hour is doing useful work. Both vary far more across the market than the hardware does. This page shows the dispersion, names both ends of it, and works the arithmetic from a rental rate to a cost per million tokens so it can be checked.

2.7x
Hyperscaler over neocloud, H100

$7.24 against $2.70 per GPU-hour on 2026-07-15. Both legs like-for-like normalized by the same index provider on the same day, which is what makes this the defensible number rather than the dramatic one.

01The accelerator lineup

Decode is a memory bandwidth problem, not a FLOPS problem

Autoregressive decode reads the whole weight set from HBM to produce one token. Arithmetic intensity is roughly one multiply-accumulate per parameter loaded, so the tensor cores idle while the memory system works, and output tokens per second per GPU tracks bandwidth rather than peak FLOPS. Prefill is the opposite: it is compute bound and batches well. A fleet tuned for cost per output token is therefore bought on TB/s, not on PFLOPS.

H100 SXM5 and H200 SXM share a die and an FP8 rating. The H200 carries 76% more memory and moves 43% more bytes per second, at identical compute, and that is the whole of the decode gain.

The inverse case is in the same table. GB300 NVL72 adds HBM capacity and 1.5x the dense FP4 rate over GB200 NVL72 while sitting at the same 8 TB/s per GPU. It is a capacity and compute upgrade that moves no additional bytes per second, which is why it changes what you can fit far more than it changes what you can decode.

AcceleratorHBMBandwidthTB/s per $/hrOn-demand / hrFloor / hr
H100 SXM5NVIDIA
80 GB3.35 TB/s0.88reported$3.82reported$2.15
H200 SXMNVIDIA
141 GB4.8 TB/s1.11reported$4.32reported$2.45
B200 (HGX, 1000 W)NVIDIAreported
180 GB8 TB/s1.11reported$7.22reported$3.60reported
B300 (Blackwell Ultra, HGX)NVIDIA
288 GB8 TB/s0.88reported$9.09reported$4.30
GB200 NVL72 (per GPU)NVIDIA
192 GB8 TB/s0.37reported$21.53reported$10.50
GB300 NVL72 (per GPU)NVIDIAest.
288 GB8 TB/s0.88est.$9.09est.$4.48est.
Instinct MI300XAMDreported
192 GB5.3 TB/s1.71reported$3.10reported$1.65reported
Instinct MI325XAMDest.
256 GB6 TB/s1.03est.$5.85est.$2.59est.
Instinct MI355XAMD
288 GB8 TB/s0.93reported$8.60reported$2.59reported
TPU v7 (Ironwood)Googlereported
192 GB7.38 TB/s4.61est.$1.60est.$1.60est.

All FLOPS figures are DENSE, in PFLOPS. Vendors headline structured 2:4 sparsity numbers that are essentially never used in production LLM serving, and where a datasheet published only a sparse figure the halved dense figure is marked as estimated. The on-demand column is a market average across many providers rather than any single rate card; named provider rates are in the next section, because the dispersion across providers is larger than the differences between these chips. Floor is the best observed spot, preemptible, or reserved rate. Bandwidth per dollar is derived from the two columns beside it and inherits the weaker confidence of the pair. Snapshot 2026-07-20.

Read the price columns carefully

  • GB300 NVL72 (per GPU): NO PUBLIC GB300 NVL72 RATE EXISTS. CoreWeave, Nebius, Lambda, AWS, Azure and GCP all say contact sales as of 2026-07-20. The HGX B300 street average is shown only so the row renders; the rack-scale SKU should be assumed to price above it, since GB200 NVL72 prices above HGX B200. Do not quote this as a GB300 price.
  • Instinct MI325X: NO PUBLIC $/GPU-HR RATE WAS FOUND for MI325X at any provider. It is listed at Vultr, TensorWave and DigitalOcean without a published rate. This is a bracket midpoint, not a price. Do not quote it.
  • TPU v7 (Ironwood): Ironwood is NOT on Google's public rate card. This is the only figure available and it is a negotiated large-account rate, not a price anyone can buy at self-serve. Published TPU rates for reference: v5e $1.20/chip-hr, v6e $2.70/chip-hr.

Runs on

Accelerators the serving stack is built and tuned for. Kernel work targets each generation directly rather than relying on a portability layer.

Vendor names in the table identify silicon we run on and tune kernels for. They are not partnerships, endorsements, or commercial relationships, and we do not describe them as any of those.

02Capacity arbitrage

An identical H100, priced across an order of magnitude

The cheapest posted on-demand rate for an H100 in this dataset is $1.79 at Vast.ai. The dearest is $12.29 at Azure. Same silicon, same rate tier, 6.9x apart. A spread that wide on a fixed, fungible input is the definition of an inefficient market, and sourcing across it is the largest single lever on cost per token that does not require touching the model.

2.7x
Standardized index spread

Silicon Data SDH100RT indices, distributed on Bloomberg, 2026-07-15. The only pair in this dataset where both legs were normalized by the same methodology on the same day.

6.9x
Cheapest to dearest, on-demand

Both legs non-interruptible and posted. This is the number that survives a like-for-like challenge.

27.3x
Cheapest to dearest, any tier

Mixes a single marketplace spot listing against a hyperscaler list price. Shown for context only. Quoting it on its own would be dishonest, so we do not.

NVIDIA H100 80GB, USD per GPU-hour by provider and rate tier

Every quote names a provider, a rate tier, and a source. The two highlighted rows are the Silicon Data neocloud and hyperscaler indices: the only pair here normalized by one methodology on one day, and therefore the only pair whose ratio is a clean like-for-like reading.

NVIDIA H100 80GB, USD per GPU-hour by provider and rate tier
ItemDetailValue ($)
Vast.aispot / marketplace$0.45
Compute Exchange (marketplace median)reserved / marketplace$1.70
Vast.aion-demand / marketplace$1.79
RunPod Communityon-demand / marketplace$1.99
Nebiuspreemptible / neocloud$2.15
SF Computeclearing / marketplace$2.16
Azurespot / hyperscaler$2.27
CoreWeavespot / neocloud$2.46
AWSspot / hyperscaler$2.62
Silicon Data neocloud indexon-demand / index$2.70
AWSreserved / hyperscaler$2.97
Nebiuson-demand / neocloud$3.85
Crusoeon-demand / neocloud$3.90
Lambdaon-demand / neocloud$3.99
Togetheron-demand / neocloud$3.99
GCPreserved / hyperscaler$4.86
AWScapacity-block / hyperscaler$5.19
GCPspot / hyperscaler$6.02
CoreWeaveon-demand / neocloud$6.16
AWSon-demand / hyperscaler$6.88
Azureon-demand / hyperscaler$6.98
Silicon Data hyperscaler indexon-demand / index$7.24
GCPreserved / hyperscaler$7.67
GCPon-demand / hyperscaler$11.06
Azureon-demand / hyperscaler$12.29

Three failure modes recur in public reporting of these figures: mixing rate tiers, mixing SKUs (H100 PCIe, SXM and NVL are three different products), and treating a posted price as a transacted one, which for the contract market it is not. Read the tier under each label before comparing any two bars.

Live provider pricing pages, published indices, and marketplace APIs. Snapshot 2026-07-20.

SiliconCheapest observedCheapest on-demandDearest observedOn-demand spread
NVIDIA H100 80GB
$0.45reportedVast.ai, spot
$1.79reportedVast.ai, on-demand
$12.29Azure, on-demand
6.9x
NVIDIA H200 141GB
$2.45Nebius, preemptible
$3.59RunPod Community, on-demand
$13.78reportedAzure, on-demand
3.8x
NVIDIA B200 180GB
$3.67AWS, spot
$3.75reportedPacket.ai, on-demand
$14.24AWS, on-demand
3.8x

On-demand spread compares the dearest observation against the cheapest posted on-demand rate, so both legs are non-interruptible and directly comparable. Raw spread compares the cheapest observation of any kind, usually a single marketplace spot listing, against the dearest, and therefore mixes rate tiers. It is shown greyed because it should never be quoted on its own, and it is the number most often reproduced in press coverage of this market.

Why the spread does not close

Contract structure. A hyperscaler list price is not a clearing price. It is an administratively set queue price sitting above a large book of capacity already sold on multi-year commitments, and it moves when a committee decides it moves rather than when supply does. A marketplace clearing price moves hourly. Two different price formation processes over identical silicon is exactly the condition under which a large spread survives.

Utilization risk. The cheap end of the ladder is interruptible. That discount is payment for absorbing eviction, and most buyers cannot absorb it, because a serving stack that loses a replica mid-decode drops requests. Checkpointing KV state and draining a node inside a preemption notice is an engineering problem, and a solved one, but it is capital expenditure in software that most teams will not make for a workload they run occasionally.

Switching costs. Egress fees, committed spend agreements, security review of a new operator, and the engineering time to re-qualify a serving image on unfamiliar drivers all raise the threshold a discount has to clear before a team will act. A buyer already inside a hyperscaler commitment often cannot move at all without forfeiting the commitment. The spread persists because moving is expensive, not because anyone thinks the dear end is a fair price.

What a GPU-hour costs to produceest.

H100, all in
$1.45 / hr
H200, all in
$1.85 / hr
B200, all in
$2.07 / hr
  • Assumes 70% fleet utilization. Across a 65-75% band the H100 hardware bucket moves only $1.04 to $1.20, so the conclusion is robust to any utilization assumption a competent operator would admit to.
  • Electricity is a rounding error. An H100 at 700W with PUE 1.12 at Northern Virginia rates burns $0.045/hr of actual electricity, about 1% of the rental price. The facility slice is 4-6x larger than the electrons.
  • Scarcity premium as a share of price: 52-64% at neoclouds, 77-88% at hyperscalers. There is no tier in this market, at any provider, where GPU-hours rent near cost.
  • Two of the source's reference prices look stale or tier-ambiguous as of July 2026 (Lambda B200 at $5.50 against a live $6.69, CoreWeave H200 at $3.89 against a live $6.31).

Silicon Analysts, GPU Rental Price Decomposition, 2026-07-10

A posted price is not a clearing price

Between 1 Mar and 20 Apr 2026 this series was unchanged on 39 of 50 daily transitions, coefficient of variation 0.47%. It is not a price; it is a posted ceiling revised administratively above a large book of already-priced reserved capacity.

Silicon Data hyperscaler index, on-demand

The discount is payment for eviction risk

Same silicon, same provider, same region as the $12.29 on-demand row below. A 5.4x spread without changing vendors. The constraint is a 30-second eviction notice.

Azure, spot

Committing can cost more than not committing

Lambda's own 1-Click Cluster reservation product prices the same GPU at $5.54 to $6.16, i.e. committing costs 39% to 54% more than not committing.

Lambda, on-demand

03Unit economics

Six inputs, one output, and no step you cannot check

A cost per token figure with no stated throughput, no stated interactivity target, and no stated utilization is uninterpretable. Here is the whole model, the provenance of every input, and three cases worked end to end. The third is a back-test against a published production bill.

cost per 1M output tokens
CPMT = 277.7778 x hourly x [ r * (1 - h) / P + 1 / D ] / U # prefill term:  r * (1 - h) / P# decode term:   1 / D# result:        USD per 1,000,000 output tokens

The first bracketed term is prefill, the second is decode. Every other technique in inference optimization is a modifier on one of those two denominators: quantization and speculative decoding raise D, prefix caching raises h, disaggregation lets P and D be served by different hardware. The constant is 1,000,000 / 3600, which converts a per second rate and a per hour price into a price per million tokens.

hourlyRate paid for the accelerator
USD / GPU-hr
DDecode throughput at a stated interactivity target
out tok/s/GPU
PPrefill throughput
in tok/s/GPU
rInput to output token ratio
ratio
hShare of input served from prefix cache
0 to 1
UFleet utilization, paid hours actually serving
0 to 1

Llama 3.3 70B

1x H200 SXM 141GB

$0.703
Cost per 1M output tokens

Chat, 1K input / 1K output, no prefix cache, 100% utilization

Geometry
70B dense, FP8 weights
Prefill share
31%
Stated result
$0.704

Inputs

hourly, rate paid
$4.39
D, decode throughput
2,500 tok/s
P, prefill throughput
5,654 tok/sest.
r, input to output ratio
1.00
h, prefix cache hit rate
0.0%
U, fleet utilization
100%

The arithmetic

277.7778 x 4.39 = 1219.44prefill 1219.44 x (1/5654) = 0.2157decode 1219.44 x (1/2500) = 0.4878total 0.7035 per 1M output tokens.

Where each input comes from

  • hourly: RunPod H200 Secure, live page updated 2026-07-17

  • D: SemiAnalysis InferenceMAX, H200 at a 50 tok/s/user interactivity SLA

    Interactivity is load bearing. The same model on the same GPU produces a different number at a different SLA, and a cost-per-token figure without an attached tok/s/user is uninterpretable.

  • P: Weights-only roofline at an assumed 40% MFUest.

    The weakest input in this case. It ignores the quadratic attention term, which becomes material above roughly 8K context, so it is optimistic. Measure P directly before relying on it.

gpt-oss-120b

1x B200 SXM 180GB

$0.061
Cost per 1M output tokens

Chat, 1K input / 1K output, no prefix cache, 100 tok/s/user SLA, 100% utilization

Geometry
117B total / 5.1B active mixture of experts, MXFP4 weights
Prefill share
10%
Stated result
$0.061

Inputs

hourly, rate paid
$5.89
D, decode throughput
30,000 tok/s
P, prefill throughput
264,700 tok/sest.
r, input to output ratio
1.00
h, prefix cache hit rate
0.0%
U, fleet utilization
100%

The arithmetic

277.7778 x 5.89 = 1636.11prefill 1636.11 x (1/264700) = 0.00618decode 1636.11 x (1/30000) = 0.05454total 0.06072 per 1M output tokens.

Where each input comes from

  • hourly: RunPod B200 Secure, live page updated 2026-07-17

  • D: NVIDIA / InferenceMAX, TensorRT-LLM with EAGLE-3 speculative decoding, at 100 tok/s/user

    The same model on the same silicon went from 6,000 to 30,000 tok/s/GPU at fixed interactivity in two months of pure software work, Aug to Oct 2025. Any cost model with a hardcoded throughput constant has a shelf life measured in weeks.

  • P: Weights-only roofline at an assumed 30% MFUest.

    Prefill scales with ACTIVE parameters, so a sparse 5.1B-active model prefills roughly 47x faster per GPU than a dense 70B. This is why long-context products have consolidated on mixture-of-experts architectures.

DeepSeek V3 / R1

H800 x8 nodes, 226.75 nodes average over 24 hours

$0.519
Cost per 1M output tokens

Real production traffic, r = 3.62, 56.3% prefix cache hit rate, 20-22 tok/s/user

Geometry
671B total / 37B active mixture of experts, FP8 weights, MLA attention
Prefill share
42%
Stated result
$0.519

Inputs

hourly, rate paid
$2.00
D, decode throughput
1,850 tok/s
P, prefill throughput
4,026 tok/s
r, input to output ratio
3.62
h, prefix cache hit rate
56.3%
U, fleet utilization
100%

The arithmetic

277.7778 x 2.00 = 555.556prefill 555.556 x (3.62 x 0.437 / 4026) = 0.21830decode 555.556 x (1/1850) = 0.30030total 0.51859 per 1M output tokens.

Where each input comes from

  • hourly: DeepSeek's own stated H800 lease assumption

  • D: Published 14,800 output tok/s per 8-GPU node, divided by 8

  • P: Derived: published 73,700 input tok/s per node INCLUDES cache hits, so the true uncached rate is 73,700 x (1 - 0.563) / 8

    The model forces this interpretation. Reading the 73.7k figure as excluding cache hits does not reconcile with the published bill.

0.06%
Model against a published bill

Modelled $0.519 against an actual $0.518 per 1M output tokens.

The back-test, and its honest limit

Case three reproduces a real published bill to within 0.06%. That is not an independent oracle: the published cost figure is itself a pure GPU-hours calculation, so the match demonstrates that their published per node throughput, applied to their published traffic mix, exactly accounts for their published node-hours. What it does establish is that the prefill plus decode decomposition is complete, with no missing term large enough to show at that tolerance.

It also prices the cache. With the 56.3% prefix cache hit rate in place, prefill is 42% of cost against a 58% decode majority. Without it, prefill would have been 63%, because prefill work scales with uncached input and decode does not. Total cost per million output tokens moves from $0.519 with the cache to roughly $0.80 without it, so the cache is worth about 36% of their total serving cost.

DeepSeek open-infra-index, DeepSeek-V3/R1 Inference System Overview, 24 hours to 2025-02-28

Utilization is a division, and it is the term most often omitted

Every figure above assumes the GPU is serving. It usually is not. Cost per token divides by fleet utilization, so a serving stack twice as efficient as a competitor’s, running at 35 percent against their 80, is more expensive per token. This is the operational variable, and it is larger than the entire sourcing spread on the previous section.

Fleet utilizationCost multiplierCost / 1M output tokens
100%1.00x$0.704
80%1.25x$0.879
65%1.54x$1.08
60%1.67x$1.17
40%2.50x$1.76
20%5.00x$3.52

Applied to the first worked case above, a Llama 3.3 70B on 1x H200 SXM 141GB. The top row is the roofline: every paid GPU-hour serving, which no fleet achieves. Measured enterprise Kubernetes GPU fleets sit an order of magnitude below it, and the difference between a well run fleet and a typical one is larger than the entire provider price spread in the section above. Cost per token divides by utilization, so this is a multiplier on whatever sourcing achieves, not an addition to it.

04What we do not claim

Silicon that is not in the fleet, and why

Absence here is a finding rather than an oversight. Each of these is either unshipped, unpriced in any public venue, or structurally impossible to place in a dollars per GPU-hour table without inventing a denominator. We would rather print the gap than fill it with a number nobody can check.

NVIDIA Vera Rubin VR200 / NVL72

Not in the fleet

Has not shipped. NVIDIA's own Vera Rubin NVL72 page is marked 'Preliminary information, subject to change'; TDP is unpublished; no provider anywhere has a rate card. Lambda merchandises VR200 on a public Superclusters page, but that is a restatement of NVIDIA platform specs and is not evidence of racked hardware. A separate report that NVIDIA lowered the HBM4 target could not be read (HTTP 403).

Published specifications

288 GB HBM4, ~22 TB/s per GPU, 35 PFLOPS dense NVFP4, 3.6 TB/s NVLink 6, 1,580 TB/s aggregate across an NVL72 rack.

Vendor figures, restated. We have not measured any of them and do not present them as our own results.

AMD Instinct MI455X / Helios

Not in the fleet

Announced for 2H 2026, not shipping. No rate card exists.

Published specifications

432 GB HBM4, 19.6 TB/s, 40 PFLOPS dense FP4, 72-GPU Helios rack.

Vendor figures, restated. We have not measured any of them and do not present them as our own results.

AWS Trainium2 / Trainium3

Not in the fleet

No defensible on-demand list price could be found for any Trainium SKU. AWS does not publish trn2/trn3 on the instance product pages, one aggregator's price field is visibly corrupt, and another paywalls it. The only observed figure is $8.5964/hr spot for trn2.48xlarge in us-east-2, which is $0.537 per chip-hour. Anyone quoting a Trainium $/hr should be asked for their source.

Published specifications

Trn2: 96 GB HBM3, 2.9 TB/s, 1.3 PFLOPS FP8. Trn3: 144 GB HBM3e, 4.9 TB/s, 2.517 PFLOPS MXFP8.

Vendor figures, restated. We have not measured any of them and do not present them as our own results.

Groq LPU, Cerebras WSE-3

Not in the fleet

Sold per token, not per hour, because per-chip capacity is tiny (230 MB SRAM per LPU, 44 GB per wafer) and the value is latency. Neither can be placed in a $/GPU-hr table without inventing a denominator. NVIDIA licensed Groq's LPU technology for roughly $20B in Dec 2025, which is the clearest signal that decode wants different silicon than prefill.

Published specifications

Groq LPU: 230 MB on-die SRAM, 80 TB/s on-die bandwidth, ridge point near zero. Cerebras WSE-3: 44 GB on-wafer SRAM, 21,000 TB/s on-wafer.

Vendor figures, restated. We have not measured any of them and do not present them as our own results.

Three further things we do not claim

  • We do not claim to own any of this hardware. We source capacity across the market described above, and where a rate is a spot or preemptible rate we say so, because the interruption behaviour is the product difference.
  • We do not claim a partnership, membership, or endorsement with any silicon vendor, cloud, or marketplace named on this page. They are hardware we run on and venues we buy from, which is a technical statement and not a commercial one.
  • We do not claim these prices are stable. Every figure here has a shelf life measured in weeks, and the current memory shortage has moved individual SKUs by more than 100 percent inside a year. The snapshot date is printed on the table for that reason.

Price your own traffic against this model.

Send the shape rather than the data: input and output token mix, prefix cache hit rate, target tokens per second per user, and peak to average ratio. We will run it through the same six inputs used above, show the arithmetic, and tell you where our number is weakest before you find it yourself.