Pricing
Per token, per million, published. No surcharge you have to discover in a bill.
Every rate on this page is the rate. Input, cached input and output are metered per million tokens at a single price across the full context window, with no long-context tier, no separate cache-write fee and no regional premium. Where a competitor charges one of those, we say which, and how much.
01The rate card
Four SKUs, six numbers each
The whole Seldon price list fits in one table. Cached input is a flat tenth of input on every SKU, and the context column is the window you are actually billed at, not a headline maximum that reprices above a threshold.
| SKU | Input / 1M | Cached input / 1M | Output / 1M | Blended 3:1 | Context |
|---|---|---|---|---|---|
Seldon Prime 1Frontier | $1.60 | $0.160 | $3.60 | $2.10 | 1.0M |
Seldon Core 1Balanced | $1.20 | $0.120 | $3.60 | $1.80 | 1.0M |
Seldon Vector 1Fast | $0.120 | $0.012 | $0.480 | $0.210 | 131K |
Seldon Horizon 1Specialist | $0.650 | $0.065 | $3.20 | $1.29 | 1.0M |
Blended assumes a 3:1 input to output ratio, the convention used for cross-vendor comparison. Real production workloads in our economics dataset run between 40:1 and 187:1, where the input and cached-input columns dominate and the blended figure is close to meaningless. Use the calculator below rather than the blended column if you know your own ratio. Every Seldon figure on this page is the published rate you are billed at, not a headline that reprices under conditions listed elsewhere.
Identical on all 4 SKUs. OpenAI, Anthropic and Google have converged on the same ratio, so this is parity rather than an advantage. It is stated because a cache read is the majority of input volume on most production workloads.
A write bills as ordinary input. There is no separate cache-write line on a Seldon invoice.
Seldon Prime 1 serves its full window at one rate. Crossing a token threshold does not change the price of the request, retroactively or otherwise.
02Cost calculator
Put your own volumes in
Pick a workload shape or type your own monthly token volumes, choose what you are paying for today, and choose a Seldon SKU. The arithmetic is the same function that backs the worked examples in our economics dataset, so anything this returns can be checked by hand.
1,000 developer seats. A terminal or IDE agent that reads a repository, plans, edits files and runs tools across long autonomous sessions. Sold per developer seat.
Share of all input tokens served from prefix cache, not the share of requests that hit.
Priced volume
- Input
- 419.3B tokens
- Output
- 2.2B tokens
- Ratio
- 187:1
- Cache hits
- 94.8%
Claude Opus 4.8
Anthropic
Seldon Prime 1
Seldon, projected
Monthly inference spend
Monthly delta
$283,243
difference in inference spend per month
Annual delta
$3,398,921
same workload, twelve months
Bill reduction
72.7%
off the comparator’s modelled bill
Assumptions behind this model
- Scale unit is 1,000 developer seats, not monthly active users. Seats are how agentic coding is sold and how Anthropic reports its own consumption guidance.
- Session token shape is taken verbatim from Anthropic's Claude Code /usage documentation example: 1,200 fresh input, 940,000 cache read, 50,000 cache write, 5,300 output per session. That example reconciles arithmetically to its stated $0.55 on Sonnet 4.6 rates, which is why it is trusted as a statement of token shape.
- 23.5 session-equivalents per active day across 18 active days per month gives 423 sessions per seat per month.
- Per seat per month: 0.51 M fresh input, 397.6 M cache read, 21.2 M cache write, 2.24 M output. At 1,000 seats that is 419.278 B total input and 2.242 B output.
- Input to output ratio is 187:1. Cache hit rate is 94.8 percent, consistent with the University of Washington TraceLab study of roughly 4,300 real Claude Code and Codex sessions, which measured a 95.7 percent aggregate prefix cache hit rate over 55 B tokens.
- Volume cross-check: the model implies 23.4 M tokens per developer per active day. Anthropic's separately published per-user rate limit guidance implies roughly 24 M. Two unrelated first-party figures agree within 3 percent.
- Cost cross-check: priced on Sonnet 4.6 this model produces $234 per seat per month, inside Anthropic's published $150 to $250 per developer per month band and close to its $215 typical enterprise starting point.
- Revenue is MODELLED at $200 per seat per month. It is not a disclosed price for any named company. It is anchored to Anthropic's published enterprise consumption starting point of $215 per developer per month for Claude Code, which is evidence of demonstrated enterprise willingness to pay at that level.
- Non-inference COGS is modelled at 8 percent of revenue, covering hosting, eval engineering, observability and support. SFAI Labs puts the new AI-specific COGS lines (eval team 1 to 2.5 percent, observability about 1.1 percent, customer success uplift 0.8 to 1.5 percent) at roughly 3 to 5 points of revenue on top of conventional hosting and support.
- At $200 per seat this archetype has an Inference Efficiency Ratio of about 1.3:1 on frontier pricing, far below Ben Murray's 3:1 warning threshold for AI-native products. That is the structural reason coding agents are the hardest margin case in AI application software.
- Seldon rates are our published list prices. Comparator rates are published vendor list prices as of the July 2026 snapshot in lib/data/models and exclude negotiated enterprise discounts, which materially change the picture at large volume.
- Cache writes are billed at 1.25x base input on Anthropicand at 1.00x on Seldon. Where a vendor publishes no cache write multiplier the figure is left at 1.00x, which understates that vendor’s bill rather than inflating it.
- Anthropic’s newer models use a tokenizer that produces roughly 30 percent more tokens for the same text. A comparison run on a fixed body of text rather than a fixed token count would move in Seldon’s favour, and this calculator does not model that.
- Per-request modifiers are excluded on both legs: long-context surcharges, batch discounts and regional premiums. Section 04 below sets out which vendors charge which.
Two controls matter more than the rest. Cache hit rate moves the bill further than model choice on any workload with a long stable prefix, and the input to output ratio decides whether you are buying prefill or decode. Both are properties of your application, not of your vendor.
03Dedicated capacity
Per GPU-hour, when you can keep the GPUs busy
Serverless per-token pricing bundles idle time, autoscaling and multi-tenant risk pooling into the rate. Dedicated capacity unbundles them: you pay for hours, you keep whatever you can fit in them, and you carry the idle. That is cheaper only above a break-even utilization, and the break-even is computable.
| Accelerator | Bandwidth | On-demand / GPU-hr | Best observed floor |
|---|---|---|---|
H100 SXM5NVIDIA | 3.35 TB/s | $3.82reported | $2.15 |
H200 SXMNVIDIA | 4.80 TB/s | $4.32reported | $2.45 |
B200 (HGX, 1000 W)NVIDIA | 8.00 TB/s | $7.22reported | $3.60reported |
B300 (Blackwell Ultra, HGX)NVIDIA | 8.00 TB/s | $9.09reported | $4.30 |
GB200 NVL72 (per GPU)NVIDIA | 8.00 TB/s | $21.53reported | $10.50 |
Instinct MI300XAMD | 5.30 TB/s | $3.10reported | $1.65reported |
Instinct MI355XAMD | 8.00 TB/s | $8.60reported | $2.59reported |
On-demand is a market average across many providers, not one vendor's posted rate; the floor is the best reserved, spot or preemptible rate observed with real inventory behind it. Prices in this market have a shelf life measured in weeks, and the snapshot is 2026-07-20. Excluded from this table because no quotable public rate exists for them: GB300 NVL72 (per GPU), Instinct MI325X, TPU v7 (Ironwood). Their rows in our fleet dataset carry deliberate stand-in values so the dataset stays complete, and quoting those as prices would be wrong.
Effective cost per million output tokens, dedicated against per token
Worked on the one case where the architectures match: gpt-oss-120b on 1x B200 SXM 180GB, against the Seldon SKU built on the same weights. Dedicated cost is the roofline model divided by utilization, which is the term most often left out of published cost claims.
| Item | Value ($) |
|---|---|
| 82% utilization | $0.074 |
| 70% utilization | $0.087 |
| 66% utilization | $0.092 |
| 49% utilization | $0.124 |
| Seldon Vector 1, per token | $0.600 |
| 5% utilization | $1.21 |
Workload held constant across both legs: Chat, 1K input / 1K output, no prefix cache, 100 tok/s/user SLA, 100% utilization. Changing the input to output ratio or the cache hit rate moves the dedicated leg and the per-token leg by different amounts, so this break-even is specific to this workload shape.
Rate, throughput and utilization figures from lib/data/fleet.ts, snapshot 2026-07-20. Utilization levels are the measured and modelled points in that file, not illustrative round numbers.
Below this, Seldon Vector 1 at its published per-token rate is cheaper than renting the same silicon. Above it, dedicated wins, and keeps winning linearly.
The uncomfortable part is where real fleets sit. Cast AI’s telemetry across roughly 23,000 enterprise Kubernetes clusters measures 5% GPU utilization before optimisation, which is under the break-even. The best single cluster in that dataset sustains 49%. Dedicated capacity is a good trade for a team that has already solved batching and scheduling, and an expensive one for a team that has not.
How we source capacityWhat the break-even is built from
Six inputs, one output. Two of them are ours to improve and four are yours. The rate we pay per GPU-hour and the throughput our serving stack extracts from it set the numerator; your ratio, cache hit rate and utilization set the rest.
- Rate
- $5.89 / GPU-hr
- Decode
- 30,000 tok/s
- Prefill
- 264,700 tok/sest.
- At full utilization
- $0.061 / 1M out
- Per-token equivalent
- $0.600 / 1M out
The prefill figure is a roofline estimate rather than a measurement, and it ignores the quadratic attention term, so it is optimistic above roughly 8K of context. The decode figure is a vendor benchmark at a stated interactivity target. Neither is a promise about your traffic, which is why a dedicated quote starts with a week of your own.
When dedicated is the wrong answer
- Spiky or seasonal trafficReserved hours are paid whether or not they serve. A workload with a 10x diurnal swing pays for its peak around the clock, and the idle hours are pure loss against a per-token rate that simply stops charging.
- Traffic below one GPUDedicated capacity is quantised to whole accelerators. Below roughly one GPU of sustained load there is nothing to optimise, and the minimum unit is larger than the workload.
- A model mix that keeps changingReserved capacity is committed to a serving configuration. If you are still routing between tiers and reworking prompts weekly, the flexibility of per-token billing is worth more than the rate.
- No engineering owner for utilizationUtilization is a division in the cost model, which makes it the largest single term. A fleet nobody is measuring will land near the bottom of the published range rather than the middle.
We quote dedicated capacity when the traffic supports it and say so when it does not. Selling reserved hours to a workload that cannot fill them is a short relationship.
04What is not charged for
The three modifiers that do not appear on our invoice
A headline rate is not a price until you know the per-request modifiers attached to it. Three of them are common in this market. We levy none, and each competitor position below is stated as that vendor publishes it, including the cases where a vendor already does the same thing we do.
No surcharge above a token threshold
OpenAI applies 2x input and 1.5x output above 272K input tokens. Google applies roughly 2x input above 200K. xAI applies 2x to every token in the request once the prompt reaches 200K, so crossing the line reprices the whole call and not just the excess.
Anthropic charges no long-context premium and bills its full 1M window at standard rates. On this specific point Anthropic and Seldon behave the same way, and it would be dishonest to present it as something only we do.
Why it matters more than it sounds: a product that lets a user upload an arbitrarily large file crosses the threshold silently, and the request that crosses it is usually the one the user cares most about.
No premium for populating the cache
Anthropic bills a cache write at 1.25x base input on the default five-minute TTL and 2.00x on the one-hour TTL. OpenAI’s GPT-5.6 generation matched the 1.25x figure, having previously made writes free. Google’s implicit caching bills writes as ordinary input, as do the open-weight serverless providers.
Seldon bills a write at 1.00x, which is to say as input. The one-hour TTL premium pays for itself at two reads per write, so it is a defensible design rather than a trap, but it is still a line on a bill that our customers do not have.
Size of the effect, and it varies more than you would expect: on the AI coding agent archetype 98% of uncached input is cache writes, so ignoring the write premium understates that bill by about 7 percent. On the High-volume classification and extraction pipeline it is 2.9% and barely registers.
No regional or residency premium
Anthropic and Google both list roughly a 10 percent uplift for regional or data-residency endpoints against their global rate. That uplift is a real cost being passed through, not an invention: constrained-region capacity genuinely prices above global capacity.
We absorb it into one published rate instead, which means the rate has to hold across the whole fleet rather than only where capacity is cheapest. It is a narrower promise than it looks: one price, not a claim about where any given request runs.
We make no certification or residency guarantee here. What is and is not in place is set out in full on the trust page rather than implied by a pricing table.
A rate card with three asterisks is not cheaper than one without them. It is harder to compare, which is usually the point.
05Batch and volume
We do not offer a batch discount, and here is what that costs you
Anthropic, OpenAI and Google all halve the rate for asynchronous batch work. We do not. So the fair comparison against a batch-eligible workload is our list rate against their batch rate, not against their headline, and on that basis we lose some of the rows below.
| Model | List blended 3:1 | Batch blended | vs Seldon Prime 1 |
|---|---|---|---|
Claude Fable 5Anthropic | $20.00 | $10.00 | 4.8x over us |
Claude Opus 4.8Anthropic | $10.00 | $5.00 | 2.4x over us |
Claude Sonnet 5Anthropic | $4.00 | $2.00 | 1.1x under us |
Claude Haiku 4.5Anthropic | $2.00 | $1.00 | 2.1x under us |
gpt-5.6-solOpenAI | $11.25 | $5.63 | 2.7x over us |
gpt-5.6-terraOpenAI | $5.63 | $2.81 | 1.3x over us |
gpt-5.6-lunaOpenAI | $2.25 | $1.13 | 1.9x under us |
Gemini 3.1 ProGoogle | $4.50 | $2.25 | 1.1x over us |
Gemini 3.5 FlashGoogle | $3.38 | $1.69 | 1.2x under us |
Grok 4.5xAI | $3.00 | n/a | no batch rate |
Batch discounts are as published by each vendor for asynchronous work, typically with a 24-hour completion window and no latency guarantee. xAI publishes a smaller discount covering only a subset of its models, reported rather than verified, so no batch figure is shown for it here rather than a guessed one. Batch pricing only applies to work that can genuinely wait; comparing it against a synchronous rate for interactive traffic is not a like-for-like comparison in either direction.
Why there is no batch tier
A batch discount is a promise about scheduling: the operator fills troughs with deferred work and shares the gain. Making that promise requires a load curve steady enough to schedule against, and ours moves too much to price one honestly. Publishing a 50 percent batch rate against a trough we cannot reliably schedule into would be a marketing number rather than a scheduling decision.
The mechanism we offer instead is dedicated capacity, which is the same trade made explicit: you take the scheduling risk and keep the saving, rather than us estimating it on your behalf.
Volume
There is no published volume schedule. Committed volume is quoted against dedicated capacity, because that is where a commitment actually changes our cost: reserved GPU-hours price below on-demand, and a customer who can fill them lets us buy them.
Two things worth knowing when you compare. The list prices on this page and on every competitor rate card exclude negotiated enterprise discounts, which move materially at large volume. And the frontier-class closed SKUs run $3.00 to $20.00 blended against $2.10 for our frontier tier, so the first question is usually which tier your traffic needs rather than which vendor. Against the cheaper closed tiers the gap is much narrower than that range suggests.
Compare the catalog by tierSend a week of traffic and we price it.
Token volumes, input to output ratio, cache hit rate and latency targets are enough to produce a priced bill on both serverless and dedicated, with the break-even utilization computed against your own numbers rather than a worked example. If dedicated does not clear it, we tell you that instead.