Skip to content
Seldon

Pricing

Per token, per million, published. No surcharge you have to discover in a bill.

Every rate on this page is the rate. Input, cached input and output are metered per million tokens at a single price across the full context window, with no long-context tier, no separate cache-write fee and no regional premium. Where a competitor charges one of those, we say which, and how much.

01The rate card

Four SKUs, six numbers each

The whole Seldon price list fits in one table. Cached input is a flat tenth of input on every SKU, and the context column is the window you are actually billed at, not a headline maximum that reprices above a threshold.

SKUInput / 1MCached input / 1MOutput / 1MBlended 3:1Context
Seldon Prime 1Frontier
$1.60$0.160$3.60$2.101.0M
Seldon Core 1Balanced
$1.20$0.120$3.60$1.801.0M
Seldon Vector 1Fast
$0.120$0.012$0.480$0.210131K
Seldon Horizon 1Specialist
$0.650$0.065$3.20$1.291.0M

Blended assumes a 3:1 input to output ratio, the convention used for cross-vendor comparison. Real production workloads in our economics dataset run between 40:1 and 187:1, where the input and cached-input columns dominate and the blended figure is close to meaningless. Use the calculator below rather than the blended column if you know your own ratio. Every Seldon figure on this page is the published rate you are billed at, not a headline that reprices under conditions listed elsewhere.

10%
Cached input, as a share of input

Identical on all 4 SKUs. OpenAI, Anthropic and Google have converged on the same ratio, so this is parity rather than an advantage. It is stated because a cache read is the majority of input volume on most production workloads.

1.00x
Cache write, as a multiple of input

A write bills as ordinary input. There is no separate cache-write line on a Seldon invoice.

1.0M
Longest window, billed flat

Seldon Prime 1 serves its full window at one rate. Crossing a token threshold does not change the price of the request, retroactively or otherwise.

02Cost calculator

Put your own volumes in

Pick a workload shape or type your own monthly token volumes, choose what you are paying for today, and choose a Seldon SKU. The arithmetic is the same function that backs the worked examples in our economics dataset, so anything this returns can be checked by hand.

1,000 developer seats. A terminal or IDE agent that reads a repository, plans, edits files and runs tools across long autonomous sessions. Sold per developer seat.

Cache hit rate94.8%

Share of all input tokens served from prefix cache, not the share of requests that hit.

Priced volume

Input
419.3B tokens
Output
2.2B tokens
Ratio
187:1
Cache hits
94.8%

Claude Opus 4.8

Anthropic

$389,586monthly inference cost
$0.924effective, per 1M tokens billed

Seldon Prime 1

Seldon, projected

$106,343monthly inference cost
$0.252effective, per 1M tokens billed

Monthly inference spend

Claude Opus 4.8
$389,586
Seldon Prime 1
$106,343

Monthly delta

$283,243

difference in inference spend per month

Annual delta

$3,398,921

same workload, twelve months

Bill reduction

72.7%

off the comparator’s modelled bill

Assumptions behind this model
  • Scale unit is 1,000 developer seats, not monthly active users. Seats are how agentic coding is sold and how Anthropic reports its own consumption guidance.
  • Session token shape is taken verbatim from Anthropic's Claude Code /usage documentation example: 1,200 fresh input, 940,000 cache read, 50,000 cache write, 5,300 output per session. That example reconciles arithmetically to its stated $0.55 on Sonnet 4.6 rates, which is why it is trusted as a statement of token shape.
  • 23.5 session-equivalents per active day across 18 active days per month gives 423 sessions per seat per month.
  • Per seat per month: 0.51 M fresh input, 397.6 M cache read, 21.2 M cache write, 2.24 M output. At 1,000 seats that is 419.278 B total input and 2.242 B output.
  • Input to output ratio is 187:1. Cache hit rate is 94.8 percent, consistent with the University of Washington TraceLab study of roughly 4,300 real Claude Code and Codex sessions, which measured a 95.7 percent aggregate prefix cache hit rate over 55 B tokens.
  • Volume cross-check: the model implies 23.4 M tokens per developer per active day. Anthropic's separately published per-user rate limit guidance implies roughly 24 M. Two unrelated first-party figures agree within 3 percent.
  • Cost cross-check: priced on Sonnet 4.6 this model produces $234 per seat per month, inside Anthropic's published $150 to $250 per developer per month band and close to its $215 typical enterprise starting point.
  • Revenue is MODELLED at $200 per seat per month. It is not a disclosed price for any named company. It is anchored to Anthropic's published enterprise consumption starting point of $215 per developer per month for Claude Code, which is evidence of demonstrated enterprise willingness to pay at that level.
  • Non-inference COGS is modelled at 8 percent of revenue, covering hosting, eval engineering, observability and support. SFAI Labs puts the new AI-specific COGS lines (eval team 1 to 2.5 percent, observability about 1.1 percent, customer success uplift 0.8 to 1.5 percent) at roughly 3 to 5 points of revenue on top of conventional hosting and support.
  • At $200 per seat this archetype has an Inference Efficiency Ratio of about 1.3:1 on frontier pricing, far below Ben Murray's 3:1 warning threshold for AI-native products. That is the structural reason coding agents are the hardest margin case in AI application software.
  • Seldon rates are our published list prices. Comparator rates are published vendor list prices as of the July 2026 snapshot in lib/data/models and exclude negotiated enterprise discounts, which materially change the picture at large volume.
  • Cache writes are billed at 1.25x base input on Anthropicand at 1.00x on Seldon. Where a vendor publishes no cache write multiplier the figure is left at 1.00x, which understates that vendor’s bill rather than inflating it.
  • Anthropic’s newer models use a tokenizer that produces roughly 30 percent more tokens for the same text. A comparison run on a fixed body of text rather than a fixed token count would move in Seldon’s favour, and this calculator does not model that.
  • Per-request modifiers are excluded on both legs: long-context surcharges, batch discounts and regional premiums. Section 04 below sets out which vendors charge which.

Two controls matter more than the rest. Cache hit rate moves the bill further than model choice on any workload with a long stable prefix, and the input to output ratio decides whether you are buying prefill or decode. Both are properties of your application, not of your vendor.

03Dedicated capacity

Per GPU-hour, when you can keep the GPUs busy

Serverless per-token pricing bundles idle time, autoscaling and multi-tenant risk pooling into the rate. Dedicated capacity unbundles them: you pay for hours, you keep whatever you can fit in them, and you carry the idle. That is cheaper only above a break-even utilization, and the break-even is computable.

What the capacity market charges per GPU-hour. A dedicated quote is a function of these rates plus our serving stack, not a separate list price we set independently of them.
AcceleratorBandwidthOn-demand / GPU-hrBest observed floor
H100 SXM5NVIDIA
3.35 TB/s$3.82reported$2.15
H200 SXMNVIDIA
4.80 TB/s$4.32reported$2.45
B200 (HGX, 1000 W)NVIDIA
8.00 TB/s$7.22reported$3.60reported
B300 (Blackwell Ultra, HGX)NVIDIA
8.00 TB/s$9.09reported$4.30
GB200 NVL72 (per GPU)NVIDIA
8.00 TB/s$21.53reported$10.50
Instinct MI300XAMD
5.30 TB/s$3.10reported$1.65reported
Instinct MI355XAMD
8.00 TB/s$8.60reported$2.59reported

On-demand is a market average across many providers, not one vendor's posted rate; the floor is the best reserved, spot or preemptible rate observed with real inventory behind it. Prices in this market have a shelf life measured in weeks, and the snapshot is 2026-07-20. Excluded from this table because no quotable public rate exists for them: GB300 NVL72 (per GPU), Instinct MI325X, TPU v7 (Ironwood). Their rows in our fleet dataset carry deliberate stand-in values so the dataset stays complete, and quoting those as prices would be wrong.

Effective cost per million output tokens, dedicated against per token

Worked on the one case where the architectures match: gpt-oss-120b on 1x B200 SXM 180GB, against the Seldon SKU built on the same weights. Dedicated cost is the roofline model divided by utilization, which is the term most often left out of published cost claims.

Effective cost per million output tokens, dedicated against per token
ItemValue ($)
82% utilization$0.074
70% utilization$0.087
66% utilization$0.092
49% utilization$0.124
Seldon Vector 1, per token$0.600
5% utilization$1.21

Workload held constant across both legs: Chat, 1K input / 1K output, no prefix cache, 100 tok/s/user SLA, 100% utilization. Changing the input to output ratio or the cache hit rate moves the dedicated leg and the per-token leg by different amounts, so this break-even is specific to this workload shape.

Rate, throughput and utilization figures from lib/data/fleet.ts, snapshot 2026-07-20. Utilization levels are the measured and modelled points in that file, not illustrative round numbers.

10%
Break-even utilization

Below this, Seldon Vector 1 at its published per-token rate is cheaper than renting the same silicon. Above it, dedicated wins, and keeps winning linearly.

The uncomfortable part is where real fleets sit. Cast AI’s telemetry across roughly 23,000 enterprise Kubernetes clusters measures 5% GPU utilization before optimisation, which is under the break-even. The best single cluster in that dataset sustains 49%. Dedicated capacity is a good trade for a team that has already solved batching and scheduling, and an expensive one for a team that has not.

How we source capacity

What the break-even is built from

Six inputs, one output. Two of them are ours to improve and four are yours. The rate we pay per GPU-hour and the throughput our serving stack extracts from it set the numerator; your ratio, cache hit rate and utilization set the rest.

Rate
$5.89 / GPU-hr
Decode
30,000 tok/s
Prefill
264,700 tok/sest.
At full utilization
$0.061 / 1M out
Per-token equivalent
$0.600 / 1M out

The prefill figure is a roofline estimate rather than a measurement, and it ignores the quadratic attention term, so it is optimistic above roughly 8K of context. The decode figure is a vendor benchmark at a stated interactivity target. Neither is a promise about your traffic, which is why a dedicated quote starts with a week of your own.

When dedicated is the wrong answer

  • Spiky or seasonal trafficReserved hours are paid whether or not they serve. A workload with a 10x diurnal swing pays for its peak around the clock, and the idle hours are pure loss against a per-token rate that simply stops charging.
  • Traffic below one GPUDedicated capacity is quantised to whole accelerators. Below roughly one GPU of sustained load there is nothing to optimise, and the minimum unit is larger than the workload.
  • A model mix that keeps changingReserved capacity is committed to a serving configuration. If you are still routing between tiers and reworking prompts weekly, the flexibility of per-token billing is worth more than the rate.
  • No engineering owner for utilizationUtilization is a division in the cost model, which makes it the largest single term. A fleet nobody is measuring will land near the bottom of the published range rather than the middle.

We quote dedicated capacity when the traffic supports it and say so when it does not. Selling reserved hours to a workload that cannot fill them is a short relationship.

04What is not charged for

The three modifiers that do not appear on our invoice

A headline rate is not a price until you know the per-request modifiers attached to it. Three of them are common in this market. We levy none, and each competitor position below is stated as that vendor publishes it, including the cases where a vendor already does the same thing we do.

Long context

No surcharge above a token threshold

OpenAI applies 2x input and 1.5x output above 272K input tokens. Google applies roughly 2x input above 200K. xAI applies 2x to every token in the request once the prompt reaches 200K, so crossing the line reprices the whole call and not just the excess.

Anthropic charges no long-context premium and bills its full 1M window at standard rates. On this specific point Anthropic and Seldon behave the same way, and it would be dishonest to present it as something only we do.

Why it matters more than it sounds: a product that lets a user upload an arbitrarily large file crosses the threshold silently, and the request that crosses it is usually the one the user cares most about.

Cache writes

No premium for populating the cache

Anthropic bills a cache write at 1.25x base input on the default five-minute TTL and 2.00x on the one-hour TTL. OpenAI’s GPT-5.6 generation matched the 1.25x figure, having previously made writes free. Google’s implicit caching bills writes as ordinary input, as do the open-weight serverless providers.

Seldon bills a write at 1.00x, which is to say as input. The one-hour TTL premium pays for itself at two reads per write, so it is a defensible design rather than a trap, but it is still a line on a bill that our customers do not have.

Size of the effect, and it varies more than you would expect: on the AI coding agent archetype 98% of uncached input is cache writes, so ignoring the write premium understates that bill by about 7 percent. On the High-volume classification and extraction pipeline it is 2.9% and barely registers.

Region

No regional or residency premium

Anthropic and Google both list roughly a 10 percent uplift for regional or data-residency endpoints against their global rate. That uplift is a real cost being passed through, not an invention: constrained-region capacity genuinely prices above global capacity.

We absorb it into one published rate instead, which means the rate has to hold across the whole fleet rather than only where capacity is cheapest. It is a narrower promise than it looks: one price, not a claim about where any given request runs.

We make no certification or residency guarantee here. What is and is not in place is set out in full on the trust page rather than implied by a pricing table.

A rate card with three asterisks is not cheaper than one without them. It is harder to compare, which is usually the point.

05Batch and volume

We do not offer a batch discount, and here is what that costs you

Anthropic, OpenAI and Google all halve the rate for asynchronous batch work. We do not. So the fair comparison against a batch-eligible workload is our list rate against their batch rate, not against their headline, and on that basis we lose some of the rows below.

Every closed frontier SKU at its batch rate, against Seldon Prime 1 at $2.10 blended with no discount applied. Red rows are the ones where a batch-eligible workload is genuinely cheaper elsewhere.
ModelList blended 3:1Batch blendedvs Seldon Prime 1
Claude Fable 5Anthropic
$20.00$10.004.8x over us
Claude Opus 4.8Anthropic
$10.00$5.002.4x over us
Claude Sonnet 5Anthropic
$4.00$2.001.1x under us
Claude Haiku 4.5Anthropic
$2.00$1.002.1x under us
gpt-5.6-solOpenAI
$11.25$5.632.7x over us
gpt-5.6-terraOpenAI
$5.63$2.811.3x over us
gpt-5.6-lunaOpenAI
$2.25$1.131.9x under us
Gemini 3.1 ProGoogle
$4.50$2.251.1x over us
Gemini 3.5 FlashGoogle
$3.38$1.691.2x under us
Grok 4.5xAI
$3.00n/ano batch rate

Batch discounts are as published by each vendor for asynchronous work, typically with a 24-hour completion window and no latency guarantee. xAI publishes a smaller discount covering only a subset of its models, reported rather than verified, so no batch figure is shown for it here rather than a guessed one. Batch pricing only applies to work that can genuinely wait; comparing it against a synchronous rate for interactive traffic is not a like-for-like comparison in either direction.

Why there is no batch tier

A batch discount is a promise about scheduling: the operator fills troughs with deferred work and shares the gain. Making that promise requires a load curve steady enough to schedule against, and ours moves too much to price one honestly. Publishing a 50 percent batch rate against a trough we cannot reliably schedule into would be a marketing number rather than a scheduling decision.

The mechanism we offer instead is dedicated capacity, which is the same trade made explicit: you take the scheduling risk and keep the saving, rather than us estimating it on your behalf.

Volume

There is no published volume schedule. Committed volume is quoted against dedicated capacity, because that is where a commitment actually changes our cost: reserved GPU-hours price below on-demand, and a customer who can fill them lets us buy them.

Two things worth knowing when you compare. The list prices on this page and on every competitor rate card exclude negotiated enterprise discounts, which move materially at large volume. And the frontier-class closed SKUs run $3.00 to $20.00 blended against $2.10 for our frontier tier, so the first question is usually which tier your traffic needs rather than which vendor. Against the cheaper closed tiers the gap is much narrower than that range suggests.

Compare the catalog by tier

Send a week of traffic and we price it.

Token volumes, input to output ratio, cache hit rate and latency targets are enough to produce a priced bill on both serverless and dedicated, with the break-even utilization computed against your own numbers rather than a worked example. If dedicated does not clear it, we tell you that instead.