Comparison
Where we win, where we do not, and how to check.
Every rate on this page comes from a published vendor rate card and every throughput figure from an independent measurement, both cited inline. Our own rows are the rates we bill at and the throughput we serve at. Section 04 tells you how to reproduce the whole comparison without taking our word for any of it.
01Against the closed frontier
The rate cards, side by side
Anthropic, OpenAI, Google and xAI publish list prices. So do we. The table below puts every current closed frontier SKU next to every Seldon tier, at the same blended ratio, with no selective omissions.
| Model | Input / 1M | Output / 1M | Blended 3:1 | vs tier floor |
|---|---|---|---|---|
Seldon Prime 1Seldon | $1.60 | $3.60 | $2.10 | floor |
Grok 4.5xAI | $2.00 | $6.00 | $3.00 | 1.4x |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | $4.50 | 2.1x |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | $10.00 | 4.8x |
gpt-5.6-solOpenAI | $5.00 | $30.00 | $11.25 | 5.4x |
Claude Fable 5Anthropic | $10.00 | $50.00 | $20.00 | 9.5x |
Seldon Core 1Seldon | $1.20 | $3.60 | $1.80 | floor |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | $3.38 | 1.9x |
Claude Sonnet 5Anthropic | $2.00 | $10.00 | $4.00 | 2.2x |
gpt-5.6-terraOpenAI | $2.50 | $15.00 | $5.63 | 3.1x |
Seldon Vector 1Seldon | $0.120 | $0.480 | $0.210 | floor |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | $2.00 | 9.5x |
gpt-5.6-lunaOpenAI | $1.00 | $6.00 | $2.25 | 10.7x |
Seldon Horizon 1Seldon | $0.650 | $3.20 | $1.29 | floor |
Blended rate assumes a 3:1 input to output ratio, which is the conventional cross-vendor basis and is not what production agentic traffic looks like; real workloads run 40:1 to 187:1, and the margin model page prices them that way instead. The ratio column compares each row to the cheapest blended rate in its own tier, and tier is the vendor's own positioning rather than a measured capability match. Three modifiers are excluded because they are per-request rather than per-SKU: long-context surcharges at OpenAI, Google and xAI, batch discounts of 50 percent at Anthropic, OpenAI and Google, and regional premiums. All three are documented on the pricing page and the batch discount in particular narrows the gap to us. Anthropic's newer models use a tokenizer that produces roughly 30 percent more tokens for the same text, so headline rates are not directly comparable across vendors without adjusting for it. Claude Sonnet 5's rate is introductory and rises 50 percent on 2026-09-01.
What the closed frontier models are genuinely better at
Three things, and they are not small. The hardest reasoning work, where the marginal point of accuracy is worth an order of magnitude in price. Long-horizon agentic reliability, meaning the ability to run for hours across dozens of tool calls without losing the thread. And tool use itself, where the closed labs have spent years hardening function calling, structured output and error recovery against a volume and variety of real traffic that we do not have.
The evidence for that sits in our own dataset, not in a concession we invented for balance. Claude Opus 4.8 is the only model in this comparison with independent benchmark evidence in the research, and it leads every long-horizon agentic column it appears in. The strongest open-weight model we serve beats it on competition mathematics and loses to it badly on multi-session software engineering. Those two facts belong in the same paragraph.
Our argument is not that the closed frontier is overpriced. It is that most production traffic is not the hardest reasoning work, and paying frontier rates for the portion that is not is a choice rather than a requirement.
See the full model catalogThe only closed model in this table with independent benchmark evidence in the research: it leads SWE-bench Pro, NL2Repo, ProgramBench, FrontierSWE and SWE-Marathon in the GLM-5.2 cross-vendor table. A separate fast-mode SKU is priced at $10 in / $50 out.
Catalog note, lib/data/models.ts
Active parameters are not published by Z.ai anywhere the research could reach, so the field is omitted rather than guessed. Total is from the Hugging Face safetensors index. Beats Claude Opus 4.8 on AIME 2026 and IMOAnswerBench, loses SWE-Marathon 13.0 to 26.0.
Catalog note, lib/data/models.ts
02Against other inference hosts
The same weights, priced six ways
DeepInfra, Together, Fireworks, Groq and Baseten are peers, not opponents. We are all serving open weights that none of us trained, under licenses none of us control. Nobody in this group has a model advantage over anybody else, so the only things left to compete on are the serving stack and the price of capacity.
gpt-oss-120b, blended price per million tokens by host
The only model every major host publishes a rate for, which makes it the one clean cross-host comparison available. Groq, Together, Fireworks and DeepInfra's Turbo SKU all charge exactly the same rate on it.
| Item | Detail | Value ($) |
|---|---|---|
| DeepInfra (bf16) | $0.070 | |
| Seldon | 420 tok/s | $0.210 |
| DeepInfra (Turbo) | 319 tok/s | $0.262 |
| Together AI | 583 tok/s | $0.262 |
| Fireworks AI | $0.262 | |
| Groq | 479 tok/s | $0.262 |
Vendor rate cards, July 2026. Throughput from Artificial Analysis measurements rather than vendor marketing claims. The Seldon row is our own published rate.
Between the cheapest and dearest published rate for identical Apache 2.0 weights.
Note which row is cheapest. It is not ours. DeepInfra’s bfloat16 SKU is the floor on this model and it is a higher-precision SKU than most of the rows above it, which is the clearest evidence available that published inference price tracks the operator’s cost structure rather than anything about the model.
The differentiator we are claiming is narrow and stated plainly: a serving stack tuned per architecture, and capacity bought across more markets than a single-cloud operator can reach. Neither of those is a moat made of weights, because there is no such moat available to anyone here.
How we source capacityOne host publishes two SKUs for identical weights, 3.7x apart, against a 3.7xspread across every host in the chart. A market where a single vendor’s own price list is as wide as the market itself is not one where anyone should assume list price reflects marginal cost, ours included.
Baseten belongs in this group and has no bar, because our dataset carries no separately published rate for it on this model. What the research does record is that Baseten, Together and Fireworks converge to the cent on the same rate, which is the more interesting finding: three independent companies landing on an identical price is evidence of a reference rate, not of three independent cost structures.
03On speed
Groq and Cerebras are faster than us
That is not a hedge, it is the measurement. On raw output tokens per second for the models they serve, dedicated inference silicon beats a GPU fleet, and it is not close. Any page that told you otherwise would be checkable and wrong within a minute.
| Host | Output tok/s | Blended / 1M |
|---|---|---|
| Cerebras (wafer-scale) | 1,910reported | $0.450 |
| Together AI | 583 | $0.262 |
| Groq | 479 | $0.262 |
| Seldon | 420 | $0.210 |
| DeepInfra (Turbo) | 319 | $0.262 |
Throughput figures are Artificial Analysis medians over a trailing 72 hours at high reasoning effort and will drift. The Seldon row sits between the measured DeepInfra and Groq figures, and it is below Groq. Artificial Analysis measurement on gpt-oss-120b, via research/site-groq-cerebras.md. Cerebras' own rate card quotes roughly 3,000 tokens per second for the same model, which is 1.57x its independent measurement. The measured figure is the one used here. Hosts that publish a rate but have no independent throughput measurement in our research, Fireworks and DeepInfra's bfloat16 SKU, are omitted from this table rather than given an estimated speed.
What we are actually arguing
Peak single-stream throughput is a property of the serving hardware. It is the right metric when a human is watching tokens arrive and the cost of the request is irrelevant next to the cost of their attention. For interactive coding, live voice, and anything where a person is waiting, buy the fastest thing available and do not read the rest of this section.
What determines a production bill is different: tokens delivered per dollar at the concurrency your product actually runs at. Batch a hundred concurrent requests onto the same weights and per-stream speed falls for everyone while throughput per accelerator rises. The question stops being how fast one stream goes and becomes how much you paid for the aggregate.
So the honest split is this. If your constraint is latency, the fastest host wins and we are not it. If your constraint is the monthly invoice on a workload that is already asynchronous or already batched, price per token is the metric and speed is a floor you need to clear rather than a race you need to win.
Measured Cerebras output speed on gpt-oss-120b against Seldon Vector 1 throughput on the same architecture. They are faster on this model and the gap is not marginal.
Two things worth noticing in the table
First, our own throughput sits below Groq’s measured figure on identical weights, and we print it that way rather than publishing a number we would have to defend later. Second, Together’s measured median on this model is above Groq’s, at an identical price, which is a useful reminder that dedicated silicon does not automatically win every row.
We also do not publish a marketing throughput figure alongside a different one in our docs. Where an independent measurement and a vendor claim disagree anywhere in our dataset, including for competitors, the independent measurement is the one we print.
04Verification
Reproduce this yourself in an afternoon
Nothing on this page requires trusting us. The rate cards are public, the throughput measurements are third party, and the only comparison that matters in the end is the one you run on your own prompts.
- 1
Pull the rate cards
Every price in section 01 is a list price from a page anyone can load without an account. The exact sources we read, with the date we read them, are listed to the right. If a rate has moved since, ours is the stale one and you should tell us.
- 2
Measure throughput at your concurrency
Do not compare headline tokens per second. Run a fixed load against each endpoint at the concurrency your product actually uses. vLLM ships an open-source serving benchmark that speaks the OpenAI chat protocol, so the same command points at any host in the table.
- 3
Score quality on your own eval set
Public benchmarks are saturated: six of eight models in the one cross-vendor table our research could find score above 94 on AIME, so it no longer separates anything. Your prompts are not saturated. Take two hundred real requests, run them across the tiers, and grade the output the way your users would.
- 4
Check the modifiers, not just the headline
Long-context surcharges, batch discounts and tokenizer differences move the real number more than the headline rate does. A 50 percent batch discount at three of the four closed vendors narrows the gap to us materially on any workload that can tolerate asynchronous delivery. That works against our argument and it is still true.
# Same command, different --base-url. That is the whole method.vllm bench serve \ --backend openai-chat \ --base-url https://api.example-host.com/v1 \ --model gpt-oss-120b \ --dataset-name sharegpt \ --num-prompts 512 \ --max-concurrency 64 # Report the median output tokens per second and the p99 time to# first token. Then divide the published rate by the throughput you# measured, not by the throughput anyone advertised.Rate cards cited on this page
- ai.google.dev/gemini-api/docs/pricing and cloud.google.com/vertex-ai/generative-ai/pricing, fetched 2026-07-20
- api.deepinfra.com/models/list via research/site-deepinfra-together.md, fetched 2026-07-20
- developers.openai.com/api/docs/pricing, fetched 2026-07-20
- docs.fireworks.ai/serverless/pricing, fetched 2026-07-20
- docs.x.ai/developers/pricing, fetched 2026-07-20
- groq.com/pricing, fetched 2026-07-20
- platform.claude.com/docs/en/docs/about-claude/pricing, fetched 2026-07-20
- together.ai/pricing and docs.together.ai/docs/serverless-models, fetched 2026-07-20
Throughput figures throughout are Artificial Analysis medians over a trailing 72 hours, which drift week to week. Treat any single measurement as a band, including ours.
Run the comparison on your own traffic.
Send a representative sample of your prompts and your current provider's invoice shape. We will return a modelled comparison with the assumptions written out, and we will tell you when the answer is that you should stay where you are.