Skip to content
Seldon

Model catalog

Four tiers. One API. No retraining your application.

Every Seldon model is a tuned serving configuration of an open-weight architecture with a permissive license. We publish which architecture sits behind each SKU, because you should know what you are running and be able to leave if you want to.

01The catalog

Pick the cheapest tier that clears your quality bar

Most production traffic does not need a frontier model. The largest single cost saving available to most teams is not switching provider, it is routing the 70 percent of requests that are genuinely easy to a model that costs a twentieth as much.

Seldon Prime 1

Frontier
$2.10
blended / 1M
Input
$1.60 / 1M
Output
$3.60 / 1M
Cached input
$0.160 / 1M
Context
1.0M
Parameters
1600B total / 49B active
Throughput
180 tok/s

Seldon-tuned serving variant of the DeepSeek V4-Pro architecture (MIT, 1.6T total / 49B active per the model card, with the parameter-count caveat noted on that entry). Blended 3:1 rate of $2.10 sits 29% above DeepInfra's $1.625 floor, marginally under the $2.175 Together and Fireworks market rate, and 79% below Claude Opus 4.8's blended $10.00. Throughput is set conservatively against the 267 t/s Artificial Analysis measures for the lighter Kimi K2.7-Code on Together.

Seldon Core 1

Balanced
$1.80
blended / 1M
Input
$1.20 / 1M
Output
$3.60 / 1M
Cached input
$0.120 / 1M
Context
1.0M
Parameters
753.3B total
Throughput
380 tok/s

Seldon-tuned serving variant of the GLM-5.2 architecture (MIT, 753.3B total; active parameters are not published by Z.ai). Blended $1.80 against a DeepInfra floor of $1.448 and a market rate of $2.15, and 55% below Claude Sonnet 5 at its introductory $2 / $10. Throughput is set below the 476 t/s Artificial Analysis measures for GLM-5.2 on Together.

Seldon Vector 1

Fast
$0.210
blended / 1M
Input
$0.120 / 1M
Output
$0.480 / 1M
Cached input
$0.012 / 1M
Context
131K
Parameters
117B total / 5.1B active
Throughput
420 tok/s

Seldon-tuned serving variant of the gpt-oss-120b architecture (Apache 2.0, 117B total / 5.1B active). Priced 20% under the $0.15 / $0.60 rate that Groq, Together, Fireworks and DeepInfra's Turbo SKU all charge, while remaining well above DeepInfra's $0.037 / $0.17 bf16 SKU, which is the true commodity floor on this model. Throughput sits between the 319 t/s Artificial Analysis measures on DeepInfra and the 479 t/s it measures on Groq.

Seldon Horizon 1

Specialist
$1.29
blended / 1M
Input
$0.650 / 1M
Output
$3.20 / 1M
Cached input
$0.065 / 1M
Context
1.0M
Parameters
397B total / 17B active
Throughput
260 tok/s

Long-context specialist. Seldon-tuned serving variant of the Qwen3.5-397B-A17B architecture (Apache 2.0, 397B total / 17B active), served at the model card's full extended 1,010,000-token window rather than the 262,144 native window that DeepInfra and Together serve. Billed at a single flat rate across the whole window: no threshold, no retroactive repricing of the request. Blended $1.29 against a DeepInfra floor of $1.088 and Together at $1.35.

02Quality against price

The axis that actually matters

Price alone is a bad argument, and so is quality alone. What a buyer needs is the frontier of the tradeoff. Up and to the left is better: same capability, lower cost.

Composite quality index against blended price per million tokens

Blended price assumes a 3:1 input to output ratio. The quality index is deliberately coarse, banded in steps of 10, because the 2026 model generation does not share a common public benchmark.

  • Seldon
  • Closed frontier APIs
  • Open-weight, other hosts
Composite quality index against blended price per million tokens
ModelCategoryBlended $/1M tokensQuality index
Seldon Prime 1Seldon$2.1080
Seldon Core 1Seldon$1.8080
Seldon Vector 1Seldon$0.2150
Seldon Horizon 1Seldon$1.2980
Claude Fable 5Closed frontier API$20.0090
Claude Opus 4.8Closed frontier API$10.0090
Claude Sonnet 5Closed frontier API$4.0080
Claude Haiku 4.5Closed frontier API$2.0060
gpt-5.6-solClosed frontier API$11.2590
gpt-5.6-terraClosed frontier API$5.6380
gpt-5.6-lunaClosed frontier API$2.2560
Gemini 3.1 ProClosed frontier API$4.5090
Gemini 3.5 FlashClosed frontier API$3.3870
Grok 4.5Closed frontier API$3.0080
DeepSeek V4-ProOpen-weight, other host$1.6380
DeepSeek V4-FlashOpen-weight, other host$0.1160
GLM-5.2Open-weight, other host$1.4580
Kimi K2.7-CodeOpen-weight, other host$1.4370
Qwen3.5-397B-A17BOpen-weight, other host$1.0980
Mistral Large 3Open-weight, other host$0.7570

Read the quality axis as a band, not a score. Several 2026 models publish no comparable public benchmarks, and every score shown is for unquantized weights while every price shown is for a quantized serving SKU.

Pricing: vendor rate cards, July 2026. Quality index: see derivation note in lib/data/models.ts

03Against the closed frontier

Same job, itemised

The closed frontier models are genuinely excellent, and for the hardest reasoning work they remain the right tool. The question is what fraction of your traffic actually needs one. Compared like for like, the cheapest frontier-class closed SKU is Grok 4.5 at $3.00 blended, which is 1.4x the blended rate of Seldon Prime 1 at $2.10. The table below carries every tier, not just the frontier ones, and the cheap closed tiers price close to ours.

ModelInput / 1MOutput / 1MBlended 3:1vs Opus 4.8
Seldon Prime 1Seldon
$1.60$3.60$2.10 4.8x cheaper
Seldon Core 1Seldon
$1.20$3.60$1.80 5.6x cheaper
Seldon Vector 1Seldon
$0.120$0.480$0.210 47.6x cheaper
Seldon Horizon 1Seldon
$0.650$3.20$1.29 7.8x cheaper
Claude Fable 5Anthropic
$10.00$50.00$20.00 2.0x dearer
Claude Opus 4.8Anthropic
$5.00$25.00$10.00baseline
Claude Sonnet 5Anthropic
$2.00$10.00$4.00 2.5x cheaper
Claude Haiku 4.5Anthropic
$1.00$5.00$2.00 5.0x cheaper
gpt-5.6-solOpenAI
$5.00$30.00$11.25 1.1x dearer
gpt-5.6-terraOpenAI
$2.50$15.00$5.63 1.8x cheaper
gpt-5.6-lunaOpenAI
$1.00$6.00$2.25 4.4x cheaper
Gemini 3.1 ProGoogle
$2.00$12.00$4.50 2.2x cheaper
Gemini 3.5 FlashGoogle
$1.50$9.00$3.38 3.0x cheaper
Grok 4.5xAI
$2.00$6.00$3.00 3.3x cheaper

Blended rate assumes a 3:1 input to output token ratio. Closed-model rates exclude per-request modifiers such as long-context surcharges, batch discounts, and regional premiums, which vary by vendor and are documented on the pricing page. Anthropic's newer models use a tokenizer that produces roughly 30 percent more tokens for the same text, so headline rates are not directly comparable across vendors without adjusting for it.

04Host against host

The same open model, priced five ways

Open weights do not mean uniform pricing. The identical model, with identical weights, trades across a wide band depending on who is serving it and on what hardware. This is the clearest available evidence that inference pricing reflects serving efficiency and capacity cost, not model cost.

gpt-oss-120b, blended price per million tokens by host

The only model every major host publishes a rate for, which makes it the one clean cross-host comparison available.

gpt-oss-120b, blended price per million tokens by host
ItemDetailValue ($)
DeepInfra (bf16)$0.070
Seldon420 tok/s$0.210
DeepInfra (Turbo)319 tok/s$0.262
Together AI583 tok/s$0.262
Fireworks AI$0.262
Groq479 tok/s$0.262

Vendor rate cards, July 2026. Throughput from Artificial Analysis measurements rather than vendor marketing claims.

3.7x
Spread, cheapest to dearest

Between the lowest and highest published rate for the identical open-weight model across major hosts.

A spread this wide on a fixed input is the definition of an inefficient market. It exists because serving cost is dominated by GPU sourcing and kernel efficiency, and those vary far more between operators than the model does.

How we source capacity

05Open-weight catalog

Everything else we serve

Beyond the tuned Seldon tiers, the full open-weight catalog is available at cost-plus rates. Licenses are the model authors' own and are listed so you can check them against your own legal constraints.

ModelContextInput / 1MOutput / 1M
DeepSeek V4-Pro1.0M$1.30$2.60
DeepSeek V4-Flash1.0M$0.090$0.180
GLM-5.21.0M$0.930$3.00
Kimi K2.7-Code262K$0.740$3.50
Qwen3.5-397B-A17B262K$0.450$3.00
Mistral Large 3262K$0.500$1.50
Llama 4 Maverick1.0M$0.200$0.800
gpt-oss-120b131K$0.037$0.170
Qwen3-235B-A22B-Instruct-2507262K$0.090$0.550
DeepSeek-V3.2164K$0.260$0.380

Rates shown are the cheapest published third-party host rate for each model as of July 2026, which establishes the market floor we price against. Parameter counts marked as reported come from model cards where the published figure conflicts with the computed weight index.

Not sure which tier your traffic needs?

Send a representative sample of your prompts. We run them across the catalog, report quality against your own evaluation set rather than a public benchmark, and tell you honestly if a cheaper tier does not clear your bar.