Model catalog
Four tiers. One API. No retraining your application.
Every Seldon model is a tuned serving configuration of an open-weight architecture with a permissive license. We publish which architecture sits behind each SKU, because you should know what you are running and be able to leave if you want to.
01The catalog
Pick the cheapest tier that clears your quality bar
Most production traffic does not need a frontier model. The largest single cost saving available to most teams is not switching provider, it is routing the 70 percent of requests that are genuinely easy to a model that costs a twentieth as much.
Seldon Prime 1
Frontier- Input
- $1.60 / 1M
- Output
- $3.60 / 1M
- Cached input
- $0.160 / 1M
- Context
- 1.0M
- Parameters
- 1600B total / 49B active
- Throughput
- 180 tok/s
Seldon-tuned serving variant of the DeepSeek V4-Pro architecture (MIT, 1.6T total / 49B active per the model card, with the parameter-count caveat noted on that entry). Blended 3:1 rate of $2.10 sits 29% above DeepInfra's $1.625 floor, marginally under the $2.175 Together and Fireworks market rate, and 79% below Claude Opus 4.8's blended $10.00. Throughput is set conservatively against the 267 t/s Artificial Analysis measures for the lighter Kimi K2.7-Code on Together.
Seldon Core 1
Balanced- Input
- $1.20 / 1M
- Output
- $3.60 / 1M
- Cached input
- $0.120 / 1M
- Context
- 1.0M
- Parameters
- 753.3B total
- Throughput
- 380 tok/s
Seldon-tuned serving variant of the GLM-5.2 architecture (MIT, 753.3B total; active parameters are not published by Z.ai). Blended $1.80 against a DeepInfra floor of $1.448 and a market rate of $2.15, and 55% below Claude Sonnet 5 at its introductory $2 / $10. Throughput is set below the 476 t/s Artificial Analysis measures for GLM-5.2 on Together.
Seldon Vector 1
Fast- Input
- $0.120 / 1M
- Output
- $0.480 / 1M
- Cached input
- $0.012 / 1M
- Context
- 131K
- Parameters
- 117B total / 5.1B active
- Throughput
- 420 tok/s
Seldon-tuned serving variant of the gpt-oss-120b architecture (Apache 2.0, 117B total / 5.1B active). Priced 20% under the $0.15 / $0.60 rate that Groq, Together, Fireworks and DeepInfra's Turbo SKU all charge, while remaining well above DeepInfra's $0.037 / $0.17 bf16 SKU, which is the true commodity floor on this model. Throughput sits between the 319 t/s Artificial Analysis measures on DeepInfra and the 479 t/s it measures on Groq.
Seldon Horizon 1
Specialist- Input
- $0.650 / 1M
- Output
- $3.20 / 1M
- Cached input
- $0.065 / 1M
- Context
- 1.0M
- Parameters
- 397B total / 17B active
- Throughput
- 260 tok/s
Long-context specialist. Seldon-tuned serving variant of the Qwen3.5-397B-A17B architecture (Apache 2.0, 397B total / 17B active), served at the model card's full extended 1,010,000-token window rather than the 262,144 native window that DeepInfra and Together serve. Billed at a single flat rate across the whole window: no threshold, no retroactive repricing of the request. Blended $1.29 against a DeepInfra floor of $1.088 and Together at $1.35.
02Quality against price
The axis that actually matters
Price alone is a bad argument, and so is quality alone. What a buyer needs is the frontier of the tradeoff. Up and to the left is better: same capability, lower cost.
Composite quality index against blended price per million tokens
Blended price assumes a 3:1 input to output ratio. The quality index is deliberately coarse, banded in steps of 10, because the 2026 model generation does not share a common public benchmark.
- Seldon
- Closed frontier APIs
- Open-weight, other hosts
| Model | Category | Blended $/1M tokens | Quality index |
|---|---|---|---|
| Seldon Prime 1 | Seldon | $2.10 | 80 |
| Seldon Core 1 | Seldon | $1.80 | 80 |
| Seldon Vector 1 | Seldon | $0.21 | 50 |
| Seldon Horizon 1 | Seldon | $1.29 | 80 |
| Claude Fable 5 | Closed frontier API | $20.00 | 90 |
| Claude Opus 4.8 | Closed frontier API | $10.00 | 90 |
| Claude Sonnet 5 | Closed frontier API | $4.00 | 80 |
| Claude Haiku 4.5 | Closed frontier API | $2.00 | 60 |
| gpt-5.6-sol | Closed frontier API | $11.25 | 90 |
| gpt-5.6-terra | Closed frontier API | $5.63 | 80 |
| gpt-5.6-luna | Closed frontier API | $2.25 | 60 |
| Gemini 3.1 Pro | Closed frontier API | $4.50 | 90 |
| Gemini 3.5 Flash | Closed frontier API | $3.38 | 70 |
| Grok 4.5 | Closed frontier API | $3.00 | 80 |
| DeepSeek V4-Pro | Open-weight, other host | $1.63 | 80 |
| DeepSeek V4-Flash | Open-weight, other host | $0.11 | 60 |
| GLM-5.2 | Open-weight, other host | $1.45 | 80 |
| Kimi K2.7-Code | Open-weight, other host | $1.43 | 70 |
| Qwen3.5-397B-A17B | Open-weight, other host | $1.09 | 80 |
| Mistral Large 3 | Open-weight, other host | $0.75 | 70 |
Read the quality axis as a band, not a score. Several 2026 models publish no comparable public benchmarks, and every score shown is for unquantized weights while every price shown is for a quantized serving SKU.
Pricing: vendor rate cards, July 2026. Quality index: see derivation note in lib/data/models.ts
03Against the closed frontier
Same job, itemised
The closed frontier models are genuinely excellent, and for the hardest reasoning work they remain the right tool. The question is what fraction of your traffic actually needs one. Compared like for like, the cheapest frontier-class closed SKU is Grok 4.5 at $3.00 blended, which is 1.4x the blended rate of Seldon Prime 1 at $2.10. The table below carries every tier, not just the frontier ones, and the cheap closed tiers price close to ours.
| Model | Input / 1M | Output / 1M | Blended 3:1 | vs Opus 4.8 |
|---|---|---|---|---|
Seldon Prime 1Seldon | $1.60 | $3.60 | $2.10 | 4.8x cheaper |
Seldon Core 1Seldon | $1.20 | $3.60 | $1.80 | 5.6x cheaper |
Seldon Vector 1Seldon | $0.120 | $0.480 | $0.210 | 47.6x cheaper |
Seldon Horizon 1Seldon | $0.650 | $3.20 | $1.29 | 7.8x cheaper |
Claude Fable 5Anthropic | $10.00 | $50.00 | $20.00 | 2.0x dearer |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | $10.00 | baseline |
Claude Sonnet 5Anthropic | $2.00 | $10.00 | $4.00 | 2.5x cheaper |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | $2.00 | 5.0x cheaper |
gpt-5.6-solOpenAI | $5.00 | $30.00 | $11.25 | 1.1x dearer |
gpt-5.6-terraOpenAI | $2.50 | $15.00 | $5.63 | 1.8x cheaper |
gpt-5.6-lunaOpenAI | $1.00 | $6.00 | $2.25 | 4.4x cheaper |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | $4.50 | 2.2x cheaper |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | $3.38 | 3.0x cheaper |
Grok 4.5xAI | $2.00 | $6.00 | $3.00 | 3.3x cheaper |
Blended rate assumes a 3:1 input to output token ratio. Closed-model rates exclude per-request modifiers such as long-context surcharges, batch discounts, and regional premiums, which vary by vendor and are documented on the pricing page. Anthropic's newer models use a tokenizer that produces roughly 30 percent more tokens for the same text, so headline rates are not directly comparable across vendors without adjusting for it.
04Host against host
The same open model, priced five ways
Open weights do not mean uniform pricing. The identical model, with identical weights, trades across a wide band depending on who is serving it and on what hardware. This is the clearest available evidence that inference pricing reflects serving efficiency and capacity cost, not model cost.
gpt-oss-120b, blended price per million tokens by host
The only model every major host publishes a rate for, which makes it the one clean cross-host comparison available.
| Item | Detail | Value ($) |
|---|---|---|
| DeepInfra (bf16) | $0.070 | |
| Seldon | 420 tok/s | $0.210 |
| DeepInfra (Turbo) | 319 tok/s | $0.262 |
| Together AI | 583 tok/s | $0.262 |
| Fireworks AI | $0.262 | |
| Groq | 479 tok/s | $0.262 |
Vendor rate cards, July 2026. Throughput from Artificial Analysis measurements rather than vendor marketing claims.
Between the lowest and highest published rate for the identical open-weight model across major hosts.
A spread this wide on a fixed input is the definition of an inefficient market. It exists because serving cost is dominated by GPU sourcing and kernel efficiency, and those vary far more between operators than the model does.
How we source capacity05Open-weight catalog
Everything else we serve
Beyond the tuned Seldon tiers, the full open-weight catalog is available at cost-plus rates. Licenses are the model authors' own and are listed so you can check them against your own legal constraints.
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
| DeepSeek V4-Pro | 1.0M | $1.30 | $2.60 |
| DeepSeek V4-Flash | 1.0M | $0.090 | $0.180 |
| GLM-5.2 | 1.0M | $0.930 | $3.00 |
| Kimi K2.7-Code | 262K | $0.740 | $3.50 |
| Qwen3.5-397B-A17B | 262K | $0.450 | $3.00 |
| Mistral Large 3 | 262K | $0.500 | $1.50 |
| Llama 4 Maverick | 1.0M | $0.200 | $0.800 |
| gpt-oss-120b | 131K | $0.037 | $0.170 |
| Qwen3-235B-A22B-Instruct-2507 | 262K | $0.090 | $0.550 |
| DeepSeek-V3.2 | 164K | $0.260 | $0.380 |
Rates shown are the cheapest published third-party host rate for each model as of July 2026, which establishes the market floor we price against. Parameter counts marked as reported come from model cards where the published figure conflicts with the computed weight index.
Not sure which tier your traffic needs?
Send a representative sample of your prompts. We run them across the catalog, report quality against your own evaluation set rather than a public benchmark, and tell you honestly if a cheaper tier does not clear your bar.