Skip to content
Seldon

Company

Inference is a cost problem, and cost problems are solved with engineering and procurement.

Seldon exists because the price of a token is the binding constraint on what can be built with language models, and because that price is set by serving efficiency and capacity sourcing rather than by anything intrinsic to the models themselves.

17.3xModelled cost reductionDerived from the published serving stack, discounted for overlap between techniques. The arithmetic is on the technology page.

01The thesis

The constraint moved, and most of the industry has not noticed

For three years the interesting question was capability. It is now largely a question of unit economics, and the companies that will still be here in five years are the ones whose gross margin survives their own success.

An application layer built on inference it cannot afford is not a business, it is a subsidy with a burn chart.

The pattern repeats across the application layer. A product works, usage grows, and inference cost grows with it linearly while revenue does not. Gross margin compresses toward the point where every additional user makes the company worse off. The usual responses are to cap usage, to move to a weaker model and hope nobody notices, or to raise a round against the assumption that token prices will fall on their own.

None of those are engineering. The observation Seldon is built on is that inference cost is not a mystery: it decomposes into four terms a provider actually controls, and every one of them responds to work. Raise decode throughput, raise fleet utilization, avoid prefill you have already paid for, and lower the hourly cost of the capacity underneath. There is no fifth lever, and any provider claiming a cost advantage is pulling one of these four whether they describe it that way or not.

See the cost identity

Two of those levers are engineering and two are procurement, and the split matters more than it sounds. The engineering half is published work: sparse architectures, speculative decoding, low-precision serving, KV cache compression, disaggregated prefill and decode. None of it is secret. What is hard is assembling all of it into one stack and keeping it working together under real traffic, which is why most providers ship three or four of these rather than all of them.

The procurement half is less glamorous and at least as valuable. The same open-weight model, with identical weights, trades across a wide band depending on who is serving it and on what hardware they bought. That spread is not a quality difference. It is the market telling you that capacity sourcing and kernel efficiency vary far more between operators than models do. Treating GPU capacity as a commodity to be sourced continuously across markets, rather than a contract signed once a year, is a cost lever that has nothing to do with machine learning at all.

So: not a model problem. We do not think the frontier labs are wrong, and for the hardest reasoning work their models remain the right tool. We think the question almost nobody asks carefully is what fraction of production traffic actually needs one, and that the honest answer is a lot less than current spending implies.

What this does to gross margin

02How we work

Six rules, each visible somewhere on this site

A values page is worth nothing if it cannot be checked against the artifact it appears on. Every principle below is followed by a link to the place where it is already being applied, or contradicted.

01

Publish the derivation, not just the number

Every headline figure on this site is accompanied by the arithmetic that produced it and the assumptions it rests on. The compounded cost reduction is not asserted, it is built stage by stage from per-technique factors, each with the published range it came from, so a reader can dispute one factor rather than having to accept or reject the whole claim.

02

Mark the provenance of every figure

Every number on this site records where it came from in the type system itself, not in a footnote. Verified against a primary source, reported by a vendor, estimated, and derived are distinct categories and are stored as distinct values. Third-party figures are copied from a dated research file rather than recalled, and where a vendor's own page disagrees with an independent measurement, we publish the independent one and say which is which.

03

Discount our own claims before anyone else has to

The techniques in our serving stack overlap: sparsity and quantization both attack memory bandwidth, and two layers of caching compete for the same repeated tokens. Multiplying their individual published gains would produce a much larger number than the one we publish. We apply an overlap discount per technique and show what it removed, because the flattering number is the one that falls apart in a technical review.

04

State what we do not have

Seldon holds no third-party certifications. Not SOC 2, not ISO 27001, not FedRAMP. SOC 2 Type II is in progress and no report has been issued. The trust center publishes that status directly and lists what is running in production next to what is not, because an empty certification column is survivable and a claimed certification that evaporates in diligence is not.

05

Describe relationships as they actually are

There is no partner logo wall on this site and no partner page. What exists is a list of hardware we run on, wire formats we are compatible with, and environments we deploy into, each labelled with the verb that describes the real relationship. Naming a company as a partner without an agreement is a false-association claim, and we would rather the accurate version.

06

Make leaving possible

We publish the open-weight architecture behind every Seldon SKU and stay wire-compatible with the APIs you already use, which means switching away is a base URL change rather than a rewrite. A customer who cannot leave is not a customer who has chosen to stay, and pricing that only holds because migration is painful is not pricing we want to defend.

Runs on

  • NVIDIA H100
  • NVIDIA H200
  • NVIDIA B200
  • NVIDIA GB200 NVL72
  • AMD MI355X
  • InfiniBand NDR

Accelerators the serving stack is built and tuned for. Kernel work targets each generation directly rather than relying on a portability layer.

Drop-in for

  • OpenAI API format
  • Anthropic Messages format
  • LangChain
  • LlamaIndex
  • Vercel AI SDK
  • OpenTelemetry

Wire-compatible surfaces. Existing application code moves over by changing a base URL and a key, with no SDK rewrite.

Deploys into

  • AWS
  • Google Cloud
  • Azure
  • Customer VPC
  • On-premises
  • Air-gapped

Target environments for dedicated and private deployments, including customer-controlled infrastructure.

Serves

  • DeepSeek
  • Qwen
  • Llama
  • Kimi
  • GLM
  • Mistral
  • gpt-oss

Open-weight model families in the catalog. Licenses are the model authors' own; we host and optimize them.

Built with

  • vLLM
  • SGLang
  • TensorRT-LLM
  • PyTorch
  • Triton
  • Ray

Open source infrastructure we run in production and contribute back to.

03Customers

Teams whose margin depends on this working

The companies running production traffic on Seldon are mostly application businesses where inference is a large enough share of cost of revenue to show up in the gross margin line. That is a demanding kind of customer to have: they read the rate card closely, they measure what we serve, and they notice a regression before we have finished writing the incident note.

It is also the only kind of customer that makes the argument on this site checkable. A buyer evaluating a cost claim wants to know what happened to someone with a comparable workload, so each quote below carries the figure it turns on rather than an adjective, together with the specific pairing that figure was computed from. All four assume a move from gpt-5.6-luna to Seldon Vector 1, applied to that customer’s published workload shape. The margin calculator holds its Seldon leg at Seldon Prime 1, so it answers a different question than these quotes do.

Run the margin model

What this page still does not publish

  • Customer namesCustomers are described by segment, stage, and workload rather than named. A logo goes up when the company it belongs to has given written permission, and not before. We will introduce you to a reference customer in your own segment on request.
  • Leadership and teamNo named individuals until each has agreed in writing to appear here.
  • Investors and fundingNo round, amount, or investor named until a round has closed and the investor has approved the reference.
  • HeadcountNot published. A number that moves weekly and cannot be verified is decoration.
  • Offices and locationsNot published until there is an address rather than a plan.

If you are evaluating us and need to know who you would be working with, ask and you will get a direct answer rather than a page.

Ask us directly

What changed

The same argument, from the other side of it

Every figure quoted below is computed, not asserted: each is the same cost function the calculators run, applied to that customer's published workload shape, for the specific pairing named beside it. All four assume a move from gpt-5.6-luna to Seldon Vector 1. The margin calculator fixes its Seldon leg to Seldon Prime 1, so it will not return these figures directly, and your own volumes on your own tier are the number that actually matters.

Inference was 40 percent of revenue and it was the only line that scaled with every seat we sold. Moving the completion path over took an afternoon because the API is wire-compatible. The margin change was the difference between raising a bridge and not needing one.

VP EngineeringDeveloper tools company, Series BMoved from a closed frontier tier, 2026-01
+35.9 ptsgross margingpt-5.6-luna to Seldon Vector 1

We were paying a long-context surcharge on every contract we ingested. Seldon does not charge one, and the prefix cache holds across a document set, so a hundred-clause review costs what the first clause used to.

Head of PlatformLegal company, Series BMoved from a closed balanced tier, 2026-03
10.4xlower inference COGSgpt-5.6-luna to Seldon Vector 1

The routing layer is the part that surprised us. Roughly seventy percent of our classification traffic never needed a frontier model, and nobody had been separating those requests out. We did not change our quality bar to get the saving.

Staff EngineerData infrastructure company, Series AMoved from a closed fast tier, 2025-11
9.1xlower cost per itemgpt-5.6-luna to Seldon Vector 1

They told us upfront that inference was only six percent of our revenue and that switching would not fix our margins, which is not what a vendor normally says in a sales call. We moved anyway for the latency, and the honesty is why we expanded.

CTOCustomer support company, Series AMoved from a closed fast tier, 2025-09
+5.7 ptsmargin, as predictedgpt-5.6-luna to Seldon Vector 1

04Careers

Four problems, described as problems rather than as postings

We do not run a jobs board, because a requisition list tells you less about the work than the work does. What follows is a description of what the four areas actually involve, and an address that a person reads.

The work divides into four areas. Treat these as a description of the problems rather than a list of vacancies: what is open in a given week changes faster than a page does. If one of them is what you already do, the conversation is worth having regardless of whether a formal role is posted that week.

Inference systems

Kernel work against current accelerator generations, quantized serving paths, speculative decoding, and KV cache management. The measurable output is tokens per second per GPU at a fixed quality bar, and the work is closer to performance engineering than to machine learning.

Distributed serving and scheduling

Disaggregated prefill and decode, request routing across heterogeneous hardware, prefix cache locality, and keeping utilization high without breaking tail latency. Most of the interesting failures live here rather than in the model.

Capacity and supply

Sourcing accelerator capacity continuously across markets, modelling its cost, and building the tooling that decides what to run where. This is an engineering job with a procurement problem attached, and it is undervalued almost everywhere.

Platform, security, and compliance

The API surface, tenancy and isolation, audit and observability, and the control and evidence work an attestation depends on. Building the thing the trust center describes, rather than the page describing it.

How to apply

There is no applicant tracking system and no careers portal. Write to the address below with what you have built and which of the four areas you want to work in. A short description of a hard problem you actually shipped is worth more here than a formatted resume.

sales@seldon.ai

Routed by hand. We do not publish a response time we cannot currently keep.

What we will not tell you

No compensation bands, equity ranges, benefits, or team size are published here. A range on a page is a worse answer than a number in a conversation, and it is the number you would end up negotiating against anyway. You will get the real figures from us directly.

The fastest way to evaluate us is to argue with the numbers.

Everything quantitative on this site has its derivation published alongside it. Pick the figure you find least believable and write to us about that one.