Skip to content
Seldon

Contact

Reach an engineer, not a form queue.

Seldon is a small company running a live service. Most enquiries are answered by someone who works on the thing you are asking about. Pick the address that matches what you need and write to it directly.

01Routes in

Four addresses, each with a job

These are separate inboxes rather than aliases onto one. Every one of them gets a first reply from a person within one business day. Sending to the right one is the difference between an answer and a forwarded thread.

Sales and evaluations

sales@seldon.ai

Pricing, committed-use and dedicated capacity, contracting, and workload evaluations. This is the right address if you are trying to work out whether moving traffic to us is worth the migration.

Monthly token volumes, the input to output ratio, and what you are paying today. That is enough to model the comparison before a call.

Technical support

support@seldon.ai

API behaviour, integration problems, error responses, and anything that looks like a fault on our side. Engineering questions about the serving stack are welcome here rather than routed through sales.

The model id, a request id if you have one, the timestamp with a timezone, and the exact error body. Redact prompt content if you need to.

Security and privacy

security@seldon.ai

Vulnerability disclosure, security questionnaires, data handling questions, and vendor due diligence. Reports are read by an engineer, not by a ticket queue.

For a disclosure: reproduction steps and the affected surface. Please do not test against another customer's data or degrade the service to prove a point.

Media enquiries and fact-checking. If you are checking a figure published on this site, quote the page and we will point you at the derivation behind it.

Your outlet, your deadline, and the specific claim you are checking. We will say when we do not know something.

02Write to us

A form that tells you the truth about itself

The form validates what you type and then hands the composed message to your own mail client. You send it yourself, from your own address, and you can see exactly what leaves.

Send a message

Required fields are marked. Submitting opens your mail client with the message composed, so it arrives from your own address and you keep a copy in your sent items.

Optional.

Token volumes, input to output ratio, and latency targets are the three things that let us answer usefully on the first reply.

No backend attached. See what happens next below.

Why it works this way

A form that posts to a server you cannot see asks you to trust a confirmation screen. This one hands the message to your mail client instead, so delivery is something you can verify in your own sent folder rather than something we assert.

Nothing you type leaves the page until you press send in your mail client. The fields hold their state in the browser and post to nothing.

Or write directly

sales@seldon.ai

The same inbox the form composes to. Skip the form entirely if you would rather write the message yourself.

03Workload evaluation

Five numbers turn a conversation into arithmetic

If you want a cost comparison rather than a pitch, send these. Each one feeds a specific term in the model, and with all five we can return a modelled bill with the assumptions written out. With none of them, any figure we quote is a guess dressed up as a quote.

Send usWhat it changes
Monthly token volumeSets which pricing regime applies. Below a few hundred million tokens a month the shared endpoint is almost always cheaper than reserved capacity, and above a certain point the reverse becomes true.
Input to output ratioOutput tokens cost several times more than input tokens on every vendor's rate card, so two workloads with identical total volume can differ by a factor of three in cost. Headline blended rates assume 3:1 and most real workloads are not 3:1.
Cache hit rate, or the prompt structure that determines itCached input is billed at ten percent of input across the Seldon catalog. On a workload with a large stable system prompt this is frequently the single largest term in the bill, and it is the one most often left on the table.
Latency targetsTime to first token and inter-token latency are set by different parts of the stack, and pricing follows from which one you are constrained by. A batch summarisation job and an interactive assistant with the same volume are different products.
Compliance and data residency requirementsRegion pinning, zero retention, and private deployment each change the achievable cost, and some combinations rule out the cheapest capacity entirely. Better to find that out before modelling a price you cannot be sold.

Approximations are fine. An order of magnitude on volume and a rough ratio are enough to produce a comparison worth arguing with, and precision beyond that rarely changes which tier is the right answer.

What you get back

A modelled monthly cost on the Seldon tier that clears your quality bar, set against your current spend, with every assumption listed so you can dispute an individual line rather than the whole figure. Where a cheaper tier does not clear your bar on your own evaluation set, we will say so. We would rather lose the traffic than have it move back three months later.

Model it yourself first

What a modelled comparison is not

The rates are what we bill and the throughput is what we serve at, but neither was measured under your traffic. A comparison we return is arithmetic on the shape you send us, so it inherits every approximation in that shape. Treat a modelled saving as a hypothesis to test on a slice of real traffic rather than a number to put in a board deck, and run the slice before you move anything that matters.

What is and is not in place