Every price below is a vendor list price we read off the vendor's own page or API, on the date in the row, at the URL in the row. It is a cost input, not a quote: it is what the model costs before anything we do to it. We are pre-launch and have no contracted wholesale rate to publish, so there is no customer price on this page and inventing one would be the exact thing this product exists to stop.
Nothing here is modelled, interpolated or remembered. Where we could not read a price, the endpoint is named below as unpriced rather than filled in — because a gap you cannot see is worse than a gap.
Most of these prices are not attached to anything you can call. Measured against a running router on 2026-09-04: it serves 20 endpoints, and ten of them have a price in this card — the other ten are the unpriced list below. The remaining cards on this page are published vendor prices for models the router does not currently route to. They are cost references with a source and a date, which is all they claim to be.
DeepSeek V4 Pro and V4 Flash, Kimi K3 and GLM-5.3 are carded here from each vendor's own pricing page, read on 2026-09-02. They are in the catalogue for the same reason everything else is: one API, one rate card, and the same signed receipt per request. A frontier-class open-weight model is a routing option, not a separate product with separate rules.
DeepSeek prices by time of day. Its published rate is roughly half off-peak. The card carries the peak number deliberately — a card quoting the cheap hour would understate the cost of a request served at the expensive one, and every margin floor in this system is computed against the card.
GLM-5.3-Flash is missing on purpose. Its list price is under a 50% promotion that ends 2026-09-09. A card whose number reverts inside the month would be wrong before its own re-read date came round, so it is not carded until the promotion resolves.
These four are retaining, and that has teeth. No
zero-retention agreement exists for them to read, so they carry the weakest posture
the schema has. A project whose floor is stricter does not quietly get one of these
instead — the router refuses the request with retention_floor_unmet and
the prompt never leaves our network. Measured, not asserted.
Ollama, llama.cpp, LM Studio and vLLM are in this catalogue as ordinary
endpoints — self_hosted · on-prem, jurisdiction
LOCAL — with the same wire, the same eligibility filter and the same
signed receipt as everything else. Filter this table to Self-hosted
to see them.
Their cards read zero, and zero is a price we know. Local
inference produces no vendor invoice, so input and output are carded at
0 and the receipt's provider cost is a computed zero. That is
deliberately not the same as the unpriced endpoints below,
which mean we could not read a price at all — a gap you can see, rather than a gap
filled in.
They are retaining like everything else, and that is the
honest answer rather than the flattering one. A model on your own laptop
plainly has the better retention story; we may not say so, because any posture
above retaining is gated on a document somebody has read, and what a
local runtime writes to its own disk is a property of software we did not audit.
What is checked is where the bytes go: the base URL must be a numeric
loopback literal, and even localhost is refused because a name is
resolved by the host. Measured 2026-09-04, starting the shipped router with
OLLAMA_BASE_URL=http://localhost:11434/v1: it exited 78
without opening a socket.
Configuration and the rest of the argument: docs → local & self-hosted runtimes.
| Model | Substrate · region | In / 1M | Out / 1M | Cache read | Cache write | Reasoning | Observed | Source |
|---|
Generated from router/data/rate-card.json, the router’s own price document — the same file the router prices against. A dash is not zero: it means that token class carries no published price on that endpoint, and the loader will not let a token be charged at a number nobody published.
Endpoints in the catalogue with no published price
The rate card refuses to load unless every catalogue endpoint is either priced or named here, so this list is not a courtesy — it is the mechanism that keeps a missing price from becoming an invented one. Each entry carries the source we checked and the date we checked it.
| Endpoint | Why there is no price | Checked |
|---|
The reasons are reproduced verbatim from the rate card, in English, because they are the record rather than prose about it.
Where each number came from
| Ref | Document | What it yielded | Fetched |
|---|
What this card does not yet carry
Counted from the card itself rather than asserted here, so this section cannot quietly go stale while the card changes underneath it.
Each of these is something the rate-card schema defines and this bundle does not yet stand behind — a field left empty, or a field filled in with no document behind it. They are listed because a schema that could carry evidence is not the same as a card that does.
The retention block the rate-card schema defines — posture,
evidence kind, evidence URL, verification state, verification date — is now on
every card above and on every unpriced endpoint below. Twenty-six blocks, and the
router refuses to start if any one of them is missing, exactly as it refuses to
start when a price is missing. The posture the routing filter enforces is read
out of that document rather than out of a table compiled into the binary.
All thirty-five read unverified. That is a state
the schema defines, not a field left blank: evidence_url and
verified_at are null, and unverified_reason records
where the posture value came from instead. We hold no DPA clause reference, no
zero-retention addendum and no verified organisation setting for any endpoint,
and there is no vendor contract for one to live in. A plausible-looking link in
those fields would be worse than the compiled table it replaced, so there is
still no posture column and no jurisdiction column on this page. They go in per
endpoint as the evidence lands, and not one row before it.
Read the taxonomy on Trust → retention postures. It is what we will publish per endpoint once we hold the documents, not what we publish today.
Under GDPR Article 28 you are entitled to know every party that processes your data and to be notified when that list changes. For a model gateway, the subprocessor list is the set of endpoints traffic can reach. Publishing them as two separate documents lets them drift; publishing one makes the routing filter and the legal commitment the same fact.
A policy envelope that restricts posture or residency restricts your subprocessor set with it, and the receipt for each request names the endpoint that actually served it. Both are specified and neither has served traffic — see status for what is running.