One endpoint,
every model.
Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.
Billing rules
What Tare charges depends on whose money bought the tokens, and nothing else.
| Route used | Token charge | Routing fee |
|---|---|---|
| A platform model | Your price list | — |
| One of your own channels | None — your provider billed you directly | Only if your contract has one |
[!NOTE] Because of this, a balance of zero does not block calls on your own channels. Those tokens are not owed to Tare, and blocking them would contradict the design premise that changing one base URL is enough to integrate.
Routing fee, if it applies
Charged per million tokens (input + output) from a published rate card, not as a percentage of what your provider charges you. The Routing page shows the rate card when the account has one, and states explicitly when it does not.
Balance and credit limits
- The account balance applies to platform models only.
- A credit limit on a key caps what that key can spend, independently of the balance. A runaway job stops there instead of draining your account.
- When you run out, calls are refused with a distinct error (see Errors), not a generic failure.
Aborted calls are still billed
A stream that breaks mid-flight has still consumed tokens upstream. Tare reconciles those after the fact rather than discarding them; otherwise the statement would fall below the provider's own invoice.
Failover covers every configured route
A model can have any number of routes, and one call tries all of them, in order, until one succeeds. There is no platform-imposed ceiling: four configured fallbacks means four attempts. Whether an additional fallback justifies the additional latency is left to the account owner.
Two outcomes do not count as an attempt, and both are listed separately in the call trace:
- parked — that channel was circuit-broken at the time, so nothing was sent to it;
- never tried — a candidate that was not reached (only possible when an attempt cap is configured for the account; there is no cap by default).
⚠️ Each real attempt counts toward the routing-fee minimum. A call that fails over once made two upstream requests — two connections opened, two timeouts waited on — so the minimum is applied twice. Parked candidates never count: nothing was sent, and the cause was on the platform side.
Failover and billing lines
One call never produces two billing lines. Failover happens before the first byte reaches you, while the attempt has consumed no tokens on your behalf; once content starts flowing there is no failover. So:
- The prompt tokens of the failed attempt are not billed to you;
- One business call produces one charged line.
⚠️ The risk sits on the platform side: a non-streaming request that times out and fails over may have completed upstream anyway, in which case the platform is charged twice. That cost is absorbed, not passed on.
One other kind of line is not a charge: when a stream breaks mid-flight the usage never arrives
(it lives in the final frame), so the upstream is queried afterwards and the real cost recorded with
cost = 0 and only an upstream cost — absorbed by the platform, not charged to you. It carries the same gen- id,
so you can correlate by id.