One endpoint,
every model.
Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.
Limits
Four tiers, counted independently
| Counted against | Default | Scope |
|---|---|---|
| Model calls (per key) | 100/minute | One end user's runaway loop |
| Model calls (account total) | 6000/minute | Ten thousand keys making a few calls each — undetectable per key, saturating in aggregate |
| Provisioning | 1200/minute | Issuing, capping and revoking child keys; as frequent as your billing events |
| Reconciliation / integration | 1200/minute | Read-only reporting. Month-end pulls hundreds of pages by day, key and model |
⚠️ Model calls are counted per key, not per account: give every end user their own child key and each gets their own 100/minute, with no crowding out.
A streaming call counts once, regardless of duration or frame count — the counter increments at authentication time.
The response when a limit is exceeded
429 with Retry-After (seconds). ⚠️ A spent cap is 403, not 429: 429 means "slow
down and retry", which clients back off on, while an exhausted cap never recovers by waiting. Full
mapping in Errors.
Raising a limit
All four can be raised per account, with no release on the platform side, and changes take effect immediately.
A request needs three things: which tier, the target, and the shape of your peak (sustained versus bursty, and roughly how many seconds a burst lasts). The limit is then set and confirmed.
The ceiling depends on capacity at the time rather than being a fixed number. Supply an expected peak and the platform states plainly whether it can be carried, rather than letting you find out by hitting 429.
Number of child keys
The number of child keys per account is not capped. Ten thousand is the measured scale
(authentication, lookup by external_id, first page and deep pagination are all index hits). There
is no ceiling to reach, so there is no error code for reaching one.
⚠️ What is capped is how many rows one list request returns: GET /v1/keys?limit= defaults to
100 and is capped at 200. A larger value is clamped, not rejected. Page with next_cursor
(see Provisioning keys).
Ejection and recovery
These are observability figures, not an availability promise — no SLA number is published, as the product page says.
| Default | |
|---|---|
| Consecutive failures before ejection | 5 |
| Cooldown after ejection | 30 s, exponential backoff, capped at 300 s |
| Active probing | every 30 s, at most 20 channels per round; first round 60 s after startup |
| Health window | 300 s, minimum 5 samples, error-rate threshold 50% |
| Probe timeouts | 5 s connect / 20 s read |
Recovery is decided by a successful probe, not by elapsed time: once the cooldown ends the prober checks, and a channel that answers returns to the candidate list. An upstream wobble therefore costs at most one cooldown plus one probe interval.
⚠️ Breaker state is shared across instances (via Redis) rather than judged per instance; otherwise every instance would hit the same bad channel five times of its own.
Concurrency and timeouts
- Concurrent connections are not capped separately; the per-minute tiers above are the effective limit.
- Connect timeout to an upstream is 10 seconds; failing to connect fails that attempt and fails over.
- Read timeout is 300 seconds, and it is the gap between frames, not the total response length: a ten-minute answer is never cut off by it, an upstream that stalls for five minutes is.
- The number of upstreams tried per request is not capped by default: a call tries as many
routes as the model has configured. A cap applies only when one is configured for the account,
and the skipped candidates are listed as
never triedin the call trace.