Docs

One endpoint,
every model.

Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.

OperationsLimits

Limits

Four tiers, counted independently

Counted againstDefaultScope
Model calls (per key)100/minuteOne end user's runaway loop
Model calls (account total)6000/minuteTen thousand keys making a few calls each — undetectable per key, saturating in aggregate
Provisioning1200/minuteIssuing, capping and revoking child keys; as frequent as your billing events
Reconciliation / integration1200/minuteRead-only reporting. Month-end pulls hundreds of pages by day, key and model

⚠️ Model calls are counted per key, not per account: give every end user their own child key and each gets their own 100/minute, with no crowding out.

A streaming call counts once, regardless of duration or frame count — the counter increments at authentication time.

The response when a limit is exceeded

429 with Retry-After (seconds). ⚠️ A spent cap is 403, not 429: 429 means "slow down and retry", which clients back off on, while an exhausted cap never recovers by waiting. Full mapping in Errors.

Raising a limit

All four can be raised per account, with no release on the platform side, and changes take effect immediately.

A request needs three things: which tier, the target, and the shape of your peak (sustained versus bursty, and roughly how many seconds a burst lasts). The limit is then set and confirmed.

The ceiling depends on capacity at the time rather than being a fixed number. Supply an expected peak and the platform states plainly whether it can be carried, rather than letting you find out by hitting 429.

Number of child keys

The number of child keys per account is not capped. Ten thousand is the measured scale (authentication, lookup by external_id, first page and deep pagination are all index hits). There is no ceiling to reach, so there is no error code for reaching one.

⚠️ What is capped is how many rows one list request returns: GET /v1/keys?limit= defaults to 100 and is capped at 200. A larger value is clamped, not rejected. Page with next_cursor (see Provisioning keys).

Ejection and recovery

These are observability figures, not an availability promise — no SLA number is published, as the product page says.

Default
Consecutive failures before ejection5
Cooldown after ejection30 s, exponential backoff, capped at 300 s
Active probingevery 30 s, at most 20 channels per round; first round 60 s after startup
Health window300 s, minimum 5 samples, error-rate threshold 50%
Probe timeouts5 s connect / 20 s read

Recovery is decided by a successful probe, not by elapsed time: once the cooldown ends the prober checks, and a channel that answers returns to the candidate list. An upstream wobble therefore costs at most one cooldown plus one probe interval.

⚠️ Breaker state is shared across instances (via Redis) rather than judged per instance; otherwise every instance would hit the same bad channel five times of its own.

Concurrency and timeouts

  • Concurrent connections are not capped separately; the per-minute tiers above are the effective limit.
  • Connect timeout to an upstream is 10 seconds; failing to connect fails that attempt and fails over.
  • Read timeout is 300 seconds, and it is the gap between frames, not the total response length: a ten-minute answer is never cut off by it, an upstream that stalls for five minutes is.
  • The number of upstreams tried per request is not capped by default: a call tries as many routes as the model has configured. A cap applies only when one is configured for the account, and the skipped candidates are listed as never tried in the call trace.
Docs last updated Sep 20, 2026, 06:31 (UTC+8)