One endpoint,
every model.
Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.
Errors
Errors follow the OpenAI shape — {"error": {"message": …, "type": …}} — because that is
what the SDKs read; anything else degrades to undefined in a client's logs.
| Status | Meaning | Action |
|---|---|---|
| 401 | The key is missing, malformed or revoked | Check the header. A 401 is about the credential, nothing else |
| 403 | Valid credential, but this call is not allowed | Wrong key type for the endpoint, a spent cap, or a source IP that is not on the key's allowlist — read metadata.refusal_reason |
| 404 | No route is configured for that model | Fix the model name, add a route, or add a catch-all |
| 402 | Account is out of balance (this and nothing else) | Top up. Access is restored automatically |
| 429 | Over the per-minute request tier | Back off using Retry-After; see Limits |
| 503 | Routes exist but every channel is failing | Retry. This is not a configuration problem |
Three kinds of insufficient funds
⚠️ A spent cap is 403, not 402. 402 means one thing here: the account is out of balance. The split between the two codes is whose problem it is — 402 is money owed by the account to the platform, 403 is a governance ceiling (a key's cap, a monthly budget) turning the call away.
The table below is the authoritative one; implement against it:
| Situation | HTTP | error.type / metadata.error_type | metadata.refusal_reason |
|---|---|---|---|
| This key's cumulative cap is spent | 403 | token_limit_exceeded | key_credit_exhausted |
| The account's monthly budget is spent | 403 | token_limit_exceeded | budget_exhausted |
| The account is out of balance | 402 | payment_required | account_balance_insufficient |
| The call came from an address not on the key's allowlist | 403 | permission_denied | source_ip_not_allowed |
If you are a platform relaying the reason to an end user, read metadata.refusal_reason, not the
message text:
refusal_reason | What happened | What to tell the end user |
|---|---|---|
key_credit_exhausted | This key's cumulative cap is spent | "You're out of credit" — topping up fixes it |
budget_exhausted | The account's monthly budget is spent | Nothing to do with them; do not send them to pay |
account_balance_insufficient | The account is in arrears | Nothing to do with them either |
source_ip_not_allowed | The call came from an address the key does not allow | Not about money. Retry from an allowed machine, or have the key's owner add this one |
Getting the middle two wrong means a user pays for something that was never about them; getting the last one wrong sends them to top up an account that has plenty of money.
{"error":{"code":403,"type":"token_limit_exceeded","message":"...","metadata":{
"error_type":"token_limit_exceeded",
"refusal_reason":"key_credit_exhausted"}}}
The two response envelopes
| Endpoints | Success | Failure |
|---|---|---|
/v1/chat/completions, /v1/keys/*, /v1/models, /v1/generation | OpenAI shape | non-2xx HTTP + {"error":{…}} |
/v1/integration/*, /v1/content-settings* | {"code","message","data","success"} | non-2xx HTTP, still that envelope (not 200 with success:false) |
⚠️ Reached through the api.jamerly.dev gateway, reconciliation responses are not unwrapped,
so the payload nests twice (data.data); reached directly on tare.jamerly.ai it nests once.
The test is whether data contains another code/success, not where you are deployed —
that holds under either deployment.
Error body shape
{"error":{
"code": 403,
"type": "token_limit_exceeded",
"message": "...",
"metadata": {
"error_type": "token_limit_exceeded",
"refusal_reason": "key_credit_exhausted",
"provider_code": "503",
"provider_raw": "{\"error\":{\"code\":\"model_not_found\",\"message\":\"...\"}}"
}}}
| Field | Guarantee |
|---|---|
error.code | Always equals the HTTP status; branching on it is safe |
error.type | Always present — the field OpenAI SDKs read |
metadata.error_type | Same value as error.type; some SDKs read this one instead |
metadata.refusal_reason | Only on the three refusals above |
metadata.provider_code | The upstream's own status code, passed through; may be absent |
metadata.provider_raw | The upstream's response body, byte for byte. Not parsed, not summarised, not filtered. Absent only when no upstream request was sent (the call was refused beforehand) |
[!NOTE]
error.messageis generated by Tare: it states which class of failure occurred and which status the upstream returned. Tare does not parse the upstream's error body to produce it — error shapes differ per provider, and inferring which field carries the explanation fails silently and yields a confidently worded but incorrect description.The upstream's own body is preserved verbatim in
metadata.provider_raw, for debugging and for forwarding to the provider unchanged. One exception: when the failure concerns the request itself (a malformed tool schema, an out-of-range parameter),messagealso carries the upstream's explanation, because "rejected as invalid" alone is not actionable.
[!NOTE] The top-level
typeandmetadata.error_typeare the same value, duplicated on purpose: the two ecosystems do not read the same field. With only the latter, a client written against the OpenAI SDK getsundefined— no error, it just renders a precise refusal as "something went wrong".
error_typevalues are unchanged; the breakdown is added onrefusal_reason.
404 versus 503
"No route configured" is fixed by changing configuration; "all upstreams are down" is fixed by waiting. Collapsing them into one status sends people to the wrong place, which is also why a catch-all route rescues the 404 case and never the 503 one.
[!NOTE] 404 used to be 402 here. 402 means "pay and you may continue", and this error never gets better if you pay — people topped up, checked their balance, and only then found a typo in a model name.
Routes exist but the modality does not match
You get a 404 with a different message: the configured upstreams do not accept images (or audio, or video). Declare the modality on the route, or use a model that takes it.