One endpoint,
every model.
Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.
Provisioning keys
When end users hold accounts and balances with you, a single shared key is not sufficient. The requirement is one key per user, each with its own cap, revocable at any time, which is what a provisioning key provides.
[!NOTE] The response nests the key as
{key, data:{…}}: the credential itself at the top level, its attributes underdata. A client already written against a provisioning API of this shape needs onlybase_urland the management credential changed.
Obtaining a provisioning key
Issue one yourself from the API Keys page in the console — pick Provisioning (master key) as the type. The plaintext appears exactly once, at creation; after that only a fingerprint is retained. Tare can also issue one on request.
⚠️ A provisioning key can only manage keys: it cannot make model calls, read billing, or touch channels and routing. A credential that can mint unlimited-quota child keys must not also be able to spend, or a leak has no ceiling.
Creating a child key
curl https://tare.jamerly.ai/v1/keys \
-H "Authorization: Bearer $TARE_PROVISIONING_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "user-8f3a2b91",
"limit": 12.50,
"external_id": "usr_8f3a2b91",
"include_byok_in_limit": false
}'
key in the response is the full credential and appears only once; data.hash is the
stable identifier you use from then on.
| Field | Meaning |
|---|---|
name | Purpose, supplied by the caller; Tare does not model your end users |
limit | Cumulative spend cap in USD, not a balance. Omit for unlimited |
external_id | Your own identifier, used for create idempotency (see below) |
include_byok_in_limit | Whether BYOK spend counts against this cap. Defaults to true |
Response fields
Create, fetch and cap-change all return the same data object:
| Field | Type | Notes |
|---|---|---|
hash | string | The stable identifier; every later operation uses it |
name | string | The purpose supplied at creation |
label | string | Prefix plus ellipsis, for display |
limit | number or null | Cumulative cap; null means uncapped, not 0 |
limit_remaining | number or null | What is left; also null when limit is null |
usage | number | Spent so far |
disabled | boolean | Whether it is deactivated |
include_byok_in_limit | boolean | Whether BYOK spend counts against the cap |
external_id | string | The external identifier; absent entirely when never set |
byok_usage | number | Always 0 today; reserved |
created_at / updated_at | string | ISO 8601 |
expires_at | string or null | Expiry |
Listing and pagination
GET /v1/keys?limit=&after=, cursor-paginated, including disabled keys (rotating your
management credential means walking every key you own).
[!NOTE] The console's API Keys page lists these too — pick the Child keys tab. It searches on name, purpose, and your own
external_id, which is usually the faster way to answer "what happened to my user usr_8f3a2b91".
{"data": [ /* the object above */ ], "next_cursor": 8123}
Where next_cursor lives | Top level, a sibling of data |
| Type | Number (an internal id) or null; null means no further page |
| How to use it | Pass it back verbatim as after |
limit default / max | 100 / 200; a larger value is clamped, not rejected |
⚠️ The cursor is an id, not an offset. With an offset, keys created mid-pagination make a row get skipped or returned twice.
When the plaintext is not available
Re-creating past the 5-minute window returns key: null plus two sibling fields:
{"key": null,
"key_unavailable_reason": "plaintext_window_expired",
"key_unavailable_message": "The plaintext is only returned within 5 minutes ...",
"data": { }}
key_unavailable_reasonis a programmable value;plaintext_window_expiredis the only one today; additions are announced in advance. Branch on this field.key_unavailable_messageis prose for humans. The wording will change; do not switch on it.
Expiry (expires_at)
Settable at creation and changeable afterwards (PATCH with expires_at; an explicit null
clears it).
- ⚠️ An expired key gets 401 on the model endpoint, indistinguishable from a revoked one. Telling them apart would confirm to someone holding an invalid credential that the key existed.
- ⚠️ Expiry does not flip
disabledtotrue; that field only means "deactivated". Readexpires_atto detect expiry.
Rotating a provisioning key
Provisioning keys are independent: one account can hold several at once and deactivate them individually, so rotation is zero-downtime:
- A new provisioning key is issued (plaintext appears once);
- Switch the configuration over so new traffic uses it;
- Confirm it works with
/whoamior anyGET /v1/keys; - Request deactivation of the old key.
⚠️ Do not deactivate the old key first. A provisioning key is your only path to issuing child keys; while it is gone every new user signup fails, and a failed signup is not recoverable.
⚠️ Rotating a provisioning key does not affect child keys already issued: they are independent credentials and do not expire with the key that created them.
Idempotent creation
When a create request times out you cannot tell whether it went through: retrying yields a second
key, not retrying leaves the user with none. Send external_id and the problem disappears —
creating again with the same identifier returns the existing key, with HTTP 200 rather
than 201.
[!WARNING] The full credential is only returned within 5 minutes of creation.
After that window, creating again still returns the key's
data, butkeyisnullwith akey_unavailable_reasonalongside it.Without the limit, anyone holding the provisioning key could retrieve any child key's plaintext by knowing or enumerating
external_id— and child keys spend money, so "the provisioning key cannot spend" would no longer cap a leak. If yourexternal_idis a user id, enumeration costs nothing.Genuine create-retries occur within seconds, so 5 minutes is ample. To obtain a fresh credential, rotate: disable the old key and create one under a new
external_id. That leaves an audit trail; silent re-retrieval does not.
[!WARNING]
external_iduniqueness includes disabled keys: one id maps to one key, permanently. To un-ban a user sendPATCH {"disabled": false}; do not create a new key with the sameexternal_id, or that user's usage history splits in two.
Addressing keys by external identifier
Instead of maintaining an external_id → hash mapping table, operate directly on your own
identifier:
curl https://tare.jamerly.ai/v1/keys/external/usr_8f3a2b91 \
-H "Authorization: Bearer $TARE_PROVISIONING_KEY"
curl -X PATCH https://tare.jamerly.ai/v1/keys/external/usr_8f3a2b91 \
-H "Authorization: Bearer $TARE_PROVISIONING_KEY" \
-H "Content-Type: application/json" \
-d '{"limit": 3.75}'
Semantics are identical to the hash variants. The path is /external/{id} rather than
letting {hash} accept both: both identifiers are free text, and sharing one position means that
the day your identifier looks like a hash, you silently operate on a different key.
Changing the cap and revoking
curl -X PATCH https://tare.jamerly.ai/v1/keys/$HASH \
-H "Authorization: Bearer $TARE_PROVISIONING_KEY" \
-H "Content-Type: application/json" \
-d '{"limit": 3.75}'
Effective on the next call, with no cache expiry to wait for. This is deliberate: lowering the cap right after deducting a user's credits makes any minute-scale delay a window of overspend.
Send {"disabled": true} to revoke, equally immediate.
Cap semantics
limit is a cumulative spend cap; usage only ever increases and never resets. Your
credit-sync logic can build directly on this:
limit = already_spent + value_of_remaining_credits
Three independent gates:
| Gate | Governs | Set by |
|---|---|---|
| Account balance | Whether the account is in arrears | Platform |
| Monthly budget | Ceiling for this month, resets monthly | Customer |
| Per-key cap | Total this key may ever spend, no reset | Customer (the limit field) |
The third is deliberately not merged into the monthly budget: credits do not grow back on the 1st.
Authorization
You can only touch keys you own. A hash that is not yours returns 404, not 403: a 403 would
confirm the hash is real.
Rate limits
Provisioning has its own tier of 1200/minute, counted separately from model traffic. Cap changes happen about as often as your billing events, and the model tier would 429 you on the first ramp.
All four tiers, how to have them raised, and the concurrency and timeout figures are in Limits.