One endpoint,
every model.
Tare is an OpenAI-compatible gateway in front of every model you use — ours and the ones you bring yourself. One base URL, one key, and every token accounted for.
Endpoints
Base URL: https://tare.jamerly.ai
Model traffic
| Method | Path | What it does |
|---|---|---|
| POST | /v1/chat/completions | Model calls. OpenAI-shaped, streaming supported |
| POST | /v1/embeddings | Text to vectors. OpenAI-shaped, no streaming |
| POST | /v1/images | Image generation. No streaming |
| POST | /v1/images/generations | The same endpoint under its other common name |
| GET | /v1/models | Models available to your account, with prices |
| GET | /v1/generation?id=gen-… | What one call cost, after the fact |
Unsupported endpoints
Model traffic is the six above and nothing else. These do not exist; requesting them returns a 404:
| Path | Status |
|---|---|
/v1/images/edits, /v1/images/variations | Not supported |
/v1/audio/* (transcription, TTS) | Not supported |
/v1/completions (legacy text completion) | Not supported. Use /v1/chat/completions |
/v1/credits | Not supported; read your balance from /v1/integration/balance — see Reconciliation API |
POST /v1/embeddings: text to vectors
OpenAI-shaped. The request body is passed through untouched; the response is the upstream's
own body plus a usage.cost field:
{"model":"text-embedding-3-small","input":"hello"}
Three differences from chat:
- ⚠️ Model names are not mapped. On chat you send the platform-side model name (the
idfromGET /v1/models); here you send the upstream's own name. Embeddings do not use the model routing table, so there is no mapping layer — and the model name in the response is not rewritten either. - ⚠️ A channel serves embeddings only once you configure it. Enter the embeddings address for that channel on the Channels page; leaving it empty means the channel does not offer the endpoint. Nothing is inferred from the dialect — speaking the OpenAI dialect for chat says nothing about having an embeddings endpoint, and many relays, aggregating gateways and self-hosted inference servers offer chat only.
- Selection order: your own channels win, the platform channel is the fallback. Every channel of yours with an embeddings address configured is a candidate, in ascending channel id. Only when you have configured none does the platform channel designated for embeddings take over.
- Your own channels fail over between themselves. Configure several and each is tried in turn; every attempt is listed in the call trace.
- ⚠️ There is no failover from your channels to the platform one. They bill differently: your own channel is metered but not charged, the platform channel is charged at list price. A silent fallback would begin charging without consent, while the only visible signal is a successful call. When all of your own channels fail, the call returns an error rather than a substitution.
- ⚠️ A model with no platform price is refused there, not served for free. Tokens on the platform channel are bought by Tare from the upstream; when the price list has no row for that model the charge computes to zero, and serving it silently for free is worse than failing. The call returns 404 with the reason. Use one of your own channels, or ask us to price the model.
No streaming. The gen- id is returned in the response header as usual, and
/v1/generation can look the call up.
[!NOTE]
⚠️ Embedding models do not appear in GET /v1/models. That catalogue comes from the
upstream's own model list, and upstreams do not list embedding models there (measured: 421
models, 0 of them embeddings). Take the model name from the upstream's documentation; it cannot be
selected from the platform catalogue.
[!NOTE] Embedding calls follow the same retention switch as chat. With the switch on, the request body and the upstream response are stored and can be expanded in the call trace; with it off, nothing is written.
⚠️ This endpoint previously stored nothing regardless of the switch, on the grounds that embedding inputs are routinely whole documents. The cost was that it showed nothing in call traces at all — a successful embedding call looked exactly like one that never finished. Retention is now your decision rather than a property of the endpoint.
POST /v1/images: image generation
{"model":"google/gemini-3-pro-image","prompt":"a red fox in the snow"}
Request
| Field | |
|---|---|
model | Required. A model you have routed on the Routing page, or a platform default |
prompt | Required. The image description |
| anything else | Passed to the upstream unchanged — size, aspect_ratio, n, input_references, provider-specific options. What is accepted depends on the model; a field the upstream does not support returns the upstream’s error, not Tare’s |
POST /v1/images/generations is the same endpoint — same request, same response.
Response
{
"id": "gen-c9d21c16eec04f968fc9bbcd08e140f9",
"created": 1787820572,
"model": "google/gemini-3-pro-image",
"data": [
{
"b64_json": "iVBORw0KGgo…",
"media_type": "image/png",
"url": "https://s.jamerly.ai/img/36/b21a5d9fa79b4135993282e1baf34999.png",
"expires_at": 1787906972
}
],
"usage": {"prompt_tokens": 7, "completion_tokens": 4175, "total_tokens": 4182, "cost": 0.000900}
}
| Field | |
|---|---|
id | The call's id. Same value as the X-Generation-Id response header; use it with /v1/generation |
model | The requested model name, not the upstream’s |
data[].b64_json | The image bytes, base64. This is the durable copy |
data[].media_type | image/png, image/jpeg, … |
data[].url | A hosted copy of the same bytes. ⚠️ Expires 24 hours after the call |
data[].expires_at | Unix seconds: when that URL stops working. Authoritative |
text | Text returned by the model alongside the image. Absent when empty. The standard image response has no field for free text, so this field is an addition; no standard field is repurposed |
finish_reason | The upstream’s stop reason. ⚠️ Read this first when data is empty — content_filter means the image was generated and then blocked by the upstream’s own policy |
usage | Tokens, and cost = the cost of this call. Images are billed on tokens, which is why completion_tokens is large |
[!WARNING]
datacan be an empty array, and that is not an error. The upstream answered normally and returned no image — the model replied with text only, or the upstream content policy blocked it. Tare does not reclassify this as a failure: the status is still 200, and the reason is carried infinish_reasonandtext.
[!WARNING] Do not persist
urlin your database. It expires the next day, and is indistinguishable from a permanent link until it breaks. Persistb64_jsoninstead, or copy the bytes to storage you control. The URL exists so an image can be handed to another system without transferring megabytes through your own service.
Behaviour
- Routing is by model, the same as
/v1/chat/completions: whichever channel you pointgoogle/gemini-3-pro-imageat is where the call goes. Image models differ in price by more than an order of magnitude, so model choice has a significant cost impact. - ⚠️ No failover. A channel failure returns that error directly; Tare does not silently retry elsewhere. A second candidate can bill differently, and falling back would incur cost on a channel that was not chosen. An image costs cents, not fractions of a cent.
- Billing, limits and call traces are identical to chat: the same balance, budget and key-credit checks run first, the call lands in call traces with its prompt, and it is billed the same way. The endpoint does not change the billing rules applied.
- No streaming.
Image generation from a chat call: modalities
Also supported: add modalities to a /v1/chat/completions call and the images are returned
in message.images. Use this when one call should produce text and images; use /v1/images
when only the image is needed.
{"model":"…","messages":[{"role":"user","content":"a red fox"}],
"modalities":["image","text"]}
The request body is passed through untouched and the response is the upstream's own body
(only model and id are rewritten). Three points to note:
- ⚠️ Routing looks at the input modality only, not at the output you asked for. If a model name has several candidate channels and only some can produce images, a failover can land on one that cannot; the symptom is "the same request sometimes returns an image and sometimes only text". Avoid it in configuration: give the image-generating model its own model name, or a single channel. Selecting candidates by output modality is on the roadmap.
- Billing follows tokens and the cost the upstream reports.
pricing.imageinGET /v1/modelsis always 0, which does not mean free — there is no per-image tier. For what a call actually cost, readusage.costfrom the response. - Whether the image is returned as base64 or a URL is decided by the upstream; Tare does not rewrite it.
Request headers
⚠️ Unknown request headers are ignored: they never cause a rejection, and they are never forwarded
upstream. The headers sent upstream are rebuilt by Tare (credential, Content-Type, whatever the
dialect requires) — no header you send appears on the upstream side.
Attribution headers such as HTTP-Referer and X-Title fall into this case: harmless to
sending them has no effect (they affect neither routing, billing nor rate limits), but they do not
reach the upstream. For attribution, use one of the following:
| Requirement | Method |
|---|---|
| Which workload a call belongs to | usage_label in the request body, stored verbatim on the usage record |
| Which end user a call belongs to | Issue that user their own child key — see Provisioning keys |
[!NOTE] CORS is not enabled. Called straight from a browser, those custom headers are blocked by the browser itself at the preflight step: the request never leaves. Server-side calls are unaffected, and server-side is where these calls belong — an inference key in a browser is a published key.
gen- ids
Every call returns an id of the form gen-xxxxx. It is ours, not the upstream's: use it to
look up cost, and as an idempotency key in your own bookkeeping.
The prefix matters. Clients commonly test id.startsWith("gen-") to decide
whether a call has an id at all; a different prefix makes that check fail silently and the whole
reconciliation chain does nothing.
Key management (provisioning)
For platforms issuing keys to their own end users. Authenticated by a provisioning key, no login session involved. See Provisioning keys.
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/keys | Create a child key (idempotent when external_id is supplied) |
| GET | /v1/keys | List your child keys — cursor-paginated, disabled ones included |
| GET | /v1/keys/{hash} | Read one key's cap and usage |
| PATCH | /v1/keys/{hash} | Change the cap, revoke, or rename |
| GET | /v1/keys/external/{externalId} | Same, addressed by your own identifier |
| PATCH | /v1/keys/external/{externalId} | Same, addressed by your own identifier |
Console and reconciliation
The console (channels, routing, billing, alerts) is a separate surface authenticated by your login. For machine-to-machine reconciliation there is a dedicated key-authenticated API — see Reconciliation API.
[!NOTE] The three key types do not overlap: inference keys only call models, integration keys only read reports, provisioning keys only manage keys. Using the wrong type returns an explicit refusal, not a vague 401.
GET /v1/models: fields and units
[!NOTE] Send your inference key and the catalogue includes the models you configured yourself: the ones on your own providers ∪ the platform ones, yours winning on a name both sides have — exactly the routing you get when you actually call. Without a key you get the platform catalogue (that path stays; a client may fetch the catalogue with no headers at all).
Both an inference key and a provisioning key work. A provisioning key is not bound to a router, so it sees the catalogue of your account's default routing; to see what one child key can call, call with that child key.
⚠️ A key that is present but invalid or expired returns 401 — you never silently get a catalogue that is missing your own models. A reconciliation key returns 403: it does not reach the model surface.
⚠️ For models on your own providers, pricing.prompt / completion are 0 —
those tokens are paid to your upstream directly and do not pass through the platform price sheet.
Read usage.cost in the response for what a call actually cost.
{"data":[{"id":"GLM 5.1","name":"GLM 5.1","description":"…","context_length":128000,
"architecture":{"modality":"text->text","input_modalities":["text"],"output_modalities":["text"]},
"reasoning":{"mandatory":false,"default_enabled":true,
"supported_efforts":["low","medium","high"],"default_effort":"medium"},
"pricing":{"prompt":"0.0000006","completion":"0.0000022",
"input_cache_read":"0.00000006","input_cache_write":"0.00000075",
"request":"0","image":"0"}}]}
| Field | Meaning |
|---|---|
id | The platform-side model name — use it to make calls |
pricing.prompt / completion | Per token (not per million), as strings |
pricing.input_cache_read / input_cache_write | The two cache tiers, per token as well |
pricing.request / image | Always 0 — see the warning below |
| Currency | USD. No conversion to other currencies today |
description | The model blurb from the upstream catalogue; absent when unavailable |
context_length | From the upstream catalogue; absent when unavailable (do not read it as 0) |
architecture | From the upstream catalogue; falls back to text-only when unavailable |
reasoning | Reasoning capabilities from the upstream catalogue; absent when unavailable — no default is invented |
⚠️ pricing.image is always 0, but images are not free. The price sheet has four token
tiers only (input / output / cache read / cache write); there is no per-image tier, and calls
carrying images are billed on the tokens and cost the upstream reports. Do not estimate image
cost from this field — read usage.cost off the response.
reasoning: reasoning capabilities
{"mandatory":false,"default_enabled":true,
"supported_efforts":["low","medium","high"],"default_effort":"medium"}
Use it to decide whether to show a reasoning-effort picker and which levels to offer:
- The
supported_effortsvocabulary comes from the upstream and is not always those three (some upstreams also offerminimal/xhigh/none). Render what the field says; do not hard-code a list client-side. mandatory: truemeans it cannot be turned off — do not render a toggle in that case.- ⚠️ An absent
reasoningmeans unknown, not "no reasoning". Self-hosted and local models have no catalogue entry, so the field is missing for them; reading absence as "unsupported" makes the picker vanish for exactly those models. No fallback is applied: inventing a default would make a promise on the upstream's behalf, and a user picking a level the upstream rejects gets a 400 that looks like a broken model.
⚠️ There is no tool-calling capability flag today; whether a model supports tool calls cannot be read from this catalogue.
usage.cost: the cost of the call
In USD:
- Platform models: the platform sell price;
- BYOK: upstream cost + service fee. BYOK tokens are paid to the vendor directly; a 0 here would make the cost report show millions of tokens as free for the month.
- ⚠️ When it cannot be computed the field is absent, not 0 — a fake 0 gets believed.
"This model has no price yet" counts as not computable, so a
0you do see is a real 0. - ⚠️ The final streaming frame carries
costtoo (that is the frame with usage). - For upstreams that report their own cost in the response, the reported value is used, so no cost-price entry is required.
- ⚠️ Platform-model responses carry no
usage.cost_details: that field is what the upstream charged the platform. On BYOK it is retained — that amount appears on your own upstream invoice and is exactly what cost attribution requires. The same rule applies toupstream_inference_costonGET /v1/generation: BYOK calls only.
⚠️ architecture is passed through from the upstream catalogue: self-hosted and local models have
no catalogue entry and fall back to text-only, so a front end filtering on input_modalities
containing image will drop them.
GET /v1/generation?id=: when the data is queryable
It reads from the usage record, and the usage record is written after the response, so querying the moment the id is received returns 404.
⚠️ A 404 means two things and shares one code today: not written yet (retry) and not your call (give up). Use time as the test: a 404 within 60 seconds of receiving the id is "not written yet", beyond that treat it as the latter. Not distinguishing them is deliberate — it would hand anyone an oracle for "does this id exist".
- Non-streaming: usually queryable within a second of the response.
- Streaming: written after the final frame — measure from the end of the stream.
- 404 is a normal path, not an error: treat it as "try again shortly" and back off a few seconds.
- ⚠️ For streams that broke mid-flight, the cost has to be recovered from the upstream afterwards, which can take minutes to appear.
Streaming: stream_options.include_usage is forced on
Without usage there is no billing, so this overrides the setting in the request (the whole
stream_options object is replaced). Handle the shape it produces:
| Upstream dialect | Final frame |
|---|---|
| OpenAI-compatible | The upstream's own final frame, passed through: choices is an empty array, with usage |
| Anthropic | Translated by Tare: choices has one entry, an empty delta, a finish_reason and usage |
⚠️ The two dialects produce different final frames. If you render per frame, write the test as
"choices is empty or the delta carries no content → skip rendering, take usage" and both
dialects are covered. The Anthropic path is not padded with an extra empty-choices frame: that
would add a frame.
OpenAPI / JSON Schema
The customer-facing surface (/v1/chat/completions, /v1/keys, /v1/models,
/v1/generation, /v1/integration/*) has a hand-maintained OpenAPI 3.1 document, kept in
step with these docs:
| View it | Single page, works offline, no external dependencies |
| Download openapi.yaml | Generate a client, or import into Postman / Insomnia |
⚠️ The auto-generated document is not published: it exposes the console, ops and health endpoints, which are not a public contract and are not open on the gateway. The hand-written one covers exactly the surface described in these docs.
⚠️ The document describes shapes only. It does not repeat the reasoning here (why a spent cap is 403, why the final streaming frame differs between dialects, …). Read both when you integrate.