API reference
What Saylek's API surface serves: the endpoints, how streaming and model ids behave, and how a call fails. For RC integrations, use the hosted Saylek API. Start with Connect your app to choose the format your client uses.
It is not a list of request-body fields. Per-field support beyond stream and the model
id is unpublished, and rather than print a list we have not established, we are leaving it
unstated. One family is published: the chat path's four thinking controls, whose handling
is stated in full under Thinking
controls.
Base URLs
| Surface | Base URL | Auth |
|---|---|---|
| Hosted, OpenAI dialect | https://api.saylek.com/v1 | Authorization: Bearer <key> |
| Hosted, Anthropic dialect | https://api.saylek.com | x-api-key or Authorization: Bearer, plus anthropic-version |
The Anthropic base URL carries no /v1, because Anthropic clients append
/v1/messages themselves. The OpenAI base URL does. Getting this backwards is the most
common cause of a 404.
Advanced: local daemon reference
The following table is retained for troubleshooting existing local installations, not as
the recommended RC application setup. The local daemon API at 127.0.0.1:8443 does not
authenticate requests. Do not expose it to a network. A loopback connection is not a
local-only guarantee; see Privacy and egress.
For hosted request examples, use OpenAI setup or Anthropic setup. Do not infer hosted endpoint support from this local table.
The daemon serves the following OpenAI-dialect routes:
| Endpoint | Purpose |
|---|---|
GET /v1/models | What this daemon can serve: models discovered on this machine, plus any upstreams you configured. It is not the list of what you can reach. For that, run saylek models --account. |
POST /v1/chat/completions | Chat completions. The main path, and the one to build on. Its handling of reasoning_effort and the three sibling thinking fields is published under Thinking controls. |
POST /v1/embeddings | Embeddings. |
POST /v1/audio/transcriptions | Audio transcription, but no released build ships a speech model, so it answers 400 rather than guessing. Pointing [audio] at an engine of your own is necessary without being sufficient: the daemon verifies that engine's own size bound instead of taking your word for it, by sending an over-budget clip and requiring a 413. An engine that accepts the clip, or refuses it any other way, leaves audio un-offered. |
POST /v1/responses | Responses-style calls, but on this daemon only against a proxy upstream you have configured. A local model or a machine shared with you returns model_not_found here. The hosted surface is a different path: it has answered a plain, non-streaming text request from a shared machine. |
POST /v1/completions | Not implemented. The route exists and answers 501. It is listed so you do not spend an afternoon wondering why the legacy path 404s differently than you expected. Use the chat path. |
The hosted Anthropic surface adds POST /v1/messages, covered on Claude Code and the
Anthropic SDK.
Streaming
stream: true is supported on the chat path and emits standard OpenAI SSE frames, so a
client that already consumes OpenAI streaming works without changes.
A streamed request that fails after the stream has started does not change the status code:
the response is already HTTP 200, so the error arrives as a frame,
data: {"error":{"code":…,"message":…,"details":{"reason":…}}}, with no [DONE] after it.
Check each frame for error rather than relying on the status.
Model ids are verbatim
Saylek carries the advertised model id, exactly as advertised. There is no
aliasing layer: gpt-4o, claude-3-5-sonnet and similar names from other providers do not
resolve, even through the Anthropic dialect. Always discover first:
curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
Rate limiting and availability
There is a rate limiter, and it is on by default. Your local daemon ships with a limit of
100 requests per second with a burst allowance of 200, and it answers HTTP 429
when you exceed it. It is a spike absorber rather than a quota: you are unlikely to meet it
by hand, and quite likely to meet it with a parallel batch job. Handle 429 and back off. The
values live under [ratelimit] in ~/.saylek/config.toml if you need to change them on
your own machine.
What does not exist yet is a published quota or token ceiling on the hosted surface. Rather than print a number we have not committed to, we are leaving that unstated until it is real.
Availability is the failure mode that will actually bite you, because capacity comes from members' machines:
- A request that nothing you can reach takes right now waits for a machine, by default up
to 60 seconds (up to 180 seconds when streamed), and then fails with a 503 that
carries the reason
wait_elapsed. It does not queue indefinitely. A model id that nothing serves anywhere fails at once withmodel_not_found. - Locally, a request for a model with nothing loaded returns
model_not_found. - A daemon built with no inference backend returns HTTP 503
no inference backend compiled in.
See Troubleshooting for the error shapes.
Limits and status codes a client has to handle
| Request body | 32 MiB on the chat, embeddings and responses paths. Larger bodies are rejected at the edge with 413. Audio transcription is bounded separately by the upload itself. |
| Rate | 100 requests/second, burst 200. Over that is 429. |
| Retry-After | Sent on the hosted API's 429 and 503 error responses. Honour it rather than retrying immediately. It tells you the earliest sensible retry, not that the retry will work, so read error.code to tell a busy moment from a condition that waiting does not fix. |
The status codes worth branching on:
| Code | What it means | What to do |
|---|---|---|
400 | The request is malformed. | Fix the request. Retrying will not help. |
401 | Key missing, wrong, or revoked. On the hosted surface only; a local daemon does not check keys. | Check the key. |
404 | The base URL or the model. A hosted 404 whose error.code is model_not_found is about the model; see Hosted error codes. Otherwise, suspect the base URL first. | Check whether your base URL should include /v1, then check the model id against GET /v1/models. |
413 | Body over 32 MiB. | Send less. |
429 | Rate limited. | Back off, honour Retry-After. |
501 | The route exists but is not implemented, for example /v1/completions. | Use the chat path. |
503 | Nothing could serve the request. On the hosted surface, error.code says which case you hit; see Hosted error codes. A local daemon built with no inference backend also answers 503. | Honour Retry-After, and read error.code before deciding whether to retry. |
Hosted error codes
On the hosted surface, branch on the error body's code and not only on the HTTP status. Several
503s exist, and a 403 or 404 is not always about your key or your URL. Match
code exactly: the casing is mixed, so some codes arrive lowercase and some uppercase.
The table covers chat completions, responses and messages requests. Hosted
/v1/embeddings and /v1/audio/transcriptions are narrower: for a model outside your reach
they answer 404 model_not_found, even when a machine you cannot reach serves it. They do
not return NOT_IN_CIRCLE, but they do return CASCADE_EXHAUSTED, machine_offline and
outside_sharing_hours.
| Status | code | What it means | What to do |
|---|---|---|---|
404 | model_not_found | Nothing anywhere serves this model id right now. | Check the id against GET /v1/models and send it exactly as listed. Retrying the same id does not help until something serves it. |
403 | NOT_IN_CIRCLE | Nobody shares with you, and no machine of your own serves this model. | Ask someone who runs this model to share a machine with you, or connect a machine of your own that runs it. Retrying on its own does not change this. |
503 | CASCADE_EXHAUSTED | Nothing you can reach took this request. Either nobody who shares with you runs this model, or the machines that run it are busy, paused or offline, or Saylek could not confirm just now who shares with you. | If nobody who shares with you runs this model, waiting does not help: ask someone who runs it to share a machine with you, or connect a machine of your own that runs it. Otherwise retry after Retry-After, and check the machines you rely on are online if it keeps happening. |
503 | CAPACITY_FULL | Machines that run this model are all busy, or the one picked for your request became unavailable before it started. | Retry after Retry-After. |
503 | machine_offline | A machine you can use runs this model, and it is offline right now. You still have access. | Retry after Retry-After. The model comes back when that machine reconnects. |
503 | outside_sharing_hours | Machines shared with you run this model, and every one that is online is outside its sharing hours right now. You still have access. | Retry later, during those machines' sharing hours. Retry-After here is the standard back-off, not the time those hours begin. |
A 503 can also say why it was refused. When the request waited for a machine and none took
it in time, the error carries the reason wait_elapsed: in error.reason on a non-streamed
request, and in error.details.reason in a streamed error frame. By default that happens after
up to 60 seconds, or up to 180 seconds when streamed.
In the OpenAI dialect, the error type follows the status rather than the cause: requests
for a 429, server_error for a 5xx, and invalid_request_error otherwise. The error also
carries a correlation_id, echoed in the X-Correlation-ID header; quote it when you report
a problem. The Anthropic dialect uses the same code values inside its own error envelope and
type names.
Keys
Member API keys are created at /applications. A key's secret is shown once, at creation.
Keys are member-scoped, so a key reaches everything shared with you in its region.
Revoke a key at /applications when it should stop working, for example if it leaked.
Where requests run
Every hosted call, and every local call your own machine cannot serve, goes through Saylek's registry to a machine that can serve it: one of your own machines, one belonging to someone who shares with you, or, if you joined a group (one an organization runs or one anyone can join), one belonging to anyone in it. Nowhere else. Build with that in mind. Privacy and egress is the page to read before sending anything sensitive through an integration.
Next steps
- Use your models from anywhere: hosted setup, end to end.
- Troubleshooting: 401s, 404s, and rejected model ids.
- Privacy and egress: what leaves the machine.
Last checked 2026-09-23 · read as markdown at /docs/api-reference.md