Skip to content

API reference

What Saylek's API surface serves: the endpoints, how streaming and model ids behave, and how a call fails. For RC integrations, use the hosted Saylek API. Start with Connect your app to choose the format your client uses.

It is not a list of request-body fields. Per-field support beyond stream and the model id is unpublished, and rather than print a list we have not established, we are leaving it unstated. One family is published: the chat path's four thinking controls, whose handling is stated in full under Thinking controls.

Base URLs

SurfaceBase URLAuth
Hosted, OpenAI dialecthttps://api.saylek.com/v1Authorization: Bearer <key>
Hosted, Anthropic dialecthttps://api.saylek.comx-api-key or Authorization: Bearer, plus anthropic-version

The Anthropic base URL carries no /v1, because Anthropic clients append /v1/messages themselves. The OpenAI base URL does. Getting this backwards is the most common cause of a 404.

Advanced: local daemon reference

The following table is retained for troubleshooting existing local installations, not as the recommended RC application setup. The local daemon API at 127.0.0.1:8443 does not authenticate requests. Do not expose it to a network. A loopback connection is not a local-only guarantee; see Privacy and egress.

For hosted request examples, use OpenAI setup or Anthropic setup. Do not infer hosted endpoint support from this local table.

The daemon serves the following OpenAI-dialect routes:

EndpointPurpose
GET /v1/modelsWhat this daemon can serve: models discovered on this machine, plus any upstreams you configured. It is not the list of what you can reach. For that, run saylek models --account.
POST /v1/chat/completionsChat completions. The main path, and the one to build on. Its handling of reasoning_effort and the three sibling thinking fields is published under Thinking controls.
POST /v1/embeddingsEmbeddings.
POST /v1/audio/transcriptionsAudio transcription, but no released build ships a speech model, so it answers 400 rather than guessing. Pointing [audio] at an engine of your own is necessary without being sufficient: the daemon verifies that engine's own size bound instead of taking your word for it, by sending an over-budget clip and requiring a 413. An engine that accepts the clip, or refuses it any other way, leaves audio un-offered.
POST /v1/responsesResponses-style calls, but on this daemon only against a proxy upstream you have configured. A local model or a machine shared with you returns model_not_found here. The hosted surface is a different path: it has answered a plain, non-streaming text request from a shared machine.
POST /v1/completionsNot implemented. The route exists and answers 501. It is listed so you do not spend an afternoon wondering why the legacy path 404s differently than you expected. Use the chat path.

The hosted Anthropic surface adds POST /v1/messages, covered on Claude Code and the Anthropic SDK.

Streaming

stream: true is supported on the chat path and emits standard OpenAI SSE frames, so a client that already consumes OpenAI streaming works without changes.

A streamed request that fails after the stream has started does not change the status code: the response is already HTTP 200, so the error arrives as a frame, data: {"error":{"code":…,"message":…,"details":{"reason":…}}}, with no [DONE] after it. Check each frame for error rather than relying on the status.

Model ids are verbatim

Saylek carries the advertised model id, exactly as advertised. There is no aliasing layer: gpt-4o, claude-3-5-sonnet and similar names from other providers do not resolve, even through the Anthropic dialect. Always discover first:

curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"

Rate limiting and availability

There is a rate limiter, and it is on by default. Your local daemon ships with a limit of 100 requests per second with a burst allowance of 200, and it answers HTTP 429 when you exceed it. It is a spike absorber rather than a quota: you are unlikely to meet it by hand, and quite likely to meet it with a parallel batch job. Handle 429 and back off. The values live under [ratelimit] in ~/.saylek/config.toml if you need to change them on your own machine.

What does not exist yet is a published quota or token ceiling on the hosted surface. Rather than print a number we have not committed to, we are leaving that unstated until it is real.

Availability is the failure mode that will actually bite you, because capacity comes from members' machines:

  • A request that nothing you can reach takes right now waits for a machine, by default up to 60 seconds (up to 180 seconds when streamed), and then fails with a 503 that carries the reason wait_elapsed. It does not queue indefinitely. A model id that nothing serves anywhere fails at once with model_not_found.
  • Locally, a request for a model with nothing loaded returns model_not_found.
  • A daemon built with no inference backend returns HTTP 503 no inference backend compiled in.

See Troubleshooting for the error shapes.

Limits and status codes a client has to handle

Request body32 MiB on the chat, embeddings and responses paths. Larger bodies are rejected at the edge with 413. Audio transcription is bounded separately by the upload itself.
Rate100 requests/second, burst 200. Over that is 429.
Retry-AfterSent on the hosted API's 429 and 503 error responses. Honour it rather than retrying immediately. It tells you the earliest sensible retry, not that the retry will work, so read error.code to tell a busy moment from a condition that waiting does not fix.

The status codes worth branching on:

CodeWhat it meansWhat to do
400The request is malformed.Fix the request. Retrying will not help.
401Key missing, wrong, or revoked. On the hosted surface only; a local daemon does not check keys.Check the key.
404The base URL or the model. A hosted 404 whose error.code is model_not_found is about the model; see Hosted error codes. Otherwise, suspect the base URL first.Check whether your base URL should include /v1, then check the model id against GET /v1/models.
413Body over 32 MiB.Send less.
429Rate limited.Back off, honour Retry-After.
501The route exists but is not implemented, for example /v1/completions.Use the chat path.
503Nothing could serve the request. On the hosted surface, error.code says which case you hit; see Hosted error codes. A local daemon built with no inference backend also answers 503.Honour Retry-After, and read error.code before deciding whether to retry.

Hosted error codes

On the hosted surface, branch on the error body's code and not only on the HTTP status. Several 503s exist, and a 403 or 404 is not always about your key or your URL. Match code exactly: the casing is mixed, so some codes arrive lowercase and some uppercase.

The table covers chat completions, responses and messages requests. Hosted /v1/embeddings and /v1/audio/transcriptions are narrower: for a model outside your reach they answer 404 model_not_found, even when a machine you cannot reach serves it. They do not return NOT_IN_CIRCLE, but they do return CASCADE_EXHAUSTED, machine_offline and outside_sharing_hours.

StatuscodeWhat it meansWhat to do
404model_not_foundNothing anywhere serves this model id right now.Check the id against GET /v1/models and send it exactly as listed. Retrying the same id does not help until something serves it.
403NOT_IN_CIRCLENobody shares with you, and no machine of your own serves this model.Ask someone who runs this model to share a machine with you, or connect a machine of your own that runs it. Retrying on its own does not change this.
503CASCADE_EXHAUSTEDNothing you can reach took this request. Either nobody who shares with you runs this model, or the machines that run it are busy, paused or offline, or Saylek could not confirm just now who shares with you.If nobody who shares with you runs this model, waiting does not help: ask someone who runs it to share a machine with you, or connect a machine of your own that runs it. Otherwise retry after Retry-After, and check the machines you rely on are online if it keeps happening.
503CAPACITY_FULLMachines that run this model are all busy, or the one picked for your request became unavailable before it started.Retry after Retry-After.
503machine_offlineA machine you can use runs this model, and it is offline right now. You still have access.Retry after Retry-After. The model comes back when that machine reconnects.
503outside_sharing_hoursMachines shared with you run this model, and every one that is online is outside its sharing hours right now. You still have access.Retry later, during those machines' sharing hours. Retry-After here is the standard back-off, not the time those hours begin.

A 503 can also say why it was refused. When the request waited for a machine and none took it in time, the error carries the reason wait_elapsed: in error.reason on a non-streamed request, and in error.details.reason in a streamed error frame. By default that happens after up to 60 seconds, or up to 180 seconds when streamed.

In the OpenAI dialect, the error type follows the status rather than the cause: requests for a 429, server_error for a 5xx, and invalid_request_error otherwise. The error also carries a correlation_id, echoed in the X-Correlation-ID header; quote it when you report a problem. The Anthropic dialect uses the same code values inside its own error envelope and type names.

Keys

Member API keys are created at /applications. A key's secret is shown once, at creation. Keys are member-scoped, so a key reaches everything shared with you in its region.

Revoke a key at /applications when it should stop working, for example if it leaked.

Where requests run

Every hosted call, and every local call your own machine cannot serve, goes through Saylek's registry to a machine that can serve it: one of your own machines, one belonging to someone who shares with you, or, if you joined a group (one an organization runs or one anyone can join), one belonging to anyone in it. Nowhere else. Build with that in mind. Privacy and egress is the page to read before sending anything sensitive through an integration.

Next steps

Last checked 2026-09-23 · read as markdown at /docs/api-reference.md