Skip to content

Use the models you already serve

Connect the runtime you already run (Ollama, LM Studio, llama.cpp or vLLM), and call its models from compatible applications with a Saylek API key, on this machine or anywhere else. Your runtime keeps loading the models and managing the GPU.

Before you begin

You needDetail
Your runtime, runningOn the machine you are connecting, listening on its own address (127.0.0.1). Default ports: Ollama 11434, LM Studio 1234, llama.cpp 8080, vLLM 8000.
A supported machineLinux x86_64, or macOS on Apple Silicon. On Windows, the Linux build inside WSL2.
A Saylek accountCreate one at saylek.com with your email.

1. Connect the machine and its runtime

On the machine that runs your runtime:

curl -fsSL https://saylek.com/install.sh | sh -s -- --connect

This signs you in through your browser, links the machine to your account, and connects a runtime it finds answering on one of the default ports. It does not share the machine with anyone. Installed earlier without --connect? Run the line again; it is safe to repeat.

You know it worked when the installer says the machine is connected, and in Machines its Models section lists your runtime's models.

If the installer did not find your runtime, for example because it uses another port, connect it on the machine with its address and port:

saylek model add http://127.0.0.1:11434   # Ollama. vLLM is usually :8000

It prints [ok] register endpoint when it connects. If the runtime needs its own API key, put the key in an environment variable and add --token-env VARIABLE_NAME. For a runtime you started after installing, you can also open the machine in Machines, then Settings, Manage sources, and choose Scan again; the machine's Models section then shows the same command, with Copy command.

2. Check that your models are ready

In Machines, open the machine. Its Models section shows each of your runtime's models as Ready, and the page says Only you can use its models. Your own keys can use them with sharing off.

From the terminal, saylek models lists the same models, each with your runtime's address.

3. Connect an application

  1. Open Applications, choose Add application, name it, and choose Create key. Copy the key now: it is shown once.

  2. List the models your key can use, and pick one of your runtime's model ids:

    curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
    

You know it worked when your runtime's models are in the list.

4. Send your first request

curl https://api.saylek.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'

You know it worked when the response carries an answer. To set up your own client, follow Connect your app: OpenAI-compatible clients use https://api.saylek.com/v1, and Claude Code and Anthropic clients use https://api.saylek.com with no /v1.

Where your prompt goes. Hosted requests pass through Saylek's servers, even when your own machine serves them. If someone who shares with you serves the same model, their machine can answer instead of yours. Privacy and egress has the details.

If something goes wrong

What you seeWhat to do
The installer ends with waiting for this machine to connectRun saylek status on the machine; Machines updates when it connects
saylek models says no models loaded yetStart the runtime, then saylek model add <address>
missing portAdd the port to the address, for example :11434
connection refused or upstream unreachableStart the runtime, or give the port it actually listens on
authentication requiredThe runtime expects a key: add --token-env VARIABLE_NAME
The machine's page says your runtime "is not answering", or no longer lists its modelsStart the runtime again, then check with saylek models
Your models are missing from the hosted listCheck the machine is online in Machines, or run step 1 again with --connect
Only some of them are missingAn earlier selection leaves them out: saylek host models --clear
The daemon is not runningStart it as Quickstart step 2 describes

Troubleshooting covers the rest.

Optional: use Saylek on this machine without a key

Saylek also runs a local API on the machine, for applications running there:

SettingValue
Base URLhttp://127.0.0.1:8443/v1
API keyAny non-empty string. The local API does not check it.

Because it has no authentication, anything that can reach that port on the machine can use it. It speaks the OpenAI format only. saylek call sends a one-line test through it and prints [ok] <model> responded.

Use it instead of your runtime's own address for two things. A request your runtime cannot serve can go to another machine you can use. And your own work goes first: while work sent through the local API needs the runtime, Saylek stops passing hosted requests to it, your own included, and those wait or are stopped part-way. Saylek does not unload the runtime's models or free its memory. Work sent to the runtime's own address bypasses Saylek.

Reference

Last checked 2026-09-24 · read as markdown at /docs/use-your-existing-runtime.md