Start without a runtime
No model runtime on your machine? Saylek can download one starter model, run it on your GPU, and answer your applications with a Saylek API key.
Before you begin
| You need | Detail |
|---|---|
| A machine that can serve | Linux x86_64 with an NVIDIA GPU and NVIDIA driver 555.42.02 or newer (CUDA 12.5); Windows with WSL2 (x86_64) and an NVIDIA GPU, with NVIDIA driver 555.85 or newer installed in Windows; or a Mac with Apple Silicon. A machine without one of these cannot run the starter model. It can still serve a runtime you already run (Use the models you already serve) and use models shared with you. |
| A terminal on that machine | The download asks for your consent there. The browser cannot start it. On Windows, use your WSL2 terminal, not PowerShell. |
| A Saylek account | Create one at saylek.com with your email. |
1. Connect the machine
curl -fsSL https://saylek.com/install.sh | sh -s -- --connect
This signs you in through your browser and links the machine to your account. It does not
download a model. With none on the machine, the installer ends with no model on this machine yet and points you to saylek gpu wizard.
2. Download the starter model
On the machine:
saylek gpu wizard
After its setup steps, and only when nothing on the machine can serve a model yet, the
wizard offers one starter model and asks download it now? [y/N]. Answer y. It picks the largest Qwen3.5 model that fits
your largest GPU, or a Mac's memory, from 0.8B (about 0.5 GB) to 27B (about 16.7 GB). The
file comes from huggingface.co and is checked against a pinned checksum. Saylek downloads
nothing without this answer.
In Machines, Add model on the machine's page shows the same command under Starting without a model?.
You know it worked when the wizard prints [ok] downloaded + registered.
3. Check that the model is ready
In Machines, open the machine. The model shows as Ready, labelled
Saylek-managed. From the terminal, saylek models lists it.
The first request loads the model, so it takes longer than the rest. On Linux, the first load also downloads Saylek's inference runtime once and verifies its signature.
4. Connect an application and send a request
In Applications, choose Add application, name it, and choose Create key. Copy the key now: it is shown once. Then send a test request, using the model id from the models list:
curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
curl https://api.saylek.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'
You know it worked when the response carries an answer. To set up your own client, follow Connect your app. Hosted requests pass through Saylek's servers; Privacy and egress says what that means for your prompt.
Add another model later
Saylek does not download models on its own. To serve another, register a model file you have:
saylek model add /absolute/path/to/model.gguf
Or copy a .gguf file into ~/.saylek/models/ and run saylek model rescan. At most
two models stay loaded at once; Make more of one GPU explains.
If something goes wrong
| What you see | What to do |
|---|---|
| The wizard does not offer a download | Run it in an interactive terminal. If you did, something on the machine can already serve: check with saylek models, or follow Use the models you already serve |
| On Linux without an NVIDIA GPU, the wizard says the machine will serve on its CPU, or offers a model to run on the CPU | Saylek cannot run a model on a Linux CPU. Answer N; connect a runtime you already run, or use models shared with you |
download did not complete | Run saylek gpu wizard again. saylek help models shows how to fetch a model yourself |
checksum mismatch or size mismatch | The partial file was removed. Run saylek gpu wizard again |
NVIDIA driver not detected | On a machine with an NVIDIA GPU, install its driver (under WSL2, in Windows), then retry. The message may start with "no free model slot" |
no published inference runtime for GPU arch | Either the driver supports too old a CUDA version, or Saylek has no runtime for this GPU model yet. Update the driver and retry; if the message stays, this GPU cannot serve yet |
the inference runtime needs glibc | The machine needs a newer Linux release: Ubuntu 24.04+, Debian 13+ or Fedora 40+ |
could not fetch the inference runtime | A network problem. Retry |
model_load_failed | The model did not load, often because it does not fit. saylek call on the machine, or ~/.saylek/log/, shows the reason |
Troubleshooting covers the rest.
Reference
- Quickstart: install and first steps.
- Connect your app: client setup.
- CLI reference: every command.
Last checked 2026-09-25 · read as markdown at /docs/start-without-a-runtime.md