Skip to content

Make more of one GPU

Serve several models from one GPU, keep your own work first, and, if you share the machine, choose how much of it other people's requests can use.

Who runs the model decides what Saylek can do

A model on your machine is run either by Saylek itself or by your own runtime (Ollama, LM Studio, llama.cpp or vLLM). The machine's page in Machines labels each model Saylek-managed or with its runtime's name.

What differsSaylek runs the modelYour runtime runs the model
LoadingOn the first request for itYour runtime decides
Models loaded at onceAt most 2Your runtime decides
UnloadingAfter 5 minutes idle, by defaultYour runtime decides
Checking that a model fitsNot checked when a model loads (the wizard only estimates): a model that does not fit fails to loadYour runtime decides
Requests at onceOne per modelOne shared request at a time, unless the runtime reports more room

1. Add the models you want

For models Saylek runs, register each model file:

saylek model add /absolute/path/to/model.gguf

Or copy .gguf files into ~/.saylek/models/ and run saylek model rescan. For your own runtime, add the models in the runtime; if they do not appear, choose Scan again under the machine's Settings, Manage sources.

You know it worked when saylek models, or the machine's page, lists every model.

2. Know how models load and unload

For models Saylek runs, a request loads its model if it is not loaded yet, so that first request takes longer. At most two models stay loaded. When a third is requested, the loaded model used least recently gives way if it is idle; if both are in use, the request gets all model slots busy. A model unloads after 5 minutes without requests.

To change that, set keep_warm_secs under [daemon] in ~/.saylek/config.toml, in seconds: higher keeps models ready longer, 0 frees memory as soon as the last request ends. Saylek reads it when it starts. For your own runtime, use its settings instead.

3. Keep your own work first

  • Work sent to Saylek's local API goes first. Requests that applications on this machine send to http://127.0.0.1:8443/v1 (set up in Use the models you already serve) are taken before waiting shared requests, and stop a shared request that is already running at its next step. That request goes to another machine if one is free; otherwise it waits, or ends with an error the application can retry. A request that has already started streaming ends with owner_preempted. With your own runtime, Saylek stops passing it shared requests; the runtime may finish one it already has.
  • Requests through the hosted API do not stop running work, even your own requests with your own key on your own machine. When requests are waiting for your machine, Saylek's servers offer your own first.
  • Application priority stops nothing. It orders one account's own requests; see Give important applications priority.
  • Work sent straight to your runtime's own address bypasses Saylek, so Saylek neither counts nor stops it.

Optional: limit what people you share with can use

These apply only when sharing is on. host models and host slots are terminal only.

ToIn the webFrom the terminal
Offer only some modelsTerminal onlysaylek host models MODEL_ID; --clear offers all again
Take fewer shared requests at onceTerminal onlysaylek host slots 1; --clear resets it
Share only at set hoursEdit schedule on the machine's pagesaylek host hours --set 09:00-17:00
Stop shared requests for nowTurn off Share from this machine on the machine's pagesaylek host stop, and saylek host start to share again

Hours use the machine's own clock. A schedule has at most 5 periods, each a set of days with one time range, or all day. A slot number above what the machine can take is capped, and saylek host slots says so.

If something goes wrong

What you seeWhat to do
all model slots busyTwo models Saylek runs are both in use. Wait, or send fewer models at once
model_load_failedThe model may not fit next to what is loaded. Lower keep_warm_secs (step 2), or use a smaller model
The wizard says a model may not fitIt is an estimate. Skip it, choose another GPU, or try anyway
Shared requests still arrive at hours you setHours use the machine's clock, not yours. Check with saylek host hours

Reference

Last checked 2026-09-29 · read as markdown at /docs/make-more-of-one-gpu.md