# Sharing your models

Sharing a machine is how its models become reachable by someone else. It is always an
explicit choice: nothing you do as a consumer ever turns your machine into one that serves
other people.

## What sharing commits you to

When you share a machine, it serves requests from the people you shared it with while
sharing is switched on, and, if you joined a group, from that group's members. Sharing yours does not share theirs with you: each direction is a separate
decision by the person whose machine it is.

Two things are worth being clear about before you start:

- **Your own work takes priority on your own GPU.** Work you send from the machine itself
  to its local Saylek API goes first. If it arrives for a model that is serving someone you
  share with, their request is stopped before it finishes, rather than immediately.
- **You will see what you serve.** Serving a request means your machine handles that
  person's prompt in the clear, because that is how it computes the answer. The other side
  of this is documented for consumers in [Privacy and
  egress](/docs/privacy-and-egress); sharing puts you on the receiving end of that trust.

The full version of what sharing a machine commits you to, including what serving does *not*
expose and who is answerable for what, is the [Acceptable Use Policy](/aup). Read it before
you run the wizard below rather than after: running it is the act that agrees to it.

## Requirements

| | |
|---|---|
| A model to serve | A GPU the daemon can serve from (NVIDIA via `cuda`, Apple Silicon via `metal`), or a model runtime already running on the machine: Ollama, LM Studio, llama.cpp or vLLM. |
| Binary | A build with an inference backend compiled in. Released binaries have one. |
| An account | You must be signed in with an active Saylek account. |

A from-source build with no `--features` flag compiles but cannot serve: the first call
returns HTTP 503 `no inference backend compiled in`. Build with `cpu`, `cuda`, or `metal`.

## Start sharing

Get a model serving first, if you have not already. If a model runtime already runs on the
machine, connect it:

```bash
saylek model add http://127.0.0.1:11434   # your runtime's address; this is Ollama's
```

Otherwise let the wizard find your GPUs and help you pick models. It also offers to connect a
runtime it finds running:

```bash
saylek gpu wizard
```

Then turn sharing on:

```bash
saylek host wizard
```

This wizard is the step that turns sharing on. It signs the machine in for sharing and keeps
Saylek running in the background, so the machine keeps serving after you close the terminal.
`saylek gpu wizard` above only gets a model serving for you; it does not share anything.

The wizard lists the models it will offer from the ones Saylek serves itself, so on a
machine that serves through a connected runtime it may find none. That is expected: until
you pick models yourself, the machine offers everything it serves, the runtime's included.
If it lists some models but not your runtime's, run `saylek host models --clear`.

Once the machine is set up, `saylek host` serves in the foreground if you want to watch it.

Confirm what your machine is offering:

```bash
saylek models     # what this machine can serve right now
saylek status     # health, and whether you are connected
```

## Stopping: three different scopes

These are easy to confuse, and one of them is much larger than people expect.

| You want to | Command | What keeps working |
|---|---|---|
| Stop **serving other people**, keep using Saylek yourself | `saylek host stop` | Your own use of the machine, including through your own keys, and who you share with |
| Stop **every Saylek request** to the machine, yours included, and stay connected | `saylek host pause` | Local use on the machine |
| Stop the **daemon** entirely | `saylek stop` | Nothing. `saylek call` will not answer until `saylek start` |
| Step away from Saylek but keep your account | `saylek settings account deactivate` | Reversible with `activate` |

`saylek host stop` is the small, targeted one, and it is what turning off **Share from this
machine** on a machine's page does. `saylek host pause` is wider than its name suggests: it
stops the requests your own keys send to that machine too.

There is no bare `saylek pause`. "Pause" means two different things at two different scopes,
so the CLI does not guess which one you meant: it answers with both commands and lets you
pick. Use `saylek host stop` when you mean serving other people, `saylek host pause` when you
mean every hosted request to the machine, and `saylek stop` when you mean the daemon.

## Checking what your machine served

```bash
saylek statement <YYYY-MM>
```

A record of where each request that went through this machine ran, as counts. It records no
money.

## Next steps

- [Sharing and access](/docs/sharing-and-access): who your machine will serve.
- [Troubleshooting](/docs/troubleshooting): when your machine shows as offline.
- [CLI reference](/docs/cli-reference): the full everyday command table.
