Sharing your models
Sharing a machine is how its models become reachable by someone else. It is always an explicit choice: nothing you do as a consumer ever turns your machine into one that serves other people.
What sharing commits you to
When you share a machine, it serves requests from the people you shared it with while sharing is switched on, and, if you joined a group, from that group's members. Sharing yours does not share theirs with you: each direction is a separate decision by the person whose machine it is.
Two things are worth being clear about before you start:
- Your own work takes priority on your own GPU. Work you send from the machine itself to its local Saylek API goes first. If it arrives for a model that is serving someone you share with, their request is stopped before it finishes, rather than immediately.
- You will see what you serve. Serving a request means your machine handles that person's prompt in the clear, because that is how it computes the answer. The other side of this is documented for consumers in Privacy and egress; sharing puts you on the receiving end of that trust.
The full version of what sharing a machine commits you to, including what serving does not expose and who is answerable for what, is the Acceptable Use Policy. Read it before you run the wizard below rather than after: running it is the act that agrees to it.
Requirements
| A model to serve | A GPU the daemon can serve from (NVIDIA via cuda, Apple Silicon via metal), or a model runtime already running on the machine: Ollama, LM Studio, llama.cpp or vLLM. |
| Binary | A build with an inference backend compiled in. Released binaries have one. |
| An account | You must be signed in with an active Saylek account. |
A from-source build with no --features flag compiles but cannot serve: the first call
returns HTTP 503 no inference backend compiled in. Build with cpu, cuda, or metal.
Start sharing
Get a model serving first, if you have not already. If a model runtime already runs on the machine, connect it:
saylek model add http://127.0.0.1:11434 # your runtime's address; this is Ollama's
Otherwise let the wizard find your GPUs and help you pick models. It also offers to connect a runtime it finds running:
saylek gpu wizard
Then turn sharing on:
saylek host wizard
This wizard is the step that turns sharing on. It signs the machine in for sharing and keeps
Saylek running in the background, so the machine keeps serving after you close the terminal.
saylek gpu wizard above only gets a model serving for you; it does not share anything.
The wizard lists the models it will offer from the ones Saylek serves itself, so on a
machine that serves through a connected runtime it may find none. That is expected: until
you pick models yourself, the machine offers everything it serves, the runtime's included.
If it lists some models but not your runtime's, run saylek host models --clear.
Once the machine is set up, saylek host serves in the foreground if you want to watch it.
Confirm what your machine is offering:
saylek models # what this machine can serve right now
saylek status # health, and whether you are connected
Stopping: three different scopes
These are easy to confuse, and one of them is much larger than people expect.
| You want to | Command | What keeps working |
|---|---|---|
| Stop serving other people, keep using Saylek yourself | saylek host stop | Your own use of the machine, including through your own keys, and who you share with |
| Stop every Saylek request to the machine, yours included, and stay connected | saylek host pause | Local use on the machine |
| Stop the daemon entirely | saylek stop | Nothing. saylek call will not answer until saylek start |
| Step away from Saylek but keep your account | saylek settings account deactivate | Reversible with activate |
saylek host stop is the small, targeted one, and it is what turning off Share from this
machine on a machine's page does. saylek host pause is wider than its name suggests: it
stops the requests your own keys send to that machine too.
There is no bare saylek pause. "Pause" means two different things at two different scopes,
so the CLI does not guess which one you meant: it answers with both commands and lets you
pick. Use saylek host stop when you mean serving other people, saylek host pause when you
mean every hosted request to the machine, and saylek stop when you mean the daemon.
Checking what your machine served
saylek statement <YYYY-MM>
A record of where each request that went through this machine ran, as counts. It records no money.
Next steps
- Sharing and access: who your machine will serve.
- Troubleshooting: when your machine shows as offline.
- CLI reference: the full everyday command table.
Last checked 2026-09-29 · read as markdown at /docs/share-your-models.md