# Start without a runtime

No model runtime on your machine? Saylek can download one starter model, run it on your GPU,
and answer your applications with a Saylek API key.

## Before you begin

| You need | Detail |
|---|---|
| A machine that can serve | Linux x86_64 with an NVIDIA GPU and NVIDIA driver 555.42.02 or newer (CUDA 12.5); Windows with WSL2 (x86_64) and an NVIDIA GPU, with NVIDIA driver 555.85 or newer installed in Windows; or a Mac with Apple Silicon. A machine without one of these cannot run the starter model. It can still serve a runtime you already run ([Use the models you already serve](/docs/use-your-existing-runtime)) and use models shared with you. |
| A terminal on that machine | The download asks for your consent there. The browser cannot start it. On Windows, use your WSL2 terminal, not PowerShell. |
| A Saylek account | Create one at [saylek.com](https://saylek.com) with your email. |

## 1. Connect the machine

```bash
curl -fsSL https://saylek.com/install.sh | sh -s -- --connect
```

This signs you in through your browser and links the machine to your account. It does not
download a model. With none on the machine, the installer ends with `no model on this
machine yet` and points you to `saylek gpu wizard`.

## 2. Download the starter model

On the machine:

```bash
saylek gpu wizard
```

After its setup steps, and only when nothing on the machine can serve a model yet, the
wizard offers one starter model and asks `download it now? [y/N]`. Answer `y`. It picks the largest Qwen3.5 model that fits
your largest GPU, or a Mac's memory, from 0.8B (about 0.5 GB) to 27B (about 16.7 GB). The
file comes from huggingface.co and is checked against a pinned checksum. Saylek downloads
nothing without this answer.

In [Machines](/machines), **Add model** on the machine's page shows the same command under
**Starting without a model?**.

**You know it worked when** the wizard prints `[ok] downloaded + registered`.

## 3. Check that the model is ready

In [Machines](/machines), open the machine. The model shows as **Ready**, labelled
**Saylek-managed**. From the terminal, `saylek models` lists it.

The first request loads the model, so it takes longer than the rest. On Linux, the first
load also downloads Saylek's inference runtime once and verifies its signature.

## 4. Connect an application and send a request

In [Applications](/applications), choose **Add application**, name it, and choose **Create
key**. Copy the key now: it is shown once. Then send a test request, using the model id from
the models list:

```bash
curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"

curl https://api.saylek.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "Say hi in one sentence."}]}'
```

**You know it worked when** the response carries an answer. To set up your own client,
follow [Connect your app](/docs/connect-your-app). Hosted requests pass through Saylek's
servers; [Privacy and egress](/docs/privacy-and-egress) says what that means for your prompt.

## Add another model later

Saylek does not download models on its own. To serve another, register a model file you have:

```bash
saylek model add /absolute/path/to/model.gguf
```

Or copy a `.gguf` file into `~/.saylek/models/` and run `saylek model rescan`. At most
two models stay loaded at once; [Make more of one GPU](/docs/make-more-of-one-gpu) explains.

## If something goes wrong

| What you see | What to do |
|---|---|
| The wizard does not offer a download | Run it in an interactive terminal. If you did, something on the machine can already serve: check with `saylek models`, or follow [Use the models you already serve](/docs/use-your-existing-runtime) |
| On Linux without an NVIDIA GPU, the wizard says the machine will serve on its CPU, or offers a model to run on the CPU | Saylek cannot run a model on a Linux CPU. Answer `N`; connect a runtime you already run, or use models shared with you |
| `download did not complete` | Run `saylek gpu wizard` again. `saylek help models` shows how to fetch a model yourself |
| `checksum mismatch` or `size mismatch` | The partial file was removed. Run `saylek gpu wizard` again |
| `NVIDIA driver not detected` | On a machine with an NVIDIA GPU, install its driver (under WSL2, in Windows), then retry. The message may start with "no free model slot" |
| `no published inference runtime for GPU arch` | Either the driver supports too old a CUDA version, or Saylek has no runtime for this GPU model yet. Update the driver and retry; if the message stays, this GPU cannot serve yet |
| `the inference runtime needs glibc` | The machine needs a newer Linux release: Ubuntu 24.04+, Debian 13+ or Fedora 40+ |
| `could not fetch the inference runtime` | A network problem. Retry |
| `model_load_failed` | The model did not load, often because it does not fit. `saylek call` on the machine, or `~/.saylek/log/`, shows the reason |

[Troubleshooting](/docs/troubleshooting) covers the rest.

## Reference

- [Quickstart](/docs/quickstart): install and first steps.
- [Connect your app](/docs/connect-your-app): client setup.
- [CLI reference](/docs/cli-reference): every command.
