Skip to content

Claude Code & the Anthropic SDK

Connect Claude Code or an Anthropic Messages client with your Saylek API key and a suitable model. The connection uses a different base URL from an OpenAI-compatible client.

Available models come from your own machines, from people who share with you, and, if you joined a group, from machines in that group. A model appearing in your list is not a reservation: a machine still has to be free to serve the request. Nor is being listed a capability claim: it does not mean the model supports every Messages feature or the tool calling an agent needs.

This guide uses the hosted Saylek API, the recommended RC integration path. If your client uses OpenAI-compatible chat completions instead, follow OpenAI setup.

The endpoint

Anthropic clients own the version segment: ANTHROPIC_BASE_URL is the origin, and the client appends /v1/messages itself. So the base URL you set must not carry a /v1 suffix:

https://api.saylek.com

(The OpenAI-compatible surface, by contrast, uses https://api.saylek.com/v1: the extra /v1 is deliberate there and absent here.)

Claude Code

If you already have Claude Code installed, one command points it at the models you can reach:

saylek claude

It starts the claude you already have, pointed at the models you can reach, with a key minted from your account and a model available to you. Your own Claude sign-in is left alone, so a plain claude keeps using it.

The key is minted the first time you run it, named for this machine (saylek-claude-cli-<host>-<timestamp>) so you can tell your machines apart when you revoke one. It is saved to ~/.secret/saylek with owner-only permissions, so you are not asked again after a reboot. Existing keys are never renamed or replaced.

Anything you type after claude is passed straight through, so saylek claude -p "..." and saylek claude --resume work as usual. For Claude Code's own help rather than this command's, use saylek claude -- --help.

It also sets Claude Code up for long prompts, the way Long prompts below describes, so a prompt that takes minutes to read still gets its answer.

Two things have to be true or nothing launches, and it tells you which one is missing: Claude Code is installed, and a model is available to you. It does not install Claude Code for you.

To point these at a pre-release environment, use the same base-URL overrides every other Saylek command takes:

export SAYLEK_REGISTRY_BASE=https://stage1-api.saylek.com   # the data plane
export SAYLEK_WORKER_BASE=https://stage1.saylek.com         # your account
export SAYLEK_IDENTITY_BASE=https://stage1-id.saylek.com    # the people who share with you

Set all three together. They are separate services, and setting one leaves the others on production, which mints a key on one environment and uses it on another.

Setting it up by hand

export ANTHROPIC_BASE_URL="https://api.saylek.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_API_KEY"
unset ANTHROPIC_API_KEY

# One id, every model slot Claude Code reaches for.
MODEL="MODEL_ID"
export ANTHROPIC_MODEL="$MODEL"
export ANTHROPIC_SMALL_FAST_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_OPUS_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_SONNET_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="$MODEL"
export CLAUDE_CODE_SUBAGENT_MODEL="$MODEL"

export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
export API_TIMEOUT_MS=600000              # self-hosted models answer slower than the stock timeout allows
export CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1   # see Long prompts below
export CLAUDE_STREAM_IDLE_TIMEOUT_MS=900000          # see Long prompts below

claude

saylek claude sets the connection variables and preserves your exported context limit.

Find an available MODEL_ID:

curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"

For Claude Code, choose a model that supports tool calling; a successful plain-text answer alone does not establish that it can complete an agent task.

Set every slot, not just ANTHROPIC_MODEL. An unset slot can select a default claude-* name that is not in your Saylek model list, causing a later call to fail.

saylek claude does not choose a context window. If you already have a verified explicit CLAUDE_CODE_MAX_CONTEXT_TOKENS override, it is preserved; otherwise leave it unset.

Every line above needs export. A bare NAME=value sets a shell variable that claude never sees.

Long prompts

A long prompt can take minutes to read before the first word comes back, and until then only a keep-alive crosses the connection. The last two lines of the block keep that working:

  • CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1 stops Claude Code from answering a failed stream by sending the request again without streaming. That second request sends nothing back until the whole answer exists, so the network between you and Saylek can close it as idle, and you see a timeout for work that was still running.
  • CLAUDE_STREAM_IDLE_TIMEOUT_MS=900000 is how long Claude Code waits for the next streamed event before it gives up; unset, it waits 5 minutes. 900000 is 15 minutes, the longest Saylek holds a streamed request, so Saylek always answers first.

Claude Code will also warn that your claude.ai connectors are disabled. That is expected: an auth token pointed at Saylek takes precedence over a claude.ai login, so connectors managed there do not load. Nothing else is affected, and your own MCP servers are untouched. Turn the notice off in .claude/settings.json:

{
  "disableClaudeAiConnectors": true
}

This stops Claude Code attempting a load that cannot succeed while you are authenticated here. To get those connectors back, remove the setting and unset ANTHROPIC_AUTH_TOKEN, which means using a claude.ai login for that session instead of Saylek.

Claude Code's built-in WebSearch and WebFetch are executed by the model provider, not by whoever serves the request, and no search engine sits behind Saylek to honor them. So Saylek refuses the request with a 400 naming the tool, rather than serving it without the search and letting the model answer as though it had one. Replace them in two steps.

  1. Add a search server to .mcp.json at the root of your project:
{
  "mcpServers": {
    "search": {
      "command": "npx",
      "args": ["-y", "mcp-searxng"],
      "env": { "SEARXNG_URL": "http://YOUR_SEARXNG_HOST:8080" }
    }
  }
}
  1. In .claude/settings.json, deny the built-ins and grant the MCP tools:
{
  "permissions": {
    "deny": ["WebSearch", "WebFetch"],
    "allow": ["mcp__search__searxng_web_search", "mcp__search__web_url_read"]
  }
}

Then run claude once in that directory and accept the workspace-trust prompt. Until you do, Claude Code honors the deny but ignores the allow, which is the worst of both: the built-ins are off and your search tool is ungranted, so the model calls it every turn and is refused every turn. A scripted claude -p in a fresh checkout hits this, because there is no prompt to accept.

Both steps are needed. Leave the built-ins enabled and the model picks them every time and never calls your MCP tool.

Both must also land in the config that session actually loads, which is the trap that looks exactly like the recipe not working. Claude Code reads settings and MCP registrations once, at startup, from the home directory the session was launched with. A session already running will not see either change, and a session launched with a different HOME reads a different config entirely and behaves as if you configured nothing. Restart, then confirm with /mcp that search reads connected; if it does not, you edited a file that session is not reading.

The allow half matters as much as the deny. Without it the model calls your search tool and the request is refused for want of permission. Interactively you get a prompt; a scripted claude -p run just fails. The names are the server's own tools prefixed with mcp__<server>__, so they follow whichever server you picked: mcp-searxng exposes searxng_web_search and web_url_read, not web_search. Check yours with /mcp in a session, or claude mcp list.

Confirm the server is live before blaming the search:

claude mcp list

✔ Connected means the MCP server started, not that the search backend answered. For that, query the backend directly. A 403 rather than a 200 is why results come back empty.

Any MCP search server works. SearXNG can be self-hosted, which keeps your queries off a third-party search API; if you run one, its JSON API and bot limiter both need attention, since the defaults refuse programmatic callers.

The Anthropic SDK

The Python and TypeScript SDKs take a base URL and an auth token directly:

from anthropic import Anthropic

client = Anthropic(base_url="https://api.saylek.com", auth_token="YOUR_API_KEY")
resp = client.messages.create(
    model="MODEL_ID",
    max_tokens=1024,
    messages=[{"role": "user", "content": "How are you doing?"}],
)
print(resp.content[0].text)
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://api.saylek.com",
  authToken: "YOUR_API_KEY",
});
const resp = await client.messages.create({
  model: "MODEL_ID",
  max_tokens: 1024,
  messages: [{ role: "user", content: "How are you doing?" }],
});
const [block] = resp.content;
if (block.type === "text") console.log(block.text);

curl

The raw Messages API endpoint accepts either x-api-key or Authorization: Bearer, plus the required anthropic-version header:

curl https://api.saylek.com/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "How are you doing?"}]
  }'

When it can answer

An eligible machine must be online, able to serve the requested model, and have capacity. If a call fails because no capacity is available, check your model list and try again when capacity returns. Adding a key does not add a machine or reserve capacity.

Hosted requests pass through Saylek and are handled by the serving machine, whether it is your own or shared with you. Read Privacy and egress before sending sensitive content.

A key's secret is shown once, at creation, so paste it in place of YOUR_API_KEY above and keep your copy safe. To be exact about where it travels: your key is sent to Saylek on every call, because that is what authenticates you. It is not passed on to the machine that serves the request. If you lose a key or a key leaks, revoke it at /applications and create another.

Next steps

Last checked 2026-09-24 · read as markdown at /docs/connect-anthropic.md