Claude Code & the Anthropic SDK
Connect Claude Code or an Anthropic Messages client with your Saylek API key and a suitable model. The connection uses a different base URL from an OpenAI-compatible client.
Available models come from your own machines, from people who share with you, and, if you joined a group, from machines in that group. A model appearing in your list is not a reservation: a machine still has to be free to serve the request. Nor is being listed a capability claim: it does not mean the model supports every Messages feature or the tool calling an agent needs.
This guide uses the hosted Saylek API, the recommended RC integration path. If your client uses OpenAI-compatible chat completions instead, follow OpenAI setup.
The endpoint
Anthropic clients own the version segment: ANTHROPIC_BASE_URL is the origin, and the
client appends /v1/messages itself. So the base URL you set must not carry a
/v1 suffix:
https://api.saylek.com
(The OpenAI-compatible surface, by contrast, uses https://api.saylek.com/v1: the extra
/v1 is deliberate there and absent here.)
Claude Code
If you already have Claude Code installed, one command points it at the models you can reach:
saylek claude
It starts the claude you already have, pointed at the models you can reach, with a key minted from
your account and a model available to you. Your own Claude sign-in is left
alone, so a plain claude keeps using it.
The key is minted the first time you run it, named for this machine
(saylek-claude-cli-<host>-<timestamp>) so you can tell your machines apart when you
revoke one. It is saved to ~/.secret/saylek with owner-only permissions, so you are
not asked again after a reboot. Existing keys are never renamed or replaced.
Anything you type after claude is passed straight through, so saylek claude -p "..."
and saylek claude --resume work as usual. For Claude Code's own help rather than this
command's, use saylek claude -- --help.
It also sets Claude Code up for long prompts, the way Long prompts below describes, so a prompt that takes minutes to read still gets its answer.
Two things have to be true or nothing launches, and it tells you which one is missing: Claude Code is installed, and a model is available to you. It does not install Claude Code for you.
To point these at a pre-release environment, use the same base-URL overrides every other Saylek command takes:
export SAYLEK_REGISTRY_BASE=https://stage1-api.saylek.com # the data plane
export SAYLEK_WORKER_BASE=https://stage1.saylek.com # your account
export SAYLEK_IDENTITY_BASE=https://stage1-id.saylek.com # the people who share with you
Set all three together. They are separate services, and setting one leaves the others on production, which mints a key on one environment and uses it on another.
Setting it up by hand
export ANTHROPIC_BASE_URL="https://api.saylek.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_API_KEY"
unset ANTHROPIC_API_KEY
# One id, every model slot Claude Code reaches for.
MODEL="MODEL_ID"
export ANTHROPIC_MODEL="$MODEL"
export ANTHROPIC_SMALL_FAST_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_OPUS_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_SONNET_MODEL="$MODEL"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="$MODEL"
export CLAUDE_CODE_SUBAGENT_MODEL="$MODEL"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
export API_TIMEOUT_MS=600000 # self-hosted models answer slower than the stock timeout allows
export CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1 # see Long prompts below
export CLAUDE_STREAM_IDLE_TIMEOUT_MS=900000 # see Long prompts below
claude
saylek claude sets the connection variables and preserves your exported context limit.
Find an available MODEL_ID:
curl https://api.saylek.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
For Claude Code, choose a model that supports tool calling; a successful plain-text answer alone does not establish that it can complete an agent task.
Set every slot, not just ANTHROPIC_MODEL. An unset slot can select a default
claude-* name that is not in your Saylek model list, causing a later call to fail.
saylek claude does not choose a context window. If you already have a verified explicit
CLAUDE_CODE_MAX_CONTEXT_TOKENS override, it is preserved; otherwise leave it unset.
Every line above needs export. A bare NAME=value sets a shell variable that claude
never sees.
Long prompts
A long prompt can take minutes to read before the first word comes back, and until then only a keep-alive crosses the connection. The last two lines of the block keep that working:
CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1stops Claude Code from answering a failed stream by sending the request again without streaming. That second request sends nothing back until the whole answer exists, so the network between you and Saylek can close it as idle, and you see a timeout for work that was still running.CLAUDE_STREAM_IDLE_TIMEOUT_MS=900000is how long Claude Code waits for the next streamed event before it gives up; unset, it waits 5 minutes. 900000 is 15 minutes, the longest Saylek holds a streamed request, so Saylek always answers first.
Claude Code will also warn that your claude.ai connectors are disabled. That is expected:
an auth token pointed at Saylek takes precedence over a claude.ai login, so connectors
managed there do not load. Nothing else is affected, and your own MCP servers are
untouched. Turn the notice off in .claude/settings.json:
{
"disableClaudeAiConnectors": true
}
This stops Claude Code attempting a load that cannot succeed while you are authenticated
here. To get those connectors back, remove the setting and unset
ANTHROPIC_AUTH_TOKEN, which means using a claude.ai login for that session instead of
Saylek.
Web search
Claude Code's built-in WebSearch and WebFetch are executed by the model provider, not
by whoever serves the request, and no search engine sits behind Saylek to honor them. So
Saylek refuses the request with a 400 naming the tool, rather than serving it without
the search and letting the model answer as though it had one. Replace them in two steps.
- Add a search server to
.mcp.jsonat the root of your project:
{
"mcpServers": {
"search": {
"command": "npx",
"args": ["-y", "mcp-searxng"],
"env": { "SEARXNG_URL": "http://YOUR_SEARXNG_HOST:8080" }
}
}
}
- In
.claude/settings.json, deny the built-ins and grant the MCP tools:
{
"permissions": {
"deny": ["WebSearch", "WebFetch"],
"allow": ["mcp__search__searxng_web_search", "mcp__search__web_url_read"]
}
}
Then run claude once in that directory and accept the workspace-trust prompt. Until you
do, Claude Code honors the deny but ignores the allow, which is the worst of both: the
built-ins are off and your search tool is ungranted, so the model calls it every turn and is
refused every turn. A scripted claude -p in a fresh checkout hits this, because there is
no prompt to accept.
Both steps are needed. Leave the built-ins enabled and the model picks them every time and never calls your MCP tool.
Both must also land in the config that session actually loads, which is the trap that
looks exactly like the recipe not working. Claude Code reads settings and MCP registrations
once, at startup, from the home directory the session was launched with. A session already
running will not see either change, and a session launched with a different HOME reads
a different config entirely and behaves as if you configured nothing. Restart, then
confirm with /mcp that search reads connected; if it does not, you edited a file
that session is not reading.
The allow half matters as much as the deny. Without it the model calls your search
tool and the request is refused for want of permission. Interactively you get a prompt;
a scripted claude -p run just fails. The names are the server's own tools prefixed
with mcp__<server>__, so they follow whichever server you picked: mcp-searxng exposes
searxng_web_search and web_url_read, not web_search. Check yours with /mcp in a
session, or claude mcp list.
Confirm the server is live before blaming the search:
claude mcp list
✔ Connected means the MCP server started, not that the search backend answered. For that,
query the backend directly. A 403 rather than a 200 is why results come back empty.
Any MCP search server works. SearXNG can be self-hosted, which keeps your queries off a third-party search API; if you run one, its JSON API and bot limiter both need attention, since the defaults refuse programmatic callers.
The Anthropic SDK
The Python and TypeScript SDKs take a base URL and an auth token directly:
from anthropic import Anthropic
client = Anthropic(base_url="https://api.saylek.com", auth_token="YOUR_API_KEY")
resp = client.messages.create(
model="MODEL_ID",
max_tokens=1024,
messages=[{"role": "user", "content": "How are you doing?"}],
)
print(resp.content[0].text)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://api.saylek.com",
authToken: "YOUR_API_KEY",
});
const resp = await client.messages.create({
model: "MODEL_ID",
max_tokens: 1024,
messages: [{ role: "user", content: "How are you doing?" }],
});
const [block] = resp.content;
if (block.type === "text") console.log(block.text);
curl
The raw Messages API endpoint accepts either x-api-key or Authorization: Bearer, plus the
required anthropic-version header:
curl https://api.saylek.com/v1/messages \
-H "x-api-key: YOUR_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "How are you doing?"}]
}'
When it can answer
An eligible machine must be online, able to serve the requested model, and have capacity. If a call fails because no capacity is available, check your model list and try again when capacity returns. Adding a key does not add a machine or reserve capacity.
Hosted requests pass through Saylek and are handled by the serving machine, whether it is your own or shared with you. Read Privacy and egress before sending sensitive content.
A key's secret is shown once, at creation, so paste it in place of YOUR_API_KEY above and
keep your copy safe. To be exact about where it travels: your key is sent to Saylek on
every call, because that is what authenticates you. It is not passed on to the machine that
serves the request. If you lose a key or a key leaks, revoke it at /applications and
create another.
Next steps
- Use your models from anywhere: the same endpoint, OpenAI dialect.
- Privacy and egress: what the serving machine sees when it serves your agent.
- Troubleshooting: 401s, 404s, and rejected model ids.
Last checked 2026-09-24 · read as markdown at /docs/connect-anthropic.md