Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Using OpenCode with a Fuzzball-Hosted Model

OpenCode is a coding agent that talks to any OpenAI-compatible API. A model served on Fuzzball publishes that kind of API, so OpenCode can use a cluster GPU while it edits your files.

On an endpoint whose scope is not public, the proxy tells the model service who you are, and the LiteLLM in front of the model admits you on that alone – see Identifying the caller. No LiteLLM key is involved, except in two cases.

You can run OpenCode in two places:

  • On Fuzzball, with the OpenCode catalog entry. The workflow reaches the model with your identity, so you hand it no credential at all. Start here.
  • On your workstation, with a local install. You paste the URL and a Fuzzball endpoint token into OpenCode’s config.

Before you start

You need a running model. Either:

  • A vllm catalog entry (or one of its presets such as GPT-OSS 20B) with its default Proxy=true. OpenCode asks for the model by its repository name: the entry’s Model value without hf:// or anything after ?, for example openai/gpt-oss-20b.
  • Several models behind the LiteLLM Model Gateway entry. See Serving an AI Model.

Get the endpoint URL

$ fuzzball workflow endpoints list

Copy the URL and endpoint ID from the row for the model workflow’s litellm service, or the gateway workflow’s gateway service.

Run OpenCode on Fuzzball

$ fuzzball workflow catalog start OpenCode \
    --values Endpoint=https://ENDPOINT-URL,ServiceScope=public,MaxContextSize=MODEL-CONTEXT-SIZE

Then attach from your workstation. Take the server URL from fuzzball workflow endpoints list and the password from the show-server job:

$ fuzzball workflow log WORKFLOW_ID show-server
$ opencode attach https://OPENCODE-URL -p PASSWORD

What the values do:

  • Endpoint is optional. Left empty, the workflow registers every gateway endpoint its own identity can reach. Set it to pin one endpoint, or to reach an API outside Fuzzball.
  • ServiceScope=public is what makes opencode attach and the web app work. At any other scope the Fuzzball endpoint proxy consumes the Authorization header they send, and only API clients that append ?auth_token=BASE64(opencode:PASSWORD) to the URL get through. Public means the password is the only barrier, and anyone who has it can run commands and edit files in the container.
  • MaxContextSize must match the model. The entry defaults to 32768, while the vllm presets serve 131072 to 1048576. Use the model workflow’s own MaxContextSize value.
  • For a model endpoint that is not public, the workflow mints its own Fuzzball token at startup and the model admits it as you, the workflow’s owner. It lasts seven days at most. Restart the workflow when it lapses.
  • It registers every model the endpoint lists at startup. Model=openai/gpt-oss-20b picks the default. A model the endpoint gains later needs a workflow restart.

Two things to know:

  • The workspace starts empty and is lost when the workflow stops. The agent works in /data/workspace on the workflow’s volume, which defaults to ephemeral. Name a persistent volume with Volume= to keep files and sessions. The image is Alpine plus the OpenCode binary, with no git or compilers. Override Image if the agent needs a toolchain.
  • The server password is readable by others. It is in the workflow definition, so anyone who can read the workflow can read it, and anyone who has it can drive the agent.

Run OpenCode on your workstation

Install OpenCode, for example brew install sst/tap/opencode.

Trust the cluster certificate

If the cluster’s API uses a private CA, every client must trust it or you get:

SSL certificate problem: unable to get local issuer certificate

Get the CA in PEM form from your administrator and point OpenCode at it:

$ export NODE_EXTRA_CA_CERTS=/path/to/cluster-ca.pem

Get a Fuzzball token

A model endpoint defaults to user scope, which needs a Fuzzball credential. Mint one for the endpoint:

$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 8h

This token is the only credential you send. A public endpoint needs no token, and Fuzzball does no authentication on it. Only use public with a strong key.

Check with curl

This separates connection problems from configuration problems. The token goes in Authorization:

$ curl -H "Authorization: Bearer YOUR-FUZZBALL-TOKEN" \
    https://ENDPOINT-URL/v1/models

A healthy response lists the model:

{"data":[{"id":"openai/gpt-oss-20b","object":"model","created":1677610602,"owned_by":"openai"}],"object":"list"}

Configure OpenCode

Export the token, then define Fuzzball as an OpenAI-compatible provider in ~/.config/opencode/opencode.json or a project’s opencode.json:

$ export FUZZBALL_TOKEN=YOUR-FUZZBALL-TOKEN
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "fuzzball": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Fuzzball",
      "options": {
        "baseURL": "https://ENDPOINT-URL/v1",
        "apiKey": "{env:FUZZBALL_TOKEN}"
      },
      "models": {
        "openai/gpt-oss-20b": {
          "name": "gpt-oss-20b on Fuzzball",
          "limit": {
            "context": 32768,
            "output": 8192
          }
        }
      }
    }
  },
  "model": "fuzzball/openai/gpt-oss-20b"
}

{env:NAME} and {file:/path} keep credentials out of the file. Set limit.context to the entry’s MaxContextSize; a larger value leads to requests the model rejects.

Use it

$ opencode models fuzzball
fuzzball/openai/gpt-oss-20b

$ opencode run --model fuzzball/openai/gpt-oss-20b "Reply with the single word: ready"
ready

Agentic work needs tool calling. Start the vllm entry with the right flags in ExtraArgs, for example --tool-call-parser openai --enable-auto-tool-choice for gpt-oss models. Without them OpenCode can chat but cannot read or edit files.

Small models sometimes invent tool names, such as apply_patch. OpenCode reports the error and usually recovers on the next step.

When the Fuzzball token expires, OpenCode fails to authenticate. Mint and export a new one.

When you still need a LiteLLM key

Two cases. A public endpoint authenticates nobody, so nothing is sent about the caller. And on a cluster whose nodes have not picked up the node extension that publishes the signing keys, LiteLLM cannot check what is sent and falls back to its own key check.

Send a key only in those two cases. A key you send decides the request, so a wrong or stale one is refused even with a valid Fuzzball token.

Get the key. If you started the model workflow with your own, reuse it: the ApiKey or ApiKeySecret you gave the vllm entry, or the gateway’s MasterKeySecret. Otherwise the entry generated one, and it is in the workflow definition:

$ fuzzball workflow get WORKFLOW_ID | grep -oE 'sk-[A-Za-z0-9_-]+' | head -1

A generated key is visible to anyone who can read the workflow, and it changes every time the workflow starts. On the gateway it is the master key, and it alone reaches management routes such as /key/generate, so mint OpenCode a virtual key with it rather than handing it over.

On your workstation, on a public endpoint, put the key in options.apiKey. Otherwise keep the Fuzzball token there and add the key, exported as LITELLM_KEY, in a header the proxy leaves alone, inside options:

"headers": {
  "x-litellm-api-key": "{env:LITELLM_KEY}"
}

On Fuzzball, store the key as a user-scoped secret of type value and name it in ApiKeySecret. If the model endpoint is public, add EndpointAuth=api-key as well:

$ printf 'sk-...' | fuzzball secret create secret://user/litellm-key --type value
$ fuzzball workflow catalog start OpenCode \
    --values Endpoint=https://ENDPOINT-URL,ApiKeySecret=secret://user/litellm-key

Fuzzball rejects group and organization secrets for environment variables, so each user stores their own copy. See Secrets in Workflows.

Troubleshooting

SymptomCause and fix
unable to get local issuer certificateThe cluster certificate is not trusted. Set NODE_EXTRA_CA_CERTS to the cluster CA.
401 with a plain Unauthorized bodyThe Fuzzball endpoint proxy rejected the request. Authorization did not carry a valid Fuzzball token.
401 with Authentication Error, No api key passed in.Fuzzball let you through, but LiteLLM was told nothing about you and wanted a key: the endpoint is public, or the cluster’s nodes have not picked up the signing keys. See When you still need a LiteLLM key.
401 with Invalid proxy server tokenThe LiteLLM key you sent is wrong or stale, and a key you send decides the request. On an endpoint that is not public, drop the key and rely on your Fuzzball token.
ApiKeySecret must be a user-scoped secret reference at startApiKeySecret names a group or organization secret, or is malformed. Use secret://user/NAME.
400 Invalid model name, and /v1/models returns {"data":[]}No replica is ready yet. Retry once /v1/models lists the model.
400 Invalid model name while /v1/models lists a modelUse exactly the id from /v1/models, such as openai/gpt-oss-20b, in both the models map and the model default.
A reply arrives with no contentA reasoning model spent its output budget on reasoning. Raise limit.output.
The model is missing from opencode modelsThe identifier is <provider key>/<model key>, both from your config file.