Serving an AI Model
Fuzzball serves large language models from the workflow catalog. One command starts a model on
a cluster GPU and gives you a base URL that speaks the OpenAI API, so any client written for
OpenAI can use it: a coding agent, an IDE plugin, a notebook, or plain curl.
You do not need to write a Fuzzfile, choose a container image, or know how vLLM, the server that runs the model on the GPU, is configured. Each model has its own catalog entry, already carrying the settings that model needs.
The only thing you have to match is GPU memory. A model needs enough of it to hold its weights plus the context it is serving:
| Catalog entry | GPU memory to run one copy |
|---|---|
GPT-OSS 20B | 16GB |
Ministral 3 8B | 24GB |
Ministral 3 14B | 48GB |
Gemma 4 31B | 80GB |
GPT-OSS 120B | 80GB |
Nemotron 3 Nano | 80GB |
Qwen3-Coder 30B | 160GB |
Qwen3-Coder Next 80B | 160GB |
Nemotron 3 Super | 640GB |
Mistral Medium 3.5 | 640GB |
Llama 4 Scout | 640GB |
Two numbers move those figures. Each entry assumes its own default context size, which is
large, so lowering MaxContextSize lets a model run on less memory. And the memory does not
have to come from one card: each entry defaults to however many 80GB-class GPUs its total
needs, which you change with GpusPerNode. Qwen3-Coder 30B defaults to two GPUs to reach
160GB, but runs on one card that has the memory, or on one 80GB card with a smaller context.
Those GPUs all come from a single node; if no node in your cluster has enough of them, see
Spanning several nodes.
To see everything your cluster offers:
$ fuzzball workflow catalog list --category ML_AND_AI
$ fuzzball workflow catalog list --tag LLMGPT-OSS 20B is the smallest and the usual starting point. If you are unsure what GPUs your
cluster has, run fuzzball node list or ask your administrator.
To serve a model that has no entry of its own, use the generalvLLMentry and set itsModelparameter to a Hugging Face URI, for example--values Model=hf://openai/gpt-oss-20b. You then choose the GPU count, memory, and context size yourself, which the per-model entries otherwise do for you.
Apart from the gated one covered below, every per-model entry runs with no parameters at all:
$ fuzzball workflow catalog start "GPT-OSS 20B"This prints a workflow ID. Copy it: the commands below need it, and fuzzball workflow list
finds it again later. Add --watch to follow the workflow until it is running:
$ fuzzball workflow catalog start "GPT-OSS 20B" --watchThe first start is slow. The model weights download from the Hugging Face Hub before the server comes up, which is tens of gigabytes for a 20B model. By default they land on an ephemeral volume, temporary storage that is cleared when the workflow stops, so the next start downloads them again. To keep them, point the entry at a persistent volume:--values Volume=my-models.
A gated model additionally needs a Hugging Face token in a Fuzzball secret. Llama 4 Scout is
the one gated entry today:
$ fuzzball secret create secret://user/hf-token --type hf '{"token":"hf_..."}'
$ fuzzball workflow catalog start "Llama 4 Scout" --values HfTokenSecret=secret://user/hf-tokenA running model holds its GPUs until you stop it. Stop it when you are done:fuzzball workflow stop WORKFLOW_ID.
The workflow is ready when its service is running and its health check has passed:
$ fuzzball workflow describe WORKFLOW_IDThen find the base URL:
$ fuzzball workflow endpoints listThe entry generates an API key at start time unless you set one with ApiKey. It is stored in
the workflow definition:
$ fuzzball workflow get WORKFLOW_ID | grep -oE 'sk-[A-Za-z0-9_-]+' | head -1Every endpoint has a Scope controlling who may call it. It defaults to group, so requests
also need a Fuzzball token bound to the endpoint:
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 8hBoth credentials travel on every request, in different headers. The Fuzzball endpoint proxy
consumes the Authorization header to authenticate you, then removes it before forwarding the
request. The API key therefore has to travel in a header the proxy leaves alone. Ask for the
model list to confirm the whole path works:
$ curl -H "Authorization: Bearer YOUR-FUZZBALL-TOKEN" \
-H "x-litellm-api-key: sk-fb-YOUR-API-KEY" \
https://ENDPOINT-URL/v1/modelsA healthy response names the model:
{"data":[{"id":"openai/gpt-oss-20b","object":"model","created":1677610602,"owned_by":"openai"}],"object":"list"}
An endpoint started with Scope=public needs only the API key, and is served without any
Fuzzball authentication, so use it only with a strong ApiKey.
Everything above puts one copy of the model on one node. When no single node has the memory,
Nodes splits a copy across several. Fuzzball starts and stops those nodes together, and the
first one serves the endpoint:
$ fuzzball workflow catalog start "Nemotron 3 Super" --values Nodes=2,GpusPerNode=4A copy then uses Nodes times GpusPerNode GPUs, and no two of its ranks share a node, so the
cluster needs that many nodes free. When they are merely busy the workflow waits for them and
reports a scheduling_blocked event. It is rejected at submission only when the copy asks for
more nodes than the cluster is configured to give it, which is a limit no amount of waiting
changes. You do not need to choose how the model is split across the nodes: ExpertParallelism
defaults to reading the model’s own configuration and picking a layout.
What you may need to set is the network. The nodes talk to each other directly, and on a
cluster with more than one network interface the collective library has to be told which one to
use. That goes in ExtraEnv as space-separated NAME=VALUE pairs:
$ fuzzball workflow catalog start "Nemotron 3 Super" \
--values 'Nodes=2,GpusPerNode=4,ExtraEnv=NCCL_SOCKET_IFNAME=eth0'The interface name is specific to your cluster, so ask your administrator which one carries node-to-node traffic. A fabric with RDMA may need more than one variable.
Multi-node copies need GPUs that can talk to each other directly. Virtualised GPUs generally cannot, and fail at startup reporting that the operation is not supported. Use passthrough or bare-metal GPUs. A copy that fits on one node is unaffected.
A copy spanning nodes fails as a unit. If any of its ranks dies, the whole copy is stopped and returned for re-placement, rather than left serving with part of the model missing, and the error names the rank, its node, and the exit status. A later scale-up replaces the copy if the pool is still above its minimum; losing a copy at the minimum ends the workflow. Multi-node serving is untested on AMD GPUs.
If one model is all you need, the default already does the right thing and you can skip to See also. Read on when you want several models behind one URL.
Each model entry runs a pool of replicas, meaning copies of the model that the cluster scales
with demand. The Proxy parameter controls how that pool is published, and decides whether
clients talk to the model directly or through a gateway. Either way the URL you are given stays
stable for the life of the workflow.
Proxy=true, the default. The workflow runs its own LiteLLM proxy, which holds the
endpoint and load-balances across whichever replicas are ready. The ApiKey is enforced there,
and the request rate through it is what triggers autoscaling. This is what the previous section
describes.
Proxy=false. Choose this when you plan to put several models behind one gateway. No
LiteLLM proxy is started, and the pool carries the endpoint itself, still on one stable URL. It
additionally publishes one address per replica, annotated so a gateway can discover the
replicas and balance across them directly. Two things change: the ApiKey is no longer
enforced, so the endpoint’s scope becomes the only access control, and request-rate autoscaling
stops, since it is the proxy that measures the rate.
$ fuzzball workflow catalog start "GPT-OSS 20B" --values Proxy=falseStart the models that way, then start the gateway. The LiteLLM Model Gateway entry polls the
endpoints API, picks up every annotated replica you have access to, and serves them all from a
single OpenAI-compatible base URL, adding and dropping them as pools scale:
$ fuzzball workflow catalog start "LiteLLM Model Gateway"Point your clients at the gateway’s URL. They ask for a model by name and never learn which workflow serves it. A request for a model whose pool has scaled to zero wakes it and succeeds on retry.
Adding the gateway to a model already running withProxy=truemeans restarting that model, not changing it in place. The gateway discovers per-replica addresses, which a pool publishes only withProxy=false.
- Using OpenCode with a Fuzzball-Hosted Model connects a terminal coding agent to the endpoint you just created, including trusting a private cluster certificate from your workstation.
- Workflow Endpoints covers endpoint scopes and access tokens in full.
- Service Autoscaling explains the replica pool the model runs in.