Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Serving an AI Model

Fuzzball serves large language models from the workflow catalog. One command starts a model on a cluster GPU and gives you a base URL that speaks the OpenAI API, so any client written for OpenAI can use it: a coding agent, an IDE plugin, a notebook, or plain curl.

You do not need to write a Fuzzfile, choose a container image, or know how vLLM, the server that runs the model on the GPU, is configured. Each model has its own catalog entry, already carrying the settings that model needs.

Choosing a model

The only thing you have to match is GPU memory. A model needs enough of it to hold its weights plus the context it is serving:

Catalog entryGPU memory to run one copy
GPT-OSS 20B16GB
Ministral 3 8B24GB
Ministral 3 14B48GB
Gemma 4 31B80GB
GPT-OSS 120B80GB
Nemotron 3 Nano80GB
Qwen3-Coder 30B160GB
Qwen3-Coder Next 80B160GB
Nemotron 3 Super640GB
Mistral Medium 3.5640GB
Llama 4 Scout640GB

Two numbers move those figures. Each entry assumes its own default context size, which is large, so lowering MaxContextSize lets a model run on less memory. And the memory does not have to come from one card: each entry defaults to however many 80GB-class GPUs its total needs, which you change with GpusPerNode. Qwen3-Coder 30B defaults to two GPUs to reach 160GB, but runs on one card that has the memory, or on one 80GB card with a smaller context. Those GPUs all come from a single node; if no node in your cluster has enough of them, see Spanning several nodes.

To see everything your cluster offers:

$ fuzzball workflow catalog list --category ML_AND_AI
$ fuzzball workflow catalog list --tag LLM

GPT-OSS 20B is the smallest and the usual starting point. If you are unsure what GPUs your cluster has, run fuzzball node list or ask your administrator.

To serve a model that has no entry of its own, use the general vLLM entry and set its Model parameter to a Hugging Face URI, for example --values Model=hf://openai/gpt-oss-20b. You then choose the GPU count, memory, and context size yourself, which the per-model entries otherwise do for you.

Starting it

Apart from the gated one covered below, every per-model entry runs with no parameters at all:

$ fuzzball workflow catalog start "GPT-OSS 20B"

This prints a workflow ID. Copy it: the commands below need it, and fuzzball workflow list finds it again later. Add --watch to follow the workflow until it is running:

$ fuzzball workflow catalog start "GPT-OSS 20B" --watch
The first start is slow. The model weights download from the Hugging Face Hub before the server comes up, which is tens of gigabytes for a 20B model. By default they land on an ephemeral volume, temporary storage that is cleared when the workflow stops, so the next start downloads them again. To keep them, point the entry at a persistent volume: --values Volume=my-models.

A gated model additionally needs a Hugging Face token in a Fuzzball secret. Llama 4 Scout is the one gated entry today:

$ fuzzball secret create secret://user/hf-token --type hf '{"token":"hf_..."}'

$ fuzzball workflow catalog start "Llama 4 Scout" --values HfTokenSecret=secret://user/hf-token
A running model holds its GPUs until you stop it. Stop it when you are done: fuzzball workflow stop WORKFLOW_ID.

Confirming it is up

The workflow is ready when its service is running and its health check has passed:

$ fuzzball workflow describe WORKFLOW_ID

Then find the base URL:

$ fuzzball workflow endpoints list

The entry generates an API key at start time unless you set one with ApiKey. It is stored in the workflow definition:

$ fuzzball workflow get WORKFLOW_ID | grep -oE 'sk-[A-Za-z0-9_-]+' | head -1

Every endpoint has a Scope controlling who may call it. It defaults to group, so requests also need a Fuzzball token bound to the endpoint:

$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 8h

Both credentials travel on every request, in different headers. The Fuzzball endpoint proxy consumes the Authorization header to authenticate you, then removes it before forwarding the request. The API key therefore has to travel in a header the proxy leaves alone. Ask for the model list to confirm the whole path works:

$ curl -H "Authorization: Bearer YOUR-FUZZBALL-TOKEN" \
    -H "x-litellm-api-key: sk-fb-YOUR-API-KEY" \
    https://ENDPOINT-URL/v1/models

A healthy response names the model:

{"data":[{"id":"openai/gpt-oss-20b","object":"model","created":1677610602,"owned_by":"openai"}],"object":"list"}

An endpoint started with Scope=public needs only the API key, and is served without any Fuzzball authentication, so use it only with a strong ApiKey.

Spanning several nodes

Everything above puts one copy of the model on one node. When no single node has the memory, Nodes splits a copy across several. Fuzzball starts and stops those nodes together, and the first one serves the endpoint:

$ fuzzball workflow catalog start "Nemotron 3 Super" --values Nodes=2,GpusPerNode=4

A copy then uses Nodes times GpusPerNode GPUs, and no two of its ranks share a node, so the cluster needs that many nodes free. When they are merely busy the workflow waits for them and reports a scheduling_blocked event. It is rejected at submission only when the copy asks for more nodes than the cluster is configured to give it, which is a limit no amount of waiting changes. You do not need to choose how the model is split across the nodes: ExpertParallelism defaults to reading the model’s own configuration and picking a layout.

What you may need to set is the network. The nodes talk to each other directly, and on a cluster with more than one network interface the collective library has to be told which one to use. That goes in ExtraEnv as space-separated NAME=VALUE pairs:

$ fuzzball workflow catalog start "Nemotron 3 Super" \
    --values 'Nodes=2,GpusPerNode=4,ExtraEnv=NCCL_SOCKET_IFNAME=eth0'

The interface name is specific to your cluster, so ask your administrator which one carries node-to-node traffic. A fabric with RDMA may need more than one variable.

Multi-node copies need GPUs that can talk to each other directly. Virtualised GPUs generally cannot, and fail at startup reporting that the operation is not supported. Use passthrough or bare-metal GPUs. A copy that fits on one node is unaffected.

A copy spanning nodes fails as a unit. If any of its ranks dies, the whole copy is stopped and returned for re-placement, rather than left serving with part of the model missing, and the error names the rank, its node, and the exit status. A later scale-up replaces the copy if the pool is still above its minimum; losing a copy at the minimum ends the workflow. Multi-node serving is untested on AMD GPUs.

One model, or several

If one model is all you need, the default already does the right thing and you can skip to See also. Read on when you want several models behind one URL.

Each model entry runs a pool of replicas, meaning copies of the model that the cluster scales with demand. The Proxy parameter controls how that pool is published, and decides whether clients talk to the model directly or through a gateway. Either way the URL you are given stays stable for the life of the workflow.

Proxy=true, the default. The workflow runs its own LiteLLM proxy, which holds the endpoint and load-balances across whichever replicas are ready. The ApiKey is enforced there, and the request rate through it is what triggers autoscaling. This is what the previous section describes.

Proxy=false. Choose this when you plan to put several models behind one gateway. No LiteLLM proxy is started, and the pool carries the endpoint itself, still on one stable URL. It additionally publishes one address per replica, annotated so a gateway can discover the replicas and balance across them directly. Two things change: the ApiKey is no longer enforced, so the endpoint’s scope becomes the only access control, and request-rate autoscaling stops, since it is the proxy that measures the rate.

$ fuzzball workflow catalog start "GPT-OSS 20B" --values Proxy=false

Start the models that way, then start the gateway. The LiteLLM Model Gateway entry polls the endpoints API, picks up every annotated replica you have access to, and serves them all from a single OpenAI-compatible base URL, adding and dropping them as pools scale:

$ fuzzball workflow catalog start "LiteLLM Model Gateway"

Point your clients at the gateway’s URL. They ask for a model by name and never learn which workflow serves it. A request for a model whose pool has scaled to zero wakes it and succeeds on retry.

Adding the gateway to a model already running with Proxy=true means restarting that model, not changing it in place. The gateway discovers per-replica addresses, which a pool publishes only with Proxy=false.

See also