Workflow Endpoints
When a workflow service exposes network endpoints (configured via the network section in your
workflow definition), the fuzzball workflow endpoints command and its subcommands let you list
the live endpoints you have access to, inspect a single endpoint, mint an access token for a
non-public endpoint, or print the URL(s) derived from a workflow specification.
fuzzball workflow endpoints list returns the workflow service endpoints currently alive on the
cluster that you have access to. Access is scoped to the groups (accounts) you belong to and to
your organization.
An endpoint is alive only while the workflow that declared it is still running. Once that workflow
reaches a terminal state — Finished, Errored, or Canceled (see
Workflow Status)
— its endpoints are removed, whether the service shut down normally or exited unexpectedly.
fuzzball workflow endpoints get rechecks the owning workflow on every call and stops returning
the endpoint immediately. The list above does not recheck each row, so an endpoint whose workflow
has just ended can still appear there briefly; get is authoritative.
$ fuzzball workflow endpoints list
ENDPOINT ID WORKFLOW ID SERVICE ENDPOINT REPLICA TYPE SCOPE URL
endpoint-1a2b3c4d <workflow> jupyter web path user https://endpoints.example.com/endpoints/...
endpoint-5e6f7a8b <workflow> api-server grpc subdomain group https://api-server.<workflow>.<account>.example.com/The REPLICA column is blank for an ordinary endpoint and for an autoscaled pool’s own endpoint.
It carries the replica number only on the per-replica endpoints described in
“Addressing individual replicas” below.
Flags:
| Flag | Description |
|---|---|
-p, --page-size | Server batch size per request |
-m, --max-pages | Hard cap on the number of pages fetched (0 = no cap) |
-o, --output | Output format (table default, or json / yaml) |
Use -o json or -o yaml for scripting; the structured output includes additional fields such as
the owning account and organization.
$ fuzzball workflow endpoints list -o jsonfuzzball workflow endpoints get shows the details of one endpoint by its ID, including its URL:
$ fuzzball workflow endpoints get ENDPOINT_IDEndpoints whose scope is not public require a bearer token. fuzzball workflow endpoints generate-token mints a token bound to the endpoint and carrying your identity:
$ fuzzball workflow endpoints generate-token ENDPOINT_IDPresent the returned token to the endpoint using the Authorization or FB-Authorization HTTP
header.
A token can only be minted while the owning workflow is still running. Once the workflow reaches a terminal state the request is refused, and tokens already issued for its endpoints stop working as soon as the endpoint is removed. Mint a fresh token against a running workflow instead of holding a long-lived one across workflow restarts.
On a federated deployment, an endpoint belonging to a workflow the federate cluster has not yet learned about can still be read and minted; the cluster running the workflow remains authoritative and stops serving the endpoint once it is gone.
The --expiration flag sets the token lifetime. It accepts the units understood by Go durations
(s, m, h) plus day, week, month, and year units — 7d, 2w, 1mo, 1y (and long forms
like 7days). When omitted, the server default lifetime is used:
# Valid for 24 hours
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 24h
# Valid for 7 days
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 7d
# Valid for 1 month
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 1moInvoked with a workflow ID (and optional service and endpoint names), fuzzball workflow endpoints prints the endpoint URL(s) derived from that workflow’s specification:
In addition to querying endpoint URLs with the CLI, the Fuzzball Web UI provides a Connect button on workflow detail pages and in the workflows list. When a workflow is running and has at least one named endpoint defined, clicking the Connect button opens the workflow’s primary endpoint in a new browser tab.
This provides a convenient shortcut for accessing web-based services without needing to copy and paste URLs. For details on how the primary endpoint is selected when multiple endpoints exist, and for information on connecting to specific non-primary endpoints or using client-script mode, see the Connecting to Workflow Services page.
$ fuzzball workflow endpoints WORKFLOW_ID [SERVICE_NAME] [ENDPOINT_NAME]Arguments:
| Argument | Required | Description |
|---|---|---|
WORKFLOW_ID | Yes | The ID of a running workflow |
SERVICE_NAME | No | Filter results to a specific service |
ENDPOINT_NAME | No | Filter results to a specific endpoint within a service |
Get all endpoints for a workflow:
$ fuzzball workflow endpoints <workflow_id>
SERVICE ENDPOINT URL
jupyter web https://endpoints.example.com/endpoints/accounts/.../jupyter/web/
api-server grpc https://endpoints.example.com/endpoints/accounts/.../api-server/grpc/Get all endpoints for a specific service, or a specific endpoint by name:
$ fuzzball workflow endpoints <workflow_id> jupyter
$ fuzzball workflow endpoints <workflow_id> jupyter webOutput the results as JSON (useful for scripting):
$ fuzzball workflow endpoints <workflow_id> -o jsonEndpoints are declared in the network section of a service definition. For example, a Jupyter
service that exposes a web endpoint:
services:
jupyter:
image:
uri: docker://jupyter/base-notebook:latest
script: |
#!/bin/bash
jupyter lab --ip=0.0.0.0 --no-browser
resource:
cpu:
cores: 2
memory:
size: 4GB
network:
ports:
- name: web
port: 8888
protocol: tcp
endpoints:
- name: web
port-name: web
protocol: https
type: path
scope: user
See the workflow syntax reference for full details on
the network configuration options.
The type field selects how an endpoint is addressed:
path(the default) serves the endpoint beneath a single cluster-wide hostname, ashttps://endpoints.<domain>/endpoints/accounts/....subdomaingives the endpoint its own hostname,https://endpoint-<id>.endpoints.<domain>/.
Prefer subdomain for applications that assume they are served from the root of a domain. Many web
applications build absolute links or set cookies in ways that break beneath a path prefix, and a
dedicated hostname avoids the problem entirely.
The trade-off is a DNS requirement. Because a new hostname is generated for each endpoint,
subdomain endpoints require the cluster domain to resolve wildcards — every name matching
*.endpoints.<domain> must reach the cluster. Managed and cloud deployments handle this already.
For a single-host Docker Compose deployment, see
DNS resolution; note
that a /etc/hosts entry alone is not sufficient, since hosts files cannot express wildcards.
If asubdomainendpoint URL does not resolve whilepathendpoints on the same cluster work, the cluster domain is missing wildcard DNS rather than the endpoint being misconfigured.
An endpoint can carry annotations — arbitrary key-value pairs that Fuzzball stores with the
endpoint and returns from fuzzball workflow endpoints list and get. Fuzzball does not interpret
them. They exist so a client that discovers endpoints through the API can tell which ones it is
meant to consume, and skip the rest:
services:
vllm:
image:
uri: docker://vllm/vllm-openai:latest
network:
ports:
- name: http
port: 8000
protocol: tcp
endpoints:
- name: openai
port-name: http
protocol: https
type: subdomain
scope: group
annotations:
example.com/api: openai
example.com/model: llama-3.1-8b
The annotations come back in the structured output:
$ fuzzball workflow endpoints get ENDPOINT_ID
annotations:
example.com/api: openai
example.com/model: llama-3.1-8b
endpointName: openai
id: endpoint-5e6f7a8b
...They are not shown in the list table; use -o yaml or -o json to see them for a whole list.
Limits. Each endpoint may declare at most 32 annotations. A key may be up to 253 bytes and a
value up to 1024, with all keys and values together limited to 8192 bytes. A key is an optional
DNS subdomain prefix followed by / and an alphanumeric name of up to 63 characters, with dashes,
underscores and dots allowed inside the name – for example example.com/api or role. The
fuzzball.io prefix, and any subdomain of it, is reserved for the platform and rejected in your own
annotations. A workflow that exceeds any of these limits is rejected at submit time.
An autoscaled service runs a pool of interchangeable replicas rather than a single container, so its endpoint works differently: one endpoint addresses the whole pool. The endpoint exists for the life of the workflow, its URL never changes, and Fuzzball forwards each request to one of the replicas that is ready at that moment, cycling through them in turn.
services:
vllm:
image:
uri: docker://vllm/vllm-openai:latest
autoscaler:
replicas:
min: 1
max: 4
network:
ports:
- name: http
port: 8000
protocol: tcp
endpoints:
- name: api
port-name: http
protocol: http
type: subdomain
scope: group
Callers of https://<endpoint-id>.endpoints.<cluster-domain>/ reach the pool without knowing how
many replicas exist. Scaling up adds replicas to the rotation as they pass their readiness probe;
scaling down removes a replica from the rotation when its drain period begins, so requests already
in flight to it can finish while new ones go elsewhere.
Two rules apply to a pool’s endpoints, and a workflow that breaks either is rejected at submission:
- The endpoint must set
type: subdomain. Apathendpoint addresses one backend and has no way to select among a pool’s replicas. - The endpoint must reference a named port. Each replica serves on a randomly assigned host port, and Fuzzball tracks those ports per named port.
A pool with replicas.min: 0 keeps its endpoint while it is idle at zero replicas. A request
arriving then cannot be served, because there is no replica to serve it. Fuzzball answers
503 Service Unavailable with a Retry-After header and, at the same time, starts one replica.
The request is not held open: starting a replica of a large model takes minutes, far longer than
most clients wait. Retry after the cold start and the request is served normally.
This works every time the pool returns to zero, not only on the first request.
Two rules apply to waking a pool:
- The request must be authenticated. Starting a replica consumes cluster resources, GPUs in the
usual case, so only a caller Fuzzball can identify may trigger it. An endpoint with
scope: publicis served without authentication and therefore never wakes its pool. Give a pool you want woken on demand auser,groupororganizationscope, and reach it with a user token or an endpoint access token. - One replica per wake. A burst of requests against an idle pool starts a single replica rather
than one per request: further wakes for that endpoint are suppressed for a minute, or until the
replica that wake asked for has served a request. A pool that scales back to zero is therefore
wakeable again as soon as it goes idle, rather than waiting out a cooldown its own replica already
satisfied. Once a replica is running, the service’s own
scale-uptriggers grow the pool to match its load, up toreplicas.max.
Each wake is recorded on the workflow, so fuzzball workflow events <workflow-id> shows when a
request woke an idle pool.
Some callers do their own load balancing – an API gateway that tracks each backend’s health and
spreads requests according to its own policy, for example. Setting per-replica: true on a pool’s
endpoint gives every live replica its own endpoint as well, so such a caller sees one target per
replica instead of one target for the pool.
endpoints:
- name: api
port-name: http
protocol: http
type: subdomain
scope: group
per-replica: true
The pool’s own endpoint and the per-replica endpoints coexist; per-replica adds endpoints and
takes nothing away. Use the pool’s URL when you want Fuzzball to balance across the replicas, and
the per-replica URLs when the caller balances for itself.
A replica’s endpoint appears when the replica passes its readiness probe and disappears when the
autoscaler begins draining it, so a caller that refreshes its list of endpoints follows the pool as
it scales, and a retiring replica leaves the rotation before it is stopped. Both kinds of endpoint
come back from fuzzball workflow endpoints list and the /endpoints API; a per-replica endpoint
reports replicaIndex (the replica it addresses, numbered from 1) and poolEndpointId (the pool’s
own endpoint), so a caller can group a pool’s replicas into one entry. replicaIndex is the
REPLICA column of the endpoint list table; poolEndpointId is available via -o json and
-o yaml.
per-replica: true requires type: subdomain and a service that defines autoscaler.replicas; a
workflow that sets it anywhere else is rejected at submission.
Every replica endpoint is a hostname of its own, created and removed as the pool scales, so a
cluster serving them needs the wildcard DNS that any subdomain endpoint needs – see
Choosing between path and subdomain.
Per-replica endpoints are local to the cluster running the workflow. In a federated deployment they are not published to the federate cluster, whose endpoint list carries the pool’s own endpoint instead – one address that always resolves to a ready replica.
On an endpoint whose scope is not public, Fuzzball authenticates every request before proxying it
and then removes the Authorization header, so a caller’s own token never reaches your service. To
let the service still know who it is serving, Fuzzball forwards a signed assertion of the caller’s
identity in the X-Fuzzball-Caller-Identity header.
This is what lets a service offer per-user behaviour – usage accounting, per-user access rules, audit trails – without asking the caller for a second credential of its own.
The header holds a JWT signed by the cluster. Its claims:
| Claim | Meaning |
|---|---|
aud | endpoint:<endpoint id> – what marks this as an assertion rather than a bearer token. Always check it against FB_ENDPOINT_ID |
sub | The calling user’s id |
email | The calling user’s email address |
account_id | For a user, the group the caller is acting in. For a caller presenting an endpoint token, the endpoint’s own account |
organization_id | The caller’s organization |
endpoint_id | The endpoint the assertion was minted for; the same id as in aud |
exp | Expiry, two minutes after minting |
iat | When it was minted |
iss | The cluster’s internal issuer, for information only. Never fetch from it – see below |
type | Always Bearer. An artefact of the shared minting path; it does not mean the assertion may be used as a credential |
cluster_id_origin | The cluster that issued it |
The assertion is signed with ES256. Fuzzball publishes the keys that verify it into the node trust store, the same directory it mounts the cluster CA into, so a service reads them from a local file rather than fetching them over the network:
/run/fuzzball-substrate/trusted-certs/signing-keys.json
The file is an ordinary JWKS document; match the assertion’s kid against it. Reject any assertion
whose signature does not verify, whose exp has passed, or whose aud is not your own endpoint.
Do not fetch keys from the URL in the assertion’s iss claim. Until the signature has been checked,
every claim is whatever the sender chose to put there, so a forged assertion would simply name a key
server the forger controls. The local file is the only source to trust.
Requireaud. An endpoint token – the credentialfuzzball workflow endpoints generate-tokenhands out – is signed by the same cluster key, names the same issuer, and carries the same identity claims and the sameendpoint_id.audis the only claim that separates the two. A service that verifies the signature without checkingaudwill accept an endpoint token as an assertion naming whoever minted it, and endpoint tokens are meant to be shared and may be minted with any lifetime.
Your endpoint’s own id arrives in the environment, so the service does not have to derive it:
| Variable | Value |
|---|---|
FB_ENDPOINT_ID | The id of the service’s first endpoint |
FB_ENDPOINT_ID_<NAME> | The id of the endpoint named <NAME>, upper-cased with punctuation replaced by _ |
FB_ENDPOINT_REPLICA_ID, FB_ENDPOINT_REPLICA_ID_<NAME> | On a per-replica endpoint of a replica pool, this replica’s own endpoint id |
A pooled service is reachable two ways – through the pool’s endpoint and through its own
per-replica endpoint – and the assertion names whichever the request arrived on. Accept both
FB_ENDPOINT_ID and FB_ENDPOINT_REPLICA_ID when you set one up.
The example uses PyJWT with its cryptography extra
(pip install "pyjwt[crypto]" – note that the PyPI package jwt is a different library):
import json
import os
import jwt
KEYS_PATH = "/run/fuzzball-substrate/trusted-certs/signing-keys.json"
MY_AUDIENCES = [
f"endpoint:{os.environ[name]}"
for name in ("FB_ENDPOINT_ID", "FB_ENDPOINT_REPLICA_ID")
if os.environ.get(name)
]
raw = request.headers.get("X-Fuzzball-Caller-Identity")
if raw is None or not MY_AUDIENCES:
# No header, or nothing to bind it to: treat the caller as anonymous
# rather than verifying against an empty audience list.
return handle_anonymous_request()
with open(KEYS_PATH) as handle:
keys = jwt.PyJWKSet.from_dict(json.load(handle))
kid = jwt.get_unverified_header(raw)["kid"]
key = next(k for k in keys.keys if k.key_id == kid).key
claims = jwt.decode(
raw,
key,
algorithms=["ES256"],
audience=MY_AUDIENCES,
options={"require": ["exp", "sub", "aud"]},
)
user_id = claims["sub"]
iss is deliberately not pinned. The key file is what establishes the signer – only the cluster
can produce a signature that verifies against it – and aud is what binds the assertion to this
endpoint, so an issuer check adds nothing. There is also no issuer value to check against inside a
container: the claim names an internal address a workflow container cannot resolve.
A rotated signing key does not reach a running node. A node writes this file once, when its Fuzzball extension starts. The cluster serves the current key immediately, but an existing node keeps the set it was given, so after rotating the cluster signing key the extension has to be restarted on every node before services can verify again. Until then a service sees assertions signed with a
kidabsent from its file and must reject them.The same applies to withdrawing a key. A node served no keys removes the file, but only at its next extension start – a running node keeps what it already has, so a withdrawn key stays trusted there until the extension restarts.
The assertion is not a credential. Fuzzball refuses it as a bearer token, so a service cannot use one to call the Fuzzball API as the caller. Treat it as a statement about who called, nothing more.
A
publicendpoint authenticates nobody, so no assertion is sent. Your service must treat a missing header as an anonymous caller rather than failing closed on a value it expects to be present.Fuzzball always discards an
X-Fuzzball-Caller-Identityheader supplied by the client, at every scope, so a caller cannot choose the identity your service sees. Never trust the header without verifying its signature: a service also listens on its node’s port, where requests arrive without passing through the endpoint proxy at all.
Three other headers are added on every proxied request – X-Fuzzball-Workflow-ID,
X-Fuzzball-Account-ID and X-Fuzzball-Service. These describe the endpoint’s own workflow, not
the caller, and are unsigned; use them for tracing, not for authorization.
When a workflow service runs, Fuzzball automatically injects environment variables for each configured endpoint. These variables allow your service code to discover the public URLs assigned to its endpoints.
For each endpoint, two environment variables are provided:
FB_ENDPOINT_PATH_<ENDPOINT_NAME>— the full URL path to the endpointFB_ENDPOINT_URL_<ENDPOINT_NAME>— the complete URL including protocol and host
The <ENDPOINT_NAME> suffix is derived from the endpoint’s name field by:
- Converting the name to uppercase
- Replacing any character that is not a letter, digit, or underscore with an underscore (
_)
This transformation ensures that the generated environment variable names are valid POSIX shell identifiers and compatible with all container runtimes.
Examples:
| Endpoint name | Environment variables |
|---|---|
web | FB_ENDPOINT_PATH_WEBFB_ENDPOINT_URL_WEB |
jupyter-one | FB_ENDPOINT_PATH_JUPYTER_ONEFB_ENDPOINT_URL_JUPYTER_ONE |
my.endpoint | FB_ENDPOINT_PATH_MY_ENDPOINTFB_ENDPOINT_URL_MY_ENDPOINT |
api_v2 | FB_ENDPOINT_PATH_API_V2FB_ENDPOINT_URL_API_V2 |
Simple alphanumeric endpoint names (containing only letters, digits, and underscores) are unaffected by this transformation beyond being converted to uppercase.
You can reference these environment variables in your service’s startup script to configure applications dynamically:
services:
web-app:
image:
uri: docker://myapp:latest
script: |
#!/bin/bash
echo "Service available at: $FB_ENDPOINT_URL_WEB"
./myapp --public-url="$FB_ENDPOINT_URL_WEB"
network:
ports:
- name: http
port: 8080
protocol: tcp
endpoints:
- name: web
port-name: http
protocol: https
type: path
scope: user
For an endpoint named jupyter-notebook, you would reference the sanitized variable names:
echo "Notebook URL: $FB_ENDPOINT_URL_JUPYTER_NOTEBOOK"