Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Workflow Endpoints

When a workflow service exposes network endpoints (configured via the network section in your workflow definition), the fuzzball workflow endpoints command and its subcommands let you list the live endpoints you have access to, inspect a single endpoint, mint an access token for a non-public endpoint, or print the URL(s) derived from a workflow specification.

Listing live endpoints

fuzzball workflow endpoints list returns the workflow service endpoints currently alive on the cluster that you have access to. Access is scoped to the groups (accounts) you belong to and to your organization.

An endpoint is alive only while the workflow that declared it is still running. Once that workflow reaches a terminal state — Finished, Errored, or Canceled (see Workflow Status) — its endpoints are removed, whether the service shut down normally or exited unexpectedly.

fuzzball workflow endpoints get rechecks the owning workflow on every call and stops returning the endpoint immediately. The list above does not recheck each row, so an endpoint whose workflow has just ended can still appear there briefly; get is authoritative.

$ fuzzball workflow endpoints list

ENDPOINT ID         WORKFLOW ID   SERVICE     ENDPOINT   REPLICA   TYPE        SCOPE   URL
endpoint-1a2b3c4d    <workflow>    jupyter     web                  path        user    https://endpoints.example.com/endpoints/...
endpoint-5e6f7a8b    <workflow>    api-server  grpc                 subdomain   group   https://api-server.<workflow>.<account>.example.com/

The REPLICA column is blank for an ordinary endpoint and for an autoscaled pool’s own endpoint. It carries the replica number only on the per-replica endpoints described in “Addressing individual replicas” below.

Flags:

FlagDescription
-p, --page-sizeServer batch size per request
-m, --max-pagesHard cap on the number of pages fetched (0 = no cap)
-o, --outputOutput format (table default, or json / yaml)

Use -o json or -o yaml for scripting; the structured output includes additional fields such as the owning account and organization.

$ fuzzball workflow endpoints list -o json

Showing a single endpoint

fuzzball workflow endpoints get shows the details of one endpoint by its ID, including its URL:

$ fuzzball workflow endpoints get ENDPOINT_ID

Generating an access token

Endpoints whose scope is not public require a bearer token. fuzzball workflow endpoints generate-token mints a token bound to the endpoint and carrying your identity:

$ fuzzball workflow endpoints generate-token ENDPOINT_ID

Present the returned token to the endpoint using the Authorization or FB-Authorization HTTP header.

A token can only be minted while the owning workflow is still running. Once the workflow reaches a terminal state the request is refused, and tokens already issued for its endpoints stop working as soon as the endpoint is removed. Mint a fresh token against a running workflow instead of holding a long-lived one across workflow restarts.

On a federated deployment, an endpoint belonging to a workflow the federate cluster has not yet learned about can still be read and minted; the cluster running the workflow remains authoritative and stops serving the endpoint once it is gone.

The --expiration flag sets the token lifetime. It accepts the units understood by Go durations (s, m, h) plus day, week, month, and year units — 7d, 2w, 1mo, 1y (and long forms like 7days). When omitted, the server default lifetime is used:

# Valid for 24 hours
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 24h

# Valid for 7 days
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 7d

# Valid for 1 month
$ fuzzball workflow endpoints generate-token ENDPOINT_ID --expiration 1mo

Printing endpoint URLs from a workflow specification

Invoked with a workflow ID (and optional service and endpoint names), fuzzball workflow endpoints prints the endpoint URL(s) derived from that workflow’s specification:

Accessing endpoints via the Web UI

In addition to querying endpoint URLs with the CLI, the Fuzzball Web UI provides a Connect button on workflow detail pages and in the workflows list. When a workflow is running and has at least one named endpoint defined, clicking the Connect button opens the workflow’s primary endpoint in a new browser tab.

This provides a convenient shortcut for accessing web-based services without needing to copy and paste URLs. For details on how the primary endpoint is selected when multiple endpoints exist, and for information on connecting to specific non-primary endpoints or using client-script mode, see the Connecting to Workflow Services page.

$ fuzzball workflow endpoints WORKFLOW_ID [SERVICE_NAME] [ENDPOINT_NAME]

Arguments:

ArgumentRequiredDescription
WORKFLOW_IDYesThe ID of a running workflow
SERVICE_NAMENoFilter results to a specific service
ENDPOINT_NAMENoFilter results to a specific endpoint within a service

Get all endpoints for a workflow:

$ fuzzball workflow endpoints <workflow_id>

SERVICE       ENDPOINT   URL
jupyter       web        https://endpoints.example.com/endpoints/accounts/.../jupyter/web/
api-server    grpc       https://endpoints.example.com/endpoints/accounts/.../api-server/grpc/

Get all endpoints for a specific service, or a specific endpoint by name:

$ fuzzball workflow endpoints <workflow_id> jupyter
$ fuzzball workflow endpoints <workflow_id> jupyter web

Output the results as JSON (useful for scripting):

$ fuzzball workflow endpoints <workflow_id> -o json

Configuring endpoints in a workflow

Endpoints are declared in the network section of a service definition. For example, a Jupyter service that exposes a web endpoint:

services:
  jupyter:
    image:
      uri: docker://jupyter/base-notebook:latest
    script: |
      #!/bin/bash
      jupyter lab --ip=0.0.0.0 --no-browser
    resource:
      cpu:
        cores: 2
      memory:
        size: 4GB
    network:
      ports:
        - name: web
          port: 8888
          protocol: tcp
      endpoints:
        - name: web
          port-name: web
          protocol: https
          type: path
          scope: user

See the workflow syntax reference for full details on the network configuration options.

Choosing between path and subdomain

The type field selects how an endpoint is addressed:

  • path (the default) serves the endpoint beneath a single cluster-wide hostname, as https://endpoints.<domain>/endpoints/accounts/....
  • subdomain gives the endpoint its own hostname, https://endpoint-<id>.endpoints.<domain>/.

Prefer subdomain for applications that assume they are served from the root of a domain. Many web applications build absolute links or set cookies in ways that break beneath a path prefix, and a dedicated hostname avoids the problem entirely.

The trade-off is a DNS requirement. Because a new hostname is generated for each endpoint, subdomain endpoints require the cluster domain to resolve wildcards — every name matching *.endpoints.<domain> must reach the cluster. Managed and cloud deployments handle this already. For a single-host Docker Compose deployment, see DNS resolution; note that a /etc/hosts entry alone is not sufficient, since hosts files cannot express wildcards.

If a subdomain endpoint URL does not resolve while path endpoints on the same cluster work, the cluster domain is missing wildcard DNS rather than the endpoint being misconfigured.

Annotating endpoints for discovery

An endpoint can carry annotations — arbitrary key-value pairs that Fuzzball stores with the endpoint and returns from fuzzball workflow endpoints list and get. Fuzzball does not interpret them. They exist so a client that discovers endpoints through the API can tell which ones it is meant to consume, and skip the rest:

services:
  vllm:
    image:
      uri: docker://vllm/vllm-openai:latest
    network:
      ports:
        - name: http
          port: 8000
          protocol: tcp
      endpoints:
        - name: openai
          port-name: http
          protocol: https
          type: subdomain
          scope: group
          annotations:
            example.com/api: openai
            example.com/model: llama-3.1-8b

The annotations come back in the structured output:

$ fuzzball workflow endpoints get ENDPOINT_ID
annotations:
  example.com/api: openai
  example.com/model: llama-3.1-8b
endpointName: openai
id: endpoint-5e6f7a8b
...

They are not shown in the list table; use -o yaml or -o json to see them for a whole list.

Limits. Each endpoint may declare at most 32 annotations. A key may be up to 253 bytes and a value up to 1024, with all keys and values together limited to 8192 bytes. A key is an optional DNS subdomain prefix followed by / and an alphanumeric name of up to 63 characters, with dashes, underscores and dots allowed inside the name – for example example.com/api or role. The fuzzball.io prefix, and any subdomain of it, is reserved for the platform and rejected in your own annotations. A workflow that exceeds any of these limits is rejected at submit time.

Endpoints on an autoscaled pool

An autoscaled service runs a pool of interchangeable replicas rather than a single container, so its endpoint works differently: one endpoint addresses the whole pool. The endpoint exists for the life of the workflow, its URL never changes, and Fuzzball forwards each request to one of the replicas that is ready at that moment, cycling through them in turn.

services:
  vllm:
    image:
      uri: docker://vllm/vllm-openai:latest
    autoscaler:
      replicas:
        min: 1
        max: 4
    network:
      ports:
        - name: http
          port: 8000
          protocol: tcp
      endpoints:
        - name: api
          port-name: http
          protocol: http
          type: subdomain
          scope: group

Callers of https://<endpoint-id>.endpoints.<cluster-domain>/ reach the pool without knowing how many replicas exist. Scaling up adds replicas to the rotation as they pass their readiness probe; scaling down removes a replica from the rotation when its drain period begins, so requests already in flight to it can finish while new ones go elsewhere.

Two rules apply to a pool’s endpoints, and a workflow that breaks either is rejected at submission:

  • The endpoint must set type: subdomain. A path endpoint addresses one backend and has no way to select among a pool’s replicas.
  • The endpoint must reference a named port. Each replica serves on a randomly assigned host port, and Fuzzball tracks those ports per named port.

Waking a pool idle at zero replicas

A pool with replicas.min: 0 keeps its endpoint while it is idle at zero replicas. A request arriving then cannot be served, because there is no replica to serve it. Fuzzball answers 503 Service Unavailable with a Retry-After header and, at the same time, starts one replica. The request is not held open: starting a replica of a large model takes minutes, far longer than most clients wait. Retry after the cold start and the request is served normally.

This works every time the pool returns to zero, not only on the first request.

Two rules apply to waking a pool:

  • The request must be authenticated. Starting a replica consumes cluster resources, GPUs in the usual case, so only a caller Fuzzball can identify may trigger it. An endpoint with scope: public is served without authentication and therefore never wakes its pool. Give a pool you want woken on demand a user, group or organization scope, and reach it with a user token or an endpoint access token.
  • One replica per wake. A burst of requests against an idle pool starts a single replica rather than one per request: further wakes for that endpoint are suppressed for a minute, or until the replica that wake asked for has served a request. A pool that scales back to zero is therefore wakeable again as soon as it goes idle, rather than waiting out a cooldown its own replica already satisfied. Once a replica is running, the service’s own scale-up triggers grow the pool to match its load, up to replicas.max.

Each wake is recorded on the workflow, so fuzzball workflow events <workflow-id> shows when a request woke an idle pool.

Addressing individual replicas

Some callers do their own load balancing – an API gateway that tracks each backend’s health and spreads requests according to its own policy, for example. Setting per-replica: true on a pool’s endpoint gives every live replica its own endpoint as well, so such a caller sees one target per replica instead of one target for the pool.

      endpoints:
        - name: api
          port-name: http
          protocol: http
          type: subdomain
          scope: group
          per-replica: true

The pool’s own endpoint and the per-replica endpoints coexist; per-replica adds endpoints and takes nothing away. Use the pool’s URL when you want Fuzzball to balance across the replicas, and the per-replica URLs when the caller balances for itself.

A replica’s endpoint appears when the replica passes its readiness probe and disappears when the autoscaler begins draining it, so a caller that refreshes its list of endpoints follows the pool as it scales, and a retiring replica leaves the rotation before it is stopped. Both kinds of endpoint come back from fuzzball workflow endpoints list and the /endpoints API; a per-replica endpoint reports replicaIndex (the replica it addresses, numbered from 1) and poolEndpointId (the pool’s own endpoint), so a caller can group a pool’s replicas into one entry. replicaIndex is the REPLICA column of the endpoint list table; poolEndpointId is available via -o json and -o yaml.

per-replica: true requires type: subdomain and a service that defines autoscaler.replicas; a workflow that sets it anywhere else is rejected at submission.

Every replica endpoint is a hostname of its own, created and removed as the pool scales, so a cluster serving them needs the wildcard DNS that any subdomain endpoint needs – see Choosing between path and subdomain.

Per-replica endpoints are local to the cluster running the workflow. In a federated deployment they are not published to the federate cluster, whose endpoint list carries the pool’s own endpoint instead – one address that always resolves to a ready replica.

Identifying the caller

On an endpoint whose scope is not public, Fuzzball authenticates every request before proxying it and then removes the Authorization header, so a caller’s own token never reaches your service. To let the service still know who it is serving, Fuzzball forwards a signed assertion of the caller’s identity in the X-Fuzzball-Caller-Identity header.

This is what lets a service offer per-user behaviour – usage accounting, per-user access rules, audit trails – without asking the caller for a second credential of its own.

The header holds a JWT signed by the cluster. Its claims:

ClaimMeaning
audendpoint:<endpoint id> – what marks this as an assertion rather than a bearer token. Always check it against FB_ENDPOINT_ID
subThe calling user’s id
emailThe calling user’s email address
account_idFor a user, the group the caller is acting in. For a caller presenting an endpoint token, the endpoint’s own account
organization_idThe caller’s organization
endpoint_idThe endpoint the assertion was minted for; the same id as in aud
expExpiry, two minutes after minting
iatWhen it was minted
issThe cluster’s internal issuer, for information only. Never fetch from it – see below
typeAlways Bearer. An artefact of the shared minting path; it does not mean the assertion may be used as a credential
cluster_id_originThe cluster that issued it

The assertion is signed with ES256. Fuzzball publishes the keys that verify it into the node trust store, the same directory it mounts the cluster CA into, so a service reads them from a local file rather than fetching them over the network:

/run/fuzzball-substrate/trusted-certs/signing-keys.json

The file is an ordinary JWKS document; match the assertion’s kid against it. Reject any assertion whose signature does not verify, whose exp has passed, or whose aud is not your own endpoint.

Do not fetch keys from the URL in the assertion’s iss claim. Until the signature has been checked, every claim is whatever the sender chose to put there, so a forged assertion would simply name a key server the forger controls. The local file is the only source to trust.

Require aud. An endpoint token – the credential fuzzball workflow endpoints generate-token hands out – is signed by the same cluster key, names the same issuer, and carries the same identity claims and the same endpoint_id. aud is the only claim that separates the two. A service that verifies the signature without checking aud will accept an endpoint token as an assertion naming whoever minted it, and endpoint tokens are meant to be shared and may be minted with any lifetime.

Your endpoint’s own id arrives in the environment, so the service does not have to derive it:

VariableValue
FB_ENDPOINT_IDThe id of the service’s first endpoint
FB_ENDPOINT_ID_<NAME>The id of the endpoint named <NAME>, upper-cased with punctuation replaced by _
FB_ENDPOINT_REPLICA_ID, FB_ENDPOINT_REPLICA_ID_<NAME>On a per-replica endpoint of a replica pool, this replica’s own endpoint id

A pooled service is reachable two ways – through the pool’s endpoint and through its own per-replica endpoint – and the assertion names whichever the request arrived on. Accept both FB_ENDPOINT_ID and FB_ENDPOINT_REPLICA_ID when you set one up.

The example uses PyJWT with its cryptography extra (pip install "pyjwt[crypto]" – note that the PyPI package jwt is a different library):

import json
import os

import jwt

KEYS_PATH = "/run/fuzzball-substrate/trusted-certs/signing-keys.json"

MY_AUDIENCES = [
    f"endpoint:{os.environ[name]}"
    for name in ("FB_ENDPOINT_ID", "FB_ENDPOINT_REPLICA_ID")
    if os.environ.get(name)
]

raw = request.headers.get("X-Fuzzball-Caller-Identity")
if raw is None or not MY_AUDIENCES:
    # No header, or nothing to bind it to: treat the caller as anonymous
    # rather than verifying against an empty audience list.
    return handle_anonymous_request()

with open(KEYS_PATH) as handle:
    keys = jwt.PyJWKSet.from_dict(json.load(handle))

kid = jwt.get_unverified_header(raw)["kid"]
key = next(k for k in keys.keys if k.key_id == kid).key

claims = jwt.decode(
    raw,
    key,
    algorithms=["ES256"],
    audience=MY_AUDIENCES,
    options={"require": ["exp", "sub", "aud"]},
)

user_id = claims["sub"]

iss is deliberately not pinned. The key file is what establishes the signer – only the cluster can produce a signature that verifies against it – and aud is what binds the assertion to this endpoint, so an issuer check adds nothing. There is also no issuer value to check against inside a container: the claim names an internal address a workflow container cannot resolve.

A rotated signing key does not reach a running node. A node writes this file once, when its Fuzzball extension starts. The cluster serves the current key immediately, but an existing node keeps the set it was given, so after rotating the cluster signing key the extension has to be restarted on every node before services can verify again. Until then a service sees assertions signed with a kid absent from its file and must reject them.

The same applies to withdrawing a key. A node served no keys removes the file, but only at its next extension start – a running node keeps what it already has, so a withdrawn key stays trusted there until the extension restarts.

The assertion is not a credential. Fuzzball refuses it as a bearer token, so a service cannot use one to call the Fuzzball API as the caller. Treat it as a statement about who called, nothing more.

A public endpoint authenticates nobody, so no assertion is sent. Your service must treat a missing header as an anonymous caller rather than failing closed on a value it expects to be present.

Fuzzball always discards an X-Fuzzball-Caller-Identity header supplied by the client, at every scope, so a caller cannot choose the identity your service sees. Never trust the header without verifying its signature: a service also listens on its node’s port, where requests arrive without passing through the endpoint proxy at all.

Three other headers are added on every proxied request – X-Fuzzball-Workflow-ID, X-Fuzzball-Account-ID and X-Fuzzball-Service. These describe the endpoint’s own workflow, not the caller, and are unsigned; use them for tracing, not for authorization.

Environment variables for endpoints

When a workflow service runs, Fuzzball automatically injects environment variables for each configured endpoint. These variables allow your service code to discover the public URLs assigned to its endpoints.

For each endpoint, two environment variables are provided:

  • FB_ENDPOINT_PATH_<ENDPOINT_NAME> — the full URL path to the endpoint
  • FB_ENDPOINT_URL_<ENDPOINT_NAME> — the complete URL including protocol and host

Environment variable naming

The <ENDPOINT_NAME> suffix is derived from the endpoint’s name field by:

  1. Converting the name to uppercase
  2. Replacing any character that is not a letter, digit, or underscore with an underscore (_)

This transformation ensures that the generated environment variable names are valid POSIX shell identifiers and compatible with all container runtimes.

Examples:

Endpoint nameEnvironment variables
webFB_ENDPOINT_PATH_WEB
FB_ENDPOINT_URL_WEB
jupyter-oneFB_ENDPOINT_PATH_JUPYTER_ONE
FB_ENDPOINT_URL_JUPYTER_ONE
my.endpointFB_ENDPOINT_PATH_MY_ENDPOINT
FB_ENDPOINT_URL_MY_ENDPOINT
api_v2FB_ENDPOINT_PATH_API_V2
FB_ENDPOINT_URL_API_V2
Simple alphanumeric endpoint names (containing only letters, digits, and underscores) are unaffected by this transformation beyond being converted to uppercase.

Using endpoint variables in startup scripts

You can reference these environment variables in your service’s startup script to configure applications dynamically:

services:
  web-app:
    image:
      uri: docker://myapp:latest
    script: |
      #!/bin/bash
      echo "Service available at: $FB_ENDPOINT_URL_WEB"
      ./myapp --public-url="$FB_ENDPOINT_URL_WEB"
    network:
      ports:
        - name: http
          port: 8080
          protocol: tcp
      endpoints:
        - name: web
          port-name: http
          protocol: https
          type: path
          scope: user

For an endpoint named jupyter-notebook, you would reference the sanitized variable names:

echo "Notebook URL: $FB_ENDPOINT_URL_JUPYTER_NOTEBOOK"