Fuzzball v4.3.0 release notes
Fuzzball v4.3.0 is a feature release.
Headline features:
- The workflow catalog provides a new set of AI catalog entries: there’s a vLLM entry, eleven pre-configured entries for specific self-hosted models, a gateway for providing access to multiple models simultaneously, a RAG document service, and two coding-agents. Models are provided as auto-scaled service pools, so they can scale up and down to meet demand.
v4.3.0 rolls up v4.2.1, v4.2.2, and v4.2.3, so these notes are complete for
anyone upgrading directly from v4.2.0; items marked (also in vX.Y.Z) already
shipped on those patches.
Upgrade notes.
- Catalog repository URIs are now restricted to HTTPS and publicly routable hosts by default. Other repositories will stop syncing and must be re-added over HTTPS. Operator-managed deployments can relax this with
spec.fuzzball.workflowCatalog.restrictions.allowPrivateAddresses, and tighten it withallowedHosts,maxCloneBytesandcloneTimeout.- Catalog entry IDs change on the first sync after upgrade. The ID is now derived from the catalog source as well as the entry’s
idfrontmatter. Saved links and scripts referring to catalog entries by UUID must be updated.- Ingress routes on the
kongandnginxclasses now require TLS and redirect plaintext callers with a 308. Other classes are unchanged.- A leftover pending-restore marker now fails the deployment instead of being silently ignored, which previously skipped the required pre-upgrade database dump. Delete the
pending-pg-restoreConfigMap to clear one.
Serving a large language model no longer requires writing a Fuzzfile or configuring vLLM from scratch. New workflow catalog entries start a model on cluster GPUs and hand back one stable base URL speaking the OpenAI API, so any OpenAI-compatible client can use it.
- vLLM entry. Serves a model from an autoscaled pool of vLLM replicas on NVIDIA or AMD GPUs. The model is downloaded once at workflow start and served offline from the workflow’s volume, and a replica runs on one node or spans several with tensor or expert parallelism. By default an in-workflow LiteLLM proxy holds the endpoint, load-balances across ready replicas, and drives scale-up from its own metrics.
- Eleven model presets. Individual catalog entries each configure the vLLM entry template with GPU count, context size, and serving arguments validated for each model. GPT-OSS 20B and 120B, Ministral 3 8B and 14B, Qwen3-Coder 30B and Next 80B, Nemotron 3 Nano and Super, Mistral Medium 3.5, Llama 4 Scout, and Gemma 4 31B are included in this release.
- LiteLLM Model Gateway. As an alternative to the in-workflow proxy, a single LiteLLM instance started from the workflow catalog can serve several models from one endpoint, so a client agent reaches all of them from a single configuration. It discovers pools through the Fuzzball endpoints API, registering the per-replica endpoints a pool publishes when its own proxy is turned off, and a request for a pool scaled to zero wakes it and succeeds on retry.
- Document retrieval. The
ragentry runs a retrieval service built on haiku.rag. It ingests PDF, DOCX, PPTX, HTML and plain text, indexes them for hybrid vector and full-text search in an embedded LanceDB on the workflow’s volume, and exposes retrieval as MCP tools on a workflow endpoint. Embeddings are served by a customer-provided OpenAI-compatible endpoint. - Coding agents.
OpenCodeandhermes-agentrun a coding agent in the cluster against a discovered, Fuzzball-hosted model or an OpenAI-compatible API outside the cluster.
Workflow service endpoints now work against an autoscaled pool, not just single-instance services. A pool’s endpoint keeps one stable URL for the life of the workflow and forwards each request to a replica ready at that moment, so callers never need to know how many replicas exist.
- Pool endpoints. An ordinary endpoint on a service declaring
autoscaler.replicasbecomes a pool endpoint. It must usetype: subdomainand reference a named port. Scaling up adds replicas to the rotation as they pass readiness, and scaling down removes one when its drain period begins, after in-flight requests finish. - Waking a pool idle at zero. A pool at
replicas.min: 0keeps its endpoint while idle. A request arriving then gets503withRetry-After: 120and starts one replica; retry after the cold start and it is served. The request must be authenticated, so ascope: publicendpoint never wakes its pool, and a burst starts one replica rather than one per request. Each wake shows up infuzzball workflow events. - Per-replica endpoints. Setting
per-replica: truegives every live replica its own endpoint alongside the pool’s, for callers doing their own load balancing.fuzzball workflow endpoints listadds aREPLICAcolumn andpoolEndpointIdis available with-o jsonor-o yaml. - Annotations. Any endpoint may declare key-value pair
annotationsto assist a client in discovering endpoints. - Multi-node replicas. A service may combine
multinodewithautoscaler.replicas, so each replica is a gang-scheduled group of ranks serving from rank 0.replicas.minandreplicas.maxcount groups, and pool capacity is total rank slots divided by group size.
fuzzball workflow delete WORKFLOW permanently removes a single finished,
failed, or canceled workflow and fuzzball workflow prune --older-than DURATION
removes them in bulk, with --limit and --dry-run. Who may delete is set per
role in the central config’s workflow.delete section, which takes
clusterAdmin, organizationOwner, groupOwner, and workflowOwner booleans.
All four default to false, so deletion is off until a cluster administrator
enables a role, and a caller may delete a workflow when any enabled role applies
to them. prune fails fast with a clear message when deletion is disabled
rather than reporting every candidate as skipped. On a federate deployment the
federate cluster and the orchestrate cluster owning the workflow each apply
their own setting.
Operators can also age workflows out on a schedule with a workflow.retention
block (spec.fuzzball.workflow.retention on the CRD), which takes enabled,
maxAge, interval, and batchSize parameters.
Deletion is permanent and takes the workflow’s stages, events, metrics,
placements, cost and egress records, volume attachments, and service endpoints
with it. It refuses while published volumes or live service DNS records remain,
and a finished workflow with no recorded end time is never prunable. Both
commands are audited and replicate to federate clusters, while the scheduled
sweep has no caller identity to audit as, so its service log is the record of
what it removed. A delete against a Federate cluster is forwarded to the owning
Orchestrate cluster. prune must run against the owning cluster directly.
A storage provisioner whose base path is per-node storage rather than the same
shared filesystem on every node declares local: true. Every volume on such a
provisioner exists on exactly one holding node and Fuzzball places every
stage that uses the volume there.
- Persistent volumes. Local provisioners now accept
createandaccesspolicies, and now support persistent volumes. The holding node is recorded when the volume is created, andfuzzball volume create --nodetargets a particular node when creating the volume. - Ephemeral volumes. An ephemeral volume on a local provisioner is created once rather than once per consuming node, allowing additional stages to access the volume later in the workflow.
- Single-node only. A local volume exists on one node, so a shape needing more than one is refused at submission. Multinode stages, concurrent task arrays, and autoscaled services that scale above a single node may not mount a local volume.
- Concurrent volumes. A stage mounting more than one local volume places them on a single node, and an ephemeral local volume mounted alongside a persistent local volume is created on the persistent volume’s holding node.
- Capacity-aware placement. Placement skips nodes without room for the
requested size, and targeting a node that lacks space fails naming the node
and its free space rather than reporting a bare capacity error. Sized
ephemeral volumes on a local provisioner setting
costPerGbHourare now billed, having previously been created on a path that recorded no usage. - Scanning.
fuzzball volume provisioner scanscans one node, selectable with--nodeand records the holding node of every volume it finds.
Organization compute policy grants shipped in v4.2.0 as an API with enforcement, but with no way to manage them short of calling the API. v4.3.0 adds CLI and web UI interfaces for managing access to node provisioners via the “compute policy.” Policies can be set at organization and group level.
Compute policy enforcement is still opt-in via computePolicy.enforce, which
defaults to false.
fuzzball workflow update WORKFLOW --priority N changes a queued workflow’s
priority in place. The value is re-stamped onto every unfinished stage, the
scheduler applies it on its next pass, and each stage records a
priority_updated entry in the activity log.
- AMD GPU health collection. ECC errors and overheating on AMD cards now feed the node reliability score, cordon, drain, and fault eviction alongside NVIDIA faults.
- Substrate nodes are no longer degraded for IO errors on disks backing no mounted filesystem or swap.
- Multi-node jobs fail as a group. Every rank now reports its own exit, so a rank that is killed, runs out of memory, or crashes fails the whole job, and the error names the rank, its node, and the exit status.
- Defaults-level resource flags inherit correctly.
exclusive,cpu.threads, andmemory.by-coreset underdefaultsnow apply to jobs and services declaring any other resource field. - Multinode services run with host networking and publish resolvable replica SRV records, fixing the startup bind failure and service discovery.
FB_GROUP_IDis now injected into job and service containers.
- New volumes default to mode 0750 instead of 0755, so they are no longer
world-readable or world-traversable without an explicit
--permissionsdefinition.
- Template reuse. A catalog entry can point at another entry’s template with
a
templatefield in itsmetadata.mdfrontmatter instead of carrying its owntemplate.yaml. References resolve within a repository and one level deep. - Workflow catalog search can now find entries by tag.
- Expanded access for the workflow token, including the ability to read job
logs, resolve cluster names, browse and render the workflow catalog, list and
create volumes, use the object cache, and explain scheduling through
fuzzball workflow whyandqueue show --mine. - Workflow containers receive a read-only bind mount of the node trust store, so the token works on clusters terminating TLS with a private CA. (also in v4.2.3)
- Docker Compose subdomain endpoints. Workflow service endpoint URLs now
resolve on a Compose deployment, with a default base domain
<host-ip>.nip.io. The stack can serve a host-only wildcard zone from the stack with--wildcard-dns, and--setup-dnsconfigures the local host to use the stack’s resolver. - Azure. New
--substrate-image-base-idand--substrate-image-nvidia-idadopt an already-registered gallery image. - GCP.
cluster gcp deploy|updatepulls images and Helm charts using the v-prefixed tags the source registry publishes. - AWS. Adds the
p5.4xlargeandp6-b300.48xlargeinstance types. - AWS SaaS deployments can give Federate its own DNS zone with
--federate-rootand--federate-root-hosted-zone-id.
- Fuzzball debug CLI.
fuzzball debugis now available to assist with cluster troubleshooting.debug dumpincludes collectors for metrics, JetStream, CRDs, ingress, certificates, storage, database health, and node health, plus a--namespaceflag for including additional underlying namespaces. - Lease-grant failures are bounded and explained. Failure to secure a
resource lease is reported in the workflow’s activity log and explained by
fuzzball workflow why. fuzzball report chargesandreport charge-summarynow correctly renders--output yaml.
- Streamlined sign-in experience. The page opens on identity-provider sign-in, looks up the organizations an address can use as you finish typing it, and remembers your last-used organization, listing it first and marked “Recent”. When the lookup finds nothing it says so and offers a local sign-in button; otherwise local sign-in stays available as a de-emphasized link. A local sign-in that fails for an address registered with an identity provider now offers a button back to the provider form.
- Volume permissions. The volume creation form has a Permissions control for the POSIX mode, and the detail panel shows an existing volume’s mode.
- Fuzzfile tab. A workflow’s detail page shows the submitted workflow definition as read-only YAML, alongside the Events tab.
- Cluster kind. The clusters list shows each cluster’s kind.
- Volume connections in the workflow editor. The graph draws a connection between a volume and the jobs and services that mount it.
- Workflow search matches email, ID, cluster ID, parent workflow ID, and error text as well as name.
- Secrets. The image URI field filters secret suggestions to those matching the registry’s domain, and the secrets table has a Private column showing whether an organization secret is private or shared.
- Services can select a node provisioner in the service form; the field is renamed from Node pool to Node provisioner.
- Cancelled workflows no longer display with error styling, and the workflow events and logs panels no longer show an unwanted horizontal scrollbar.
- Fixed unsafe SQL construction on the catalog, usage-stats, and billing endpoints by hardening how those queries are built.
- Fixed an issue where a catalog repository URI could make the server issue requests on the caller’s behalf (server-side request forgery). Repository URIs are now restricted to HTTPS and publicly routable hosts by default.
- Fixed an issue where object cache TTL configuration could reach across
organizations.
fuzzball org set-object-ttlandget-object-ttlnow correctly restrict to the organization scope. - Fixed an issue where the password change workflow could be bypassed. Setting a password during signup now requires the correct short-lived signup token.
- Fixed an open redirect on sign-in, where a login
returnUrlcontaining a backslash could send users to an external site or pack in additional request parameters. The return target is now validated before the redirect. - Fixed an issue where the endpoint proxy forwarded Fuzzball credentials to service endpoints. They are now stripped from the forwarded request.
- Fixed an issue where public ingress routes on the
kongandnginxclasses accepted plaintext traffic. TLS is now required on those classes.
- Fixed several issues in
cluster azure update: it attempted a PostgreSQL downgrade when--postgres-versionwas omitted, cleared configured substrate image IDs when--substrate-image-basewas omitted, and dropped the kubelet gallery Reader grant, which made provisioning fail withLinkedAuthorizationFailed. Also fixed fresh deployments, where the first context login failed withForbiddenByRbacandazure_filesvolume creation failed withAuthorizationPermissionMismatch.
- Fixed an issue where a Degraded Orchestrate or Federate stayed Degraded after a failed configuration check or deployment. It now retries automatically.
- Fixed an issue where the operator never re-ran a pending database restore, leaving a stale marker that silently skipped the required pre-upgrade database dump on later version upgrades.
- Fixed an issue where a federate database restore might erroneously consume an
orchestrate cluster’s dump, because both used the same
pg-migratedump path. The operator’s dump path is now qualified per cluster. - Fixed an operator panic when the ingress proxy type is
NodePortwithout both an http and a tls node port. The CRD now rejects that combination.
- Fixed an issue where Nginx held onto a resolved upstream address, so a service
recreated by
docker compose upkept receiving traffic at its old address. Nginx now re-resolves its upstreams. - Fixed an issue where a missing Keycloak realm or UI client silently broke Web UI login. Now such issues generate an error at startup.
- Fixed an issue where a Compose deployment advertised its default storage provisioner to nodes that cannot reach the shared Docker volume. The provisioner is now scoped to the compose-managed nodes on new deployments; an upgraded deployment keeps its original unsegmented provisioner, which has to be retired by hand.
- Fixed an issue where
cluster gcp updatedropped--substrate-image-project. The flag is now honored.
- Fixed an issue where the ECS task runner’s launch request exceeded the container-overrides size limit, causing stack updates on parameter-heavy deployments to be rejected. The runner now keeps the request under the limit (also in v4.2.2).
- Fixed an issue where parallel checkpointing of targeted EKS control-plane steps could persist a corrupt Pulumi snapshot. Those steps are now serialized (also in v4.2.1).
- Fixed several issues that left jobs queued indefinitely: whole-node-exclusive
jobs were placed onto nodes that already had other work, wedging the workflow
in READY forever; the scheduler could pick a provisioner definition holding
fewer nodes than the job needed, leaving it queued with no error; and
workflows could remain Pending after the scheduler had already finished their
allocations. Undersized definitions are no longer selected, and
fuzzball workflow whynow reportsINSUFFICIENT_RESOURCESnaming the definition and both node counts. - Fixed an issue where
workflow whymisreported a capacity-blocked service replica as a dependency wait. A replica the autoscaler is holding back, or one whose pool shrank beneath it, now reportsSERVICE_POOL_AT_CAPACITYnaming the provisioner definition and how many replicas the pool can hold. - Fixed three scheduler leadership issues: a leader that hit a JetStream timeout during a leader start following the first election kept leadership with node health tracking off, and now steps down so the start is retried on re-election; a demoted replica kept its service event consumer and took service health results from the new leader, and now releases it; and a crash during leader-election churn left the scheduler unavailable until the pod restarted.
- Fixed several cases where nodes and instances leaked. Cloud VMs left behind by a failed create or an interrupted provisioner are now recorded and deleted on all four backends; Azure instance deletion no longer reports success without deleting anything when a node reported a hostname that was not its VM name; and a node registration whose address has been reused by a different node is now purged, so the scheduler stops counting the departed node’s capacity as free.
- Fixed an issue where
fuzzball workflow start --priorityrejected a positive value from an organization owner. It now accepts one, matchingupdate. - Fixed an issue where
FB_JOB_IMAGEcarried auri:prefix. - Fixed an issue where cyclic
dependsOnworkflow graphs were accepted at submission and left unschedulable. They are now refused at submission.
- Fixed an issue where autoscaler scale-up released replicas before their image and other dependency stages had completed, causing failed replicas and pool flapping, and where a service with no dependents was stopped before its container finished starting.
- Fixed an issue where stopping or canceling a workflow running an autoscaled
pool left the pool’s autoscaler, SRV, and per-replica DNS records in place.
Endpoint rows are now removed when the owning workflow ends even if its
service never ran teardown, and getting an endpoint or minting an endpoint
token for an ended workflow now fails with
FAILED_PRECONDITION. - Fixed an issue where endpoints created or deleted during a page walk produced duplicated and skipped rows.
- Fixed an issue where ingress re-executed endpoint requests running past 60 seconds and buffered streamed responses, and where the embedded endpoint servers used by Compose and single-binary deployments cut responses, streams, and uploads off after 15 seconds.
- Fixed an issue where volume ownership defaulted to root on LDAP-federated deployments. UID now resolves from the caller’s JWT claims before the database, and GID from the selected account before the claim. (also in v4.2.3)
- Fixed three Web UI defects: the storage provisioner form rejected every local provisioner by always sending create and access policies, the workflow editor crashed while an environment variable was being added in the YAML view, and saving a template with an empty YAML body reported nothing rather than the validation error.
- Fixed an issue where an image build stage’s Web UI panel had no Logs tab, so
the post script’s output was unreachable there although the CLI could already
read it with
fuzzball workflow log. The panel now carries one. - Fixed an issue where the nodes list did not show its cluster column.
- Fixed an issue where the workflow editor dropped volumes, jobs, and services whose YAML map entry had an empty body.
- Fixed an issue where deleted federate organizations could not be recreated. Also fixed a missing default volume provisioner per registered cluster, fan-out errors for federated local users that did not name what failed, and cluster administrators being unable to leave an assumed session without re-authenticating.
- Fixed an issue where listing a large catalog over REST failed because the OpenAPI gateway’s gRPC receive limit was smaller than the one used by the CLI and backend services. The limits now match.
- Fixed an issue where a malformed
values.yamlaborted the whole catalog sync run. Sync now skips the offending entry and continues. - Fixed filtering the workflow catalog by tag.
- Fixed an unhandled exception when starting a workflow from a catalog template ID that does not exist.