Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Fuzzball v4.2.0 release notes

Fuzzball v4.2.0 is a large minor release built around cluster reliability and multi-tenancy. Compute nodes now report hardware health, earn a reliability score, and can be cordoned, drained, or evacuated by a configurable policy; storage restrictions and per-organization default provisioners keep tenants apart; and organization-owned compute policy grants gate which provisioner definitions each tenant may place work on. Alongside that, workflow containers receive a scoped API token and address, a new fuzzball mcp serve speaks Model Context Protocol over stdio, GPU nodes advertise amd.com/gpu for ROCm-capable workloads, and fuzzball report charges prices completed workflows against per-organization price adjustments. The AWS upgrade path is now async, heartbeat-monitored, and steps EKS through minor versions. v4.2.0 rolls up everything in v4.1.1 and v4.1.2, so these notes are complete for anyone upgrading directly from v4.1.0; items marked (also in v4.1.1) or (also in v4.1.2) were already available on those patches.

Upgrade notes.

  • Update the fuzzball-substrate-orchestrate extension on every compute node to match Orchestrate. Node heartbeats, hardware health collection, the substrate version reported by fuzzball node, and the aarch64 RPM split all require matching extension code on the node. fuzzball version still does not report the extension version, so skew is otherwise invisible; do not use a bare dnf update.
  • Node health defaults to observe. The health engine is on for every cluster but the policy mode is observe and nodeHealth.evictOnFaultClasses is empty by default, so scores and events are recorded and nothing is cordoned, drained, or evicted without explicit configuration. Turn on cordon, drain, or replace per definition once the collected evidence makes sense for your fleet.
  • Multi-tenant storage restrictions are inferred from SaaS mode. On a single-tenant on-prem or private-cloud install nothing changes; on a SaaS-shaped install, hostpath and NFS provisioner creation is now restricted to cluster administrators and each organization receives its own default provisioner. Operators can set storageRestrictions and defaultStorage.perOrganization explicitly to override the inferred defaults.
  • Compute policy grants apply cluster-wide with backfill. Existing provisioner definitions are backfilled into per-organization grants at upgrade so placement continues unchanged, but subsequent additions of a definition to a specific organization must go through the grant API.
  • AWS upgrades return before the operator has finished reconciling. The CloudFormation call returns as soon as the stack update is accepted, and the operator reconciles asynchronously with a heartbeat. This removes the race against CloudFormation’s one-hour timeout on cross-major upgrades, which now also require typing the stack name to confirm. Watch operator status for readiness, not the CFN return code. Clusters upgrading from v3.x should also read the AWS Upgrade Path section under Enhancements – Kong IngressClass, EKS stepped upgrades, and Pulumi refresh self-heal are all v3.x -> v4.x blockers this release addresses.
  • Removing a self-provisioned EFS provisioner tears down its filesystem. Two independent ownership confirmations (definition marked self_provisioned plus a matching creation token and ManagedBy=fuzzball tag on the filesystem) gate deletion of the filesystem, mount targets, and security group. Bring-your-own filesystems (a filesystemId set at creation) are never removed.
  • --output=json and node show --output=yaml include zero-valued scalar fields for a stable schema. Scripts checking field presence (if .priority) must switch to explicit-value checks (if .priority > 0).
  • Workflow ingress and egress URIs are scheme-checked at submission. Only http, https, file, s3, fuzzball, fb, hf, and huggingface are accepted, file must be exactly file://, and the undocumented local:// scheme is rejected. (also in v4.1.2)
  • AWS --version accepts both bare and v-prefixed values. A leading v is now stripped, so --version v4.2.0 and --version 4.2.0 both work. (also in v4.1.1)

New Features

Node Health and Fault-Tolerant Scheduling

Compute nodes now heartbeat, report hardware health, earn a per-node reliability score, and can be cordoned, drained, or replaced automatically by a configurable policy. The engine is advisory by default – it collects evidence, records events, and produces a score – and only takes action once an operator opts a definition into one of the enforcement modes.

  • Substrate node heartbeat. Nodes report every 30 seconds by default, and the control plane records a node as no longer reporting after it misses three consecutive reports (roughly 90 seconds); both the cadence and the miss threshold are configurable. Stale registrations are reaped automatically and fuzzball node remove deletes one immediately.
  • Host and GPU health collection. Substrate nodes surface machine-check exceptions, EDAC memory errors, disk SMART and IO errors, and thermal state; NVIDIA nodes additionally report ECC errors, Xid faults, and thermal or power throttling. Nodes without an NVIDIA GPU report GPU health as unavailable. A read-only block device with no filesystem is not treated as a disk fault.
  • Reliability score and events. Each node carries a score derived from active conditions and attributed job failures, with configurable penalties and a full breakdown of every contribution. The 90-day node-scoped event stream records health conditions, state transitions, score threshold crossings, and attributed failures; fuzzball node events NODE prints it with --since, repeatable --type, --latest, and standard pagination.
  • Policy engine and eviction. A nodeHealth block (cluster-wide or per provisioner definition) selects an observe (default), cordon, drain, or replace action tier, bounded by maxUnavailablePercent and by cordonSuppressionMinutes (default 30), the window during which automation leaves a manually uncordoned node alone. nodeHealth.evictOnFaultClasses names conditions that evacuate a node the instant they are raised; jobs may set policy.restart: on-node-fault to be restarted elsewhere. Only Fuzzball-provisioned nodes are replaced, and only after running work finishes.
  • Scheduling integration. The scheduler no longer places queued work on a node whose reliability score has fallen below drainBelow (45 by default); degraded nodes (95-46 by default) still receive work, and never-scored nodes are not gated. Work stranded on a node that goes offline is released immediately, and job endings caused by an unresponsive node now name the node and the specific fault.
  • CLI, UI, metrics, and webhooks. Node status, health state and score appear in the node API, the CLI, and the web UI (with a per-node reliability trend and a per-cluster dashboard rollup); per-node Prometheus series export score, health state, active conditions, and attributed failures, dropping when a node is decommissioned; and node events can be POSTed to configured endpoints as signed CloudEvents 1.0 JSON with HMAC-SHA256 signatures and per-endpoint type filters. Node health flows through Federate, so a federated cluster view shows the same health as the owning cluster.
  • Provisioner-time abort for nodes that never register. The scheduler aborts allocations whose provisioned nodes fail to register within a configurable deadline, across every cloud backend, and cleans up the orphaned instances.

Multi-Tenant Storage

Storage provisioners can now be scoped per organization and restricted so that privileged drivers stay in cluster-admin hands, with volume paths isolated per allocation and a per-directory sub_path confining a provisioner’s volumes to a subtree.

  • Per-organization default provisioner. When defaultStorage.perOrganization is set, each new organization receives a default storage provisioner at creation. Drivers that cannot scope storage per tenant are refused. On operator-managed deployments both defaultStorage.perOrganization and hostpath allowlists now take effect (the setting was previously written to a config file the storage services did not read).
  • Admin-only hostpath and NFS on shared clusters. With storageRestrictions enabled, hostpath and NFS provisioner creation is restricted to cluster administrators, with configurable path and server allowlists. Single-tenant deployments are unaffected. Both restrictions are inferred from SaaS mode by default; operators can set them explicitly.
  • Individual users in access policies. Provisioner access, create, and ephemeral policies may name individual users as well as groups. Users and groups are additive in both list and map forms.
  • Per-allocation publish paths. Volume publish paths are scoped per allocation instead of by volume name, preventing collisions between organizations or jobs sharing a same-named volume. Substrate nodes reclaim stale mounts left by ended allocations, orchestrator crashes, or substrate restarts.
  • sub_path for hostpath and Filestore. Provisioners can confine their volumes to a subdirectory via sub_path, immutable once set. The DETAIL column now reports a status message for non-READY provisioners.
  • Self-provisioned EFS lifecycle. New EFS filesystems are created encrypted at rest and in Elastic throughput mode; two provisioners in one organization no longer collide on creation tokens or security group names; and removing a self-provisioned EFS provisioner now deletes its filesystem, mount targets, and security group (with two independent ownership confirmations, so bring-your-own filesystems are never removed).
  • GCP Filestore fix. GCP Filestore provisioners can now be created; the driver had no case in the definition validator and every attempt was rejected as an unknown driver type.
  • Housekeeping. AWS account limit and throttling errors are surfaced with the specific limit hit or a retry hint, soft-deleted provisioner rows are purged after 30 days once no volumes reference them (volume tombstones stay for accounting), and provisioner listing returns stable pages.

Compute Policy Grants

Organizations now own signed compute policy grants that gate which provisioner definitions each tenant may place work on, enforced at placement. Existing definitions are backfilled into grants at upgrade so placement continues unchanged, and cross-cluster materialization is handled by the admin AppendProvisionerDefinitions RPC so a Federate can push a definition into every downstream cluster in one call. A companion warning is emitted whenever the policy fallback places an internal or system job because no definition’s policy admitted it.

Resource Accounting

fuzzball report charges and fuzzball report charge-summary compute per-workflow charges for compute, volumes, and network egress, applying per-organization price adjustments. charges workflow, charges volume, charge-summary workflow, charge-summary list, and charge-summary aggregate give tenants a per-workflow view and cluster admins an aggregated view across a time window.

In-Workflow API Access

Every workflow container now receives a workflow-scoped bearer token and API address so scripts and tools inside the container can call the Fuzzball API without a login step. FB_TOKEN is a token minted for the workflow owner that expires 1 hour after the job’s walltime, or 7 days for services and untimed jobs; FB_API_HOST is the gRPC host:port that the CLI picks up automatically; FB_OPENAPI_URL is the REST endpoint suitable for curl. An API endpoint exchanges a valid workflow token for a fresh one so long-running services can renew their credential, and workflow tokens can now list workflow service endpoints and mint endpoint access tokens scoped to the workflow owner’s access.

Model Context Protocol Server

fuzzball mcp serve runs a Model Context Protocol server over stdio, exposing the cluster’s read operations as tools that LLM agents can call directly. Read tools are always available; state changes, arbitrary compute, and destructive operations are gated behind --allow-write, --allow-exec, and --allow-destructive.

AMD (ROCm) GPU Support

The GPU substrate image now bundles both the NVIDIA and AMD device plugins and advertises amd.com/gpu on AMD hosts. Container images are pulled architecture-matched per node, so a workflow can request amd.com/gpu alongside nvidia.com/gpu and land on the right host. AMD GPUs are inventory-only in v4.2 – Fuzzball reports their presence and count, but GPU-specific health signals (ECC, Xid) are collected on NVIDIA nodes only.

Enhancements

Deployment and Infrastructure

  • AWS upgrade path. EKS upgrades now step one minor version at a time and manage coredns, kube-proxy, vpc-cni, and EBS CSI as EKS addons. AWS stack updates return while the operator reconciles asynchronously so upgrades no longer race CloudFormation’s one-hour limit, cross-major upgrades require typing the stack name to confirm, and updates print a heartbeat when CloudFormation is silent for over a minute. In-place v3.x to v4.x upgrades work again on Kong ingress by dropping the legacy IngressClass and applying the newer Kong CRDs before the chart rolls. The operator tolerates Pulumi stack refresh failures on update and delete, so rotated credentials self-heal.
  • Automated internal certificate renewal. In-cluster PostgreSQL and substrate mTLS certificates are renewed automatically by cert-manager with metrics and alerts on certificate expiry; ingress and JetStream certificates continue to reload in place.
  • Optional monitoring stack. Setting monitoring.enabled deploys Prometheus with Alertmanager, Grafana with a node reliability dashboard, and node health alert rules. Off by default.
  • Egress proxy and custom CAs. spec.proxy and spec.trustedCACerts on FuzzballOrchestrate and FuzzballFederate configure HTTP(S) egress through a proxy with additional trusted CAs.
  • Object cache config. The ObjectCache storage configuration now uses a nested storage block (class, size, path); the flat storageClassName, storageSize, and storagePath fields are deprecated. Central config also gains a flag to disable object cache write-back for image pulls.
  • Substrate v2.6.0. Substrate package, container, and AMI updated to v2.6.0. The substrate bridge now bundles and serves orchestrate extension RPMs for both x86_64 and aarch64 nodes, and cloud-init selects the URL by node architecture. (also in v4.1.2)
  • Docker Compose stack. A new scripts/docker-stack.sh interactive installer deploys a single-host compose stack; docker-compose info prints the browser endpoint. Substrate SIF object-cache uploads are routed through nginx by aliasing api.DOMAIN on the compose network. (also in v4.1.2)
  • Fixes. AWS deployments with a custom --stack-name are now visible to list, delete, and cleanup via a CreatedBy=FuzzballCLI tag; Azure auto-tagging is applied to deployed resources and federated identity and resource-group bugs are fixed; the CLI context created after cluster aws/gcp/oci deploy is usable for login; GCP deployments work when substrate images are hosted outside the CIQ image project and destroy deletes the deployment’s Secret Manager secrets; EKS addon version lookup no longer fails on stacks that disable default Pulumi providers; manual routes no longer prevent Pulumi update; workflow endpoints no longer intermittently return 401 from an ingress rule matching any hostname; orchestrator deployment uses the Recreate strategy so rolling updates cannot wedge scheduler leader election; deprecated Keycloak proxy settings are corrected; the operator ignores credential rotations after the passwordRef migration ran (it now reconciles them), and no longer leaves a stale Ready status when a pg-migrate dump failure aborts reconcile (it reports Degraded); provisioner ServiceAccount has permission to pull the substrate image; and Azure Files bootstrap creates a single default provisioner instead of separate persistent and ephemeral ones.

Cluster Validation and Diagnostics

  • fuzzball context validate runs a three-tier cluster readiness check with a sectioned report.
  • fuzzball cluster {aws,gcp,azure,oci} validate runs the same operator-level Kubernetes checks against a cloud session.
  • fuzzball debug report collects a support bundle covering the control plane and, optionally, compute nodes over the substrate channel (no SSH). Debug bundles now redact Kubernetes secret values: key names, type, and value length are kept; the value is not.

Organization Management

  • fuzzball organization create/delete for cluster administrators (Create shows the generated password once; Delete removes the Keycloak realm unless --keep-realm is given) and fuzzball-admin organization list.
  • Volume API calls with a Keycloak-issued token lacking account selection now act as the caller’s personal group instead of failing with an internal error.
  • --require-update-password now forces a password change on first login for new members and owners.

Command-Line Interface

  • --values accepts inline key=value pairs on workflow catalog start and render, repeatable and mixable with values files; later values override earlier ones.
  • --cluster on volume create targets a specific Orchestrate cluster through Federate.
  • Substrate version in node output. fuzzball node show and fuzzball node list -o yaml/json include the running fuzzball-substrate version.
  • Stable JSON schema. --output=json and node show --output=yaml include zero-valued scalar fields, so consumers should check for zero values rather than absent fields.
  • CLI workflow log backoff retries are interrupted for unknown stages.
  • fuzzball run --dry-run no longer emits the defer-start annotation, which caused the generated YAML to hang when submitted. Defer-start workflows now require exactly one job.
  • Homebrew tap. Fuzzball CLI Homebrew artifacts now build a single multi-architecture formula suitable for a tap; the macOS install instructions are corrected.
  • Federate cluster admins can run fuzzball queue show and queue stats without a permission error.
  • --version accepts a leading v on fuzzball cluster aws deploy and update. (also in v4.1.1)

Web UI

  • Node health panel on the node detail view, plus a reliability score trend and per-cluster health rollup on the dashboard, with the nodes table filterable by health state.
  • Node provisioner details in the node view.
  • Workflow graph now links an image to the jobs and services that use it in every workflow state, and the service Connect button appears without a page refresh (staying visible but disabled until the service is ready).
  • Workflow editor offers Running and Finished dependency states, hints the default shell used when a script has no shebang, and no longer produces blank volume references that blocked submission.
  • Object uploads now use a folder-relative destination with the object name auto-filled from the selected file.
  • Back button on the workflow detail view returns to the filtered workflow list.

Federation and Scheduling

  • Federate central config. Central configuration now supports Federate clusters.
  • Federate cluster registration verifies the remote cluster’s API certificate against the caCert from the registration config instead of system roots only.
  • Priority-aware federation routing. Workflow score reports the scored cluster, its queue depths, and its capacity; workflow why names the receiving cluster; and preempted events carry preempted_in_cluster_id.
  • Federate priority replication. Group, organization, and user scheduling priorities and group preemption lists now replicate from Federate deployments to Orchestrate clusters.
  • Scoring uses segmentation and priority. A workflow’s score reflects whether the receiving cluster could actually run it.
  • Scheduler robustness. Leadership is re-acquired after stepping down when the in-memory leader stream is reset; a workflow service or job that depends on another job’s RUNNING status starts once the dependency transitions; CPU core allocation no longer picks duplicate cores on multi-NUMA-socket nodes (AMD NPS, Intel SNC), and resources_allocated and container_started events now include projected and kernel-enforced cpusets; substrate nodes' bucket size is configurable via fuzzball.substrate.nodesBucketMaxBytes with new Prometheus series; the scheduler tiebreak no longer prefers never-provisioned definitions over ones holding warm nodes; and workflows carrying a fuzzball.io/connect/<service> annotation no longer fail with “no suitable provision definition found”.
  • Provisioner and substrate fixes. CSI substrate ListVolumes paginates all pages so large provisioners do not incorrectly prune volumes; provisioner scan re-scan reports already-tracked volumes; autoscaler drain deregisters replicas from DNS before stopping them; GCP and Azure provisioner definitions pick up nvidia.com/gpu annotations from nodes; OCI live pricing keeps shape cost estimates on published rates; Azure VM size seeding no longer times out on large SKU catalogs and no longer holds a transaction across the cloud catalog walk; dynamic node registration finds wildcard-expanded definitions when no instance type catalog is available; and substrate node annotation updates are continuously reconciled with the provisioner and scheduler.
  • Backported from the v4.1 patch line. Preemption fit math now counts task-array victims running multiple ranks on one node correctly, preempted task arrays no longer re-dispatch out-of-range task ids or re-place ahead of the evictor, autoscaled service replica pools are checked against pool capacity (min rejected, max capped with a warning), a failed autoscaled replica above replicas.min is returned to the pending pool instead of erroring the whole workflow, a definition whose nodes are fully held by services no longer stalls scheduling for the whole definition (bounded-walltime work backfills), uncordoned nodes return to the scheduler immediately, provisioner definition policies apply to internal jobs (with fallback to unrestricted placement), and non-admin users can create volumes from provisioners in the web UI. (also in v4.1.1) The v4.1.1 host-port regression on services is reversed: non-pool services publish on the declared port again and autoscaled pools keep random assignments. (also in v4.1.2)

Workflow Definitions and Data Movement

  • Autoscaler validation. Workflow submission rejects autoscaler metrics queries with invalid PromQL or with rate-style windows shorter than twice the metrics collection interval.
  • URI validation. Workflow validation now rejects unsupported ingress and egress URI schemes; the undocumented local:// scheme, which previously could hang fuzzball-substrate, is no longer accepted. (also in v4.1.2)
  • Multi-node helper. A new helper script simplifies use of the multinode generic implementation, distributed as a workflow example.
  • Volume ingress from previously running jobs. Segment-aware dynamic node provisioning ensures storage operations land on nodes that can reach the target segment.
  • Deprecation. WorkflowVolumeState.driver_id is deprecated in favor of a provisioner reference.
  • Autoscaler scale-down drain. Autoscaled replicas are deregistered from DNS before being stopped, over a configurable drain period.

Security

  • JWT issuer URL path validation. The full issuer URL path is validated before JWKS fetch, preventing unauthenticated internal SSRF and topology disclosure.
  • Path traversal remediation. File handling in the deployment path uses the os.Root API.
  • IMDS restricted to root. Access to the instance metadata service on cloud instances is restricted to the root user.

Bug Fixes and Stability

Workflow Engine

  • Multi-GB object-cache blob transfers no longer abort mid-stream: the Orchestrate HTTP server’s per-request read and write deadlines are cleared for streaming transfers, so large container image layers and other multi-second blob transfers complete instead of being closed at 15 seconds. (also in v4.1.2)
  • Kernels without XFRM support. RouteGetSourceAddress no longer fails on such kernels; the netlink handle now subscribes only to the route family it needs. (also in v4.1.2)
  • Substrate extension version. The extension no longer reports a hardcoded version instead of the build version.
  • Volume publish retries. Volume publish now retries transient CSI errors, fixing intermittent “no such file or directory” failures on shared filesystems.

Storage

  • Hostpath volume unmount at filesystem root. The refusal that blocked ephemeral volume creation and leaked job volume mounts is fixed.
  • Provisioner list order. Storage provisioner listing no longer returns unstable pages that could repeat one provisioner and omit another.

Autoscaler and Scheduler

  • Autoscaler tear-down race. Jobs no longer fail with “Job stopped responding” when the autoscaler tears down a node the scheduler was placing onto.

JetStream and Substrate

  • Lease consumer ack wait increased to 5 minutes so substrate nodes tolerate short interruptions.