Fuzzball v4.2.0 release notes
Fuzzball v4.2.0 is a large minor release built around cluster reliability and
multi-tenancy. Compute nodes now report hardware health, earn a reliability
score, and can be cordoned, drained, or evacuated by a configurable policy;
storage restrictions and per-organization default provisioners keep tenants
apart; and organization-owned compute policy grants gate which provisioner
definitions each tenant may place work on. Alongside that, workflow containers
receive a scoped API token and address, a new fuzzball mcp serve speaks
Model Context Protocol over stdio, GPU nodes advertise amd.com/gpu for
ROCm-capable workloads, and fuzzball report charges prices completed
workflows against per-organization price adjustments. The AWS upgrade path is
now async, heartbeat-monitored, and steps EKS through minor versions. v4.2.0
rolls up everything in v4.1.1 and v4.1.2, so these notes are complete for
anyone upgrading directly from v4.1.0; items marked (also in v4.1.1) or
(also in v4.1.2) were already available on those patches.
Upgrade notes.
- Update the
fuzzball-substrate-orchestrateextension on every compute node to match Orchestrate. Node heartbeats, hardware health collection, the substrate version reported byfuzzball node, and the aarch64 RPM split all require matching extension code on the node.fuzzball versionstill does not report the extension version, so skew is otherwise invisible; do not use a barednf update.- Node health defaults to observe. The health engine is on for every cluster but the policy mode is
observeandnodeHealth.evictOnFaultClassesis empty by default, so scores and events are recorded and nothing is cordoned, drained, or evicted without explicit configuration. Turn oncordon,drain, orreplaceper definition once the collected evidence makes sense for your fleet.- Multi-tenant storage restrictions are inferred from SaaS mode. On a single-tenant on-prem or private-cloud install nothing changes; on a SaaS-shaped install, hostpath and NFS provisioner creation is now restricted to cluster administrators and each organization receives its own default provisioner. Operators can set
storageRestrictionsanddefaultStorage.perOrganizationexplicitly to override the inferred defaults.- Compute policy grants apply cluster-wide with backfill. Existing provisioner definitions are backfilled into per-organization grants at upgrade so placement continues unchanged, but subsequent additions of a definition to a specific organization must go through the grant API.
- AWS upgrades return before the operator has finished reconciling. The CloudFormation call returns as soon as the stack update is accepted, and the operator reconciles asynchronously with a heartbeat. This removes the race against CloudFormation’s one-hour timeout on cross-major upgrades, which now also require typing the stack name to confirm. Watch operator status for readiness, not the CFN return code. Clusters upgrading from v3.x should also read the AWS Upgrade Path section under Enhancements – Kong IngressClass, EKS stepped upgrades, and Pulumi refresh self-heal are all v3.x -> v4.x blockers this release addresses.
- Removing a self-provisioned EFS provisioner tears down its filesystem. Two independent ownership confirmations (definition marked
self_provisionedplus a matching creation token andManagedBy=fuzzballtag on the filesystem) gate deletion of the filesystem, mount targets, and security group. Bring-your-own filesystems (afilesystemIdset at creation) are never removed.--output=jsonandnode show --output=yamlinclude zero-valued scalar fields for a stable schema. Scripts checking field presence (if .priority) must switch to explicit-value checks (if .priority > 0).- Workflow ingress and egress URIs are scheme-checked at submission. Only
http,https,file,s3,fuzzball,fb,hf, andhuggingfaceare accepted,filemust be exactlyfile://, and the undocumentedlocal://scheme is rejected. (also in v4.1.2)- AWS
--versionaccepts both bare andv-prefixed values. A leadingvis now stripped, so--version v4.2.0and--version 4.2.0both work. (also in v4.1.1)
Compute nodes now heartbeat, report hardware health, earn a per-node reliability score, and can be cordoned, drained, or replaced automatically by a configurable policy. The engine is advisory by default – it collects evidence, records events, and produces a score – and only takes action once an operator opts a definition into one of the enforcement modes.
- Substrate node heartbeat. Nodes report every 30 seconds by default,
and the control plane records a node as no longer reporting after it
misses three consecutive reports (roughly 90 seconds); both the cadence
and the miss threshold are configurable. Stale registrations are reaped
automatically and
fuzzball node removedeletes one immediately. - Host and GPU health collection. Substrate nodes surface machine-check exceptions, EDAC memory errors, disk SMART and IO errors, and thermal state; NVIDIA nodes additionally report ECC errors, Xid faults, and thermal or power throttling. Nodes without an NVIDIA GPU report GPU health as unavailable. A read-only block device with no filesystem is not treated as a disk fault.
- Reliability score and events. Each node carries a score derived from
active conditions and attributed job failures, with configurable
penalties and a full breakdown of every contribution. The 90-day
node-scoped event stream records health conditions, state transitions,
score threshold crossings, and attributed failures;
fuzzball node events NODEprints it with--since, repeatable--type,--latest, and standard pagination. - Policy engine and eviction. A
nodeHealthblock (cluster-wide or per provisioner definition) selects anobserve(default),cordon,drain, orreplaceaction tier, bounded bymaxUnavailablePercentand bycordonSuppressionMinutes(default 30), the window during which automation leaves a manually uncordoned node alone.nodeHealth.evictOnFaultClassesnames conditions that evacuate a node the instant they are raised; jobs may setpolicy.restart: on-node-faultto be restarted elsewhere. Only Fuzzball-provisioned nodes are replaced, and only after running work finishes. - Scheduling integration. The scheduler no longer places queued work
on a node whose reliability score has fallen below
drainBelow(45 by default); degraded nodes (95-46 by default) still receive work, and never-scored nodes are not gated. Work stranded on a node that goes offline is released immediately, and job endings caused by an unresponsive node now name the node and the specific fault. - CLI, UI, metrics, and webhooks. Node status, health state and score appear in the node API, the CLI, and the web UI (with a per-node reliability trend and a per-cluster dashboard rollup); per-node Prometheus series export score, health state, active conditions, and attributed failures, dropping when a node is decommissioned; and node events can be POSTed to configured endpoints as signed CloudEvents 1.0 JSON with HMAC-SHA256 signatures and per-endpoint type filters. Node health flows through Federate, so a federated cluster view shows the same health as the owning cluster.
- Provisioner-time abort for nodes that never register. The scheduler aborts allocations whose provisioned nodes fail to register within a configurable deadline, across every cloud backend, and cleans up the orphaned instances.
Storage provisioners can now be scoped per organization and restricted so
that privileged drivers stay in cluster-admin hands, with volume paths
isolated per allocation and a per-directory sub_path confining a
provisioner’s volumes to a subtree.
- Per-organization default provisioner. When
defaultStorage.perOrganizationis set, each new organization receives a default storage provisioner at creation. Drivers that cannot scope storage per tenant are refused. On operator-managed deployments bothdefaultStorage.perOrganizationand hostpath allowlists now take effect (the setting was previously written to a config file the storage services did not read). - Admin-only hostpath and NFS on shared clusters. With
storageRestrictionsenabled, hostpath and NFS provisioner creation is restricted to cluster administrators, with configurable path and server allowlists. Single-tenant deployments are unaffected. Both restrictions are inferred from SaaS mode by default; operators can set them explicitly. - Individual users in access policies. Provisioner access, create, and ephemeral policies may name individual users as well as groups. Users and groups are additive in both list and map forms.
- Per-allocation publish paths. Volume publish paths are scoped per allocation instead of by volume name, preventing collisions between organizations or jobs sharing a same-named volume. Substrate nodes reclaim stale mounts left by ended allocations, orchestrator crashes, or substrate restarts.
sub_pathfor hostpath and Filestore. Provisioners can confine their volumes to a subdirectory viasub_path, immutable once set. The DETAIL column now reports a status message for non-READY provisioners.- Self-provisioned EFS lifecycle. New EFS filesystems are created encrypted at rest and in Elastic throughput mode; two provisioners in one organization no longer collide on creation tokens or security group names; and removing a self-provisioned EFS provisioner now deletes its filesystem, mount targets, and security group (with two independent ownership confirmations, so bring-your-own filesystems are never removed).
- GCP Filestore fix. GCP Filestore provisioners can now be created; the driver had no case in the definition validator and every attempt was rejected as an unknown driver type.
- Housekeeping. AWS account limit and throttling errors are surfaced with the specific limit hit or a retry hint, soft-deleted provisioner rows are purged after 30 days once no volumes reference them (volume tombstones stay for accounting), and provisioner listing returns stable pages.
Organizations now own signed compute policy grants that gate which
provisioner definitions each tenant may place work on, enforced at
placement. Existing definitions are backfilled into grants at upgrade so
placement continues unchanged, and cross-cluster materialization is
handled by the admin AppendProvisionerDefinitions RPC so a Federate can
push a definition into every downstream cluster in one call. A companion
warning is emitted whenever the policy fallback places an internal or
system job because no definition’s policy admitted it.
fuzzball report charges and fuzzball report charge-summary compute
per-workflow charges for compute, volumes, and network egress, applying
per-organization price adjustments. charges workflow, charges volume,
charge-summary workflow, charge-summary list, and charge-summary aggregate give tenants a per-workflow view and cluster admins an
aggregated view across a time window.
Every workflow container now receives a workflow-scoped bearer token and
API address so scripts and tools inside the container can call the
Fuzzball API without a login step. FB_TOKEN is a token minted for the
workflow owner that expires 1 hour after the job’s walltime, or 7 days
for services and untimed jobs; FB_API_HOST is the gRPC host:port that
the CLI picks up automatically; FB_OPENAPI_URL is the REST endpoint
suitable for curl. An API endpoint exchanges a valid workflow token
for a fresh one so long-running services can renew their credential,
and workflow tokens can now list workflow service endpoints and mint
endpoint access tokens scoped to the workflow owner’s access.
fuzzball mcp serve runs a Model Context Protocol server over stdio,
exposing the cluster’s read operations as tools that LLM agents can
call directly. Read tools are always available; state changes,
arbitrary compute, and destructive operations are gated behind
--allow-write, --allow-exec, and --allow-destructive.
The GPU substrate image now bundles both the NVIDIA and AMD device
plugins and advertises amd.com/gpu on AMD hosts. Container images
are pulled architecture-matched per node, so a workflow can request
amd.com/gpu alongside nvidia.com/gpu and land on the right host.
AMD GPUs are inventory-only in v4.2 – Fuzzball reports their presence
and count, but GPU-specific health signals (ECC, Xid) are collected on
NVIDIA nodes only.
- AWS upgrade path. EKS upgrades now step one minor version at a
time and manage
coredns,kube-proxy,vpc-cni, and EBS CSI as EKS addons. AWS stack updates return while the operator reconciles asynchronously so upgrades no longer race CloudFormation’s one-hour limit, cross-major upgrades require typing the stack name to confirm, and updates print a heartbeat when CloudFormation is silent for over a minute. In-place v3.x to v4.x upgrades work again on Kong ingress by dropping the legacy IngressClass and applying the newer Kong CRDs before the chart rolls. The operator tolerates Pulumi stack refresh failures on update and delete, so rotated credentials self-heal. - Automated internal certificate renewal. In-cluster PostgreSQL and substrate mTLS certificates are renewed automatically by cert-manager with metrics and alerts on certificate expiry; ingress and JetStream certificates continue to reload in place.
- Optional monitoring stack. Setting
monitoring.enableddeploys Prometheus with Alertmanager, Grafana with a node reliability dashboard, and node health alert rules. Off by default. - Egress proxy and custom CAs.
spec.proxyandspec.trustedCACertson FuzzballOrchestrate and FuzzballFederate configure HTTP(S) egress through a proxy with additional trusted CAs. - Object cache config. The ObjectCache storage configuration now
uses a nested
storageblock (class,size,path); the flatstorageClassName,storageSize, andstoragePathfields are deprecated. Central config also gains a flag to disable object cache write-back for image pulls. - Substrate v2.6.0. Substrate package, container, and AMI updated to v2.6.0. The substrate bridge now bundles and serves orchestrate extension RPMs for both x86_64 and aarch64 nodes, and cloud-init selects the URL by node architecture. (also in v4.1.2)
- Docker Compose stack. A new
scripts/docker-stack.shinteractive installer deploys a single-host compose stack;docker-compose infoprints the browser endpoint. Substrate SIF object-cache uploads are routed through nginx by aliasingapi.DOMAINon the compose network. (also in v4.1.2) - Fixes. AWS deployments with a custom
--stack-nameare now visible tolist,delete, andcleanupvia aCreatedBy=FuzzballCLItag; Azure auto-tagging is applied to deployed resources and federated identity and resource-group bugs are fixed; the CLI context created after clusteraws/gcp/ocideploy is usable for login; GCP deployments work when substrate images are hosted outside the CIQ image project anddestroydeletes the deployment’s Secret Manager secrets; EKS addon version lookup no longer fails on stacks that disable default Pulumi providers; manual routes no longer prevent Pulumi update; workflow endpoints no longer intermittently return 401 from an ingress rule matching any hostname; orchestrator deployment uses the Recreate strategy so rolling updates cannot wedge scheduler leader election; deprecated Keycloak proxy settings are corrected; the operator ignores credential rotations after thepasswordRefmigration ran (it now reconciles them), and no longer leaves a stale Ready status when a pg-migrate dump failure aborts reconcile (it reports Degraded); provisioner ServiceAccount has permission to pull the substrate image; and Azure Files bootstrap creates a singledefaultprovisioner instead of separatepersistentandephemeralones.
fuzzball context validateruns a three-tier cluster readiness check with a sectioned report.fuzzball cluster {aws,gcp,azure,oci} validateruns the same operator-level Kubernetes checks against a cloud session.fuzzball debug reportcollects a support bundle covering the control plane and, optionally, compute nodes over the substrate channel (no SSH). Debug bundles now redact Kubernetes secret values: key names, type, and value length are kept; the value is not.
fuzzball organization create/deletefor cluster administrators (Create shows the generated password once; Delete removes the Keycloak realm unless--keep-realmis given) andfuzzball-admin organization list.- Volume API calls with a Keycloak-issued token lacking account selection now act as the caller’s personal group instead of failing with an internal error.
--require-update-passwordnow forces a password change on first login for new members and owners.
--valuesaccepts inline key=value pairs on workflow catalogstartandrender, repeatable and mixable with values files; later values override earlier ones.--clusteronvolume createtargets a specific Orchestrate cluster through Federate.- Substrate version in node output.
fuzzball node showandfuzzball node list -o yaml/jsoninclude the runningfuzzball-substrateversion. - Stable JSON schema.
--output=jsonandnode show --output=yamlinclude zero-valued scalar fields, so consumers should check for zero values rather than absent fields. - CLI workflow log backoff retries are interrupted for unknown stages.
fuzzball run --dry-runno longer emits the defer-start annotation, which caused the generated YAML to hang when submitted. Defer-start workflows now require exactly one job.- Homebrew tap. Fuzzball CLI Homebrew artifacts now build a single multi-architecture formula suitable for a tap; the macOS install instructions are corrected.
- Federate cluster admins can run
fuzzball queue showandqueue statswithout a permission error. --versionaccepts a leadingvonfuzzball cluster aws deployandupdate. (also in v4.1.1)
- Node health panel on the node detail view, plus a reliability score trend and per-cluster health rollup on the dashboard, with the nodes table filterable by health state.
- Node provisioner details in the node view.
- Workflow graph now links an image to the jobs and services that use it in every workflow state, and the service Connect button appears without a page refresh (staying visible but disabled until the service is ready).
- Workflow editor offers Running and Finished dependency states, hints the default shell used when a script has no shebang, and no longer produces blank volume references that blocked submission.
- Object uploads now use a folder-relative destination with the object name auto-filled from the selected file.
- Back button on the workflow detail view returns to the filtered workflow list.
- Federate central config. Central configuration now supports Federate clusters.
- Federate cluster registration verifies the remote cluster’s API
certificate against the
caCertfrom the registration config instead of system roots only. - Priority-aware federation routing. Workflow score reports the
scored cluster, its queue depths, and its capacity;
workflow whynames the receiving cluster; and preempted events carrypreempted_in_cluster_id. - Federate priority replication. Group, organization, and user scheduling priorities and group preemption lists now replicate from Federate deployments to Orchestrate clusters.
- Scoring uses segmentation and priority. A workflow’s score reflects whether the receiving cluster could actually run it.
- Scheduler robustness. Leadership is re-acquired after stepping
down when the in-memory leader stream is reset; a workflow service
or job that depends on another job’s RUNNING status starts once the
dependency transitions; CPU core allocation no longer picks
duplicate cores on multi-NUMA-socket nodes (AMD NPS, Intel SNC),
and
resources_allocatedandcontainer_startedevents now include projected and kernel-enforced cpusets; substrate nodes' bucket size is configurable viafuzzball.substrate.nodesBucketMaxByteswith new Prometheus series; the scheduler tiebreak no longer prefers never-provisioned definitions over ones holding warm nodes; and workflows carrying afuzzball.io/connect/<service>annotation no longer fail with “no suitable provision definition found”. - Provisioner and substrate fixes. CSI substrate
ListVolumespaginates all pages so large provisioners do not incorrectly prune volumes; provisioner scan re-scan reports already-tracked volumes; autoscaler drain deregisters replicas from DNS before stopping them; GCP and Azure provisioner definitions pick upnvidia.com/gpuannotations from nodes; OCI live pricing keeps shape cost estimates on published rates; Azure VM size seeding no longer times out on large SKU catalogs and no longer holds a transaction across the cloud catalog walk; dynamic node registration finds wildcard-expanded definitions when no instance type catalog is available; and substrate node annotation updates are continuously reconciled with the provisioner and scheduler. - Backported from the v4.1 patch line. Preemption fit math now
counts task-array victims running multiple ranks on one node
correctly, preempted task arrays no longer re-dispatch
out-of-range task ids or re-place ahead of the evictor, autoscaled
service replica pools are checked against pool capacity (min
rejected, max capped with a warning), a failed autoscaled replica
above
replicas.minis returned to the pending pool instead of erroring the whole workflow, a definition whose nodes are fully held by services no longer stalls scheduling for the whole definition (bounded-walltime work backfills), uncordoned nodes return to the scheduler immediately, provisioner definition policies apply to internal jobs (with fallback to unrestricted placement), and non-admin users can create volumes from provisioners in the web UI. (also in v4.1.1) The v4.1.1 host-port regression on services is reversed: non-pool services publish on the declared port again and autoscaled pools keep random assignments. (also in v4.1.2)
- Autoscaler validation. Workflow submission rejects autoscaler metrics queries with invalid PromQL or with rate-style windows shorter than twice the metrics collection interval.
- URI validation. Workflow validation now rejects unsupported
ingress and egress URI schemes; the undocumented
local://scheme, which previously could hangfuzzball-substrate, is no longer accepted. (also in v4.1.2) - Multi-node helper. A new helper script simplifies use of the multinode generic implementation, distributed as a workflow example.
- Volume ingress from previously running jobs. Segment-aware dynamic node provisioning ensures storage operations land on nodes that can reach the target segment.
- Deprecation.
WorkflowVolumeState.driver_idis deprecated in favor of a provisioner reference. - Autoscaler scale-down drain. Autoscaled replicas are deregistered from DNS before being stopped, over a configurable drain period.
- JWT issuer URL path validation. The full issuer URL path is validated before JWKS fetch, preventing unauthenticated internal SSRF and topology disclosure.
- Path traversal remediation. File handling in the deployment
path uses the
os.RootAPI. - IMDS restricted to root. Access to the instance metadata service on cloud instances is restricted to the root user.
- Multi-GB object-cache blob transfers no longer abort mid-stream: the Orchestrate HTTP server’s per-request read and write deadlines are cleared for streaming transfers, so large container image layers and other multi-second blob transfers complete instead of being closed at 15 seconds. (also in v4.1.2)
- Kernels without XFRM support.
RouteGetSourceAddressno longer fails on such kernels; the netlink handle now subscribes only to the route family it needs. (also in v4.1.2) - Substrate extension version. The extension no longer reports a hardcoded version instead of the build version.
- Volume publish retries. Volume publish now retries transient CSI errors, fixing intermittent “no such file or directory” failures on shared filesystems.
- Hostpath volume unmount at filesystem root. The refusal that blocked ephemeral volume creation and leaked job volume mounts is fixed.
- Provisioner list order. Storage provisioner listing no longer returns unstable pages that could repeat one provisioner and omit another.
- Autoscaler tear-down race. Jobs no longer fail with “Job stopped responding” when the autoscaler tears down a node the scheduler was placing onto.
- Lease consumer ack wait increased to 5 minutes so substrate nodes tolerate short interruptions.