Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Fuzzball v4.0.0 release notes

Fuzzball v4.0.0 is a major release. It rearchitects storage around a simpler two-tier provisioner-and-volume model, promotes the external API to v4 as a coordinated breaking change, and ships a substantially reworked command-line interface built around noun-verb commands, natural-name identifiers, and consistent list output. This release also adds Oracle Cloud Infrastructure (OCI) as a fully supported cloud target, introduces the Fuzzball Object Cache for sharing files and container images, brings GCP Filestore, Azure Files, and OCI File Storage to parity with AWS EFS, and delivers a rebuilt web UI along with self-service organization signup.

v4.0.0 is a coordinated breaking change. The external API base path moves from /v3 to /v4 and gRPC services move to the fuzzball.api.v4 package, so the CLI and server must be upgraded together — a v3 CLI cannot talk to a v4 server and vice versa. Existing v1 workflow definitions continue to run and are upgraded automatically (see Workflows).

Storage rearchitecture

Fuzzball storage moves from the previous three-tier model (Driver, Class, Volume) to a two-tier model built on storage provisioners and volumes. A storage provisioner pairs a backend driver with policy; volumes are created against a provisioner and managed through their own lifecycle.

  • fuzzball volume manages volumes directly: create, list, info, update, enable, disable, and delete.
  • fuzzball volume provisioner manages provisioners: add, edit, remove, list, list-drivers, info, scan, and migrate.
  • Built-in driver types are NFS, hostpath, Kubernetes PVC, AWS EFS, Azure Files, GCP Filestore, and OCI File Storage.
  • The legacy class-based commands (storage class, storage driver) are removed; storage is now managed entirely under volume provisioner.
  • fuzzball volume provisioner scan detects volumes that have been removed from backing storage and cleans up the stale database records. Use --dry-run to preview which volumes would be removed before committing.
  • Migration from a v3 cluster is supported. The operator storage-migrate migrate command auto-detects the cloud storage class and accepts backend-specific flags for AWS EFS, Azure, GCP Filestore, and OCI File Storage. The migration connects as the per-service fuzzball database role that owns the storage schema, so it works on external Postgres deployments (such as GCP Cloud SQL) where the bootstrap admin role lacks privileges on those tables.

V4 API

The external API is promoted to v4. The OpenAPI schema now reports version 4.0, the REST base path moves from /v3 to /v4, and all gRPC service paths use the fuzzball.api.v4 package. This is a coordinated breaking change: CLI and server versions must match.

  • gRPC-gateway no longer silently drops unknown JSON fields. Previously, service endpoint port-name fields sent in camelCase were discarded before validation, causing requests that looked correct to fail; those fields are now preserved and validated.

Workflows

  • Workflow v4 with automatic upgrade. Workflow definitions gain a v4 format. Existing v1 Fuzzfiles continue to work and are upgraded automatically on submission; fuzzball workflow upgrade converts a v1 Fuzzfile to v4 on demand.
  • Top-level and service defaults. Workflow definitions support defaults.env, defaults.mounts, and defaults.resource at the top level, plus a defaults.service block. Top-level defaults propagate to both jobs and services unless overridden by a more specific block or by the job/service itself.
  • Host networking and multi-node services. Workflow services can now request host networking and a multi-node service type.
  • fuzzball workflow connect connects to a service endpoint of a running workflow, either by opening a web browser or by executing a client locally.
  • Jobs may now declare a dependency on another job that is in the running state, not only on completed jobs.

Oracle Cloud Infrastructure (OCI)

OCI joins AWS, GCP, and Azure as a supported cloud target, with both deployment and workflow scheduling.

  • fuzzball cluster oci manages the full deployment lifecycle: deploy, update, destroy, status, list, info, logs, cleanup, and run-pulumi. A deployment provisions an OKE cluster, PostgreSQL, file storage, DNS, and Let’s Encrypt wildcard certificates via an OCI DNS webhook, with optional Depot registry integration and parallel resource cleanup.
  • OCI provisioner backend. Workflows can be scheduled onto OCI compute instances. Flex shapes are supported with a shape:ocpus:memory_gb encoding (for example VM.Standard.E4.Flex:4:64) that produces accurate resource metadata and per-unit cost estimates from the OCI pricing API, with hourly live pricing updates. GPU shapes (A10) are supported.
  • Completed-job logs on OCI are served from a native OCI Object Storage backend that authenticates through OKE Workload Identity, matching the credential-free experience already available on GCP.

Object Cache

The Fuzzball Object Cache is a built-in store for sharing files, data, and container images across workflows and clusters, with TTL-based lifecycle management.

  • fuzzball object manages cached objects: put, get, list, delete, and describe. Objects are addressed with fb:// URIs and organized into two namespaces — fb://group/... for objects shared within a group and fb://user/... for a user’s private objects. Objects placed in the user namespace are keyed to the user and no longer appear in the group listing, and vice versa. Uploads support a --ttl and recursive directory upload, and references may be pinned to a digest (fb://group/image.sif@sha256:...).
  • Automatic SIF image caching. When a workflow pulls an external container image (for example via docker:// or oras://), the converted SIF image is cached to the object cache automatically. Subsequent pulls of the same image are served from the cache instead of the external registry.
  • The web UI includes an Object Cache browser for uploading, searching, and inspecting objects across the group and user namespaces.

Multi-cloud storage drivers

The cloud file-storage drivers are brought to parity with AWS EFS.

  • GCP Filestore. A Filestore driver runs in directory mode (one shared share, one subdirectory per volume) through the GKE Filestore CSI driver. The operator provisions the Filestore PV/PVC and mounts it into the orchestrator and storage deployments when the GCP provisioner is enabled.
  • Azure Files NFS. Storage provisioners can use Azure file shares as volume backends, with each volume as a subdirectory in the share mounted over NFSv3. Requires an Azure Storage account with the NFS protocol enabled on the share.
  • OCI File Storage (FSS). A File Storage driver with parity to AWS EFS supports both bring-your-own mode (existing filesystem and mount target) and self-provisioned mode (the driver creates the filesystem and mount target). Per-volume POSIX ownership is enforced server-side via an OCI FSS identity squash policy. Volumes mount over NFSv3.
  • NFS sync mode. EFS and Azure Files provisioners accept an nfs_sync_mode option (auto, sync, or async) controlling how NFS volumes are mounted. The default auto uses sync for task-array workloads — preventing write loss across ranks — and async everywhere else for maximum write performance. The option appears in fuzzball volume provisioner list-drivers output.

Command-line interface

The CLI surface is substantially reworked for v4. The themes are noun-verb command structure, natural-name identifiers, consistent and complete list output, and discoverable filtering and sorting. Wherever a command was renamed or moved, the previous path is retained as a hidden, deprecated alias that shares behavior with the canonical command, so existing scripts keep working and print a one-line deprecation notice on first use.

Command structure

  • Deprecated top-level subtrees are replaced: fuzzball account → fuzzball group, fuzzball application → fuzzball workflow catalog, fuzzball cloud → fuzzball cluster, and the entire fuzzball admin subtree moves to top-level equivalents (for example fuzzball context, fuzzball storage, and fuzzball node provisioner).
  • Membership and ownership commands move to noun-verb form: group member {add,list,remove,update}, organization member {add,list,remove,update}, and workflow owner {add,list,remove}. On groups and organizations, --owner on member add/member update promotes or demotes a member.
  • Aggregate reports move into a new fuzzball report scope, split by audience: report group ... for the caller’s selected group (group-owner only) and report cluster ... for cluster-wide reports (cluster-admin only).
  • fuzzball provisioner and fuzzball resource-defs become fuzzball node provisioner, and fuzzball node get becomes fuzzball node show. The help text adopts the formalized “node provisioner” terminology for the configuration that tells Fuzzball how to obtain a class of compute nodes for a given backend.
  • delete is renamed to remove on operations that only unregister rather than destroy: cluster delete → cluster remove and workflow catalog source delete → workflow catalog source remove.
  • fuzzball login and fuzzball logout are added as top-level aliases for fuzzball context login/logout.

Identifiers

Commands that previously required a UUID now also accept the natural identifier where one exists — groups by name, users and members by email, and workflow templates, secrets, clusters, volumes, and catalog sources by name. UUIDs continue to work, and an ambiguous name returns an error with a hint to use the UUID. As part of this work, fuzzball context login --group now correctly selects the requested group; previously the flag was parsed but never applied, so login fell through to the user’s personal account.

Output, lists, filtering, and sorting

  • --output/-o selects the machine output format and accepts json or yaml; omit it for each command’s native rendering. The previous --json/-j flag is retained but deprecated in favor of --output json.
  • List commands now return the complete result set. Most list commands previously issued a single request and silently dropped everything past the first server page; they now walk all pages automatically. --page-size remains a server batch-size hint, --max-pages is an escape hatch against runaway fetches, and the old --page-token is hidden, deprecated, and ignored.
  • Default columns are trimmed to the Id plus the human-meaningful identifying fields; the full field set remains available via -o yaml/-o json.
  • --filter is replaced by discoverable per-field flags on the list commands. String filters accept * wildcards and time filters accept either RFC3339 timestamps or an “ago” duration such as 7d or 24h. The opaque --filter is retained, hidden and deprecated, for advanced expressions.
  • --order-by is typed, accepting friendly field names (listed in each command’s --help), multiple sort keys, and several direction syntaxes.
  • fuzzball node list gains --resources/-r and --available-resources/-a view modes showing total and currently-free cores, memory, and devices.
  • context show machine output is curated. context show -o json|yaml now emits a curated view and, importantly, no longer serializes stored OIDC tokens — the previous raw output leaked the refresh and access tokens. This is a breaking change for scripts that parsed the old keys.

Secrets

  • fuzzball secret create NAME [VALUE] and secret update SECRET [VALUE] take the value as a positional argument; when omitted, the value is read from standard input. --from-file becomes a boolean toggle that reinterprets the positional as a path to a file.
  • A --type flag declares the secret type explicitly, --edit opens an editor against a typed scaffold, and secret://scope/name references are accepted as the identifier by create, get, show, update, and delete.

Other CLI changes

  • The --password flag on organization add-member, add-owner, and update-member no longer accepts a bare form to trigger an interactive prompt; use the new --password-prompt flag instead. Omitting both continues to auto-generate a secure password.
  • User-facing account terminology is relabeled to group in context show, context list, and login messages.
  • Fixed group member list and organization member list so that omitting a role filter returns everyone instead of only non-owner members. A new --relationship=all|owner|member flag (default all) selects which role to list; the previous --owner boolean is retained as a hidden, deprecated alias.

Self-service signup & email

  • Self-service organization signup. Users can sign up with an email address and organization name, receive a verification link, and are logged in automatically after verifying. A password-creation step follows verification so they can sign in on future visits without administrator intervention.
  • SMTP email infrastructure supports plaintext, STARTTLS, and direct TLS connection modes with retry logic, and falls back to log-only delivery when SMTP is not configured. SMTP is configured on the FuzzballOrchestrate CRD (host, port, from address, TLS mode, and a credentials secret reference). In local Kind environments, Mailpit is deployed automatically for development.
  • Organization invites. Organization owners can invite new users to their organization by email.

Cluster administration

  • User impersonation. Cluster admins can assume any user’s identity in any organization for troubleshooting. Assumed-user sessions are audited with both the impersonated user and the impersonating admin recorded.
  • Federate clusters can now view node information for their registered clusters.

Web UI

  • The web UI is rebuilt on a modern React/TypeScript stack for v4, retaining the Monaco-based workflow editor and adding the new Object Cache browser.
  • The workflow editor no longer offers the deprecated retry.attempts field in its templates and validation.
  • Static UI assets are now served with caching and compression, reducing cold-load time over slow or distant connections.
  • Catalog volume picks now carry structured provisioner-name and volume-name fields, producing well-formed three-part volume URIs that upgrade cleanly to v4 use/name syntax.
  • Fixed shorthand syntax handling on workflow rerun.

Deployment & Infrastructure

  • Kubernetes scheduling controls. nodeSelector, tolerations, and affinity are supported on all orchestrate and federate components.
  • Operator Helm chart overrides. The fuzzball-operator chart supports name and fullname overrides.
  • Operator lifecycle observability is added to the operator.
  • GPU substrate image. A fuzzball-substrate-orchestrate-gpu container image bundles the NVIDIA GPU device plugin, built for both linux/amd64 and linux/arm64 substrate nodes.
  • Local docker-compose deployments. fuzzball cluster docker-compose manages local docker-compose Fuzzball deployments with interactive configuration prompts, including a --gpu flag on deploy/update that selects the GPU-enabled substrate image.
  • SaaS mode. Fuzzball can determine whether it is running in SaaS mode, with a preliminary cloudProvider value for GCP and AWS Pulumi wiring that passes FUZZBALL_CLOUD_PROVIDER to the operator.
  • AWS deployment options. fuzzball cluster aws deploy/update gain a --permissions-boundary-arn flag that attaches a managed policy as the PermissionsBoundary on every IAM role the stack creates, unblocking deployments under accounts that enforce IAM boundary guardrails. An opt-in non-Marketplace mode is also added, and the AWS CFN template parameterizes the engineering registry (with new EngRegistryAccountId/EngRegistryRegion parameters) so images can be mirrored to a different ECR without changing customer installs.
  • Stripping the IAM path from SSO role ARNs when writing EKS aws-auth mapRoles so aws-iam-authenticator matches the canonicalized caller and SSO cluster-admin roles are no longer denied.

Scheduler & Provisioner

  • Provisioner definition permissions can now be granted at organization and group scope.
  • A provisioning_blocked stage event is emitted once per allocation when a workflow hits the per-definition node cap, making the pending state visible to the user instead of being silently retried each scheduling tick.
  • Node readiness is now used as a tiebreaker for static provisioner definitions.
  • PBS provisioner validation now requires the select directive to specify a chunk count of 1.

Bug Fixes & Stability

Image & data handling

  • Fixed a panic that could leave the image cache and its database out of sync, leading to image pull failures.
  • Fixed a panic when using data egress with an empty file.

Federate

  • Fixed federate service proxying. (also in v3.4.2)
  • Fixed federate-supplied secrets being overwritten with nil during the preflight check, so encrypted secrets data is now populated correctly. (also in v3.4.2)
  • Federate agent cross-cluster proxy calls now log per-cluster failures through the agent’s JSON logger so TLS and connectivity errors appear in pod logs, and the federate application syncer no longer skips TLS verification.

Cluster & discovery

  • Fixed a concurrent map panic in cluster discovery when iterating cluster IDs during a refresh.
  • Fixed a cluster’s update time not advancing when its status, name, or endpoint changed via upsert.

Security

  • Federate-to-orchestrate communication no longer uses insecure TLS.
  • GetAccount now returns on a denied authorization check, preventing unauthorized cross-account reads.
  • Updated the NATS server dependency to address two vulnerabilities: an MQTT plaintext password disclosure through monitoring endpoints and a credential leak through a debug endpoint.
  • Pulled a patched SQLite version to address a known vulnerability. (also in v3.4.2)