Fuzzball v4.1.0 release notes
Fuzzball v4.1.0 is a minor release built around scheduling. Queued work is now
ranked by a cluster-configurable priority expression, higher-priority work can
preempt lower-priority preemptible workloads, and a backfill pass – on by
default – puts idle capacity to work while a large job waits, all backed by new
queue-inspection commands, workflow events, and a Prometheus scrape endpoint.
The release also adds storage segments for clusters whose nodes cannot all reach
the same filesystem, an Oracle Cloud Managed Lustre storage driver, HuggingFace
Hub volume ingress, and early-access autoscaling for workflow services. v4.1.0
rolls up everything from v4.0.1 and v4.0.2, so these notes are complete for
anyone upgrading directly from v4.0.0; items marked (also in v4.0.1) or
(also in v4.0.2) were already available on those patches.
Upgrade notes. Three changes alter behavior on an existing cluster with no configuration edit on your part:
- Backfill is enabled by default (
scheduler.backfillEnabled: true). Node pools that previously scheduled strictly in priority order will now run shorter, lower-priority work in gaps behind a blocked job. Set it tofalseto restore the old behavior.recognizedAnnotationsis replaced byignoredAnnotations, with the sense inverted: it was an allowlist of annotation keys to match, and is now a list of keys to skip. A central config still carryingscheduler.recognizedAnnotationsparses without error but the key is ignored and matching becomes strict across every key. Move any annotation handled by a provisioner-definitionpolicy:expression intoscheduler.ignoredAnnotations.definition.nodescounts usable nodes, not free ones. For a static definition it now includes fully allocated nodes, so policies that relied on it dropping to zero while a pool was busy will see jobs queue instead of falling through to another definition.Two operational notes carry forward. Update the
fuzzball-substrate-orchestrateextension on every compute node to match the Orchestrate version –fuzzball versiondoes not report the extension version, so skew is otherwise invisible and an older extension can silently misbehave; do not use a barednf update. And clusters that were manually forced into a running v4.0.1 state after hitting the v4.0.1 NATS JetStream failure still carry the old JetStream cluster name and hit the same conflict here; contact CIQ Support for the NATS cluster name upgrade runbook and follow it first. Every other cluster upgrades with a normal rolling update.
The scheduler now ranks, preempts, and backfills work against a fully configurable notion of priority, so a busy cluster can keep urgent work moving without leaving capacity idle.
Each allocation’s priority is recomputed on every scheduling pass from a
cluster-wide expression, scheduler.priority in the central configuration,
instead of a fixed formula. Unset, it defaults to the sum of the organization,
account, user, and workflow priority inputs plus one unit per hour the
allocation has waited, so long-queued work rises over time.
- Cluster admins set organization priority with
fuzzball organization update <ID> --priority; organization owners set account and user priority withfuzzball group update <GROUP> --priorityandfuzzball user update <USER> --priority. A newfuzzball organization list(cluster-admin) finds the organization ID. Entity priority changes now propagate to queued and running allocations, not just new submissions. fuzzball workflow start --prioritytakes a signed value that must be less than or equal to zero (default0), so users can deprioritize their own work but never raise it above their entitlement.- Priority is visible throughout: a Priority column in
fuzzball group list,fuzzball org member list, andfuzzball group member list; sorting viafuzzball workflow list --order-by priority; a full per-stage breakdown fromfuzzball workflow describe <WORKFLOW> <STAGE>; and readable events –Priority set (workflow: 0, organization: 5, group: 2, user: 1)on submission,Priority updated: organization 5 -> 10when a change is applied to a live allocation.
A blocked higher-priority allocation can evict a running, lower-priority
allocation that has opted in to being preemptible. Preemption is disabled by
default; set scheduler.preemptionEnabled: true to turn it on.
- Workloads opt in with a top-level
preemptible: truein the workflow definition (applying to every job and service), a per-job or per-servicepreemptiblethat overrides it, orfuzzball workflow start --preemptible. Nothing that has not opted in is ever evicted. - Eviction is deliberately conservative. The victim must be running,
preemptible, past an anti-thrash minimum runtime
(
scheduler.minPreemptionRuntime, default60s), and below the blocked allocation by at leastscheduler.preemptionGap(default10). The pass only fires once pool utilization reachesscheduler.preemptionThreshold(default75), and only when the freed capacity actually seats the blocked allocation – undersized victims are not evicted for nothing. - At most one victim is evicted per blocked allocation per pass. Opt-in
multi-victim preemption (
scheduler.preemptionMultiVictimEnabled, bounded byscheduler.maxEvictionsPerTick, default8) evicts a set in one pass when no single victim frees enough. - Organization owners control which groups a group may preempt with
fuzzball group update <GROUP> --preempt <TARGET>and--no-preempt. Lists are empty by default, so enabling preemption without configuring them preempts nothing, and deleted groups drop out automatically. fuzzball workflow eventsshows both sides of an eviction:Preempted by workflow X job Yon the victim andPreempted workflow X job Yon the preemptor. Evicted work is re-queued and restarts cleanly on its next placement.
While a high-priority allocation waits for capacity, one-level backfill runs
shorter, lower-priority allocations in the capacity it does not need yet,
without pushing out its predicted start time. Backfill is enabled by default
(scheduler.backfillEnabled: true).
- A job qualifies when its runtime is bounded by an execute timeout
(
policy.timeout.execute); the workflow DSL default already satisfies this. Services, which legitimately run without a timeout, are never candidates. - Image and data staging jobs are backfilled whenever they fit capacity the blocked allocation does not need, so staging no longer starves behind a large queued job.
- Backfill applies to static pools and to dynamic definitions already at their
maxNodescap, and placements appear infuzzball workflow events.
Three new commands and a metrics endpoint expose what the scheduler is doing and why, replacing guesswork about a job that has not started.
fuzzball queue showlists allocations ordered by effective priority, with columns for queue, status, priority, weight, node, preemptible, queue time, and owner, and--queue,--priority-min, and--ownerfilters. Allocation IDs print in full, so one can be copied straight intofuzzball workflow why. The cluster-wide view is cluster-admin only; any user can see their own allocations with--mine, enforced server-side.fuzzball queue stats(cluster-admin) reports queue depths, scheduler counters, tick timings, and the current backfill reservation.fuzzball workflow why <WORKFLOW> <ALLOCATION>returns the typed reason an allocation is in its current state. Workflow owners can run it on their own allocations without admin rights; anyone else needs the newexplain_allocationpermission on the owning account.- On Federate deployments,
queue showandqueue statsfan out across every Orchestrate cluster andworkflow whyresolves the owning cluster automatically.--cluster <ID>(repeatable or comma-separated) scopes to specific clusters. - A Prometheus scrape endpoint at
/metricson every service’s monitoring HTTP server and on the single binary. Scheduler and gRPC metrics that were registered but unreachable are now scrapeable: tick and interval metrics showing whether the configurablescheduler.intervalbackstop is delaying work, preemption metrics (misses, time blocked before an eviction, near-miss priority gaps, preempted-allocation outcomes), and backfill metrics (reservations held, predicted head wait, gap-fill size, share of each tick).fuzzball_scheduler_ready_wait_seconds, which previously recorded roughly zero for every placement, now measures real ready-queue wait.
Clusters where different groups of nodes mount different shared filesystems can now describe that topology, so the scheduler only places work on nodes that can reach the volumes it needs.
- A
segmentfield on a storage provisioner restricts its volumes to nodes in the same segment; the same key on a node provisioner definition marks the nodes it provisions as belonging to that segment. Names are lowercase DNS labels, and provisioners without a segment stay globally accessible, so existing deployments are unaffected. - Invalid combinations are caught at submission rather than run time: a workflow referencing a volume from a segment no node provides, or a job mounting volumes whose segments conflict, is rejected with a clear error.
- Image stages are segment-aware. When jobs in different segments reference the same image, a separate image stage runs in each so the converted SIF is cached on the right filesystem, and the distributed pull lock is scoped per segment so pulls across segments run concurrently.
- The Volume Provisioners page in the web UI exposes the segment field, and
recognizes the OCI Managed Lustre driver and the new
Provisioningstatus.
A new oci_lustre storage driver backs a provisioner with a single Oracle Cloud
File Storage with Lustre filesystem, aimed at HPC and AI/ML workloads. Each
volume is a subdirectory of that filesystem, owned by the volume’s POSIX
UID/GID.
- Bring-your-own mode uses an existing filesystem via
filesystemId. Self-provisioned mode omits it and suppliessubnetIds; the driver builds the filesystem from the compartment, availability domain, and subnet, sized bylustreCapacityGbsandlustrePerformanceTier(default: Oracle’s 31.2 TB minimum at the 125 MBps-per-TB tier). That minimum applies to every tier, so a self-provisioned filesystem is a large billable resource – prefer bring-your-own where one filesystem backs several provisioners. - Volumes mount with the native Lustre client, so every substrate node running these workloads needs the Lustre client kernel modules installed and an LNet configuration bound by subnet rather than NIC name.
A workflow volume can now pull a model, dataset, or space – or a single file from one – straight from the HuggingFace Hub.
- The source URI takes the form
hf://<org>/<repo>[/<path>][?revision=&type=&include=&exclude=]. Single-name repositories such ashf://gpt2are accepted, andhuggingface://is an alias. revisionselects a branch, tag, or commit (defaultmain);typeselectsmodel(default),dataset, orspace;includeandexcludeare repeatable globs that filter a whole-repository download. The workflow editor offers HuggingFace as a first-class ingress source with structured fields for all of these.- A new
hfsecret type stores a HuggingFace user access token, required for gated or private repositories and useful for avoiding anonymous rate limits. Public repositories need no credentials.
A workflow service can now declare an autoscaler block and run as a pool of
replicas that grows and shrinks automatically, suiting inference servers, queue
workers, and other demand-driven services. This is early access – the syntax
may change in a future release.
- Fuzzball pre-allocates capacity for up to
replicas.maxcopies and startsreplicas.minof them, which may be0. - Scaling decisions are PromQL expressions evaluated against the service’s own
scraped metrics (
metrics.enabled, with configurable path, port, and interval) or against container CPU and memory metrics Fuzzball collects automatically. Each action is rate-limited by itscooldown-period. - Replicas are reached through a single DNS name that always resolves to the currently-ready replicas, so clients never need to know the pool size.
fuzzball cluster aws|azure|gcp|oci generate-substrate-config extracts the
configuration an external substrate node needs to join the cluster – for
compute nodes that cannot mount the cluster’s shared filesystem, such as nodes
attached over a VPN. Matching the existing docker-compose subcommand, it reads
the operator-rendered provisioner secrets, adjusts the output for static nodes,
and accepts --image-dir to override the image cache directory. The
deployment’s dynamic provisioner must be enabled.
The fuzzball cluster docker-compose stack now runs on Podman as well as
Docker. The CLI auto-detects the runtime, preferring Docker; pass --runtime podman or set FUZZBALL_CONTAINER_RUNTIME=podman to force it. Podman is for
local testing – production should stay on Docker. See the README.md written
into the deployment directory for host setup.
- Scripts no longer need a shebang. Job and service
scriptfields assume#!/bin/shwhen no interpreter line is present, so a plain shell script runs without error. Scripts starting with#!still use the interpreter they name. - Optional secret and volume template inputs. Application-template inputs of
type
secretandvolumecan be left unset: an empty reference is treated as unset rather than rejected, so a credential needed only by some run configurations no longer has to be permanently required or downgraded to a plain string. Non-empty references are still format-validated, and the web UI pickers gain a “(None)” option.
fuzzball object putno longer silently replaces an existing object. Uploading over an object whose content differs fails; pass--forceto replace it. Re-uploading identical content is still a no-op success.--vpc-cidronfuzzball cluster aws deploychooses the VPC’s CIDR block (for example172.16.0.0/16) to avoid overlapping existing corporate networks. It defaults to10.0.0.0/16and is deploy-only, because changing an existing deployment’s CIDR would destroy it. AWS deploy also gains--version, the version pin that Azure, GCP, and OCI already supported.fuzzball cluster aws cleanupis scoped to one deployment by default. Pass--allfor every deployment in the region, as it previously did.fuzzball cluster oci logsmatches the other providers. It uses the shared log streaming, defaults to thefuzzballnamespace (the services) rather thanfuzzball-system(the operator), prompts for a container when a pod has several instead of erroring, and gains--componentand--since-time.fuzzball group updatetakes the group as a positional argument, with new--priority,--preempt, and--no-preemptflags. The old--group/-gflag is hidden and deprecated but still works.--owneris canonical onorg member listandgroup member list;--relationshipis hidden and deprecated but accepted.--ownerlists owners,--owner=falselists non-owners, and omitting it lists everyone.fuzzball object putandgetinfer names from the source when the destination is a directory-style path. (also in v4.0.1)fuzzball run --volumeexpresses the v4 volume model. The spec accepts[KEY=][REF:]MOUNT[,size=SIZE][,ANNOTATION=VALUE...]– a bare mount path creates an ephemeral volume, a bare name references a persistent volume, and aprovisioner/volumeform selects a provisioner. (also in v4.0.1)fuzzball runstreaming stdin errors propagate instead of being swallowed.
- Workflow editor. A stage selector and a downloadable events-and-logs view on the workflow detail page, an Edit Workflow button, an option to open a catalog template in the editor, a task array card, suggestions for the devices and annotations fields, a node provisioner field, and annotations on persistent volumes.
- Catalog run forms group values by category, under their own
display_categoryheadings instead of lumping every categorized value under “Advanced”. Catalog descriptions render Markdown. - Unified member lists across the organization and group views, the group membership menu item moved somewhere more discoverable, and a missing GCP volume driver added to the provisioner picker.
- Download CLI page. A page in the Links menu lets logged-in users download a CLI build for macOS (Apple Silicon and Intel), Linux (x86_64 and ARM64), and Windows (x86_64), preselected for the browser’s platform and always matching the deployed Fuzzball version. (also in v4.0.1)
- Filter button uses a funnel icon rather than a sliders icon, workflow list filters intersect rather than union, the Objects label was renamed, and the password reset modal no longer clips its text.
- Workflow stage events panel on the workflow detail page, and more fields in the workflow defaults section. (also in v4.0.1)
- Self-provisioning is asynchronous. AWS EFS, OCI File Storage, and OCI
Managed Lustre provisioners that create their own backing infrastructure
return immediately from
provisioner addin aProvisioningstatus and move toReadyorErrorwhen the cloud resource is done – Lustre creation takes 10 to 15 minutes. Check withfuzzball volume provisioner info <NAME>. - Image cache entries are architecture-aware, so the same image URI no longer collides in the object cache across architectures on mixed-architecture clusters.
- Five improvements carried forward from the patch line, all
(also in v4.0.1):
fuzzball object listauto-paginates, with--page-size/-psetting the per-request batch and a new--max-pages/-mcapping the total;fuzzball volume listreturns the full result set in a stable order instead of stopping at 25, and excludes system-generated ephemeral volumes; the legacyreferencefield is gone from volume responses, since the v3-shapevolume://URI was unusable under v4; price estimates include GPU cost; and the object browser shows inherited effective TTLs.
maxNodesis a real per-definition pool cap. It previously only clamped a single allocation’s node spread, leaving the pool bounded by a hardcoded limit, so a configuredmaxNodeshad no effect on total node count. At the cap, an allocation that fits waits for a node to free rather than over-provisioning, and one needing more nodes than the cap is rejected.- Annotation matching is strict by default, configured with
scheduler.ignoredAnnotations– see the upgrade note above. The<KEY>.must-havesuffix, never wired up, has been removed. definition.nodescounts usable nodes. For a static definition it now includes fully allocated and exclusively held nodes rather than dropping to zero at full occupancy, which previously made nodes-gated policies reject a pool exactly when it was busiest with “no suitable provision definition found”. A companiondefinition.max_nodesexposes the configured cap.- A missing device no longer aborts scheduling. A node lacking a requested device is skipped instead of aborting the allocation’s pass.
- AWS deploy is faster after optimization of the Pulumi program, and its Pulumi runner now runs inside its own VPC.
- Every Helm chart is mirrored to the depot registries. Some were never
published, so
fuzzball cluster gcp deploy --use-depotfailed roughly 30 minutes into the Pulumi phase withMANIFEST_UNKNOWN. - metrics-server is installed on OCI (OKE) deployments, which do not bundle it, so HorizontalPodAutoscalers scale under load instead of sitting at baseline.
- Additional trusted issuers in the CRD. The FuzzballOrchestrate and
FuzzballFederate CRDs accept an
additionalTrustedIssuerslist of JWT issuer URLs, so endpoints can be trusted using tokens from a separate Keycloak. (also in v4.0.1) - Default cluster names derive from the deployment domain rather than
unset-clusteror a timestamped stack name, with Federate clusters prefixedfederate.. Existing clusters are renamed on the next reconcile. (also in v4.0.1) - Certificates are issued before ingress resources, closing a window in which Kong could serve a default or expired certificate. (also in v4.0.1)
- The Docker Compose object cache persists and self-heals. It now survives a restart, and a blob whose backing file has gone missing is purged from the metadata so it repopulates instead of returning 404.
- Docker Compose deployment fixes. Endpoint subdomains work, the cluster domain is set correctly, and job containers no longer fail with “no space left on device”. (also in v4.0.1)
fuzzball cluster gcp deployandupdatefail when the Pulumi runner job fails instead of reporting success;destroytears down provisioned resources and DNS records before removing the deployment record. (also in v4.0.1)
- Workflow events propagate to Federate, so a federated view shows the same
event stream as the owning cluster.
fuzzball workflow why,queue show, andqueue statsalso work on Federate, where they previously returned NotFound. - Five federation fixes carried forward from the patch line, all (also in v4.0.1): service-scoped secret references are materialized when forwarding score requests downstream, so secrets referenced by workflow services resolve; a cached cluster list is served when a refresh fails instead of an error; a cluster’s own signing key is kept in the JWKS on refresh, preventing token validation failures after a rotation; cluster registration upserts the record, so re-registering a removed cluster no longer fails; and skipping a duplicate storage provisioner during sync rolls back the aborted transaction cleanly.
- Cross-tenant workflow listing. A caller-supplied workflow list filter could escape the account scope and return other tenants’ workflows. The account scope is now a separate clause and filter fields are allowlisted.
- Application template rendering is sandboxed. The
env,expandenv, andgetHostByNametemplate functions are removed from the rendering environment, so a caller-supplied application template can no longer read the Orchestrate process environment. fuzzball workflow whyis authorized. The underlying ExplainAllocation call let any authenticated user explain any allocation cluster-wide. Non-owners now require theexplain_allocationpermission on the allocation owner’s account.- Scheduling fields require an organization owner. Group priority and preemption lists could previously be set by any group owner, and permission denials now explain what is required.
- Scheduler leader takeover hung indefinitely after an orchestrator pod was replaced on a cluster where the central configuration had never been set, leaving nothing able to schedule.
- Workflow summary filters were not applied to the underlying queries, so filtered lists came back unfiltered.
- Service autoscaler minimum replicas were not honored on startup, so a pool
could start below its configured
replicas.min. - Workflow catalogs using the deprecated v1 mount syntax are parsed server-side and upgraded to v4 format instead of being rejected by the gateway. (also in v4.0.1)
fuzzball node listshowed jobs as running after they had completed. (also in v4.0.1)
- Annotation and backend-spec edits to existing provisioner definitions were silently dropped when central config was re-applied; they now persist and appear in resource-defs output.
- Removing or renaming a provisioner definition failed with “currently used by N active instance(s)” when none were running. (also in v4.0.1)
- The GCP provisioner picked an incompatible boot disk for the G4, A4, A4X,
H4D, X4, and Z4 machine families, which reject
pd-balanced. It now selects hyperdisk-balanced boot disks for those families and routes G4, A4, and A4X to the NVIDIA substrate image. GCP Filestore creation, which timed out before the filesystem was ready, now allows 60 minutes. - Object cache background jobs raced on shutdown. Purge, GC, and reconcile now build their cancellation context up front, startup fails on an invalid cluster ID instead of silently purging nothing, and every cycle is counted so one that finds nothing is distinguishable from one that never ran.
- SIF image caching failed silently on Docker Compose deployments using a custom domain. Orchestrate was not given the domain, so its certificate did not cover the object cache upload host and the upload’s TLS handshake failed – images were re-converted on every run instead of served from cache.
fuzzball node deprovisionsilently no-opped when given the node ID (theip:portvalue fromfuzzball node list); only the hostname form terminated the instance. Both forms now work. (also in v4.0.1)
- AWS deployments failed in accounts whose 12-digit account ID begins with a zero. The ID was parsed as an integer, dropping the leading zero and producing invalid IAM principal ARNs: the KMS key policy was rejected 30 minutes into the deploy, and EKS auth-map entries silently received invalid user ARNs. The account ID is now a string end to end.
fuzzball cluster aws logsandgenerate-substrate-configcould not reach the EKS cluster when its name did not match the stack name. The name is now read from the stack tag or the CloudFormation output, and the bearer token includes the expiry parameter the authenticator requires.fuzzball cluster oci logsandgenerate-substrate-configfailed with “no kubernetes endpoint” on VCN-native OKE clusters, which only read the legacy endpoint field. They now prefer the public endpoint, falling back to the private and legacy fields.fuzzball cluster oci logscrashed with a flag shorthand panic, because its--containerflag claimed-c, colliding with the global--colorflag. The flag no longer has a shorthand.fuzzball cluster azure listignored--location; it now filters deployments to that location, case-insensitively. The interactive selector used bystatus,logs, andgenerate-substrate-configlabels the resource group explicitly, since that is what--resource-grouptakes and it can differ from the name shown bylist.
- S3 secrets built from temporary credentials could not authenticate.
fuzzball secret create,update, andeditsilently dropped a suppliedsessionTokenons3secrets, so secrets made from temporary or SSO credentials (ASIA...keys) were stored incomplete and every workflow S3 ingress or egress using them failed to authenticate. - Filtering
member listby--email,--name, or other fields returned no matches. Field filters and email-based resolution now work. - Rendering or starting a catalog entry crashed when its prompts included an optional secret or volume input left unset, and opening a workflow carrying a value the client could not represent produced a rendering error.
- Workflow endpoints were unreachable on operator-deployed clusters. The
WorkflowEndpointsService had no ingress route, so
fuzzball workflow endpointsand the endpoint display in the web UI could not reach it. - Bare ephemeral volumes were stripped when re-opening a workflow in the editor or rerunning a previous run. (also in v4.0.1)
- Organization data-access audit events were recorded under the workflow
startWorkflowsubject method, so they could not be queried or filtered by operation. Each handler now emits under an organization-specific method. - Copying to S3-compatible object stores could lose data silently. Failed egress uploads now surface as errors instead of reporting success, and non-AWS S3 endpoints no longer receive the default AWS request checksum encoding that they reject. (also in v4.0.1)
- Malformed
cluster_refsrows are repaired before the owner CHECK constraint is applied, preventing migration failures on clusters carrying older data. (also in v4.0.1) - The NATS JetStream cluster name is decoupled from the display cluster
name, so renaming or defaulting a deployment’s name no longer stalls a
rolling update with the newest pod in
CrashLoopBackOff. See the upgrade note above. (also in v4.0.2)