Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Fuzzball v4.1.1 release notes

Fuzzball v4.1.1 is a patch release that hardens the scheduling and service work introduced in v4.1.0. Services now publish their declared ports on randomly assigned host ports so same-port services can share a node, with DNS SRV records and dynamic-config ip:port arrays carrying the assignment to in-cluster clients. Autoscaled service replica pools are now checked against provisioner capacity at submission and at scale-up, and a failed replica above replicas.min no longer tears down a healthy workflow. The scheduler picks up fixes for a definition-wide queue stall behind service-held nodes, preemption of task arrays, and uncordoned nodes never returning to scheduling. Provisioner definition policies now also apply to the internal jobs Fuzzball creates for image pulls and data staging.

Upgrade notes.

  • Update the fuzzball-substrate-orchestrate extension on every compute node before, or together with, Orchestrate. The service host-port change spans both sides: an Orchestrate that expects assigned host ports paired with an older node extension will not resolve services correctly. fuzzball version does not report the extension version, so skew is otherwise invisible; do not use a bare dnf update.
  • Clients must no longer assume <node>:<declared-port> reachability for services. A service on an isolated network now listens on its declared port inside the container but is reachable on the node at a randomly assigned host port. Endpoints and fuzzball workflow connect follow the change automatically; anything that hard-coded a node address plus the declared port must move to the SRV records or the dynamic-config ip:port arrays described below. Services using network.host: true are unaffected.
  • Provisioner definition policies now evaluate for internal jobs. On a cluster whose definitions carry policy: expressions, image pulls and data staging are now subject to those policies rather than bypassing them. Placement falls back to unrestricted when no definition’s policy accepts an internal job, so no existing configuration becomes unschedulable, but placement of internal jobs can change.

Otherwise v4.1.1 upgrades from v4.1.0 with a normal rolling update.

Enhancements

Service Networking & Discovery

Services that run on their own isolated network no longer bind their declared port directly on the node. Each declared port is published at a randomly assigned host port, which removes the port collision that previously stopped two services declaring the same port – including two replicas of one autoscaled service – from sharing a node. Inside the container nothing changes: the service still listens on the port it declared.

  • Transparent for endpoints and workflow connect. External network endpoints and fuzzball workflow connect are wired to the assigned host port automatically, so the common access paths need no changes.
  • DNS SRV records for in-cluster discovery. Each service port is published as an SRV record named _<port-name>._<protocol>.<service>.svc.<workflow-id>.<account-id>.fuzzball, also exported at the account and cluster scopes, resolving to the replica’s host and its assigned port. Autoscaled replica pools publish one record per ready replica under _<port-name>._tcp.<service>.autoscaler.<workflow-id>.fuzzball, so an SRV-capable client or load balancer can address every replica in the pool.
  • ip:port arrays in dynamic-config. A dynamic-config script now receives a SERVICES_<NAME>_<PORT-NAME> bash array of ip:port pairs for each named TCP port of every watched service, alongside the existing IP-only SERVICES_<NAME> array. This is the recommended way to render a proxy or load balancer backend list, since the assigned port differs per replica. Existing scripts that build backends from SERVICES_<NAME> plus a hard-coded port need to switch to the port array.
  • Host-network services unchanged. With network.host: true a service still uses its ports as declared, so two host-network services binding the same port still cannot share a node.

Service Autoscaling

The early-access service autoscaler introduced in v4.1.0 now understands the capacity of the provisioner pool its replicas land on, and treats a failure of an optional replica as recoverable rather than fatal.

  • Replica pools are validated against pool capacity. At submission, Fuzzball computes how many replicas of the requested size a definition’s pool can host (flooring independently on cores, memory, and devices, across the pool’s usable nodes or a dynamic definition’s maxNodes). A workflow whose replicas.min cannot fit is rejected with an error naming the service, the definition, and the shortfall, instead of being accepted and leaving replicas pending forever. If the baseline fits but replicas.max does not, the workflow is accepted, a warning is logged, and a service_capacity_capped stage event records the effective maximum.
  • Scale-up is skipped when the pool is full. A scale-up trigger that fires against an at-capacity pool no longer queues a replica that can never place; it emits a scale_up_capacity_exhausted event and stamps the cooldown so the pool is not re-evaluated on every tick. fuzzball workflow why reports such a replica as blocked on pool capacity.
  • Capacity-aware definition selection. When several definitions could host a replica set, they are now ranked by capacity tier – fits replicas.max, then fits replicas.min, then fits neither – before cost, so a marginally cheaper pool that is too small no longer wins. Note that the check is per service against a pool’s nominal capacity, not cross-tenant admission control: several services sharing one definition are each validated against the full pool, so size shared pools for their combined peak demand.
  • A failed replica above replicas.min is no longer fatal. A start, readiness, or lease failure of a replica above the guaranteed baseline previously errored the whole workflow and tore down its healthy replicas. Such a replica is now marked stopped, returned to the pending pool for a later scale-up retry, and reported through a replica_failed stage event, while the workflow and its other replicas keep running. Baseline replicas at or below replicas.min still fail fast.

Scheduler & Provisioner

  • Definition policies apply to internal jobs. Provisioner definition policy: expressions previously skipped the internal jobs Fuzzball creates implicitly for container image pulls and data staging, so those jobs could land anywhere. They are now policy-gated like user work and match on request.job_kind == "internal", which makes it possible to route transfers to a dedicated transfer node and keep them off compute nodes. Internal jobs keep a safety net user jobs do not have: if no definition’s policy accepts one – or a policy expression fails to evaluate – placement falls back to the unrestricted choice, so a policy can never make image pulls or data staging unschedulable. Because a definition without a policy accepts every job kind, attracting internal jobs to one definition also requires repelling them (request.job_kind != "internal") from the others. The configuration reference documents the full pattern, including the requirement that a transfer node share a storage segment with the compute nodes that consume the images it pulls.

Bug Fixes & Stability

Scheduler

  • Definition-wide scheduling stall behind service-held nodes. When every usable node in a definition’s pool was held by work with no walltime – a service, for example – the scheduler could not estimate a start time for the blocked head of the queue, produced no backfill reservation, and consequently stopped placing anything scored to that definition for as long as the condition held. The scheduler now holds an unestimable reservation in that case: with no bounded start to protect, any candidate with a bounded walltime is admitted, so shorter work keeps flowing while the head waits. Work with no walltime is still never backfilled, so backfill cannot make the condition permanent. fuzzball workflow why distinguishes the case, reporting node availability as unestimable rather than offering a start estimate it cannot compute.
  • Uncordoned nodes never returning to the scheduler. After fuzzball node uncordon, the node’s status changed but its scheduling view was not refreshed. An idle node emits no further substrate resource events, so the node stayed invisible to scheduling until its substrate restarted. Cordon and uncordon now reproject the node immediately, for both the single-node and --all forms, so the status change takes effect at once.
  • Preemption under-counting task-array victims. The fit calculation that decides whether preempting a set of victims frees enough capacity counted a task array running multiple ranks on one node incorrectly, and credited ranks in a terminal state that had already released their footprint. Preemption could therefore evict work and still not fit the job it evicted for. The accounting is now per rank and ignores terminal ranks.
  • Preempted task arrays re-dispatching incorrectly. A preempted task array could re-dispatch task IDs outside its declared range, and could re-place onto the nodes it had just been evicted from ahead of the allocation that evicted it – so the preempting job still did not start. Evicted ranks now roll back to a well-defined orphaned state, preserve their completion records, and are re-dispatched behind the preempting allocation. A rank whose container start fails on a transient container-ID collision with its own tearing-down predecessor is now rolled back and retried on its own, within a bounded retry budget, instead of failing the workflow.

Command-Line Interface

  • --version with a leading v on AWS cluster commands. fuzzball cluster aws deploy and fuzzball cluster aws update passed the flag value through to the CloudFormation FuzzballVersion parameter verbatim, so a natural --version v4.1.0 produced a version string that matched no ECR image tag and failed the deployment. The leading v is now stripped; both bare and v-prefixed forms work.

Web UI

  • Non-admin users could not create volumes from provisioners. The identity context did not expose the signed-in user’s group memberships, so the volume creation form found no groups to create a volume in for anyone who was not an administrator. Group membership is now populated from the user’s accounts.