Fuzzball v4.1.1 release notes
Fuzzball v4.1.1 is a patch release that hardens the scheduling and service work
introduced in v4.1.0. Services now publish their declared ports on randomly
assigned host ports so same-port services can share a node, with DNS SRV
records and dynamic-config ip:port arrays carrying the assignment to
in-cluster clients. Autoscaled service replica pools are now checked against
provisioner capacity at submission and at scale-up, and a failed replica above
replicas.min no longer tears down a healthy workflow. The scheduler picks up
fixes for a definition-wide queue stall behind service-held nodes, preemption of
task arrays, and uncordoned nodes never returning to scheduling. Provisioner
definition policies now also apply to the internal jobs Fuzzball creates for
image pulls and data staging.
Upgrade notes.
- Update the
fuzzball-substrate-orchestrateextension on every compute node before, or together with, Orchestrate. The service host-port change spans both sides: an Orchestrate that expects assigned host ports paired with an older node extension will not resolve services correctly.fuzzball versiondoes not report the extension version, so skew is otherwise invisible; do not use a barednf update.- Clients must no longer assume
<node>:<declared-port>reachability for services. A service on an isolated network now listens on its declared port inside the container but is reachable on the node at a randomly assigned host port. Endpoints andfuzzball workflow connectfollow the change automatically; anything that hard-coded a node address plus the declared port must move to theSRVrecords or thedynamic-configip:portarrays described below. Services usingnetwork.host: trueare unaffected.- Provisioner definition policies now evaluate for internal jobs. On a cluster whose definitions carry
policy:expressions, image pulls and data staging are now subject to those policies rather than bypassing them. Placement falls back to unrestricted when no definition’s policy accepts an internal job, so no existing configuration becomes unschedulable, but placement of internal jobs can change.Otherwise v4.1.1 upgrades from v4.1.0 with a normal rolling update.
Services that run on their own isolated network no longer bind their declared port directly on the node. Each declared port is published at a randomly assigned host port, which removes the port collision that previously stopped two services declaring the same port – including two replicas of one autoscaled service – from sharing a node. Inside the container nothing changes: the service still listens on the port it declared.
- Transparent for endpoints and
workflow connect. External network endpoints andfuzzball workflow connectare wired to the assigned host port automatically, so the common access paths need no changes. - DNS
SRVrecords for in-cluster discovery. Each service port is published as anSRVrecord named_<port-name>._<protocol>.<service>.svc.<workflow-id>.<account-id>.fuzzball, also exported at the account and cluster scopes, resolving to the replica’s host and its assigned port. Autoscaled replica pools publish one record per ready replica under_<port-name>._tcp.<service>.autoscaler.<workflow-id>.fuzzball, so an SRV-capable client or load balancer can address every replica in the pool. ip:portarrays indynamic-config. Adynamic-configscript now receives aSERVICES_<NAME>_<PORT-NAME>bash array ofip:portpairs for each named TCP port of every watched service, alongside the existing IP-onlySERVICES_<NAME>array. This is the recommended way to render a proxy or load balancer backend list, since the assigned port differs per replica. Existing scripts that build backends fromSERVICES_<NAME>plus a hard-coded port need to switch to the port array.- Host-network services unchanged. With
network.host: truea service still uses its ports as declared, so two host-network services binding the same port still cannot share a node.
The early-access service autoscaler introduced in v4.1.0 now understands the capacity of the provisioner pool its replicas land on, and treats a failure of an optional replica as recoverable rather than fatal.
- Replica pools are validated against pool capacity. At submission, Fuzzball
computes how many replicas of the requested size a definition’s pool can host
(flooring independently on cores, memory, and devices, across the pool’s usable
nodes or a dynamic definition’s
maxNodes). A workflow whosereplicas.mincannot fit is rejected with an error naming the service, the definition, and the shortfall, instead of being accepted and leaving replicas pending forever. If the baseline fits butreplicas.maxdoes not, the workflow is accepted, a warning is logged, and aservice_capacity_cappedstage event records the effective maximum. - Scale-up is skipped when the pool is full. A
scale-uptrigger that fires against an at-capacity pool no longer queues a replica that can never place; it emits ascale_up_capacity_exhaustedevent and stamps the cooldown so the pool is not re-evaluated on every tick.fuzzball workflow whyreports such a replica as blocked on pool capacity. - Capacity-aware definition selection. When several definitions could host a
replica set, they are now ranked by capacity tier – fits
replicas.max, then fitsreplicas.min, then fits neither – before cost, so a marginally cheaper pool that is too small no longer wins. Note that the check is per service against a pool’s nominal capacity, not cross-tenant admission control: several services sharing one definition are each validated against the full pool, so size shared pools for their combined peak demand. - A failed replica above
replicas.minis no longer fatal. A start, readiness, or lease failure of a replica above the guaranteed baseline previously errored the whole workflow and tore down its healthy replicas. Such a replica is now marked stopped, returned to the pending pool for a later scale-up retry, and reported through areplica_failedstage event, while the workflow and its other replicas keep running. Baseline replicas at or belowreplicas.minstill fail fast.
- Definition policies apply to internal jobs. Provisioner definition
policy:expressions previously skipped the internal jobs Fuzzball creates implicitly for container image pulls and data staging, so those jobs could land anywhere. They are now policy-gated like user work and match onrequest.job_kind == "internal", which makes it possible to route transfers to a dedicated transfer node and keep them off compute nodes. Internal jobs keep a safety net user jobs do not have: if no definition’s policy accepts one – or a policy expression fails to evaluate – placement falls back to the unrestricted choice, so a policy can never make image pulls or data staging unschedulable. Because a definition without apolicyaccepts every job kind, attracting internal jobs to one definition also requires repelling them (request.job_kind != "internal") from the others. The configuration reference documents the full pattern, including the requirement that a transfer node share a storage segment with the compute nodes that consume the images it pulls.
- Definition-wide scheduling stall behind service-held nodes. When every
usable node in a definition’s pool was held by work with no walltime – a
service, for example – the scheduler could not estimate a start time for the
blocked head of the queue, produced no backfill reservation, and consequently
stopped placing anything scored to that definition for as long as the
condition held. The scheduler now holds an unestimable reservation in that
case: with no bounded start to protect, any candidate with a bounded walltime
is admitted, so shorter work keeps flowing while the head waits. Work with no
walltime is still never backfilled, so backfill cannot make the condition
permanent.
fuzzball workflow whydistinguishes the case, reporting node availability as unestimable rather than offering a start estimate it cannot compute. - Uncordoned nodes never returning to the scheduler. After
fuzzball node uncordon, the node’s status changed but its scheduling view was not refreshed. An idle node emits no further substrate resource events, so the node stayed invisible to scheduling until its substrate restarted. Cordon and uncordon now reproject the node immediately, for both the single-node and--allforms, so the status change takes effect at once. - Preemption under-counting task-array victims. The fit calculation that decides whether preempting a set of victims frees enough capacity counted a task array running multiple ranks on one node incorrectly, and credited ranks in a terminal state that had already released their footprint. Preemption could therefore evict work and still not fit the job it evicted for. The accounting is now per rank and ignores terminal ranks.
- Preempted task arrays re-dispatching incorrectly. A preempted task array could re-dispatch task IDs outside its declared range, and could re-place onto the nodes it had just been evicted from ahead of the allocation that evicted it – so the preempting job still did not start. Evicted ranks now roll back to a well-defined orphaned state, preserve their completion records, and are re-dispatched behind the preempting allocation. A rank whose container start fails on a transient container-ID collision with its own tearing-down predecessor is now rolled back and retried on its own, within a bounded retry budget, instead of failing the workflow.
--versionwith a leadingvon AWS cluster commands.fuzzball cluster aws deployandfuzzball cluster aws updatepassed the flag value through to the CloudFormationFuzzballVersionparameter verbatim, so a natural--version v4.1.0produced a version string that matched no ECR image tag and failed the deployment. The leadingvis now stripped; both bare andv-prefixed forms work.
- Non-admin users could not create volumes from provisioners. The identity context did not expose the signed-in user’s group memberships, so the volume creation form found no groups to create a volume in for anyone who was not an administrator. Group membership is now populated from the user’s accounts.