Fuzzball v4.1.2 release notes
Fuzzball v4.1.2 is a patch release that corrects the service host-port behavior introduced in v4.1.1, unblocks multi-GB object-cache transfers such as large container image layers, and tightens workflow validation of ingress and egress URIs. The substrate bridge now delivers orchestrate extension RPMs for both x86_64 and aarch64 nodes, with provisioning selecting the right one by node architecture. v4.1.2 also includes everything shipped in v4.1.1; if you are upgrading from v4.1.0, read the v4.1.1 release notes alongside these, since the scheduler, autoscaling, and provisioner-policy changes described there apply to you as well.
Upgrade notes.
- Service host ports are back to their declared values. v4.1.1 published every service port on a randomly assigned host port; v4.1.2 restores publishing on the declared port for ordinary services and keeps random assignment only for autoscaled replica pools. If you changed clients or
dynamic-configscripts for v4.1.1,<service-fqdn>:<declared-port>works again for non-pool services. As before v4.1.1, two non-pool services that declare the same port cannot share a node. Coming from v4.1.0, the only change is for autoscaled replica pools.- Update the
fuzzball-substrate-orchestrateextension on nodes that host autoscaled replica pools. Assigned host ports are still required there, so a node running an extension that predates v4.1.1 will fail replica startup with an explicit error. Ordinary services no longer depend on the port report, so this constraint is narrower than the blanket requirement in the v4.1.1 upgrade notes.- Workflow ingress and egress URIs are now scheme-checked at validation time. Only
http,https,file,s3,fuzzball,fb,hf, andhuggingfaceare accepted, andfilemust be spelled exactlyfile://. Workflows using the undocumentedlocal://scheme, or variants such asfile:/pathorFILE://path, are now rejected at submission instead of failing or hanging at transfer time. Move them tofile://.- Docker Compose deployments no longer set
TLS_DOMAINorOBJECTCACHE_HOST. The object-cache hostname is now handled by anapi.${DOMAIN}network alias on the nginx service. If you copied the shippeddocker-compose.yamland added those variables locally, take the updated file.
Services that run on their own isolated network publish each declared port on
the node at that same port again, so an in-cluster client can reach a service at
<service>.svc.<workflow-id>.<account-id>.fuzzball:<declared-port> (the service
name alone also resolves through the workflow’s DNS search path). Autoscaled
replica pools remain the exception and keep substrate-assigned host ports, which
is what lets several replicas of one service pack onto a single node.
- Pool clients still use
SRVrecords ordynamic-config. Because a replica’s host port is assigned, clients of a pool must not connect to<node>:<declared-port>. Discover each replica through the DNSSRVrecords named_<port-name>._tcp.<service>.autoscaler.<workflow-id>.fuzzball, or through theip:portarrays adynamic-configscript receives. - Endpoints and
fuzzball workflow connectare unaffected. Both route to whatever host port a service is published on, so neither needs a change. - Host-network services are unchanged. With
network.host: truea service still binds its ports exactly as declared.
- Extension RPMs for aarch64 nodes. The substrate bridge image now bundles
and serves the
fuzzball-substrate-orchestrateextension for both x86_64 and aarch64, at architecture-suffixed well-known download paths, and the operator routes both paths on the bridge ingress. AWS and Azure cloud-init pick the extension URL from the CPU architecture advertised by the provision definition selected for the allocation, falling back to the x86_64 URL when a definition advertises no architecture or when a configuration rendered before this release supplies no aarch64 URL. The legacy unsuffixed x86_64 download path keeps working, and a request for an architecture the bridge does not carry returns 404 rather than the wrong RPM. A neworchestrateExtensionURLArm64provisioner setting holds the aarch64 URL. Note that the depot client install step in cloud-init is still x86_64-only, since no aarch64 depot client is published. - Object-cache uploads from substrate in Docker Compose. Substrate uploads
converted SIF images to the object cache at
https://api.${DOMAIN}, from inside its own container, so the name has to resolve on the Compose network. The nginx service now carries anapi.${DOMAIN}alias on that network, which makes caching of externally sourced images work. This replaces the earlierTLS_DOMAINandOBJECTCACHE_HOSTenvironment overrides, which have been removed along with thelocal://path rewriting they depended on.
- Unsupported ingress and egress URI schemes are rejected at submission.
Workflow validation previously accepted any parseable URI, so a scheme the
data mover cannot materialize was only discovered when the transfer ran, and
the undocumented
local://scheme could hangfuzzball-substrateoutright.sourceanddestinationare now checked against the schemes that can actually be copied –http,https,file,s3,fuzzball,fb,hf, andhuggingface– and anything else fails validation with a message naming the supported set.fileis accepted only in its exactfile://form, because that is the only spelling joined onto the volume mount downstream;file:/path,file:path, and uppercase variants would otherwise reach the data mover as paths outside any mount.
- Multi-GB blob transfers aborting mid-stream. The object-cache blob endpoints stream their bodies on the main Orchestrate HTTP server, which carries a 15 second per-request read and write timeout appropriate for JSON API and UI responses. Any blob that took longer than 15 seconds to transfer was therefore closed mid-stream, surfacing as a stream error on the substrate and an “upstream prematurely closed connection” in nginx, with a partial byte count that varied run to run. Small blobs succeeded and large container image layers always failed, blocking image pulls and cross-node object reuse. The blob upload and download handlers now clear the per-request deadlines for the duration of the transfer, leaving the 15 second guard in place for every other request.
- Source-address lookup failing on kernels without XFRM support. Orchestrate determines the host IP it advertises by asking the kernel which source address reaches a given destination. That lookup opened a netlink handle subscribed to every netlink family, which fails outright on a kernel built without XFRM, so startup failed on those hosts with a netlink handle error rather than resolving the address. The lookup now subscribes only to the route family it actually needs.