Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Fuzzball v4.1.2 release notes

Fuzzball v4.1.2 is a patch release that corrects the service host-port behavior introduced in v4.1.1, unblocks multi-GB object-cache transfers such as large container image layers, and tightens workflow validation of ingress and egress URIs. The substrate bridge now delivers orchestrate extension RPMs for both x86_64 and aarch64 nodes, with provisioning selecting the right one by node architecture. v4.1.2 also includes everything shipped in v4.1.1; if you are upgrading from v4.1.0, read the v4.1.1 release notes alongside these, since the scheduler, autoscaling, and provisioner-policy changes described there apply to you as well.

Upgrade notes.

  • Service host ports are back to their declared values. v4.1.1 published every service port on a randomly assigned host port; v4.1.2 restores publishing on the declared port for ordinary services and keeps random assignment only for autoscaled replica pools. If you changed clients or dynamic-config scripts for v4.1.1, <service-fqdn>:<declared-port> works again for non-pool services. As before v4.1.1, two non-pool services that declare the same port cannot share a node. Coming from v4.1.0, the only change is for autoscaled replica pools.
  • Update the fuzzball-substrate-orchestrate extension on nodes that host autoscaled replica pools. Assigned host ports are still required there, so a node running an extension that predates v4.1.1 will fail replica startup with an explicit error. Ordinary services no longer depend on the port report, so this constraint is narrower than the blanket requirement in the v4.1.1 upgrade notes.
  • Workflow ingress and egress URIs are now scheme-checked at validation time. Only http, https, file, s3, fuzzball, fb, hf, and huggingface are accepted, and file must be spelled exactly file://. Workflows using the undocumented local:// scheme, or variants such as file:/path or FILE://path, are now rejected at submission instead of failing or hanging at transfer time. Move them to file://.
  • Docker Compose deployments no longer set TLS_DOMAIN or OBJECTCACHE_HOST. The object-cache hostname is now handled by an api.${DOMAIN} network alias on the nginx service. If you copied the shipped docker-compose.yaml and added those variables locally, take the updated file.

Enhancements

Service Networking & Discovery

Services that run on their own isolated network publish each declared port on the node at that same port again, so an in-cluster client can reach a service at <service>.svc.<workflow-id>.<account-id>.fuzzball:<declared-port> (the service name alone also resolves through the workflow’s DNS search path). Autoscaled replica pools remain the exception and keep substrate-assigned host ports, which is what lets several replicas of one service pack onto a single node.

  • Pool clients still use SRV records or dynamic-config. Because a replica’s host port is assigned, clients of a pool must not connect to <node>:<declared-port>. Discover each replica through the DNS SRV records named _<port-name>._tcp.<service>.autoscaler.<workflow-id>.fuzzball, or through the ip:port arrays a dynamic-config script receives.
  • Endpoints and fuzzball workflow connect are unaffected. Both route to whatever host port a service is published on, so neither needs a change.
  • Host-network services are unchanged. With network.host: true a service still binds its ports exactly as declared.

Deployment & Infrastructure

  • Extension RPMs for aarch64 nodes. The substrate bridge image now bundles and serves the fuzzball-substrate-orchestrate extension for both x86_64 and aarch64, at architecture-suffixed well-known download paths, and the operator routes both paths on the bridge ingress. AWS and Azure cloud-init pick the extension URL from the CPU architecture advertised by the provision definition selected for the allocation, falling back to the x86_64 URL when a definition advertises no architecture or when a configuration rendered before this release supplies no aarch64 URL. The legacy unsuffixed x86_64 download path keeps working, and a request for an architecture the bridge does not carry returns 404 rather than the wrong RPM. A new orchestrateExtensionURLArm64 provisioner setting holds the aarch64 URL. Note that the depot client install step in cloud-init is still x86_64-only, since no aarch64 depot client is published.
  • Object-cache uploads from substrate in Docker Compose. Substrate uploads converted SIF images to the object cache at https://api.${DOMAIN}, from inside its own container, so the name has to resolve on the Compose network. The nginx service now carries an api.${DOMAIN} alias on that network, which makes caching of externally sourced images work. This replaces the earlier TLS_DOMAIN and OBJECTCACHE_HOST environment overrides, which have been removed along with the local:// path rewriting they depended on.

Validation

  • Unsupported ingress and egress URI schemes are rejected at submission. Workflow validation previously accepted any parseable URI, so a scheme the data mover cannot materialize was only discovered when the transfer ran, and the undocumented local:// scheme could hang fuzzball-substrate outright. source and destination are now checked against the schemes that can actually be copied – http, https, file, s3, fuzzball, fb, hf, and huggingface – and anything else fails validation with a message naming the supported set. file is accepted only in its exact file:// form, because that is the only spelling joined onto the volume mount downstream; file:/path, file:path, and uppercase variants would otherwise reach the data mover as paths outside any mount.

Bug Fixes & Stability

Object Cache & Image Pulls

  • Multi-GB blob transfers aborting mid-stream. The object-cache blob endpoints stream their bodies on the main Orchestrate HTTP server, which carries a 15 second per-request read and write timeout appropriate for JSON API and UI responses. Any blob that took longer than 15 seconds to transfer was therefore closed mid-stream, surfacing as a stream error on the substrate and an “upstream prematurely closed connection” in nginx, with a partial byte count that varied run to run. Small blobs succeeded and large container image layers always failed, blocking image pulls and cross-node object reuse. The blob upload and download handlers now clear the per-request deadlines for the duration of the transfer, leaving the 15 second guard in place for every other request.

Substrate & Networking

  • Source-address lookup failing on kernels without XFRM support. Orchestrate determines the host IP it advertises by asking the kernel which source address reaches a given destination. That lookup opened a netlink handle subscribed to every netlink family, which fails outright on a kernel built without XFRM, so startup failed on those hosts with a netlink handle error rather than resolving the address. The lookup now subscribes only to the route family it actually needs.