Updating Fuzzball
You can (and should!) update your Fuzzball cluster when new versions are released. The update process differs somewhat depending on the environment.
Database migrations run before the Orchestrate pods are replaced, so when upgrading to v4.3, the release that adds endpoints on autoscaled service pools, every workflow endpoint read fails on the Orchestrate pods still running the old version until the rollout finishes.
If your deployment includes a Federate cluster, upgrade Federate before (or at the same time as) the Orchestrate clusters it fronts. An older Federate rejects the pool endpoint rows a newer Orchestrate replicates to it, and replication is not retried after the workflow starts, so those endpoints stay unreachable through Federate for the life of the workflow.
Updating Fuzzball in an on-prem environment consists of updating the Substrate installation running on compute nodes, the Orchestrate cluster, and ensuring that client CLIs are up-to-date.
Fuzzball Substrate is installed on each compute node from the Depot yum/dnf repository. When you
update Orchestrate you should also update the substrate packages on every compute node — in
particular the fuzzball-substrate-orchestrate extension — pinned to the same version as
Orchestrate.
Do not simply run a barednf update. That pulls the newest available packages, which can overshoot your Orchestrate version (for example installing a 4.0.x extension against a 3.4.x Orchestrate) and conflicts with the recommendation to avoid skipping major or minor releases. Thefuzzball-substrate-orchestrateversion must match your Orchestrate version.
Set VERSION to match your Orchestrate version (without the leading v) and install the pinned
packages, exactly as you did when installing Substrate on the compute
node:
# VERSION="4.3.0" # matches orchestrate version without the leading v
# dnf install -y fuzzball-substrate fuzzball-substrate-orchestrate-${VERSION}-1 fuzzball-cli-${VERSION}-1After the packages are updated, restart Substrate on the compute node so the new binaries take effect:
# systemctl restart fuzzball-substrate.serviceYou can confirm the extension version installed on the node with rpm:
# rpm -q fuzzball-substrate-orchestratefb versionreports only the CLI and Agent versions, not thefuzzball-substrate-orchestrateextension, so version skew between Orchestrate and a compute node is not otherwise visible. The base Substrate platform is backward compatible with Orchestrate, but thefuzzball-substrate-orchestrateextension carries Orchestrate-matched functionality (for example the data-copy / S3 mover), so an older extension can silently misbehave against a newer Orchestrate. Keep the extension pinned to the Orchestrate version, and don’t forget to update Substrate whenever you update Orchestrate.
It’s easy to update your Fuzzball Orchestration cluster using the same K8s Custom Resource Definition (CRD) that installs Fuzzball. Simply repeat the steps that you used to install the Fuzzball Operator with a new version.
You can safely update from one patch release to another non-sequentially. (e.g Updating from v2.1.8 to v2.1.10 is supported.) But you should avoid skipping major or minor releases when updating.
First, make sure that you are logged into CIQ Depot.
# DEPOT_USER="" # populate with your username for CIQ Depot
# ACCESS_KEY="" # populate with the Depot key obtained from the CIQ sales/support team# helm registry login depot.ciq.com --username "${DEPOT_USER}" --password "${ACCESS_KEY}"Now simply update the version number and run the same commands that you originally ran to install the Fuzzball Operator the first time.
# VERSION="" # insert the updated version number
# CHART="oci://depot.ciq.com/fuzzball/fuzzball-images/helm/fuzzball-operator"# helm upgrade --install fuzzball-operator "${CHART}" \
--namespace fuzzball-system --create-namespace \
--version "${VERSION}" \
--set "image.tag=${VERSION}" \
--set "imagePullSecrets.name=repository-ciq-com" \
--set "imagePullSecrets.inline.registry=depot.ciq.com" \
--set "imagePullSecrets.inline.username=${DEPOT_USER}" \
--set "imagePullSecrets.inline.password=${ACCESS_KEY}" \
--set "storageClassName=local-path"Running this command will automatically re-deploy the Orchestrate cluster with updated images. You can monitor the progress of the update with the same command that you ran when first installing the cluster.
# kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system -f --tail=-1Under unusual circumstances, the update may not proceed. For instance, if you incorrectly specify
your Depot username or password, the deployment fails and the FuzzballOrchestrate custom resource
is marked Degraded.
The operator retries on a timer, so once you correct the mistake the next attempt picks it up with no further action from you. It rechecks a missing storage class about once a minute, and retries a failed deployment every few minutes. A failed database pre-upgrade dump waits longer after each attempt, from about 30 seconds up to ten minutes.
If you would rather not wait for the next attempt, you can start the redeployment process immediately. Here is an example.
# kubectl get deployments -A
NAMESPACE NAME READY UP-TO-DATE AVAILABLE AGE
cert-manager cert-manager 3/3 3 3 26h
cert-manager cert-manager-cainjector 1/1 1 1 26h
cert-manager cert-manager-webhook 1/1 1 1 26h
cert-manager trust-manager 1/1 1 1 26h
fuzzball-system fuzzball-operator-controller-manager 1/1 1 1 26h
[...snip]
# kubectl rollout restart deployment/fuzzball-operator-controller-manager -n fuzzball-system
deployment.apps/fuzzball-operator-controller-manager restarted
# kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system -f --tail=-1
[...snip]
Resources:
+ 7 created
265 unchanged
Duration: 10s
2025-05-29T21:05:00Z DEBUG events Resources have been deployed successfully {"type": "Normal", "object": {"kind":"FuzzballOrchestrate","name":"fuzzball-orchestrate","uid":"22ca168d-a7a7-46f8-a144-33b42937f50f","apiVersion":"deployment.ciq.com/v1alpha1","resourceVersion":"11472"}, "reason": "DeploymentSucceeded"}
2025-05-29T21:05:00Z INFO Updated Fuzzball status to ReconciliationComplete - Reconciliation completed successfully {"controller": "fuzzballorchestrate", "controllerGroup": "deployment.ciq.com", "controllerKind": "FuzzballOrchestrate", "FuzzballOrchestrate": {"name":"fuzzball-orchestrate"}, "namespace": "", "name": "fuzzball-orchestrate", "reconcileID": "162d9224-5d75-49a7-829b-cc6ad8381ebe"}In the example above, we first determine the name of the deployment that we want to restart
(fuzzball-operator-controller-manager). We then use the proper kubectl command to restart the
deployment process. Finally, we use the appropriate command to monitor the logs as the deployment
proceeds to successful completion.
The exact method that you use to update the Fuzzball CLI will depend on the way in which you originally installed it. See the CLI installation documentation for more information.
Older versions of the Fuzzball CLI are not guaranteed to work with newer versions of Orchestrate, so it’s a good practice to keep it up to date. This may require announcing updates to your users so that they can update the CLI on their personal machines too!
AWS deployments are created using the Fuzzball CLI, and can also be updated via the update
command. The CLI applies the change to the underlying CloudFormation stack for you, so there is no
need to create and execute a change set by hand in the AWS console.
Pass--versionto select the target Fuzzball version. When the flag is omitted, the target version is the version embedded in the Fuzzball CLI itself (the CLI ships a CloudFormation template stamped with its own build version).
Before updating your Fuzzball Cluster take steps to ensure that no workflows are running and no users are accessing Orchestrate.
Fuzzball can be deployed in AWS using either Marketplace or CIQ Portal mode. Once a cluster has been deployed, you cannot switch between modes via update; destroy the stack and redeploy if you need to change modes. Theupdatecommand detects the existing mode and enforces this automatically.
Ensure that you have access to 10 or more elastic IP addresses in the region where you are performing the Fuzzball update. See the requirements page for more information.
Much like the interactive deployment method, you can update Fuzzball on AWS with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of confirming the parameters. Values from the previous deployment are pre-populated, so in most cases you can accept them as-is and let the update apply the new version embedded in your CLI.
Include the--dry-runflag to see what will happen before you actually execute the command.
$ fuzzball cluster aws updateYou can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs
and using the --non-interactive flag like so:
$ fuzzball cluster aws update \
--region "$REGION" \
--stack-name "$STACK_NAME" \
--non-interactivePass --version to select the target Fuzzball version; when omitted, the update applies the
Fuzzball version embedded in the CLI you are running. For a full list of available update options,
use the fuzzball cluster aws update --help command.
If your cluster was deployed in CIQ Portal mode, the update must re-hydrate the updated images into your account’s ECR, so you also need to supply Depot credentials (--depot-user+--depot-access-token, or theDEPOT_USERandDEPOT_ACCESS_TOKENenvironment variables) or--container-images-dirfor offline updates. See the CIQ Portal deployment section of the AWS deployment guide for details.
After the stack applies, update waits for the operator to report the cluster ready. The operator
marks the FuzzballOrchestrate custom resource Degraded when it hits a problem it can recover
from – a storage class that is not available yet, for example – and the command keeps waiting
while the operator retries. The command stops waiting once the resource has been Degraded for
several minutes across all of those retries, or as soon as a deployment itself fails.
The resource reports only that something went wrong. For the underlying error, read the operator log or the resource’s events:
$ kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system --tail=-1
$ kubectl describe fuzzballorchestrateInterrupting the command with Ctrl+C is safe: the operator continues applying the update in the
cluster. Re-run fuzzball cluster aws update, or run kubectl get fuzzballorchestrate -w, to
follow the rest of the progress.
When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.
The operator will not upgrade over a database it could not safely dump or restore: if the
pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded
and retries instead of continuing. Three conditions have a specific remedy:
pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm thedatabase-postgresqlsecret is present in the database namespace and the database pod isRunning/Ready; the next retry then proceeds.refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s/pulumi-statevolume, then let the retry continue.- A restore is reported pending but no dump exists – a leftover
pending-pg-restoreConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-databaseGCP deployments are created using the Fuzzball CLI, and can also be updated via the update
command.
Much like the interactive deployment method, you can update Fuzzball on GCP with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of filling in all necessary parameters.
Include the--dry-runflag to get an idea of what will happen before you actually execute the command.
$ fuzzball cluster gcp updateYou can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs
and using the --non-interactive flag like so:
$ fuzzball cluster gcp update \
--project "$PROJECT_ID" \
--region "$REGION" \
--deployment-name "unique-name" \
--version "$VERSION" \
--non-interactiveFor a full list of available update options, use the fuzzball cluster gcp update --help command.
When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.
The operator will not upgrade over a database it could not safely dump or restore: if the
pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded
and retries instead of continuing. Three conditions have a specific remedy:
pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm thedatabase-postgresqlsecret is present in the database namespace and the database pod isRunning/Ready; the next retry then proceeds.refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s/pulumi-statevolume, then let the retry continue.- A restore is reported pending but no dump exists – a leftover
pending-pg-restoreConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-databaseAzure deployments are created using the Fuzzball CLI, and can also be updated via the update
command.
Much like the interactive deployment method, you can update Fuzzball on Azure with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of filling in all necessary parameters.
Include the--dry-runflag to get an idea of what will happen before you actually execute the command.
$ fuzzball cluster azure updateYou can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs
and using the --non-interactive flag like so:
$ fuzzball cluster azure update \
--subscription "$SUBSCRIPTION_ID" \
--location "$LOCATION" \
--resource-group "$RESOURCE_GROUP" \
--version "$VERSION" \
--non-interactiveFor a full list of available update options, use the fuzzball cluster azure update --help command.
An update reuses the configuration recorded on the deployment’s resource group, so options you omit keep their deployed values instead of reverting to defaults. The recorded settings are:
- Fuzzball version and domain
- PostgreSQL engine version
- Custom substrate image IDs
- Keycloak realm name, username, and owner email
- Let’s Encrypt email
- Instance types and cluster administrators
- DNS zone resource group
- Log Analytics setting and Azure AD admin group ID
A routine version bump therefore needs only --version and leaves everything else as deployed.
Passing an option overrides the recorded value and records the new one for subsequent updates. Use
--dry-run to confirm which values an update will apply before you run it.
Omitting an option leaves its current value unchanged — it does not reset the value to its default. To change a setting, pass its option explicitly.
The engine version is the one setting the CLI reads back from the running server rather than from a
tag, so an update stops if it cannot determine it: if the resource group holds other PostgreSQL
servers and the CLI cannot tell which one is Fuzzball’s, or if the lookup itself fails. Applying a
default in either case risks upgrading a running server in place, so the update asks you to pass
--postgres-version explicitly instead. The exception is --dry-run, which changes nothing and so
is never blocked: it reports that the version could not be determined, leaves it unresolved in the
output, and notes that a real update will refuse until you pass --postgres-version. You can still
see what an update would do while you sort out the cause.
Resource groups deployed before the CLI began recording this configuration do not carry every value. To check what yours has recorded, inspect its tags:
$ az group show --name "$RESOURCE_GROUP" --query tagsThe CLI reads the PostgreSQL engine version from the running server, so that setting needs no action either way. A custom substrate image does need one action: nothing is recorded for it yet, so the next update would clear it.
If your deployment uses a custom substrate image and thesubstrate-image-base-idtag is absent, record the image on your next update. Otherwise the update clears it and falls back to marketplace discovery, and in regions where the marketplace offer is not published, workflow provisioning then fails withno substrate image available for VM size.
Pass the image’s gallery resource ID to record it without re-uploading anything:
$ fuzzball cluster azure update \
--subscription "$SUBSCRIPTION_ID" \
--location "$LOCATION" \
--resource-group "$RESOURCE_GROUP" \
--version "$VERSION" \
--substrate-image-base-id "/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$IMAGE_RG/providers/Microsoft.Compute/galleries/$GALLERY/images/$IMAGE/versions/$IMAGE_VERSION" \
--non-interactive$VERSION is the Fuzzball version to update to, which --non-interactive always requires;
$IMAGE_VERSION is the gallery image’s own version, the last segment of its resource ID.
Use --substrate-image-nvidia-id for the GPU image. Both flags take the resource ID of an
already-registered image – either a managed image or a gallery image version – and record it like
any other setting, so later updates inherit it. A deployment whose image was uploaded from a VHD
has a managed image ID recorded, which is the form to pass back here. Their
counterparts --substrate-image-base and --substrate-image-nvidia take a local VHD file instead,
which the CLI uploads and registers for you; pass the path form or the ID form for a given image,
not both.
When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.
The operator will not upgrade over a database it could not safely dump or restore: if the
pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded
and retries instead of continuing. Three conditions have a specific remedy:
pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm thedatabase-postgresqlsecret is present in the database namespace and the database pod isRunning/Ready; the next retry then proceeds.refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s/pulumi-statevolume, then let the retry continue.- A restore is reported pending but no dump exists – a leftover
pending-pg-restoreConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-databaseDocker Compose deployments are updated in place with fuzzball cluster docker-compose update. The
command operates on a single deployment, selected with --deployment-name (defaulting to
default).
$ fuzzball cluster docker-compose update --helpAdd--dry-runto anyupdateinvocation to preview the changes without applying them.
If the stack is running, update prompts before recreating the affected containers to apply the
change. Pass -y to skip the prompt in scripts.
Use --version to move the deployment to a specific Fuzzball version:
$ fuzzball cluster docker-compose update --version v4.1.0Alternatively, use --upgrade to bump the deployment to the Fuzzball image tag that matches the
version of the fuzzball CLI you are running:
$ fuzzball cluster docker-compose update --upgradeThe two flags are mutually exclusive.
Older versions of the Fuzzball CLI are not guaranteed to work with newer versions of Orchestrate, so keep the CLI up to date. See the CLI installation documentation for how to update it.
The --reset-* flags rotate the deployment’s generated credentials. Each flag takes the new value
as its argument:
$ fuzzball cluster docker-compose update --reset-database-password "$NEW_DB_PASSWORD"
$ fuzzball cluster docker-compose update --reset-keycloak-password "$NEW_KEYCLOAK_PASSWORD"
$ fuzzball cluster docker-compose update --reset-owner-password "$NEW_OWNER_PASSWORD"The--reset-*flags require the stack to be running. Start it first withfuzzball cluster docker-compose up.
You can switch an existing deployment between single- and multi-node, and between the regular and GPU-enabled substrate images, without redeploying:
| Flag | Effect |
|---|---|
--substrate-nodes N | Set the number of substrate nodes (default 1, max 255). |
--gpu | Switch substrate nodes to the GPU-enabled image variant. |
--no-gpu | Switch substrate nodes back to the regular (non-GPU) image variant. |
$ fuzzball cluster docker-compose update --gpu--gpu and --no-gpu are mutually exclusive – passing both is an error.
The older --multi-node and --single-node flags still work, for scripts that already use them,
but they are deprecated in favour of --substrate-nodes and print a warning when used. They are
mutually exclusive with --substrate-nodes and with each other: passing more than one of the three
is an error.
These nodes all run on the host the stack is deployed on, so raising the count over-provisions one machine rather than adding hardware. It is useful for exercising multi-node scheduling in development, not for adding capacity. To attach a node on another machine, see Adding an external substrate node.
| Flag | Effect |
|---|---|
--tls-renew | Renew the deployment’s wildcard TLS leaf certificate. |
--sync-files | Refresh the deployment’s generated files from the CLI’s embedded copies. |
--dry-run | Print what the update would do, then exit without changing anything. |
--tls-renew regenerates the leaf certificate under the deployment’s certs/ directory and
reloads nginx so the new certificate is served immediately:
$ fuzzball cluster docker-compose update --tls-renewIf the CA key is missing fromcerts/, Fuzzball generates a new CA as well as a new leaf. Every external substrate node then trusts the wrong CA and must be reconfigured withfuzzball cluster docker-compose generate-substrate-config. The command warns when this happens.
--sync-files rewrites the deployment’s generated files – docker-compose.yaml, the nginx
configuration, and the rest of the embedded tree – from the copies built into the CLI. Your .env
is preserved:
$ fuzzball cluster docker-compose update --sync-filesRun it after upgrading the CLI. A deployment keeps using the files it was created with until you
do, so a fix that ships in one of those files does not arrive through --version or --upgrade
alone.
It also overwrites any edits made directly to those files. Keep local changes in
docker-compose.override.yaml instead, which Fuzzball never writes and always applies last. See
Keeping local changes to the stack.
--dry-run prints the same plan the command would otherwise carry out and stops there. It combines
with any other flag, so you can check what a rotation or a node-count change would do first:
$ fuzzball cluster docker-compose update --substrate-nodes 3 --dry-runIf an update does not behave as expected, see the Troubleshooting section in the Docker Compose Stack guide for how to inspect the stack files directly.