Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Updating Fuzzball

You can (and should!) update your Fuzzball cluster when new versions are released. The update process differs somewhat depending on the environment.

Database migrations run before the Orchestrate pods are replaced, so when upgrading to v4.3, the release that adds endpoints on autoscaled service pools, every workflow endpoint read fails on the Orchestrate pods still running the old version until the rollout finishes.

If your deployment includes a Federate cluster, upgrade Federate before (or at the same time as) the Orchestrate clusters it fronts. An older Federate rejects the pool endpoint rows a newer Orchestrate replicates to it, and replication is not retried after the workflow starts, so those endpoints stay unreachable through Federate for the life of the workflow.

Please select your environment.

Updating Fuzzball in an on-prem environment consists of updating the Substrate installation running on compute nodes, the Orchestrate cluster, and ensuring that client CLIs are up-to-date.

Update Substrate on the compute nodes

Fuzzball Substrate is installed on each compute node from the Depot yum/dnf repository. When you update Orchestrate you should also update the substrate packages on every compute node — in particular the fuzzball-substrate-orchestrate extension — pinned to the same version as Orchestrate.

Do not simply run a bare dnf update. That pulls the newest available packages, which can overshoot your Orchestrate version (for example installing a 4.0.x extension against a 3.4.x Orchestrate) and conflicts with the recommendation to avoid skipping major or minor releases. The fuzzball-substrate-orchestrate version must match your Orchestrate version.

Set VERSION to match your Orchestrate version (without the leading v) and install the pinned packages, exactly as you did when installing Substrate on the compute node:

# VERSION="4.3.0" # matches orchestrate version without the leading v

# dnf install -y fuzzball-substrate fuzzball-substrate-orchestrate-${VERSION}-1 fuzzball-cli-${VERSION}-1

After the packages are updated, restart Substrate on the compute node so the new binaries take effect:

# systemctl restart fuzzball-substrate.service

You can confirm the extension version installed on the node with rpm:

# rpm -q fuzzball-substrate-orchestrate
fb version reports only the CLI and Agent versions, not the fuzzball-substrate-orchestrate extension, so version skew between Orchestrate and a compute node is not otherwise visible. The base Substrate platform is backward compatible with Orchestrate, but the fuzzball-substrate-orchestrate extension carries Orchestrate-matched functionality (for example the data-copy / S3 mover), so an older extension can silently misbehave against a newer Orchestrate. Keep the extension pinned to the Orchestrate version, and don’t forget to update Substrate whenever you update Orchestrate.

Updating the Orchestrate cluster

It’s easy to update your Fuzzball Orchestration cluster using the same K8s Custom Resource Definition (CRD) that installs Fuzzball. Simply repeat the steps that you used to install the Fuzzball Operator with a new version.

You can safely update from one patch release to another non-sequentially. (e.g Updating from v2.1.8 to v2.1.10 is supported.) But you should avoid skipping major or minor releases when updating.

First, make sure that you are logged into CIQ Depot.

# DEPOT_USER="" # populate with your username for CIQ Depot

# ACCESS_KEY="" # populate with the Depot key obtained from the CIQ sales/support team
# helm registry login depot.ciq.com --username "${DEPOT_USER}" --password "${ACCESS_KEY}"

Now simply update the version number and run the same commands that you originally ran to install the Fuzzball Operator the first time.

# VERSION="" # insert the updated version number

# CHART="oci://depot.ciq.com/fuzzball/fuzzball-images/helm/fuzzball-operator"
# helm upgrade --install fuzzball-operator "${CHART}" \
  --namespace fuzzball-system --create-namespace \
  --version "${VERSION}" \
  --set "image.tag=${VERSION}" \
  --set "imagePullSecrets.name=repository-ciq-com" \
  --set "imagePullSecrets.inline.registry=depot.ciq.com" \
  --set "imagePullSecrets.inline.username=${DEPOT_USER}" \
  --set "imagePullSecrets.inline.password=${ACCESS_KEY}" \
  --set "storageClassName=local-path"

Running this command will automatically re-deploy the Orchestrate cluster with updated images. You can monitor the progress of the update with the same command that you ran when first installing the cluster.

# kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system -f --tail=-1

Under unusual circumstances, the update may not proceed. For instance, if you incorrectly specify your Depot username or password, the deployment fails and the FuzzballOrchestrate custom resource is marked Degraded.

The operator retries on a timer, so once you correct the mistake the next attempt picks it up with no further action from you. It rechecks a missing storage class about once a minute, and retries a failed deployment every few minutes. A failed database pre-upgrade dump waits longer after each attempt, from about 30 seconds up to ten minutes.

If you would rather not wait for the next attempt, you can start the redeployment process immediately. Here is an example.

# kubectl get deployments -A
NAMESPACE            NAME                                   READY   UP-TO-DATE   AVAILABLE   AGE
cert-manager         cert-manager                           3/3     3            3           26h
cert-manager         cert-manager-cainjector                1/1     1            1           26h
cert-manager         cert-manager-webhook                   1/1     1            1           26h
cert-manager         trust-manager                          1/1     1            1           26h
fuzzball-system      fuzzball-operator-controller-manager   1/1     1            1           26h
[...snip]

# kubectl rollout restart deployment/fuzzball-operator-controller-manager -n fuzzball-system
deployment.apps/fuzzball-operator-controller-manager restarted

# kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system -f --tail=-1
[...snip]
Resources:
    + 7 created
    265 unchanged

Duration: 10s

2025-05-29T21:05:00Z	DEBUG	events	Resources have been deployed successfully	{"type": "Normal", "object": {"kind":"FuzzballOrchestrate","name":"fuzzball-orchestrate","uid":"22ca168d-a7a7-46f8-a144-33b42937f50f","apiVersion":"deployment.ciq.com/v1alpha1","resourceVersion":"11472"}, "reason": "DeploymentSucceeded"}
2025-05-29T21:05:00Z	INFO	Updated Fuzzball status to ReconciliationComplete - Reconciliation completed successfully	{"controller": "fuzzballorchestrate", "controllerGroup": "deployment.ciq.com", "controllerKind": "FuzzballOrchestrate", "FuzzballOrchestrate": {"name":"fuzzball-orchestrate"}, "namespace": "", "name": "fuzzball-orchestrate", "reconcileID": "162d9224-5d75-49a7-829b-cc6ad8381ebe"}

In the example above, we first determine the name of the deployment that we want to restart (fuzzball-operator-controller-manager). We then use the proper kubectl command to restart the deployment process. Finally, we use the appropriate command to monitor the logs as the deployment proceeds to successful completion.

Update the Fuzzball CLI on clients

The exact method that you use to update the Fuzzball CLI will depend on the way in which you originally installed it. See the CLI installation documentation for more information.

Older versions of the Fuzzball CLI are not guaranteed to work with newer versions of Orchestrate, so it’s a good practice to keep it up to date. This may require announcing updates to your users so that they can update the CLI on their personal machines too!

AWS deployments are created using the Fuzzball CLI, and can also be updated via the update command. The CLI applies the change to the underlying CloudFormation stack for you, so there is no need to create and execute a change set by hand in the AWS console.

Pass --version to select the target Fuzzball version. When the flag is omitted, the target version is the version embedded in the Fuzzball CLI itself (the CLI ships a CloudFormation template stamped with its own build version).
Before updating your Fuzzball Cluster take steps to ensure that no workflows are running and no users are accessing Orchestrate.
Fuzzball can be deployed in AWS using either Marketplace or CIQ Portal mode. Once a cluster has been deployed, you cannot switch between modes via update; destroy the stack and redeploy if you need to change modes. The update command detects the existing mode and enforces this automatically.
Ensure that you have access to 10 or more elastic IP addresses in the region where you are performing the Fuzzball update. See the requirements page for more information.

Interactive update

Much like the interactive deployment method, you can update Fuzzball on AWS with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of confirming the parameters. Values from the previous deployment are pre-populated, so in most cases you can accept them as-is and let the update apply the new version embedded in your CLI.

Include the --dry-run flag to see what will happen before you actually execute the command.
$ fuzzball cluster aws update

Non-interactive update

You can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs and using the --non-interactive flag like so:

$ fuzzball cluster aws update \
  --region "$REGION" \
  --stack-name "$STACK_NAME" \
  --non-interactive

Pass --version to select the target Fuzzball version; when omitted, the update applies the Fuzzball version embedded in the CLI you are running. For a full list of available update options, use the fuzzball cluster aws update --help command.

If your cluster was deployed in CIQ Portal mode, the update must re-hydrate the updated images into your account’s ECR, so you also need to supply Depot credentials (--depot-user + --depot-access-token, or the DEPOT_USER and DEPOT_ACCESS_TOKEN environment variables) or --container-images-dir for offline updates. See the CIQ Portal deployment section of the AWS deployment guide for details.

Waiting for the operator

After the stack applies, update waits for the operator to report the cluster ready. The operator marks the FuzzballOrchestrate custom resource Degraded when it hits a problem it can recover from – a storage class that is not available yet, for example – and the command keeps waiting while the operator retries. The command stops waiting once the resource has been Degraded for several minutes across all of those retries, or as soon as a deployment itself fails.

The resource reports only that something went wrong. For the underlying error, read the operator log or the resource’s events:

$ kubectl logs -l app.kubernetes.io/name=fuzzball-operator -n fuzzball-system --tail=-1

$ kubectl describe fuzzballorchestrate

Interrupting the command with Ctrl+C is safe: the operator continues applying the update in the cluster. Re-run fuzzball cluster aws update, or run kubectl get fuzzballorchestrate -w, to follow the rest of the progress.

PostgreSQL major version upgrade

When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.

The operator will not upgrade over a database it could not safely dump or restore: if the pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded and retries instead of continuing. Three conditions have a specific remedy:

  • pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm the database-postgresql secret is present in the database namespace and the database pod is Running/Ready; the next retry then proceeds.
  • refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s /pulumi-state volume, then let the retry continue.
  • A restore is reported pending but no dump exists – a leftover pending-pg-restore ConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-database

GCP deployments are created using the Fuzzball CLI, and can also be updated via the update command.

Interactive update

Much like the interactive deployment method, you can update Fuzzball on GCP with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of filling in all necessary parameters.

Include the --dry-run flag to get an idea of what will happen before you actually execute the command.
$ fuzzball cluster gcp update

Non-interactive update to a new version

You can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs and using the --non-interactive flag like so:

$ fuzzball cluster gcp update \
  --project "$PROJECT_ID" \
  --region "$REGION" \
  --deployment-name "unique-name" \
  --version "$VERSION" \
  --non-interactive

For a full list of available update options, use the fuzzball cluster gcp update --help command.

PostgreSQL major version upgrade

When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.

The operator will not upgrade over a database it could not safely dump or restore: if the pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded and retries instead of continuing. Three conditions have a specific remedy:

  • pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm the database-postgresql secret is present in the database namespace and the database pod is Running/Ready; the next retry then proceeds.
  • refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s /pulumi-state volume, then let the retry continue.
  • A restore is reported pending but no dump exists – a leftover pending-pg-restore ConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-database

Azure deployments are created using the Fuzzball CLI, and can also be updated via the update command.

Interactive update

Much like the interactive deployment method, you can update Fuzzball on Azure with an interactive approach by running update without any options. The Fuzzball CLI will guide you through the process of filling in all necessary parameters.

Include the --dry-run flag to get an idea of what will happen before you actually execute the command.
$ fuzzball cluster azure update

Non-interactive update to a new version

You can also update Fuzzball non-interactively by supplying the parameters in option/argument pairs and using the --non-interactive flag like so:

$ fuzzball cluster azure update \
  --subscription "$SUBSCRIPTION_ID" \
  --location "$LOCATION" \
  --resource-group "$RESOURCE_GROUP" \
  --version "$VERSION" \
  --non-interactive

For a full list of available update options, use the fuzzball cluster azure update --help command.

Configuration preserved across updates

An update reuses the configuration recorded on the deployment’s resource group, so options you omit keep their deployed values instead of reverting to defaults. The recorded settings are:

  • Fuzzball version and domain
  • PostgreSQL engine version
  • Custom substrate image IDs
  • Keycloak realm name, username, and owner email
  • Let’s Encrypt email
  • Instance types and cluster administrators
  • DNS zone resource group
  • Log Analytics setting and Azure AD admin group ID

A routine version bump therefore needs only --version and leaves everything else as deployed. Passing an option overrides the recorded value and records the new one for subsequent updates. Use --dry-run to confirm which values an update will apply before you run it.

Omitting an option leaves its current value unchanged — it does not reset the value to its default. To change a setting, pass its option explicitly.

The engine version is the one setting the CLI reads back from the running server rather than from a tag, so an update stops if it cannot determine it: if the resource group holds other PostgreSQL servers and the CLI cannot tell which one is Fuzzball’s, or if the lookup itself fails. Applying a default in either case risks upgrading a running server in place, so the update asks you to pass --postgres-version explicitly instead. The exception is --dry-run, which changes nothing and so is never blocked: it reports that the version could not be determined, leaves it unresolved in the output, and notes that a real update will refuse until you pass --postgres-version. You can still see what an update would do while you sort out the cause.

Deployments created by an earlier CLI

Resource groups deployed before the CLI began recording this configuration do not carry every value. To check what yours has recorded, inspect its tags:

$ az group show --name "$RESOURCE_GROUP" --query tags

The CLI reads the PostgreSQL engine version from the running server, so that setting needs no action either way. A custom substrate image does need one action: nothing is recorded for it yet, so the next update would clear it.

If your deployment uses a custom substrate image and the substrate-image-base-id tag is absent, record the image on your next update. Otherwise the update clears it and falls back to marketplace discovery, and in regions where the marketplace offer is not published, workflow provisioning then fails with no substrate image available for VM size.

Pass the image’s gallery resource ID to record it without re-uploading anything:

$ fuzzball cluster azure update \
  --subscription "$SUBSCRIPTION_ID" \
  --location "$LOCATION" \
  --resource-group "$RESOURCE_GROUP" \
  --version "$VERSION" \
  --substrate-image-base-id "/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$IMAGE_RG/providers/Microsoft.Compute/galleries/$GALLERY/images/$IMAGE/versions/$IMAGE_VERSION" \
  --non-interactive

$VERSION is the Fuzzball version to update to, which --non-interactive always requires; $IMAGE_VERSION is the gallery image’s own version, the last segment of its resource ID.

Use --substrate-image-nvidia-id for the GPU image. Both flags take the resource ID of an already-registered image – either a managed image or a gallery image version – and record it like any other setting, so later updates inherit it. A deployment whose image was uploaded from a VHD has a managed image ID recorded, which is the form to pass back here. Their counterparts --substrate-image-base and --substrate-image-nvidia take a local VHD file instead, which the CLI uploads and registers for you; pass the path form or the ID form for a given image, not both.

PostgreSQL major version upgrade

When updating to a release that includes a PostgreSQL major version upgrade (for example, upgrading to v3.3.0 which moves to PostgreSQL 16), the Fuzzball operator handles the in-place database migration automatically during the update process. No manual database migration is required.

The operator will not upgrade over a database it could not safely dump or restore: if the pre-upgrade dump or restore cannot proceed, it marks the FuzzballOrchestrate resource Degraded and retries instead of continuing. Three conditions have a specific remedy:

  • pg-migrate dump failed: ... database secret not found, or a PostgreSQL connection error. The database pod exists but its secret is missing or PostgreSQL is not yet reachable. Confirm the database-postgresql secret is present in the database namespace and the database pod is Running/Ready; the next retry then proceeds.
  • refusing to restore: ... the consumed dump ... could not be removed and the restore would repeat. The operator’s dump volume is not writable, so restoring would loop and roll the database back each cycle. Restore write access to the operator’s /pulumi-state volume, then let the retry continue.
  • A restore is reported pending but no dump exists – a leftover pending-pg-restore ConfigMap from an earlier interrupted upgrade. Delete it in the database namespace to clear the signal:
# kubectl delete configmap pending-pg-restore -n fuzzball-database
Update instructions coming soon!

Docker Compose deployments are updated in place with fuzzball cluster docker-compose update. The command operates on a single deployment, selected with --deployment-name (defaulting to default).

$ fuzzball cluster docker-compose update --help
Add --dry-run to any update invocation to preview the changes without applying them.

If the stack is running, update prompts before recreating the affected containers to apply the change. Pass -y to skip the prompt in scripts.

Upgrading the Fuzzball version

Use --version to move the deployment to a specific Fuzzball version:

$ fuzzball cluster docker-compose update --version v4.1.0

Alternatively, use --upgrade to bump the deployment to the Fuzzball image tag that matches the version of the fuzzball CLI you are running:

$ fuzzball cluster docker-compose update --upgrade

The two flags are mutually exclusive.

Older versions of the Fuzzball CLI are not guaranteed to work with newer versions of Orchestrate, so keep the CLI up to date. See the CLI installation documentation for how to update it.

Rotating credentials

The --reset-* flags rotate the deployment’s generated credentials. Each flag takes the new value as its argument:

$ fuzzball cluster docker-compose update --reset-database-password "$NEW_DB_PASSWORD"

$ fuzzball cluster docker-compose update --reset-keycloak-password "$NEW_KEYCLOAK_PASSWORD"

$ fuzzball cluster docker-compose update --reset-owner-password "$NEW_OWNER_PASSWORD"
The --reset-* flags require the stack to be running. Start it first with fuzzball cluster docker-compose up.

Changing the node and GPU configuration

You can switch an existing deployment between single- and multi-node, and between the regular and GPU-enabled substrate images, without redeploying:

FlagEffect
--substrate-nodes NSet the number of substrate nodes (default 1, max 255).
--gpuSwitch substrate nodes to the GPU-enabled image variant.
--no-gpuSwitch substrate nodes back to the regular (non-GPU) image variant.
$ fuzzball cluster docker-compose update --gpu

--gpu and --no-gpu are mutually exclusive – passing both is an error.

The older --multi-node and --single-node flags still work, for scripts that already use them, but they are deprecated in favour of --substrate-nodes and print a warning when used. They are mutually exclusive with --substrate-nodes and with each other: passing more than one of the three is an error.

These nodes all run on the host the stack is deployed on, so raising the count over-provisions one machine rather than adding hardware. It is useful for exercising multi-node scheduling in development, not for adding capacity. To attach a node on another machine, see Adding an external substrate node.

Other maintenance operations

FlagEffect
--tls-renewRenew the deployment’s wildcard TLS leaf certificate.
--sync-filesRefresh the deployment’s generated files from the CLI’s embedded copies.
--dry-runPrint what the update would do, then exit without changing anything.

Renewing the TLS certificate

--tls-renew regenerates the leaf certificate under the deployment’s certs/ directory and reloads nginx so the new certificate is served immediately:

$ fuzzball cluster docker-compose update --tls-renew
If the CA key is missing from certs/, Fuzzball generates a new CA as well as a new leaf. Every external substrate node then trusts the wrong CA and must be reconfigured with fuzzball cluster docker-compose generate-substrate-config. The command warns when this happens.

Refreshing the generated files

--sync-files rewrites the deployment’s generated files – docker-compose.yaml, the nginx configuration, and the rest of the embedded tree – from the copies built into the CLI. Your .env is preserved:

$ fuzzball cluster docker-compose update --sync-files

Run it after upgrading the CLI. A deployment keeps using the files it was created with until you do, so a fix that ships in one of those files does not arrive through --version or --upgrade alone.

It also overwrites any edits made directly to those files. Keep local changes in docker-compose.override.yaml instead, which Fuzzball never writes and always applies last. See Keeping local changes to the stack.

Previewing an update

--dry-run prints the same plan the command would otherwise carry out and stops there. It combines with any other flag, so you can check what a rotation or a node-count change would do first:

$ fuzzball cluster docker-compose update --substrate-nodes 3 --dry-run

If an update does not behave as expected, see the Troubleshooting section in the Docker Compose Stack guide for how to inspect the stack files directly.