Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Deleting and Retaining Workflows

A workflow record outlives the work it describes. Once a workflow finishes, its row and everything recorded alongside it – stages, events, job metrics, placements, cost snapshots – stay in the database so you can still inspect what ran. On a busy cluster, that data accumulates indefinitely, and fuzzball workflow list fills with runs nobody is looking at anymore.

This page covers deleting finished workflows, one at a time, in bulk, or automatically. Deletion is off by default: a cluster administrator chooses which roles may delete workflows in the central config (see Who can delete), and configures the scheduled sweep.

Deleting a workflow is permanent and takes its cost and usage records with it. Anything that bills or reports from per-workflow data must aggregate before the workflow is removed – there is no way to recover the detail afterwards.

What gets deleted

Deleting a workflow removes everything recorded about it:

CategoryRecords removed
Execution historyStages, stage events
MetricsJob metrics, including egress classification
SchedulingPlacement records
Cost accountingCost snapshots, egress snapshots, volume attachments
RoutingService endpoints
LeasesObject-cache leases
AnalyticsApplication usage events
DefinitionThe submitted workflow specification

Object-cache leases are easy to overlook but have real consequences. While a workflow holds a lease on a cached container image layer, the object cache cannot reclaim that layer’s disk space. Deleting finished workflows releases those leases, so the cache can free the space.

Two things a workflow holds are deliberately not on that list: its published storage volumes and the DNS records for its services. Those are live handles other subsystems still act on, not records of what happened, and deletion refuses while they exist rather than removing them – see below.

Deleting a single workflow

Only workflows that have finished, failed, or been canceled can be deleted. This page calls all three states “finished”. A running workflow still has allocations and node state pointing at it, so Fuzzball refuses the request and tells you to cancel the workflow first.

A finished workflow is also refused while its teardown is still in progress – while it holds a published storage volume or a service DNS record. Finishing and being torn down are not the same moment: a workflow whose volume cleanup failed, or one stopped with --force, reaches a terminal state with those handles still held. Deleting the workflow row then would strand a mount on a node or an A-record pointing at an address that can be reassigned, with nothing left to tie either back.

The error names the workflow and what it still holds. Wait for teardown to finish and retry.

A prune or retention sweep stops at the first workflow it cannot delete and reports how many it removed before that point. Because candidates are taken oldest-first, a workflow that stays stuck mid-teardown holds up everything behind it, so the count will not grow between runs. If a sweep keeps reporting the same workflow, resolve its teardown before expecting the rest of the backlog to clear.
$ fuzzball workflow delete 7bc0c8b4-ea91-4c42-8ce2-2bcd708d5223
Deleted workflow 7bc0c8b4-ea91-4c42-8ce2-2bcd708d5223

Deleting requires one of the roles enabled under workflow.delete in the central config, see Who can delete. With none enabled, delete and prune are refused with a message saying deletion is disabled on this cluster.

Deletion is irreversible and takes the workflow’s cost and usage history with it, so both delete and prune record audit events on clusters with auditing enabled: the authorization decision, and the deletion itself with the workflow IDs involved. Those events are the only record of what was removed once the rows are gone.

Who can delete

Nobody can delete workflows until a cluster administrator enables at least one role in the central config:

workflow:
  delete:
    clusterAdmin: true
    organizationOwner: true
    groupOwner: false
    workflowOwner: false
RoleWho it covers
clusterAdminCluster administrators
organizationOwnerOwners of the organization the workflow’s group belongs to
groupOwnerOwners of the group the workflow was submitted to, including the owner of a personal group
workflowOwnerThe users listed as the workflow’s owners, normally the submitter

A caller may delete a workflow when any enabled role applies to them. Every role is off unless set to true, and a misspelled role within delete is rejected when the config is applied rather than silently ignored. The setting governs fuzzball workflow delete and fuzzball workflow prune; the retention sweep below is an administrator-configured process and does not consult it.

On a federate deployment the federate cluster and the orchestrate cluster that owns the workflow each apply their own workflow.delete, so enable it on both.

Deleting in bulk

fuzzball workflow prune removes every finished workflow older than a given age. Running workflows are never touched, whatever their age.

Check what a cutoff would remove before you run it for real:

$ fuzzball workflow prune --older-than 720h --dry-run
Would delete 12 workflow(s) that finished before 2026-07-25T09:14:02Z:
  7bc0c8b4-ea91-4c42-8ce2-2bcd708d5223
  9e9566d2-3d26-4860-b1f7-0e1f2a3b4c5d
  ...

Then run the same command without --dry-run:

$ fuzzball workflow prune --older-than 720h
Deleted 12 workflow(s) that finished before 2026-07-25T09:14:02Z:
  7bc0c8b4-ea91-4c42-8ce2-2bcd708d5223
  ...

When nothing matches, the command says so and exits:

$ fuzzball workflow prune --older-than 8760h
No workflows finished before 2025-08-24T09:14:02Z

--older-than takes a Go duration string. There is no day unit, so express days in hours: 720h is thirty days, 168h is a week. The flag is required – a prune with no cutoff would match every finished workflow on the cluster.

Each run deletes at most 100 workflows. --limit raises or lowers that.

The age is measured from when a workflow ended, not when it started, so a workflow that ran for a week is not pruned early on the basis of its submission time.

A finished workflow with no recorded end time is never prunable, whatever its age – there is nothing to measure the cutoff against. Delete those individually if they accumulate.

Working through a backlog

A cluster with a long history is cleared over several passes rather than in one long database-heavy burst. If more workflows match the cutoff than the limit allows, the command says so:

$ fuzzball workflow prune --older-than 720h --limit 500
Deleted 500 workflow(s) that finished before 2026-07-25T09:14:02Z:
  ...

More workflows match this cutoff than the limit allowed. Run again to continue.

Prune only removes workflows you have permission to delete. Others are skipped rather than failing the run, so the command is usable on a shared cluster without needing rights over everyone else’s work, and the count of skipped workflows is reported alongside the deletions:

$ fuzzball workflow prune --older-than 720h
Deleted 3 workflow(s) that finished before 2026-07-25T09:14:02Z:
  ...

Skipped 24 workflow(s) you do not have permission to delete.

On a busy shared cluster your own oldest workflows can sit behind a long stretch of other people’s, and prune works past them to reach yours. If none of the matches are yours, the count is the whole answer – which is a different message from an empty cutoff, because it means the cutoff was fine and there was simply nothing of yours to remove:

$ fuzzball workflow prune --older-than 720h
No workflows deleted. Skipped 24 workflow(s) you do not have permission to delete.

Run prune against the Orchestrate cluster that owns the workflows. A Federate cluster holds mirrored metadata rather than the workflows themselves, so pruning there is refused – it would delete the mirror and leave the originals in place. The same applies to the scheduled sweep below: retention is ignored on a Federate cluster, and enabling it there logs a warning rather than deleting anything.

Deleting a single workflow through a Federate cluster does work: the request is forwarded to the Orchestrate cluster that owns it, and the Federate copy is removed once that succeeds, so the workflow stops appearing in Federate listings.

Deleting on a schedule

Fuzzball can delete finished workflows on a schedule instead. The retention sweep is disabled by default, so nothing is deleted automatically. It destroys data, and a cluster that silently began deleting history on upgrade would be a worse surprise than one that keeps accumulating it.

On a Kubernetes deployment, set it on the FuzzballOrchestrate resource:

spec:
  fuzzball:
    workflow:
      retention:
        enabled: true
        maxAge: 720h
        interval: 1h
        batchSize: 500

On a Docker Compose deployment, set the same fields under the retention key of the workflow service configuration (fuzzball-workflow.yaml).

FieldTypeDescription
enabledbooleanTurns the sweep on. Default false.
maxAgedurationHow long a finished workflow is kept, measured from when it ended. Required when enabled.
intervaldurationHow often the sweep runs. Defaults to 1h.
batchSizeintegerMaximum workflows removed per sweep. Zero uses the server default.

The sweep waits one interval before its first pass, so a mistaken maxAge is not acted on the instant the service restarts. It also refuses to start if maxAge is zero or negative, which would otherwise put the cutoff at the present moment and delete every finished workflow on the cluster.

A sweep that deletes something logs what it removed and whether more are still waiting, so you can tell a backlog from a steady state. Sweeps that find nothing to delete log at debug level.

Retention applies to every workflow on the cluster, not just your own. Decide the age against your longest reporting or billing cycle, not against how long anyone expects to browse their own run history.

Do not delete workflow rows directly

Do not delete workflow rows directly in PostgreSQL. Several tables hold rows keyed on a workflow without a foreign key to it, so a manual DELETE FROM workflows leaves them behind with nothing left to associate them with. Use fuzzball workflow delete or prune, which remove those rows in the same transaction.