Debugging
This section provides documentation for debugging and troubleshooting Fuzzball clusters. These tools are designed for administrators to diagnose issues, monitor service behavior, and perform advanced maintenance tasks.
Debug commands are useful for:
- Viewing service logs to diagnose issues
- Creating debug dumps to share with the support team
You will need to install the Fuzzball CLI (see the CLI installation documentation for details), set up a Fuzzball admin context and authenticate like so:
$ fuzzball context create default-admin-context <api_url>
Configuration for "default-admin-context" created.
Configuration for "default-admin-context" now in use.
$ FUZZBALL_ADMIN_USER="admin_username"
$ read -rs FUZZBALL_ADMIN_PASSWORD # read interactively to set variable without echoing to screen
$ export FUZZBALL_ADMIN_USER FUZZBALL_ADMIN_PASSWORD
$ fuzzball context login
Logging into current cluster context...The fuzzball debug logs command allows administrators to view
logs from Fuzzball services running in the cluster. This is useful for debugging
service issues, monitoring service behavior, and tracking down errors.
$ fuzzball debug logs <SERVICE_NAME> [flags]The following Fuzzball services can be queried for logs:
agent- Fuzzball agent serviceaudit- Audit logging serviceauth- Authentication servicebilling- Billing servicecluster-admin- Cluster administration servicejetstream- JetStream message brokeropenapi- OpenAPI serviceorchestrator- Orchestrator serviceprovision- Provisioning serviceschedule- Scheduling servicestorage- Storage servicesubstrate-bridge- Substrate bridge serviceui- Web UI serviceworkflow- Workflow service
| Flag | Short | Description | Default |
|---|---|---|---|
--since | -s | Show logs since a specific time (RFC3339 or YYYY-MM-DD format) | - |
--since-duration | -d | Show logs since a specific duration (e.g., ‘1h’, ‘30m’) | - |
--tail | -t | Number of lines to show from the end of the logs | 100 |
--follow | -f | Follow the log output (stream logs) | false |
--min-log-level | - | Minimum log level to display (1: Debug, 2: Info, 3: Warning, 4: Error, 5: Fatal) | 0 |
--trace-id | - | Filter logs by trace ID | - |
--ignore-errors | - | Ignore errors when fetching logs | false |
--json | -j | Output in JSON format | false |
View the last 100 lines of logs from the workflow service:
$ fuzzball debug logs workflow
[fuzzball-workflow-0]: 2026-01-13T10:15:23Z INFO Starting workflow service
[fuzzball-workflow-0]: 2026-01-13T10:15:24Z INFO Connected to database
[fuzzball-workflow-0]: 2026-01-13T10:15:25Z INFO Workflow service readyView logs from multiple services:
$ fuzzball debug logs workflow storage
[fuzzball-workflow-0]: 2026-01-13T10:15:23Z INFO Starting workflow service
[fuzzball-storage-0]: 2026-01-13T10:15:22Z INFO Storage service initializedFollow logs in real-time:
$ fuzzball debug logs workflow --follow
[fuzzball-workflow-0]: 2026-01-13T10:15:23Z INFO Starting workflow service
[fuzzball-workflow-0]: 2026-01-13T10:15:24Z INFO Processing workflow requestShow logs since a specific time:
$ fuzzball debug logs workflow --since 2026-01-13T10:00:00ZShow logs from the last hour:
$ fuzzball debug logs workflow --since-duration 1hFilter by minimum log level (errors and above):
$ fuzzball debug logs workflow --min-log-level 4Show the last 500 lines:
$ fuzzball debug logs workflow --tail 500Output logs in JSON format:
$ fuzzball debug logs workflow --json
{"pod_name":"fuzzball-workflow-0","message":"2026-01-13T10:15:23Z INFO Starting workflow service","level":2}
{"pod_name":"fuzzball-workflow-0","message":"2026-01-13T10:15:24Z INFO Connected to database","level":2}- The
--sinceand--since-durationflags cannot be used together - When using
--follow, logs will stream continuously until interrupted (Ctrl+C) - The
--jsonflag is useful for programmatic parsing and integration with log analysis tools - Log levels: 0 (Unspecified), 1 (Debug), 2 (Info), 3 (Warning), 4 (Error), 5 (Fatal)
The fuzzball debug dump command creates a comprehensive dump of the
Fuzzball cluster state, including configuration, logs, and diagnostic
information. This is useful for gathering all relevant information when
reporting issues or performing detailed troubleshooting.
$ fuzzball debug dump [flags]| Flag | Short | Description | Default |
|---|---|---|---|
--dest | -d | Destination path to save the dump | Current directory |
--namespace | -n | Additional namespaces to include in the dump; repeat the flag or pass a comma-separated list | - |
The dump always covers Fuzzball’s own namespaces and those of its
dependencies:
fuzzball, fuzzball-federate, fuzzball-system, fuzzball-database,
fuzzball-identity, cert-manager, and metallb-system. A namespace that
does not exist on your deployment contributes nothing beyond an empty
events.txt.
Create a cluster dump in the current directory:
$ fuzzball debug dump
$ ls -lhd fuzzball-cluster-dump-*
drwxr-xr-x@ 10 user group 320B Jan 15 12:11 fuzzball-cluster-dump-20260115-121100
$ tree -L2 fuzzball-cluster-dump-20260115-121100
fuzzball-cluster-dump-20260115-121100
├── cert-manager
│ ├── certificates
│ ├── configmaps
│ ├── deployments
│ ├── events.txt
│ ├── metrics
│ ├── pods
│ ├── rbac
│ ├── secrets
│ └── services
├── cluster
│ ├── clusterissuers
│ ├── collector-diagnostics.txt
│ ├── crds
│ ├── ingressclasses
│ ├── jetstream
│ ├── metrics
│ ├── nodes
│ ├── rbac
│ └── storage
├── cluster_info.txt
├── errors.txt
├── fuzzball
│ ├── certificates
│ ├── configmaps
│ ├── deployments
│ ├── events.txt
│ ├── ingresses
│ ├── metrics
│ ├── pods
│ ├── rbac
│ ├── scheduler
│ ├── secrets
│ ├── services
│ └── storage
├── fuzzball-database
│ ├── certificates
│ ├── configmaps
│ ├── deployments
│ ├── events.txt
│ ├── health
│ ├── metrics
│ ├── pods
│ ├── rbac
│ ├── secrets
│ ├── services
│ └── storage
├── fuzzball-federate
│ ├── certificates
│ ├── configmaps
│ ├── deployments
│ ├── events.txt
│ ├── ingresses
│ ├── metrics
│ ├── pods
│ ├── rbac
│ ├── secrets
│ ├── services
│ └── storage
├── fuzzball-identity
│ ├── certificates
│ ├── configmaps
│ ├── events.txt
│ ├── ingresses
│ ├── metrics
│ ├── pods
│ ├── secrets
│ └── services
├── fuzzball-system
│ ├── configmaps
│ ├── deployments
│ ├── events.txt
│ ├── metrics
│ ├── pods
│ ├── rbac
│ ├── secrets
│ ├── services
│ └── storage
├── manifest.json
└── metallb-system
└── events.txt
$ du -sh fuzzball-cluster-dump-20260115-121100
23M fuzzball-cluster-dump-20260115-121100 # small dump for a test deployment
$ tar -czf fuzzball-cluster-dump-20260115-121100.tgz \
fuzzball-cluster-dump-20260115-121100 # tar or zip for easy sharingThis creates a directory named fuzzball-cluster-dump-YYYYMMDD-HHMMSS containing the cluster state.
Include an additional namespace beyond the defaults:
$ fuzzball debug dump -n my-workload-namespaceThe cluster dump includes:
- Service logs and pod descriptions from all Fuzzball components, plus an
aggregated
errors.txtcollecting every log line that mentionserror,fatal, orpanic(each pod also gets its ownerrors.txt) - Kubernetes resource configurations (deployments, services, configmaps, RBAC, events; secret values are redacted)
- Node and pod CPU/memory usage from metrics-server
(
cluster/metrics/nodes.txt,<namespace>/metrics/pods.txt); when metrics-server is unavailable, raw kubelet stats summaries are written tocluster/metrics/kubelet-summary/instead - JetStream stream and consumer reports, including per-stream message counts
(
cluster/jetstream/) - The
FuzzballOrchestrateandFuzzballFederatecustom resources (cluster/crds/) - Ingresses and ingress classes (
<namespace>/ingresses/,cluster/ingressclasses/) - Certificates and cluster issuers from cert-manager
(
<namespace>/certificates/,cluster/clusterissuers/) - Storage classes, CSI drivers, persistent volumes, and persistent volume
claims (
cluster/storage/,<namespace>/storage/) - Postgres health checks for in-cluster databases – settings, activity,
connection counts, locks, and long-running queries
(
fuzzball-database/health/) - The scheduler’s view of the connected compute nodes, including health
state, reliability score, and each node’s recent health event history
(
fuzzball/scheduler/) cluster/collector-diagnostics.txt, recording which Kubernetes APIs the dump could reach and therefore which optional collectors could not run; a collector that is skipped also leaves aSKIPPED.txtin its own directorymanifest.json, indexing every artifact and naming any source that failed (see Reading the Manifest under Support Bundle – the same format applies to a plain dump)
- The dump process may take several minutes depending on cluster size and log volume
- Ensure sufficient disk space is available for the dump
- The dump directory is timestamped to prevent overwrites
- Review the dump contents before sharing, as it may contain sensitive information
- Use this command when reporting issues to CIQ support for comprehensive diagnostics
A job that ends with a message naming a node died because that node stopped responding, not because of anything in the workflow:
node 10.0.0.42/9100: memory errors; job stopped responding on this node
Where a fault is named, the node was carrying that condition when the job died. Where only the node is named, the node had no active condition – the job may have hung, been killed for memory, or deadlocked, and the node is reported so you know where it was running.
Start with the node rather than the workflow:
$ fuzzball node show 10.0.0.42/9100
$ fuzzball node events 10.0.0.42/9100 --latestSee Node Health Monitoring for how attribution is decided.
fuzzball debug report collects the cluster dump above and diagnostics
from the compute nodes, into one directory to send to support.
Node diagnostics travel over the substrate channel the nodes already hold, so no SSH access and no inbound port on the node is required. Each node contributes its health report and the recent kernel fault records behind it – the machine checks and GPU Xid records that its reliability score is derived from.
$ fuzzball debug report [flags]Press Ctrl+C to stop a collection in progress. Whatever has already been
written stays in the destination directory, but the bundle is left
unfinished: its manifest.json is not amended, and node data is only
written once every node has answered, so an interrupted collection usually
has no node directories at all.
| Flag | Short | Description | Default |
|---|---|---|---|
--dest | -d | Directory to write the bundle into | Current directory |
--nodes | Which nodes to collect from: none, unhealthy, all | none | |
--node | Collect from a single node, whatever its health | ||
--max-fault-lines | Maximum recent kernel fault lines per node | Node default |
--nodes and --node are mutually exclusive.
Collect the control plane and the nodes worth looking at – those that are degraded, unknown, cordoned or offline:
$ fuzzball debug report --nodes unhealthy
bundle written to fuzzball-debug-report-20260812-104233Collect from one specific node, whatever its health:
$ fuzzball debug report --node 10.0.0.42/9100fuzzball-debug-report-20260812-104233
├── cluster/ # the cluster dump described above
├── nodes/
│ ├── gpu-14/
│ │ ├── health.json # signals, conditions and score inputs
│ │ └── kernel-faults.txt # the machine check and Xid records behind them
│ └── gpu-17/
│ └── error.txt # this node did not answer, and why
└── manifest.json
Every bundle carries a manifest.json listing each source, how many artifacts
it produced, and for anything that failed, the reason:
{
"generatedAt": "2026-08-12T10:42:33Z",
"scrubbed": true,
"sources": [
{ "name": "pod-logs", "artifacts": 42 },
{ "name": "nodes", "artifacts": 2 },
{ "name": "node-diagnostics", "artifacts": 0,
"error": "1 of 2 nodes did not answer: 10.0.0.17/7334" }
],
"errors": ["node-diagnostics: 1 of 2 nodes did not answer: 10.0.0.17/7334"]
}
scrubbed reports whether secret values were removed. Every Kubernetes secret
in the bundle keeps its name, type and key names, but each value is replaced
with its length – redacted (7 bytes) rather than the credential. That is
enough to answer whether a key is present, whether it is empty, and whether a
certificate is plausibly a certificate, without the bundle carrying anything
worth stealing. Credential-bearing fields in the dumped FuzzballOrchestrate
and FuzzballFederate custom resources (passwords, access tokens, private
keys) are redacted the same way, and the last-applied-configuration
annotation, which re-embeds the originally applied values, is removed.
Scrubbing covers secret values, not everything a bundle can contain. Pod logs, events and ConfigMaps are collected verbatim, and an application that logs a token puts that token in the bundle. Review a bundle before sending it outside your organization.
Collection is best effort by design. A node that does not answer is recorded in the bundle rather than failing it, because during an incident an unreachable node is frequently the fault being investigated – and a tool that returned nothing in that case would be useless exactly when it is needed.
Check
errorsbefore sending a bundle. An empty list means everything asked for was collected. Every source is listed insourceswhether or not it succeeded, so anode-diagnosticsentry with noerrormeans the nodes were collected – not that they were skipped. Note the two distinct sources:nodesis the cluster’s Kubernetes Node objects, whilenode-diagnosticsis the per-node health snapshots gathered by--nodes.node-diagnosticsalways reports0artifacts; read itserrorfield, not its count.
Collecting a bundle does not change what the nodes report next. Some node signals are counted as growth since the previous health report. A bundle reads those counters without resetting them, so collecting one during an incident will not hide a condition from the next health report.
When the control plane itself is unreachable, fuzzball debug report cannot
run. Fuzzball ships an sos report plugin
for that case, installed with the fuzzball-substrate-orchestrate package:
$ sos report -o fuzzballIt collects the same host-local sources a node would have contributed –
substrate configuration, service logs, EDAC counters, GPU state and kernel
fault records – and falls back to kubectl for cluster state on a control
plane host. Credentials in configuration files are scrubbed automatically.
When a job is placed, the scheduler publishes a resources_allocated workflow event that
shows both what the job requested and the CPU placement the scheduler computed for it:
$ fuzzball workflow events WORKFLOW
...
Resources allocated (node: 10.0.0.3/7331, requested cpu: 28, requested memory: 28 GB, projected cpus: 28 (ids: 0-6,8-14,16-22,24-30), projected numa nodes: 4 (ids: 0-3))requested cpu,requested memory, andrequested devicesecho the resource request from the workflow definition. When the request setssockets,requested cpuis thecoresvalue, which counts cores per socket;requested devicesis the total device count across all device types.projected cpusandprojected numa nodesare the CPUs the scheduler selected for the job and the NUMA nodes they span, each rendered as a count followed by the specific IDs (count (ids: ...)). When the job requests hardware threads or multiple sockets, the projected CPU count is larger thanrequested cpubecause it counts every schedulable CPU the job receives. For multi-node jobs, per-node values are separated by semicolons in the same order as the comma-separated node list, and each per-node count compares againstrequested cpuindividually.
The projected values are the scheduler’s placement decision, not a reading from the node: the node’s substrate service independently computes the enforced CPU set from the same lease. The two normally match, but the projection is not a guarantee — during upgrades, transient differences in scheduler and node resource accounting, or if placement behavior ever diverges between scheduler and node, the CPUs a job actually runs on can differ from the projection.
When available, the kernel-enforced values appear on the container_started event, read
from the container process on the node when it starts. They use the same count-and-IDs
format, so enforced cpus compares directly against projected cpus and
enforced numa nodes against projected numa nodes:
Container started on node: 10.0.0.3/7331 (enforced cpus: 28 (ids: 0-6,8-14,16-22,24-30), enforced numa nodes: 4 (ids: 0-3))The enforced values can be omitted when the node’s substrate version predates the feature,
when the container process exits before the values can be read, or when the node cannot be
reached. A container_started event without enforced fields is normal during upgrades or
transient node-communication failures.
Comparing the resources_allocated and container_started events for a stage shows
requested, projected, and enforced values side by side. For multi-node jobs the
container_started values cover the first node only. When the enforced values are
missing from the event, the same values can always be read from inside a job:
$ grep -E 'Cpus_allowed_list|Mems_allowed_list' /proc/self/statusA projected CPU count lower than requested cpu (without hardware threads or sockets in
play) indicates the scheduler could not compute a full placement and should be reported.
The scheduler publishes a scheduling_blocked workflow event when a job’s allocation is
waiting for node reuse because its provisioner definition’s node pool has reached the
maxNodes cap. The event means the job is queued, not rejected: it proceeds once a node
in the pool becomes free. This is expected when the cluster is at capacity and workflows
are competing for resources.
- The definition is dynamic and its pool has reached the effective cap
(
min(maxNodes, 128)). - The allocation fits within the cap, but no node is currently free for reuse.
- Two companion events are published, once per blocked allocation:
provisioning_blocked, which carries thedefinition_id,total_nodes, andmax_nodesattributes, and the genericscheduling_blockedevent. - With backfill enabled (the default), an at-cap allocation may instead hold a
reservation and publish a
reservation_heldevent.
Because the events are published once per blocked allocation, judge persistence by whether the allocation stays unscheduled, not by counting events.
Follow the workflow’s events and look for
provisioning_blocked(itsdefinition_idattribute names the affected definition):$ fuzzball workflow events WORKFLOW --followAsk the scheduler why the allocation is not scheduled. A
PROVISIONING_BLOCKEDreason indicates the pool cap; see Queue Observability for the full reasons table:$ fuzzball workflow why WORKFLOW ALLOCATIONReview the definition’s current
maxNodesvalue:$ fuzzball node provisioner get compute-pool
Increase maxNodes: Raise the
maxNodesvalue on the provisioner definition and apply the updated configuration:definitions: - id: compute-pool provisioner: pbs ttl: 3600 maxNodes: 64 provisionerSpec: cpu: 8 memory: "32GiB" queue: "workq"$ fuzzball cluster config set updated-config.yamlWait for node reuse: Transient blocking resolves on its own as running jobs complete and free their nodes.
Reduce workflow node requirements: If feasible, request fewer nodes so allocations fit within the available pool capacity.
Review provisioner policies: Ensure policy expressions are not concentrating load on a single definition.
A scheduling_blocked event whose reason attribute is node_local_volume_pin means
something different from the pool cap above: the job mounts a volume from a provisioner
marked local,
so it can only run on the one node that holds that volume. The event carries the node in
its required_node attribute, and its detail attribute distinguishes the two cases:
- The node has no capacity yet. Ordinary contention. The job runs once the node frees up, exactly like any other queued allocation.
- The node is not available to the scheduler. The node is down, drained, or has been removed. This does not resolve on its own – the volume exists only on that node, so no other node and no newly provisioned node can run the job.
A related event, reason: node_local_volume_lost, is terminal rather than a wait: it means
the holding node was deprovisioned, so the volume’s data is gone with its disk. The
workflow fails naming the node instead of queueing.
The scheduler does not provision new nodes for a pinned allocation, because a node created now cannot hold a volume created earlier somewhere else. If the node is gone for good, recreate the volume on an available node or restore the original node.
Only volumes on a provisioner withlocal: truepin jobs this way, and both persistent and ephemeral volumes do. A provisioner whose path is genuinely a shared mount across all nodes should leave the field unset, so that jobs are placed without constraint.
A scheduling_blocked event whose reason attribute is lease_grant_failed means the
scheduler picked a node for the job, but the node refused to grant the lease. The
placement is rolled back and the allocation returns to the queue, so the job stays
queued rather than failing immediately. The event carries the node in its node
attribute, the affected rank in rank, how many consecutive attempts have failed in
attempts, and the reason the node gave in detail.
Most refusals clear on their own. A lease from a previous attempt is still being
cleaned up, a node is briefly unreachable, or the node’s actual free capacity no longer
matches what the scheduler projected when it chose the node. The scheduler retries on the
next pass. How many attempts this takes is not fixed — a job with many ranks retried on
every tick can rack up a great many refusals in a short time and still recover — so a
rising attempts count on its own is not a sign of trouble.
A refusal that persists fails the workflow. Once the same allocation has been refused continuously for more than ten minutes and over at least ten attempts, the scheduler stops retrying and fails the workflow with the error the node returned. Both conditions must hold, so on a cluster running the default one-minute scheduler interval it is the elapsed window that decides, and a wedged allocation fails on the first tick past ten minutes — a little past eleven. A longer interval pushes that out proportionally, and past roughly a one-minute loop it is the attempt floor that decides instead. The bound is deliberately generous, never eager. The error appears in the workflow’s status, for example:
failed to grant lease for rank 0 on compute node 10.0.0.97:7331: could not set exclusive mode, other resources in use
“Continuously” is what matters. An allocation that is refused, then is not retried for a while because it is queued behind other work, starts a fresh streak when it is next tried; the run of attempts only counts as continuous while they arrive within three scheduler intervals of each other.
An autoscaled service replica abovereplicas.minis exempt. It is optional capacity, so a persistent refusal returns it to the queue to be retried indefinitely rather than failing the service — the replicas withinreplicas.minkeep running. Thescheduling_blockedevents andworkflow whystill report it as blocked.
Persistent refusal is not the only route to failure. A multi-rank job fails immediately, on the first refusal, if the scheduler also cannot release the leases it had already granted to the job’s other ranks while rolling the placement back — it fails the workflow then rather than retry and leave those leases stranded on their nodes.
To investigate while the job is still queued:
Read the event and the reason the node gave:
$ fuzzball workflow events WORKFLOWAsk the scheduler why the allocation is not running. A
LEASE_GRANT_FAILINGreason names the node and the refusal; see Queue Observability for the full reasons table:$ fuzzball workflow why WORKFLOW ALLOCATIONCheck the state of the node named in the event:
$ fuzzball node show NODE
A node that looks idle but keeps refusing may still be holding a lease stranded by an earlier attempt. The scheduler handles the common form of this on its own: when a node reports that the lease is already granted, it revokes the stale lease before the next attempt so that retry can grant. That revoke is best-effort, so it does not guarantee the next attempt succeeds. If refusals continue, the node’s view of its own resources has diverged from the scheduler’s and needs an administrator — restarting the Fuzzball agent on that node clears its lease state.
Do not drain the node to clear a refusal. Draining cordons it, which takes it out of service for new placements — including the job you are trying to start.
A scheduling_blocked event whose reason attribute is static_definition_no_nodes
means the allocation is pinned to a static provisioner definition that currently has no
nodes registered at all. The event carries the definition in its definition_id
attribute.
This one does not clear on its own. A static definition never provisions, so no node appears unless an operator brings one back, and the allocation keeps its definition for its whole life – the scheduler chooses a definition once, at submission, and never re-picks it. The job waits indefinitely until a node registers with that definition or the workflow is cancelled.
The scheduler will not select an empty static definition for a new submission, so seeing this event means either the allocation was submitted before the nodes went away, or it predates the version that added that check. To resolve it:
Confirm which nodes the definition should have, and whether any are registered:
$ fuzzball node listBring the missing nodes back – the usual cause is that the Fuzzball Substrate daemon is not running on them, or the definition’s matching condition no longer matches the nodes it was written for. See Hardware Grouping Requirements.
If the nodes are gone for good, cancel the workflow and resubmit it once a suitable definition has nodes. Resubmitting is what re-runs definition selection.
A scheduling_blocked event whose reason attribute is
definition_too_few_nodes means the allocation is assigned to a provisioner
definition that has fewer nodes than the job needs. The event carries the
definition in definition_id and both counts in required_nodes and
pool_nodes.
This only affects multi-node jobs, which are gang-scheduled onto a single definition: all of their nodes have to come from one pool, so a pool smaller than the job can never satisfy it however long the job waits. It does not clear on its own, and the allocation keeps the definition it was assigned at submission.
A pool that is merely busy is not reported this way. Occupied, cordoned, and
rebooting nodes all count toward pool_nodes, so a job waiting for a large
enough pool to drain stays quiet and simply runs when capacity frees up.
The scheduler rejects such a job at submission, so seeing this event means the
pool shrank after the job was submitted – nodes deregistered, or the definition
was assigned before that check existed. To resolve it, return nodes to the named
definition, or cancel and resubmit the job with a multinode.nodes the cluster
can host.
The scheduler publishes a provisioning_registration_timeout workflow event when nodes it
provisioned did not register within the definition’s
registration deadline.
Unlike scheduling_blocked, this event is terminal: the scheduler deletes the instances the
provisioning request created and fails the workflow with an error of the form
only 1/3 requested nodes registered within the registration deadline.
The usual cause is a bootstrap failure that the backend never reports. The instance is created and running, but the Fuzzball Substrate on it never comes up – a cloud-init error, a failed package install, an NFS mount that does not appear, or a network path that blocks the Substrate from reaching the cluster.
The event carries the following attributes. The transaction_ attributes refer to the provisioning
request: Fuzzball records each request as a provisioning transaction internally, and that is the
name that appears in event attributes and log messages.
| Attribute | Description |
|---|---|
definition_id | The provisioner definition the nodes were requested from |
transaction_id | The provisioning request whose nodes failed to register |
registered_nodes | How many nodes registered before the deadline expired |
required_nodes | How many nodes the provisioning request asked for |
transaction_age | How long the request had been outstanding when it expired |
A registered_nodes value greater than zero is a partial bootstrap failure: some instances came up
and others did not, which usually points at a per-instance problem (a bad zone, exhausted capacity,
a per-node mount) rather than a broken definition.
Read the event and its attributes from the failed workflow:
$ fuzzball workflow events WORKFLOWCheck whether the definition’s deadline is too short for this workload. Large images, GPU instances, and slow regions can all need longer than the
15mdefault. The deadline lives in the central configuration and is not part offuzzball node provisioner getoutput. This command prints the entire configuration as it was last set and requires cluster-admin credentials; find your definition underdefinitions:. A definition with noregistrationDeadlinekey is on the15mdefault:$ fuzzball cluster config getReproduce the failure with a single-node workflow if you need a live instance to inspect. The instances from the expired request are deleted by the time the workflow fails, so inspect the backend while the reproduction is still running.
Check the Orchestrate logs for the instance cleanup that followed the timeout. A
failed to delete provisioner instances for expired registration transactionmessage means the cleanup did not complete:$ kubectl logs -n fuzzball-system deployment/fuzzball-orchestrator --tail=200
Raise the deadline: If the nodes do eventually register, just later than the deadline allows, raise
registrationDeadlineon the definition and apply the updated configuration:definitions: - id: gpu-pool provisioner: aws ttl: 7200 registrationDeadline: 30m provisionerSpec: instanceType: p5.48xlarge$ fuzzball cluster config set updated-config.yamlFix the bootstrap failure: If the nodes never register no matter how long the deadline, raising it only delays the failure. Investigate the image, the cloud-init or package configuration, and the network path from the node to the cluster.
Account for backend queue time: On Slurm and PBS the deadline starts at
sbatchorqsubsubmission, so time the batch job spends queued in the backend counts against the deadline. See Workflow Stuck in PENDING in the Slurm and PBS troubleshooting guide, which covers both backends in identical sections.
If instance cleanup fails repeatedly, the scheduler stops retrying after three consecutive failures and fails the workflow anyway rather than leaving the workflow hung. Each instance it could not delete is recorded in the Orchestrate logs asfailed to delete instance during transaction cleanup, and those instances may need to be cleaned up manually in the backend.
For collecting some Fuzzball debugging information, administrators may need to
interact with the Kubernetes cluster directly using
kubectl. In this section we
assume that you have the kubectl tool installed following the general
installation
instructions or
when you configured your local Kubernetes
cluster and configured to access
your Kubernetes cluster.
A full tutorial about Kubernetes and kubectl are beyond the scope of this documentation. Below are some commands useful for debugging.
Direct access to the Kubernetes cluster can be potentially dangerous. Proceed with caution.
Inspect the resource consumption of your pods if you suspect that pods are approaching their limits, are overloaded, or use significantly fewer resources than they have been allocated:
$ kubectl top pods --all-namespaces --sort-by=memory --sum
NAMESPACE NAME CPU(cores) MEMORY(bytes)
fuzzball-identity identity-keycloakx-2 3m 5081Mi
fuzzball-identity identity-keycloakx-1 4m 5057Mi
fuzzball-identity identity-keycloakx-0 3m 4990Mi
ingress kong-kong-7bbf7cc6d7-6wq6b 7m 470Mi
...
fuzzball fuzzball-admin-ui-74f84b6757-m2fwt 1m 5Mi
fuzzball fuzzball-admin-0 1m 0Mi
fuzzball postgres-debug 0m 0Mi
________ ________
2243m 24839MiTo ensure that your Kubernetes nodes are not running out of resources you can check with the following:
$ kubectl top nodes
NAME CPU(cores) CPU(%) MEMORY(bytes) MEMORY(%)
node1 922m 11% 15603Mi 50%
node2 1077m 13% 18847Mi 61%
node3 667m 8% 15156Mi 49%In this example the three node Kubernetes cluster is operating below capacity.
You can check that your fuzzball pods are all in a Running state and not frequently
restarting like so:
$ kubectl get -n fuzzball pods
NAME READY STATUS RESTARTS AGE
fuzzball-admin-0 1/1 Running 0 64m
fuzzball-admin-ui-74f84b6757-m2fwt 1/1 Running 0 64m
fuzzball-agent-7467f6484b-kjwv6 1/1 Running 0 61m
fuzzball-agent-7467f6484b-lkdmf 1/1 Running 0 62m
fuzzball-agent-7467f6484b-qxwfq 1/1 Running 0 62m
fuzzball-audit-68f75847bb-5blmp 1/1 Running 0 64m
...Any states other than Running or Completed may signal problems, as are
pods with fewer than expected Ready replicas.
More detailed information about a particular pod can be obtained with:
$ kubectl describe -n fuzzball pod/fuzzball-agent-7467f6484b-kjwv6
Name: fuzzball-agent-7467f6484b-kjwv6
Namespace: fuzzball
Priority: 0
Service Account: fuzzball-agent
Node: ip-10-0-123-57.us-west-2.compute.internal/10.0.123.57
Start Time: Fri, 16 Jan 2026 19:56:49 -0500
Labels: app.kubernetes.io/instance=fuzzball-agent
app.kubernetes.io/managed-by=pulumi
...A stateful set of 3 or more pods form a
Jetstream cluster used for
message passing between Fuzzball services. A buildup of messages for certain topics
like the workflow-engine may indicate that Orchestrate is overloaded and processing
events like state-changes of workflow stages too slowly. This can be checked with the
following command:
$ kubectl exec -n fuzzball fuzzball-jetstream-0 -- /app/nats stream ls
╭─────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ Streams │
├────────────────────────────────────────┬─────────────┬─────────────────────┬──────────┬──────┬──────────────┤
│ Name │ Description │ Created │ Messages │ Size │ Last Message │
├────────────────────────────────────────┼─────────────┼─────────────────────┼──────────┼──────┼──────────────┤
│ job-metrics-stream │ │ 2026-01-06 17:00:57 │ 0 │ 0 B │ 56m2s │
│ provisioner-aws-delete-placement-group │ │ 2026-01-06 16:49:54 │ 0 │ 0 B │ never │
│ provisioner-core │ │ 2026-01-06 16:49:54 │ 0 │ 0 B │ 50m14s │
│ substrate-activity-stream │ │ 2026-01-06 17:00:57 │ 0 │ 0 B │ 56m0s │
│ substrate-conn-stream │ │ 2026-01-06 16:48:44 │ 0 │ 0 B │ 50m12s │
│ substrate-lease-stream │ │ 2026-01-06 17:00:56 │ 0 │ 0 B │ 56m0s │
│ substrate-resource-discovery-stream │ │ 2026-01-06 16:48:44 │ 0 │ 0 B │ 56m36s │
│ substrate-resource-stream │ │ 2026-01-06 16:48:44 │ 0 │ 0 B │ 55m27s │
│ substrate-service-stream │ │ 2026-01-06 17:00:57 │ 0 │ 0 B │ 1d5h52m43s │
│ workflow-execution │ │ 2026-01-06 16:49:54 │ 0 │ 0 B │ 56m0s │
│ leader │ │ 2026-01-17 00:26:15 │ 1 │ 39 B │ 4h14m37s │
│ provisioner-aws-pricing-ticker │ │ 2026-01-06 16:49:54 │ 1 │ 60 B │ 10d9h16m27s │
╰────────────────────────────────────────┴─────────────┴─────────────────────┴──────────┴──────┴──────────────╯