Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Removing Node Registrations

A compute node appears in Fuzzball because its agent registers with the control plane and reports what hardware it has. That registration is what makes the node schedulable, and it is what the scheduler counts when deciding whether the cluster has capacity for queued work.

Removing a registration is different from cordoning or draining. A cordoned node is still a node: it stops accepting new work, but it remains in fuzzball node list and the control plane still expects it to be there. Removal deletes the registration outright, so the node and the resources it was advertising leave the scheduler’s view entirely.

Prerequisites

Removing a registration requires an authorized Fuzzball admin context.

$ fuzzball context create <context_name> <api_url>

$ fuzzball context login -u <admin_username> -p '<admin_password>'

When a registration outlives its node

Normally you never need to think about this. A node whose host shuts down cleanly deregisters itself, and one that disappears without doing so is deregistered automatically once the control plane finds it unreachable.

A registration can still outlive its node when the whole cluster goes down at once – a host power cycle, or a Docker Compose deployment where the messaging service and the compute node are stopped together. The signal that would have triggered deregistration is lost along with the service that was meant to act on it, and the registration is left behind.

Until it is cleaned up, such a registration still advertises its node’s cores, memory and GPUs as free capacity, so the scheduler may try to place work on a machine that no longer exists. Cordoning does not release that capacity – only removal does.

Automatic cleanup

The control plane re-checks every registration on a recurring sweep and deregisters any node it cannot reach across several consecutive attempts. A registration left behind by a restart is normally gone within a few minutes, without any operator action.

A single failed check is never enough to deregister a node, so a node that is briefly unreachable – restarting a service, or a short network interruption – keeps its registration.

Each automatic removal is recorded in the fuzzball_substrate_node_registrations_reaped_total metric. Occasional increments are expected. A rate that grows with the size of your cluster means the control plane cannot reach nodes that are in fact healthy, which is worth investigating rather than ignoring.

When a different node takes over an address

Reachability alone cannot settle every case. If a node disappears and its address is then given to a different machine – a cloud provider reassigning a private IP to a replacement instance, or a container recreated on the same network address – the departed node’s registration still answers every check, because the machine that inherited the address answers on its behalf. Left alone, that registration survives indefinitely, and any workflow still assigned to it keeps running and keeps accruing charges.

Fuzzball identifies a registration by hostname and address together, so a replacement registers under its own entry rather than overwriting the one before it. Two registrations sharing an address cannot both be live, since an address hosts a single node, so the older one is removed.

Which registration counts as older is decided by which was updated most recently, and the one being removed must also have stopped being updated for several minutes. Both conditions are required, because an instance being torn down keeps reporting for a short while after it is deleted, so having been written to more recently is not on its own proof of being alive. In practice the stale registration is removed within a few minutes of the replacement appearing, rather than after the several consecutive failed checks the reachability sweep would need. A registration still present shortly after a replacement appears is expected, not a sign that something is stuck.

Removing the registration stops the scheduler counting the departed node’s CPUs, memory and GPUs as free capacity. Work that was still assigned to that node is released when its leases expire, which can take a few minutes longer.

The removal is recorded separately, in the fuzzball_substrate_node_registrations_superseded_total metric, so that address reuse stays distinguishable from nodes going quiet. A sustained rate means instances are being replaced at the same addresses repeatedly, which is routine on a Docker Compose network and worth understanding on a cloud provisioner.

Removing a registration manually

To remove a registration immediately rather than waiting for the sweep, use fuzzball node remove with the identifier shown in the ID column of fuzzball node list:

$ fuzzball node remove <node-id>

That identifier is the node’s address, such as 172.18.0.4/7333. Control plane logs refer to the same node by a longer key that also carries its hostname, such as node-a/172.18.0.4/7333; pass the address form from node list, not the longer key.

Example:

$ fuzzball node remove 172.18.0.4/7333
Removed the registration for node 172.18.0.4/7333

Any work the scheduler still believed was running on the node is released and requeued, rather than being left to time out one allocation at a time.

Nodes that still have running work

Because removal releases running work, a node the scheduler still shows work on is refused:

$ fuzzball node remove 172.18.0.4/7333
Error: node 172.18.0.4/7333 still has 2 running allocation(s); removing
its registration would evacuate them. Drain the node first, or pass force to
remove it anyway

A node with running jobs is far more likely to be alive than to be a leftover registration, so this usually means the identifier is not the one you meant. If you are certain, --force removes it and reports what that cost:

$ fuzzball node remove 172.18.0.4/7333 --force
Removed the registration for node 172.18.0.4/7333
Evacuating 2 allocation(s) that were running on it

Effect on work that needs the removed node

The scheduler will not assign a job to a static provisioner definition that does not have enough nodes registered to run it, so removing registrations changes what the cluster accepts, not just what it can place.

This applies to static definitions only. A dynamic definition is judged on the nodes it may provision – its maxNodes – rather than on the nodes registered right now, so removing its registrations does not make it unselectable: the scheduler picks it and provisions replacements. Only a static pool, which cannot grow, loses capacity when a registration goes away.

Once a static definition has no registered nodes left, a workflow that has no other definition to fall back on is rejected when it is submitted rather than queued:

$ fuzzball workflow start ./job.yaml
no suitable provision definition found for job analyze (allocation ID
c7c8f5a5-1582-5231-83c7-92e59cac67ff); skipped gpu-nodes for having too few nodes

A multi-node job is gang-scheduled onto a single definition, so it needs that many nodes registered with one definition. Removing registrations below its node count rejects it the same way, and the error reports the largest pool available:

$ fuzzball workflow start ./trainer.yaml
multinode job trainer requires 3 nodes in a single provisioner definition
(allocation ID bdeefe8e-578c-5faf-b812-fa3b5f602113) but the largest pool in the
cluster has 2; reduce multinode.nodes or add nodes to a definition that fits the job

This is why a workflow can start one day and be refused the next on a cluster whose registrations have been cleaned up in between. The remedy is to restore registrations – restart the agents, or let the node rejoin – or to point the job at a definition that has nodes.

Only absent nodes reduce the count. A node that is cordoned, draining, or simply busy is still registered, so its definition still accepts work and queues it until the node is available again. Submissions are refused only when the nodes are gone from fuzzball node list altogether.

Work that was already queued when the nodes went away keeps the definition it was assigned at submission – the scheduler chooses a definition once and does not re-pick it – so it waits rather than moving to another definition. Such an allocation reports a scheduling_blocked event with the reason static_definition_no_nodes, and fuzzball workflow why reports it as provisioning-blocked:

$ fuzzball workflow why <workflow-id> <allocation-id>
Workflow: 4d65d56b-3406-4424-b6df-8f89701cd80d
Allocation: 2d76dfa7-9ffa-5870-b709-40b602683983
Received by cluster: fuzzball-local-dev
Reason: provisioning blocked
Detail: no nodes available for static definition "gpu-nodes"

See Troubleshooting scheduling_blocked Events for how to work through that state, and Queue Observability for the full set of reasons workflow why can report.

Do not use removal to decommission a live node

Removal is for registrations whose node is gone. It is not a way to take a working node out of service, and it is not cheap to undo:

  • A node reports its resources when they change, not on a timer, so removing a live node’s registration takes that node out of the cluster until its resources change or its agent restarts. It does not come back on its own.
  • While it is out, a static definition is short a node, so workflows that needed it are refused at submission rather than queued – see Effect on work that needs the removed node. Cordoning has no such effect, which is another reason to prefer it.
  • To stop new work reaching a node, use fuzzball node cordon.
  • To take a node out of service and decide what happens to the work already on it, use fuzzball node drain.
  • To tear down a cloud instance provisioned by Fuzzball, use fuzzball node deprovision, which releases the underlying instance as well as the registration.

Docker Compose deployments

In a Docker Compose deployment a node’s identity includes its container address, so recreating the compute node container on a new address registers a new node rather than updating the existing one. This is not damaging – the old registration is reaped automatically, and fuzzball node remove clears one immediately, while the recreated container registers itself on startup – but you can avoid the churn:

  • Run docker compose down before bringing the stack back up when the compute node container will be recreated, rather than a bare docker compose up -d. A clean stop deregisters the node while the control plane is still running to notice it.
  • Recreate only the service that changed (docker compose up -d --no-deps <service>) so the compute node keeps its address.