Install Fuzzball on an NVIDIA DGX Spark
This guide covers the NVIDIA DGX Spark specifics for bringing up a single-host Fuzzball cluster with GPU passthrough. It builds on the platform-agnostic Docker Compose Stack reference — install, DNS, login, lifecycle, upgrading, and troubleshooting all live there. This page covers only what is different on the DGX Spark.
The DGX Spark is a Grace Blackwell GB10 personal supercomputer: linux/arm64, 128 GB of unified
LPDDR5x, an integrated Blackwell GPU, and DGX OS (Ubuntu-derived). Because every Fuzzball image is
published as a multi-arch manifest, no custom builds are required — the operator just runs the
standard CLI. Once finished, you’ll have an orchestrator, web UI, auth (Keycloak), persistent state
(PostgreSQL), and a substrate node that schedules GPU workflows onto the Blackwell GPU.
The GPU-enabled substrate image (selected with --gpu) passes the GPU through to workflow containers
via the NVIDIA Container Toolkit on the host.
DGX OS or an Ubuntu-derived distribution shipping the NVIDIA Blackwell driver. Confirm with:
$ nvidia-smiExpected output: at least one GPU listed with a non-empty driver version.
Docker Engine with the compose plugin (preinstalled on DGX OS).
Use the Quick Install script, or install the
CLI manually (use the arm64 .deb package). When
prompted for GPU support, answer y to select the GPU substrate image for the DGX Spark’s
Blackwell GPU.
To deploy by hand instead, pass --gpu:
$ fuzzball cluster docker-compose deploy \
--gpu \
--update-etc-hosts \
--upSee Deploy for what the flags do and DNS resolution for domain options, then Log in with the CLI.
$ fuzzball node list --available-resources
NODE ID | HOSTNAME | CPU TYPE | AVAILABLE CORES | AVAILABLE MEMORY (GB) | AVAILABLE DEVICES | RUNNING JOBS | CLUSTER
172.21.0.5/7333 | substrate-localnode-1 | cpu/arm64 | 20 | 127.9 | nvidia.com/gpu:1 | 0 | local-devIf AVAILABLE DEVICES shows nvidia.com/gpu:N, the substrate-side GPU plumbing is working.
For registry authentication (pull access denied), certificate trust, and inspecting the stack
directly, see the Docker Compose troubleshooting
section. DGX-specific issues:
nvidia.com/gpu doesn’t appear in AVAILABLE DEVICES
The substrate couldn’t see the GPU. On the host, confirm the Blackwell driver is loaded and
nvidia-smi lists the GPU, then re-run the deploy.
High memory pressure or OOM during workflow runs
The DGX Spark’s 128 GB is unified between CPU and GPU — a large GPU-resident model competes with the rest of the stack. The base Fuzzball stack rests at ~2-3 GB. Leave at least 4-8 GB headroom for the control plane when sizing GPU workflow memory budgets.
Lifecycle, upgrading, and teardown are the same as any Docker Compose deployment:
- Lifecycle commands — start, stop, status, logs
- Upgrading — bump to a newer image tag
- Tear down — stop or remove the deployment
- Beyond a single host — attach remote substrate nodes and production hardening
For GPU scheduling details (requesting nvidia.com/gpu in a workflow, GPU health monitoring), see
GPUs.