Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Install Fuzzball on an NVIDIA DGX Spark

This guide covers the NVIDIA DGX Spark specifics for bringing up a single-host Fuzzball cluster with GPU passthrough. It builds on the platform-agnostic Docker Compose Stack reference — install, DNS, login, lifecycle, upgrading, and troubleshooting all live there. This page covers only what is different on the DGX Spark.

The DGX Spark is a Grace Blackwell GB10 personal supercomputer: linux/arm64, 128 GB of unified LPDDR5x, an integrated Blackwell GPU, and DGX OS (Ubuntu-derived). Because every Fuzzball image is published as a multi-arch manifest, no custom builds are required — the operator just runs the standard CLI. Once finished, you’ll have an orchestrator, web UI, auth (Keycloak), persistent state (PostgreSQL), and a substrate node that schedules GPU workflows onto the Blackwell GPU.

The GPU-enabled substrate image (selected with --gpu) passes the GPU through to workflow containers via the NVIDIA Container Toolkit on the host.

Prerequisites

  • DGX OS or an Ubuntu-derived distribution shipping the NVIDIA Blackwell driver. Confirm with:

    $ nvidia-smi

    Expected output: at least one GPU listed with a non-empty driver version.

  • Docker Engine with the compose plugin (preinstalled on DGX OS).

Install and deploy

Use the Quick Install script, or install the CLI manually (use the arm64 .deb package). When prompted for GPU support, answer y to select the GPU substrate image for the DGX Spark’s Blackwell GPU.

To deploy by hand instead, pass --gpu:

$ fuzzball cluster docker-compose deploy \
    --gpu \
    --update-etc-hosts \
    --up

See Deploy for what the flags do and DNS resolution for domain options, then Log in with the CLI.

Verify the substrate registered the GPU

$ fuzzball node list --available-resources
NODE ID         | HOSTNAME              | CPU TYPE  | AVAILABLE CORES | AVAILABLE MEMORY (GB) | AVAILABLE DEVICES | RUNNING JOBS | CLUSTER
172.21.0.5/7333 | substrate-localnode-1 | cpu/arm64 | 20              | 127.9                 | nvidia.com/gpu:1  | 0            | local-dev

If AVAILABLE DEVICES shows nvidia.com/gpu:N, the substrate-side GPU plumbing is working.

Troubleshooting

For registry authentication (pull access denied), certificate trust, and inspecting the stack directly, see the Docker Compose troubleshooting section. DGX-specific issues:

nvidia.com/gpu doesn’t appear in AVAILABLE DEVICES

The substrate couldn’t see the GPU. On the host, confirm the Blackwell driver is loaded and nvidia-smi lists the GPU, then re-run the deploy.

High memory pressure or OOM during workflow runs

The DGX Spark’s 128 GB is unified between CPU and GPU — a large GPU-resident model competes with the rest of the stack. The base Fuzzball stack rests at ~2-3 GB. Leave at least 4-8 GB headroom for the control plane when sizing GPU workflow memory budgets.

Managing the stack

Lifecycle, upgrading, and teardown are the same as any Docker Compose deployment:

For GPU scheduling details (requesting nvidia.com/gpu in a workflow, GPU health monitoring), see GPUs.