Install Fuzzball on an AMD Ryzen AI Max (Strix Halo)
This guide covers the AMD Ryzen AI Max (“Strix Halo”) specifics for bringing up a single-host Fuzzball cluster with GPU passthrough. It builds on the platform-agnostic Docker Compose Stack reference for install, DNS, login, lifecycle, upgrading, and troubleshooting. This page covers only what is different on this host.
The GPU-enabled substrate image (selected with --gpu) bundles the device plugins for AMD; on this
host the AMD plugin passes the GPU device nodes (/dev/kfd and /dev/dri/renderD*) and the host’s
ROCm driver libraries through to your workflow containers. See GPUs for how AMD ROCm devices are requested and
scheduled.
Confirm the
amdgpukernel drivers are loaded and the ROCm runtime installed:$ sudo rocminfoIf needed, follow AMD’s ROCm installation guide.
Docker Engine with the compose plugin. On a stock Ubuntu host this is often not present yet — the Quick Install script installs it from Docker’s official apt repository if it is missing and adds your user to the
dockergroup. To install it yourself first, follow the Docker Engine install guide.
Use the Quick Install script, or install the
CLI manually (use the amd64 .deb package). When
prompted for GPU support, answer y to select the GPU substrate image for the host’s Radeon GPU.
To deploy by hand instead, pass --gpu:
$ fuzzball cluster docker-compose deploy \
--gpu \
--update-etc-hosts \
--upSee Deploy for what the flags do and DNS resolution for domain options, then Log in with the CLI.
$ fuzzball node list --available-resources
NODE ID | HOSTNAME | CPU TYPE | AVAILABLE CORES | AVAILABLE MEMORY (GB) | AVAILABLE DEVICES | RUNNING JOBS | CLUSTER
172.21.0.5/7333 | substrate-localnode-1 | cpu/amd64 | 30 | 120.0 | amd.com/gpu:1 | 0 | local-devIf AVAILABLE DEVICES shows amd.com/gpu:N, the substrate-side GPU plumbing is working. Submit a
GPU workflow that requests amd.com/gpu to confirm end to end. The amd-gpu-test.fz example reads the GPU’s attributes and checks that
/dev/kfd is present inside the job.
For registry authentication (pull access denied), certificate trust, and inspecting the stack
directly, see the Docker Compose troubleshooting
section. AMD-specific issues:
amd.com/gpu doesn’t appear in AVAILABLE DEVICES
The substrate couldn’t see the GPU. On the host, confirm the amdgpu driver is loaded and the ROCm
device nodes exist (ls /dev/kfd /dev/dri/renderD*) and that rocminfo lists the GPU. The device
nodes are created by the driver at boot; if they’re missing, the driver isn’t loaded. Then re-run the
deploy.
High memory pressure or OOM during workflow runs
The Ryzen AI Max’s memory is unified between the CPU and the integrated GPU — a large GPU-resident model competes with the rest of the stack. The base Fuzzball stack rests at ~2-3 GB. Leave at least 4-8 GB headroom for the control plane when sizing GPU workflow memory budgets, and note that the amount available to the GPU also depends on the UMA/GTT split configured in firmware.
Lifecycle, upgrading, and teardown are the same as any Docker Compose deployment:
- Lifecycle commands — start, stop, status, logs
- Upgrading — bump to a newer image tag
- Tear down — stop or remove the deployment
- Beyond a single host — attach remote substrate nodes and production hardening