Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Install Fuzzball on an AMD Ryzen AI Max (Strix Halo)

This guide covers the AMD Ryzen AI Max (“Strix Halo”) specifics for bringing up a single-host Fuzzball cluster with GPU passthrough. It builds on the platform-agnostic Docker Compose Stack reference for install, DNS, login, lifecycle, upgrading, and troubleshooting. This page covers only what is different on this host.

The GPU-enabled substrate image (selected with --gpu) bundles the device plugins for AMD; on this host the AMD plugin passes the GPU device nodes (/dev/kfd and /dev/dri/renderD*) and the host’s ROCm driver libraries through to your workflow containers. See GPUs for how AMD ROCm devices are requested and scheduled.

Prerequisites

  • Confirm the amdgpu kernel drivers are loaded and the ROCm runtime installed:

    $ sudo rocminfo

    If needed, follow AMD’s ROCm installation guide.

  • Docker Engine with the compose plugin. On a stock Ubuntu host this is often not present yet — the Quick Install script installs it from Docker’s official apt repository if it is missing and adds your user to the docker group. To install it yourself first, follow the Docker Engine install guide.

Install and deploy

Use the Quick Install script, or install the CLI manually (use the amd64 .deb package). When prompted for GPU support, answer y to select the GPU substrate image for the host’s Radeon GPU.

To deploy by hand instead, pass --gpu:

$ fuzzball cluster docker-compose deploy \
    --gpu \
    --update-etc-hosts \
    --up

See Deploy for what the flags do and DNS resolution for domain options, then Log in with the CLI.

Verify the substrate registered the GPU

$ fuzzball node list --available-resources
NODE ID         | HOSTNAME              | CPU TYPE  | AVAILABLE CORES | AVAILABLE MEMORY (GB) | AVAILABLE DEVICES | RUNNING JOBS | CLUSTER
172.21.0.5/7333 | substrate-localnode-1 | cpu/amd64 | 30              | 120.0                 | amd.com/gpu:1     | 0            | local-dev

If AVAILABLE DEVICES shows amd.com/gpu:N, the substrate-side GPU plumbing is working. Submit a GPU workflow that requests amd.com/gpu to confirm end to end. The amd-gpu-test.fz example reads the GPU’s attributes and checks that /dev/kfd is present inside the job.

Troubleshooting

For registry authentication (pull access denied), certificate trust, and inspecting the stack directly, see the Docker Compose troubleshooting section. AMD-specific issues:

amd.com/gpu doesn’t appear in AVAILABLE DEVICES

The substrate couldn’t see the GPU. On the host, confirm the amdgpu driver is loaded and the ROCm device nodes exist (ls /dev/kfd /dev/dri/renderD*) and that rocminfo lists the GPU. The device nodes are created by the driver at boot; if they’re missing, the driver isn’t loaded. Then re-run the deploy.

High memory pressure or OOM during workflow runs

The Ryzen AI Max’s memory is unified between the CPU and the integrated GPU — a large GPU-resident model competes with the rest of the stack. The base Fuzzball stack rests at ~2-3 GB. Leave at least 4-8 GB headroom for the control plane when sizing GPU workflow memory budgets, and note that the amount available to the GPU also depends on the UMA/GTT split configured in firmware.

Managing the stack

Lifecycle, upgrading, and teardown are the same as any Docker Compose deployment: