AMD Strix Halo setup
Serve Gemma-4 31B as an OpenAI-compatible VLM endpoint on the AMD Strix Halo platform using llama.cpp (via the AMD ryzers framework on ROCm). The server
listens on http://localhost:8080/v1, which the benchmark’s vlm_params.yaml already targets out of the box.
Benchmark note: All three platforms (AMD, Jetson Orin, Jetson Thor) run the identical model
ggml-org/gemma-4-31B-it-GGUF:Q4_K_Mthroughllama-server. Keep the model and quant identical across platforms for a valid 1:1 comparison.
1. OS & Kernel
Section titled “1. OS & Kernel”After a fresh Ubuntu install, reboot, log in, and update:
sudo apt-get update && sudo apt-get upgradeInstall the OEM kernel and reboot:
sudo apt update && sudo apt install linux-oem-24.04csudo reboot2. ROCm Drivers & GPU Memory
Section titled “2. ROCm Drivers & GPU Memory”Install and configure the ROCm drivers for Ryzen following AMD’s guide:
In particular, configure shared memory in BIOS and Linux. The BIOS may call this “Graphics Buffer Memory” (the docs call it VRAM). AMD recommends setting this to the lowest possible value (512M-2G depending on BIOS version). The GPU draws the rest from shared system memory (GTT):
3. Docker
Section titled “3. Docker”Install Docker Engine (https://docs.docker.com/engine/install/ubuntu/):
# Add Docker's official GPG key:sudo apt updatesudo apt install ca-certificates curlsudo install -m 0755 -d /etc/apt/keyringssudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.ascsudo chmod a+r /etc/apt/keyrings/docker.asc
# Add the repository to Apt sources:sudo tee /etc/apt/sources.list.d/docker.sources <<EOFTypes: debURIs: https://download.docker.com/linux/ubuntuSuites: $(. /etc/os-release && echo "${UBUNTU_CODENAME:-$VERSION_CODENAME}")Components: stableSigned-By: /etc/apt/keyrings/docker.ascEOF
sudo apt updatesudo apt install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
sudo groupadd dockersudo usermod -aG docker $USERnewgrp dockerdocker run hello-world4. Install ryzers & Build llama.cpp
Section titled “4. Install ryzers & Build llama.cpp”Install git and a Python venv, then install the AMD ryzers framework:
sudo apt-get install vim gitsudo apt install python3-venv python3-pippython3 -m venv ~/venvsource ~/venv/bin/activate
mkdir -p ~/projects && cd ~/projectsgit clone https://github.com/amdresearch/ryzerscd ryzers/pip install -e .Build the llama.cpp image with ryzers, tagging the final image llamacpp:
ryzers build --name llamacpp llamacppThis builds a two-stage chain (ryzer_env, then llama.cpp from the
rocm/pytorch:rocm7.2.2_ubuntu24.04_py3.12_pytorch_release_2.10.0 base, compiled for the
gfx1151 Strix Halo iGPU) and tags the final image llamacpp. Confirm it:
docker images | grep llamacppWe use ryzers build only to produce this image. Everything below uses plain docker build
/ docker run (not ryzers run).
Finally, create a location for the llama cache
mkdir ~/llamacpp_cacheModel Notes
Section titled “Model Notes”- Model:
ggml-org/gemma-4-31B-it-GGUF(multimodal, text + image). Default quantQ4_K_M(~18.7 GB) fits the EVO-X2’s unified memory; do not change it unless you also change it on the Jetsons. - llama.cpp version: the
gemma4architecture requires a recent llama.cpp build (the GGUF was produced with llama.cpp release ≈b8778). If the server errorsunknown model architecture: 'gemma4', the bundled llama.cpp is too old. Rebuild the ryzers image against currentllama.cpp. llama-server -hfdownloads the GGUF and its vision projector (mmproj) intoLLAMA_CACHE(set to/root/.cache/huggingfacein the image). The run command’s-v $PWD/llamacpp_cache:/root/.cachemount (section 7) persists these across restarts.