Jetson Thor setup
Serve Gemma-4 31B as an OpenAI-compatible VLM endpoint on the Jetson Thor using
llama.cpp (NVIDIA’s prebuilt Jetson container, CUDA backend). The server listens on
http://localhost:8080/v1, which the benchmark’s
vlm_params.yaml
already targets out of the box.
Benchmark note: All three platforms (AMD, Jetson Orin, Jetson Thor) run the identical model
ggml-org/gemma-4-31B-it-GGUF:Q4_K_Mthroughllama-server. Keep the model and quant identical across platforms for a valid 1:1 comparison.
1. Flash JetPack via NVIDIA SDK Manager
Section titled “1. Flash JetPack via NVIDIA SDK Manager”Flash JetPack from a host Ubuntu machine using the USB Debug/Flashing port:
-
https://docs.nvidia.com/sdk-manager/install-with-sdkm-jetson/index.html#step-03-installation
-
Connect the Jetson to the host via the USB Debug/Flashing port.
-
In SDK Manager, select Jetson Thor as the target.
-
Choose the JetPack version and include CUDA, cuDNN, and TensorRT.
-
Complete flashing and component installation.
Boot the Jetson, finish initial OS setup, then:
sudo apt-get update && sudo apt-get upgradeThor uses a newer stack than Orin (newer JetPack/L4T, PyTorch
nv25.08vs Orin’snv25.02in these notes). Use the Thor devkit guides for Docker + CUDA bring-up: setup_docker, setup_cuda.
2. Install NVIDIA Container Toolkit
Section titled “2. Install NVIDIA Container Toolkit”sudo apt-get install -y nvidia-container-toolkitsudo nvidia-ctk runtime configure --runtime=dockersudo systemctl restart dockerDocs: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/index.html
Verify GPU access from a container:
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smiModel Notes
Section titled “Model Notes”- Model:
ggml-org/gemma-4-31B-it-GGUF(multimodal, text + image). Default quantQ4_K_M(~18.7 GB) fits the Thor’s large unified memory with plenty of headroom. (Only change the quant if you change it on every platform.) - llama.cpp version: the
gemma4architecture requires a recent llama.cpp build (the GGUF was produced with llama.cpp release ≈b8778). If the server errorsunknown model architecture: 'gemma4', the container’s llama.cpp is too old. Pull a newerllama_cppJetson image. (The technician hit this exact error running Gemma-4 on Thor via an older Ollama/llama.cpp build.) llama-server -hfdownloads the GGUF and its vision projector (mmproj) intoLLAMA_CACHE(set to/root/.cache/huggingfacein the image). Mount-v $HOME/.cache/huggingface:/root/.cache/huggingfaceso the download persists.
Reference for Jetson LLM/VLM containers: https://www.jetson-ai-lab.com/models/