Jetson Orin setup
Serve Gemma-4 31B as an OpenAI-compatible VLM endpoint on the Jetson Orin AGX 64G
using llama.cpp (NVIDIA’s prebuilt Jetson container, CUDA backend). The server listens
on http://localhost:8080/v1, which the benchmark’s
vlm_params.yaml
already targets out of the box.
Benchmark note: All three platforms (AMD, Jetson Orin, Jetson Thor) run the identical model
ggml-org/gemma-4-31B-it-GGUF:Q4_K_Mthroughllama-server. Keep the model and quant identical across platforms for a valid 1:1 comparison.
Storage: The Orin AGX ships with limited onboard eMMC (64 GB). An additional NVMe SSD is required to hold the model and container storage; see step 2. Install and setup the NVME before any of these instructions or you’ll have a very bad time recovering.
1. Flash JetPack via NVIDIA SDK Manager
Section titled “1. Flash JetPack via NVIDIA SDK Manager”Flash JetPack from a host Ubuntu machine using the USB Debug/Flashing port:
-
https://docs.nvidia.com/sdk-manager/install-with-sdkm-jetson/index.html#step-03-installation
-
Connect the Jetson to the host via the USB Debug/Flashing port.
-
In SDK Manager, select Jetson Orin AGX as the target.
-
Choose the JetPack version and include CUDA, cuDNN, and TensorRT. These notes were produced with JetPack 6.2.2 (CUDA 12.6, cuDNN 9, TensorRT) on the AGX Orin.
-
Complete flashing and component installation.
Boot the Jetson, finish initial OS setup, then:
sudo apt-get update && sudo apt-get upgrade2. Provision NVMe Storage
Section titled “2. Provision NVMe Storage”The onboard eMMC is too small for a 31B model. Mount an NVMe SSD and relocate Docker’s data directory onto it:
Key steps from the guide:
- Format and mount the NVMe SSD (e.g. at
/ssd). - Relocate the Docker data root (
/var/lib/docker) onto the SSD mount. - Verify Docker writes to the SSD before proceeding.
Create the data and HuggingFace cache directories on the SSD:
mkdir -p /ssd/docker/data /ssd/docker/.cache/huggingface3. Install NVIDIA Container Toolkit
Section titled “3. Install NVIDIA Container Toolkit”sudo apt-get install -y nvidia-container-toolkitsudo nvidia-ctk runtime configure --runtime=dockersudo systemctl restart dockerDocs: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/index.html
Verify GPU access from a container:
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smiModel Notes
Section titled “Model Notes”- Model:
ggml-org/gemma-4-31B-it-GGUF(multimodal, text + image). Default quantQ4_K_M(~18.7 GB) fits the Orin’s 64 GB comfortably. (Q8_0 ~32.6 GB also fits, but only change it if you change it on every platform.) - llama.cpp version: the
gemma4architecture requires a recent llama.cpp build (the GGUF was produced with llama.cpp release ≈b8778). If the server errorsunknown model architecture: 'gemma4', the container’s llama.cpp is too old. Pull a newerllama_cppJetson image. (The technician hit this exact error running Gemma-4 via an older Ollama/llama.cpp build.) llama-server -hfdownloads the GGUF and its vision projector (mmproj) intoLLAMA_CACHE(set to/root/.cache/huggingfacein the image). Mount the SSD cache (-v /ssd/docker/.cache/huggingface:/root/.cache/huggingface) so the download persists.
Reference for Jetson LLM/VLM containers: https://www.jetson-ai-lab.com/models/