Skip to content

Jetson Orin setup

Serve Gemma-4 31B as an OpenAI-compatible VLM endpoint on the Jetson Orin AGX 64G using llama.cpp (NVIDIA’s prebuilt Jetson container, CUDA backend). The server listens on http://localhost:8080/v1, which the benchmark’s vlm_params.yaml already targets out of the box.

Benchmark note: All three platforms (AMD, Jetson Orin, Jetson Thor) run the identical model ggml-org/gemma-4-31B-it-GGUF:Q4_K_M through llama-server. Keep the model and quant identical across platforms for a valid 1:1 comparison.

Storage: The Orin AGX ships with limited onboard eMMC (64 GB). An additional NVMe SSD is required to hold the model and container storage; see step 2. Install and setup the NVME before any of these instructions or you’ll have a very bad time recovering.


Flash JetPack from a host Ubuntu machine using the USB Debug/Flashing port:

Boot the Jetson, finish initial OS setup, then:

Terminal window
sudo apt-get update && sudo apt-get upgrade

The onboard eMMC is too small for a 31B model. Mount an NVMe SSD and relocate Docker’s data directory onto it:

Key steps from the guide:

  • Format and mount the NVMe SSD (e.g. at /ssd).
  • Relocate the Docker data root (/var/lib/docker) onto the SSD mount.
  • Verify Docker writes to the SSD before proceeding.

Create the data and HuggingFace cache directories on the SSD:

Terminal window
mkdir -p /ssd/docker/data /ssd/docker/.cache/huggingface

Terminal window
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

Docs: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/index.html

Verify GPU access from a container:

Terminal window
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
  • Model: ggml-org/gemma-4-31B-it-GGUF (multimodal, text + image). Default quant Q4_K_M (~18.7 GB) fits the Orin’s 64 GB comfortably. (Q8_0 ~32.6 GB also fits, but only change it if you change it on every platform.)
  • llama.cpp version: the gemma4 architecture requires a recent llama.cpp build (the GGUF was produced with llama.cpp release ≈ b8778). If the server errors unknown model architecture: 'gemma4', the container’s llama.cpp is too old. Pull a newer llama_cpp Jetson image. (The technician hit this exact error running Gemma-4 via an older Ollama/llama.cpp build.)
  • llama-server -hf downloads the GGUF and its vision projector (mmproj) into LLAMA_CACHE (set to /root/.cache/huggingface in the image). Mount the SSD cache (-v /ssd/docker/.cache/huggingface:/root/.cache/huggingface) so the download persists.

Reference for Jetson LLM/VLM containers: https://www.jetson-ai-lab.com/models/