Methodology
How this is measured
Everything here is reproducible from the repository. If a definition on this page is not what you would have chosen, the analysis code is public and you can compute it your way from the same raw logs.
Robotics workload
Nav2 autonomously navigates a forklift material handling robot within a 200,000 sqft (18,600 m²) industrial warehouse environment, using its advanced, built-in planning, control, behavior modeling, and perception. It moves pallets from shipping/receiving to shelving units while processing multiple 3D lidars, 2D safety lidars, RGBD cameras, and internal sensors. This workload is representative of dozens of companies and tens of thousands of robots deployed today in production environments.
The mission loop runs from a charging dock: pick a pallet from a block stack or loading dock, drop it at racking, repeat. A wait is inserted at each pick and place to simulate the physical handling time, and the robot periodically returns to the dock. Maximum robot speed is 2 m/s. Runs are 900 seconds, with system metrics sampled at 1 Hz.
Sensor suite
Ten sensors. Nine feed the AMR autonomy workload; the RGB camera feeds the VLM only.
| Sensor | Count | Rate | Configuration |
|---|---|---|---|
| 3D LiDAR | 3 | 10 Hz | 30 m range, 32 beams at 512 resolution |
| RGBD camera | 3 | 10 Hz | 320x240 |
| 2D safety LiDAR | 2 | 30 Hz | 25 m range |
| RGB camera | 1 | 5 Hz | 640x480, feeds the VLM only |
| IMU | 1 | 100 Hz | Internal |
Autonomy stack
Five subsystems, all resident at once. Optimization-based control at 30 Hz is the part of this stack most sensitive to CPU contention, which is why control-loop misses are the most diagnostic single number in the results.
Planning
Hybrid-A* with full-footprint SE2 collision checking. This is the expensive, correct option, not a point-robot approximation.
Control
MPPI, an optimization-based trajectory planner, at a 30 Hz target. This is the part most sensitive to CPU contention.
Perception
Spatio-Temporal Voxel Layer for the 3D lidars; obstacle and voxel layers for the 2D lidars and RGBD cameras.
Safety
Velocity smoother and collision monitor, running alongside everything else.
Localization
AMCL with simulated odometry.
AI workload
Gemma 4.0 (31B, Q4_K_M) served by a llama.cpp server on the platform under test, integrated into the navigation behavior tree for scene understanding so that its answers affect the robot's decisions rather than running beside them. Queries are re-issued immediately after each one completes, so the GPU stays saturated for the whole run. VLMs are intensive workloads that robotics-targeted embedded platforms are being built with in mind, which makes them a good way to fully leverage the capabilities of modern platforms.
The AI workload is optional in the pipeline. Every result published here includes it.
Sensor-driver load
Real sensor drivers are not attached, because the sensors are simulated. But driver CPU cost is a real and platform-dependent expense, so it is measured on real hardware (Orbbec Gemini 355 depth cameras and Ouster OS-1 32 lidars) and then replayed on each platform as a synthetic load. The coefficients, as a fraction of one core per sensor instance:
| Sensor | X100 / Strix Halo | Jetson Thor | Jetson AGX Orin |
|---|---|---|---|
| 3D LiDAR | 0.160 | 0.220 | 0.670 |
| 2D safety LiDAR | 0.014 | 0.040 | 0.120 |
| RGBD camera | 0.520 | 1.300 | 1.550 |
Test harness
A developer computer runs the Gazebo simulation of the warehouse and the robot's sensors in a provided Docker container. This is run on another machine to isolate potential effects from resource scheduling the simulation on utilization metrics. DDS is configured to only operate on the wired ethernet interface to minimize the impacts due to discovery traffic or other interference.
On the platform under test: the AMR autonomy container, the AI workload container, the simulated sensor-driver load, and the metrics capture script. All four contend for the same CPU, power, and accelerators, which is the entire point.
The same offered sensor load reaches every platform: measured network receive rate is approximately 66 MB/s on all three, so no platform is being fed less work than another.
Metric definitions
System counters are sampled at 1 Hz for 900 seconds. Application metrics are parsed from the ROS logs. Derived metrics:
- Free core equivalents
- Available CPU percentage multiplied by the core count. What is left, expressed in the unit you budget new nodes in.
- Absolute compute headroom (GHz-cores)
- Core count multiplied by sustained clock, minus the portion the workload consumed. Counts core count and clock speed together, so a platform cannot look good on one while hiding the other.
- Single-thread headroom (GHz)
- The median across the run of the least-loaded core, converted to available clock. This is the budget for one new real-time control thread that cannot be parallelized.
- Thermal runway (°C)
- Margin between the mean measured temperature and a 100 °C throttle point.
- Control loop misses
- Occurrences of the Nav2 controller server’s "Control loop missed its desired rate" warning, counted against its 30 Hz target and normalized per second of run time.
- Completed missions
- Count of "[bt_navigator]: Goal succeeded" in the ROS logs, meaning navigation goals the robot actually finished.
- Planner cycle time
- Derived as 1 / actual rate from the planner server loop-rate warnings, so it reports the real time to produce one plan.
- VLM query outcome
- Parsed from VLM_QUERY log lines emitted by the VLM node: success, error, cancelled, retries exhausted, no image, stale image, or encode failed.
- Platform balance radar
- Six dimensions: CPU headroom, CPU capability, GPU capability, memory capability, memory headroom, and clock speed. Each is min-max normalized against the strongest platform in that category, so 1.0 means "best of this group", not "perfect".
Power is read at the board where the platform exposes it. On the Jetsons this is the board power rail; on Strix Halo it is the GPU/SoC power counter reported by the driver. Both are the figure the platform itself reports, and neither is a wall-socket measurement.
Limitations and Concerns
Stated plainly, because a benchmark that hides these is not one you should base a purchase on.
Power is not the same on every platform
The platforms are compared using max power profiles which are not the same to compare the highest performance modes possible on each. A balanced TDP comparison is also made to compare as close to the same as possible but are not exactly identical due to vendor configuration limitations.
Optimized AI runtime are not identical
The results from unoptimized AI workloads use the same model, quantization, and source to give a 1:1 comparison between platforms. The results from optimized AI workloads showcase optimized paths on each vendor recommended instructions to showcase performance using hardware specialized methods, quantizations, or libraries.
A 15-minute window
Long-duration thermal behavior, memory fragmentation, and multi-hour drift are outside what this measures.
Simulated sensor-driver load
Driver cost is replayed from coefficients measured on real hardware rather than produced by real drivers. Difference sensor configurations will vary these metrics.
Vendor agnostic Independence
This is an independently managed, vendor-agnostic benchmark. Open Navigation LLC, the maintainers of Nav2, designed the workload, ran it, and wrote the analysis.
What makes the disclosure meaningful is that the methodology, the raw logs from every run, and the analysis code are all public. Any result on this site can be recomputed from the published logs, and the whole pipeline can be re-run on your own hardware. Results that are unflattering to any platform, including in the balanced-power category, are published alongside the rest.
