Draft — not published. This page is noindex and not in the sitemap until it's approved.
Guide · Computer vision
Running YOLO on a Jetson: ONNX and TensorRT in Practice
Key takeaways
- Export to ONNX first; build the TensorRT engine on the target device, not your workstation.
- FP16 is usually free accuracy-wise; INT8 needs calibration data and a sanity check.
- Engine files are device- and version-specific — rebuild, don't copy.
- Report measured streams × fps × latency on the target hardware, never spec-sheet numbers.
By Prasanna Patil · updated
Why the workstation number is a lie
A YOLO model that runs at 30 fps on a workstation GPU can drop below real time on a Jetson — different architecture, different memory bandwidth, no eGPU. The CCTV build I did for Smart India Hackathon ran four streams on one machine at 5 fps per stream after batching; the named next step for that pipeline was exactly this export path, to buy the frame rate back on the same hardware.
The export path
Ultralytics ships model.export(format="onnx"), and TensorRT engines are built from the ONNX file with trtexec or the Python API — on the Jetson itself. The engine encodes the specific GPU, TensorRT version and precision, so it is a build artefact, not a file you copy around. Treat it like a compiled binary: rebuilt by CI or a script whenever the model changes.
Precision: FP16 vs INT8
- FP16 is the default choice: typically ~2× over FP32 on Jetson, and for detection the mAP drop is usually noise. Check it anyway on your validation set.
- INT8 buys more speed but needs representative calibration data (a few hundred frames from your actual cameras), and accuracy can slide on classes that were already marginal. Measure per class, not just overall.
Benchmarking you can hand to a client
The deliverable is a table on the target hardware: stream count, frames per second per stream, end-to-end latency (camera to alert), and detection quality on your footage. That's the report I produce for edge deployments — it turns "should be fast enough" into a number you can commit to.
Limitations of this guide
It covers single-model detection pipelines. Multi-model graphs (detection + classification + tracking heuristics on-device), DLA offload, and dynamic shapes each need their own treatment.
Frequently asked questions
Which Jetson do I need?
Depends on stream count and frame rate, which is why the benchmark comes first. An Orin Nano covers one or two streams at modest fps for person detection; more streams or higher resolution move you up the line. The honest answer comes from measuring your workload, not from a spec sheet.
Can I just run the PyTorch model on the Jetson?
It works for a demo. You pay the framework overhead on every frame, and memory pressure shows up exactly when you add a second stream. The ONNX/TensorRT path is the difference between a demo and a deployment.
How do I update the model in the field?
Rebuild the engine from the new ONNX on the same device, verify against the validation set, then swap. Containerise the runtime so the device-side change is a deploy, not a surgery session over SSH.
Model too slow for the edge device?
Tell me the device, the model and the frame rate you need.