Export Basics¶
This page covers the everyday export path: install an extra, call model.export(), locate the output files, and run inference. For the full parameter reference and less common options, see Advanced Export. For a comparison of formats and measured performance, see the Overview.
Installation¶
Install the export dependencies you need:
Basic Export¶
Export your trained model to ONNX format:
This command saves the ONNX model to the output directory by default.
Choose a Format¶
Pass format= to model.export() to pick another target; the extra from Installation must be installed first. Which format is fastest depends on the hardware — measured numbers are on the Overview.
| Deploying to | format= |
Guide |
|---|---|---|
| Anywhere ONNX Runtime runs | "onnx" (default) |
ONNX Inference |
| NVIDIA GPU | "tensorrt" (alias "trt") |
TensorRT |
| Intel CPU, GPU or NPU | "openvino" |
OpenVINO |
| Android, embedded, edge CPU | "tflite" or "litert" |
TFLite, LiteRT |
| On-device PyTorch runtime | "executorch" (alias "pte") |
ExecuTorch |
| Apple platforms (Xcode) | "coreml" or "coreai" |
Native CoreML, Core AI |
Formats marked experimental in Advanced Export emit a warning when constructed.
Check the Export¶
An exported model returns raw tensors — box and logit decoding (sigmoid, background slot, box format) is left to your inference code. The ONNX guide spells out the decoding rules and two pitfalls that apply to every format. To confirm an export matches PyTorch, run both on the same image and compare detections; the export cookbooks do this per hardware class, next to their latency numbers.
Output Files¶
Filenames are built from the model's variant name (e.g. rfdetr-medium, falling back to inference_model when no variant or output_name is set, or backbone_model when backbone_only=True in that same case) plus a detail suffix whenever a detail materially changes the artifact — even at its default value, since the file needs to say what it actually is:
| Format | Filename pattern | Detail encoded |
|---|---|---|
onnx |
{variant}.onnx (or {variant}-backbone.onnx if backbone_only=True); without a variant or output_name, inference_model.onnx (or backbone_model.onnx if backbone_only=True) |
none — -backbone is structural, not a precision detail |
coreml |
{variant}_fp32.mlpackage / {variant}_fp16.mlpackage (or {variant}_fp32-backbone.mlpackage if backbone_only=True); without a variant or output_name, backbone_model_fp32.mlpackage if backbone_only=True |
coreml_precision, plus -backbone when named |
coreai |
{variant}_fp32.aimodel / {variant}_fp16.aimodel; without a variant or output_name, inference_model_fp32.aimodel |
coreai_precision |
executorch |
{variant}_xnnpack.pte / {variant}_coreml.pte / {variant}_qnn_{soc}.pte (or {variant}_xnnpack-backbone.pte if backbone_only=True); without a variant or output_name, backbone_model_xnnpack.pte if backbone_only=True |
backend (+ soc for qnn), plus -backbone when named |
tensorrt |
{variant}_fp16.trt / {variant}_fp32.trt (or {variant}-backbone_fp16.trt / {variant}-backbone_fp32.trt if backbone_only=True) |
fp16, plus -backbone when named |
tflite |
{variant}_gs_patched_fp32.tflite + {variant}_gs_patched_fp16.tflite (+ {variant}_gs_patched_dynamic_range_quant.tflite for quantization="int8"); the _gs_patched part appears whenever the graph has GridSample nodes, as RF-DETR's do |
precision / quantization mode |
openvino |
{variant}.xml + {variant}.bin (or {variant}-backbone.xml/.bin if backbone_only=True); without a variant or output_name, inference_model.xml/.bin (or backbone_model.xml/.bin if backbone_only=True) |
none — openvino_precision controls IR weight compression, not the filename |
Pass output_name="my-model" to override the variant name and write {output_name}.{ext} verbatim — this suppresses the detail suffix for every format except tflite, which always writes multiple files and so keeps its _fp32/_fp16/_dynamic_range_quant suffix even with a custom name ({output_name}_fp32.tflite, etc.).
With backbone_only=True, ONNX, CoreML, ExecuTorch, TensorRT, and OpenVINO retain a -backbone marker before the extension even when output_name is set, for example my-model-backbone.onnx. This distinguishes the backbone artifact from the full detector exported with the same name.
Run Inference with inference-models¶
inference-models is the recommended library for running RF-DETR inference. It supports multiple backends — PyTorch, ONNX, and TensorRT — with automatic backend selection and a unified API.
Installation¶
# CPU / PyTorch only
pip install inference-models
# With TensorRT support (NVIDIA GPU required)
pip install "inference-models[trt10]" # TensorRT 10
See the inference-models installation guide for all installation options including Jetson and CUDA 11.x.
Load a Pre-trained RF-DETR Model¶
import cv2
from inference_models import AutoModel
# Automatically selects the best available backend for your environment
model = AutoModel.from_pretrained("rfdetr-small")
image = cv2.imread("image.jpg")
predictions = model(image)
# Convert to supervision Detections
detections = predictions[0].to_supervision()
print(detections)
Load a Local RF-DETR Checkpoint¶
import cv2
from inference_models import AutoModel
# Load from a local .pth checkpoint (same file used by rfdetr for training)
model = AutoModel.from_pretrained(
"/path/to/checkpoint.pth",
model_type="rfdetr-small", # specify the architecture variant
)
image = cv2.imread("image.jpg")
predictions = model(image)
Force TensorRT Backend¶
import cv2
from inference_models import AutoModel, BackendType
# Explicitly request TensorRT — requires TRT to be installed
model = AutoModel.from_pretrained("rfdetr-small", backend=BackendType.TRT)
image = cv2.imread("image.jpg")
predictions = model(image)
AutoModel.from_pretrained accepts backend="onnx", backend="torch", or backend="trt" to override automatic backend selection.
Using the Exported Model¶
Once exported, you can use the ONNX model with various inference frameworks. See ONNX Inference for a complete example, or the format-specific pages for other runtimes.