Skip to content

Export Basics

This page covers the everyday export path: install an extra, call model.export(), locate the output files, and run inference. For the full parameter reference and less common options, see Advanced Export. For a comparison of formats and measured performance, see the Overview.

Installation

Install the export dependencies you need:

pip install "rfdetr[onnx]"
pip install "rfdetr[openvino]"
pip install "rfdetr[tflite]"
pip install "rfdetr[litert]"
pip install "rfdetr[executorch]"
pip install "rfdetr[coreml]"
pip install "rfdetr[coreai]"

Basic Export

Export your trained model to ONNX format:

from rfdetr import RFDETRSmall

model = RFDETRSmall(pretrain_weights="<path/to/checkpoint.pth>")

model.export()
from rfdetr import RFDETRSegMedium

model = RFDETRSegMedium(pretrain_weights="<path/to/checkpoint.pth>")

model.export()

This command saves the ONNX model to the output directory by default.

Choose a Format

Pass format= to model.export() to pick another target; the extra from Installation must be installed first. Which format is fastest depends on the hardware — measured numbers are on the Overview.

Deploying to format= Guide
Anywhere ONNX Runtime runs "onnx" (default) ONNX Inference
NVIDIA GPU "tensorrt" (alias "trt") TensorRT
Intel CPU, GPU or NPU "openvino" OpenVINO
Android, embedded, edge CPU "tflite" or "litert" TFLite, LiteRT
On-device PyTorch runtime "executorch" (alias "pte") ExecuTorch
Apple platforms (Xcode) "coreml" or "coreai" Native CoreML, Core AI

Formats marked experimental in Advanced Export emit a warning when constructed.

Check the Export

An exported model returns raw tensors — box and logit decoding (sigmoid, background slot, box format) is left to your inference code. The ONNX guide spells out the decoding rules and two pitfalls that apply to every format. To confirm an export matches PyTorch, run both on the same image and compare detections; the export cookbooks do this per hardware class, next to their latency numbers.

Output Files

Filenames are built from the model's variant name (e.g. rfdetr-medium, falling back to inference_model when no variant or output_name is set, or backbone_model when backbone_only=True in that same case) plus a detail suffix whenever a detail materially changes the artifact — even at its default value, since the file needs to say what it actually is:

Format Filename pattern Detail encoded
onnx {variant}.onnx (or {variant}-backbone.onnx if backbone_only=True); without a variant or output_name, inference_model.onnx (or backbone_model.onnx if backbone_only=True) none — -backbone is structural, not a precision detail
coreml {variant}_fp32.mlpackage / {variant}_fp16.mlpackage (or {variant}_fp32-backbone.mlpackage if backbone_only=True); without a variant or output_name, backbone_model_fp32.mlpackage if backbone_only=True coreml_precision, plus -backbone when named
coreai {variant}_fp32.aimodel / {variant}_fp16.aimodel; without a variant or output_name, inference_model_fp32.aimodel coreai_precision
executorch {variant}_xnnpack.pte / {variant}_coreml.pte / {variant}_qnn_{soc}.pte (or {variant}_xnnpack-backbone.pte if backbone_only=True); without a variant or output_name, backbone_model_xnnpack.pte if backbone_only=True backend (+ soc for qnn), plus -backbone when named
tensorrt {variant}_fp16.trt / {variant}_fp32.trt (or {variant}-backbone_fp16.trt / {variant}-backbone_fp32.trt if backbone_only=True) fp16, plus -backbone when named
tflite {variant}_gs_patched_fp32.tflite + {variant}_gs_patched_fp16.tflite (+ {variant}_gs_patched_dynamic_range_quant.tflite for quantization="int8"); the _gs_patched part appears whenever the graph has GridSample nodes, as RF-DETR's do precision / quantization mode
openvino {variant}.xml + {variant}.bin (or {variant}-backbone.xml/.bin if backbone_only=True); without a variant or output_name, inference_model.xml/.bin (or backbone_model.xml/.bin if backbone_only=True) none — openvino_precision controls IR weight compression, not the filename

Pass output_name="my-model" to override the variant name and write {output_name}.{ext} verbatim — this suppresses the detail suffix for every format except tflite, which always writes multiple files and so keeps its _fp32/_fp16/_dynamic_range_quant suffix even with a custom name ({output_name}_fp32.tflite, etc.).

With backbone_only=True, ONNX, CoreML, ExecuTorch, TensorRT, and OpenVINO retain a -backbone marker before the extension even when output_name is set, for example my-model-backbone.onnx. This distinguishes the backbone artifact from the full detector exported with the same name.

Run Inference with inference-models

inference-models is the recommended library for running RF-DETR inference. It supports multiple backends — PyTorch, ONNX, and TensorRT — with automatic backend selection and a unified API.

Installation

# CPU / PyTorch only
pip install inference-models

# With TensorRT support (NVIDIA GPU required)
pip install "inference-models[trt10]"  # TensorRT 10

See the inference-models installation guide for all installation options including Jetson and CUDA 11.x.

Load a Pre-trained RF-DETR Model

import cv2
from inference_models import AutoModel

# Automatically selects the best available backend for your environment
model = AutoModel.from_pretrained("rfdetr-small")

image = cv2.imread("image.jpg")
predictions = model(image)

# Convert to supervision Detections
detections = predictions[0].to_supervision()
print(detections)

Load a Local RF-DETR Checkpoint

import cv2
from inference_models import AutoModel

# Load from a local .pth checkpoint (same file used by rfdetr for training)
model = AutoModel.from_pretrained(
    "/path/to/checkpoint.pth",
    model_type="rfdetr-small",  # specify the architecture variant
)

image = cv2.imread("image.jpg")
predictions = model(image)

Force TensorRT Backend

import cv2
from inference_models import AutoModel, BackendType

# Explicitly request TensorRT — requires TRT to be installed
model = AutoModel.from_pretrained("rfdetr-small", backend=BackendType.TRT)

image = cv2.imread("image.jpg")
predictions = model(image)

AutoModel.from_pretrained accepts backend="onnx", backend="torch", or backend="trt" to override automatic backend selection.

Using the Exported Model

Once exported, you can use the ONNX model with various inference frameworks. See ONNX Inference for a complete example, or the format-specific pages for other runtimes.