Skip to content

Advanced Export

This page is the reference for every model.export() parameter, plus less common export options. New to exporting? Start with Export Basics.

Export Parameters

The export() method accepts several parameters to customize the export process:

Parameter Default Description
output_dir "output" Directory where the exported model will be saved.
format "onnx" Export format: "onnx", "tflite", "tensorrt" (alias: "trt"), "executorch", "openvino", "coreml", "coreai" or "litert".
quantization None TFLite quantization mode: None/"fp32", "fp16", or "int8". Only used when format="tflite"; format="litert" accepts only None/"fp32" and raises NotImplementedError otherwise.
calibration_data None Optional image directory, .npy file path, NumPy array, or None. Not consumed when building the generated .tflite models.
max_images 100 Maximum number of images to load from a calibration_data directory. Ignored for other calibration data formats.
infer_dir None Optional directory of sample images for inference validation during export tracing. If not provided, a random dummy image is generated.
backbone_only False Export only the backbone feature extractor instead of the full model.
opset_version 17 ONNX opset version to use for export. Higher versions support more operations.
verbose True Whether to print verbose export information.
shape None Input shape as tuple (height, width). Each dimension must be divisible by the selected model's block size (patch_size * num_windows). If not provided, uses the model's default resolution.
batch_size 1 Batch size for the exported model. With dynamic_batch=True and format="tensorrt", also the batch the engine's optimization profile is tuned for.
dynamic_batch False If True, export with a dynamic batch dimension so the model accepts variable batch sizes at runtime. Supported for format="onnx" and format="tensorrt" (which then needs max_batch_size) — TFLite, ExecuTorch, CoreML, Core AI, OpenVINO and LiteRT bake a fixed batch size (TFLite because onnx2tf fails on the dynamic-batch graph).
patch_size None Backbone patch size override. Defaults to the value from model_config.patch_size. Must match the instantiated model's patch size when provided.
backend None Backend for ExecuTorch: "xnnpack" (CPU, fp32), "coreml" (Apple, fp16), or "qnn" (Qualcomm HTP, fp16). Required when format="executorch".
soc None Target SoC chip identifier for the "qnn" backend (e.g. "SM8650" for Snapdragon 8 Gen 3). Required when backend="qnn".
fp16 True Build the TensorRT engine with FP16 precision (only used when format="tensorrt"). TensorRT 11+ removed the FP16 builder flag, so there the engine is built from an FP16-cast graph instead; engine inputs and outputs stay FP32 either way. On strongly typed TensorRT (11+), this graph cast requires onnx/onnxconverter-common — install rfdetr[tensorrt] for the complete set, or export raises ImportError. A lean/partial TensorRT < 11 wheel lacking the FP16 builder flag falls back to an FP32 engine with a warning instead. Pass False for an FP32 engine.
max_batch_size None Largest batch a dynamic TensorRT engine accepts. Required when format="tensorrt" and dynamic_batch=True: the engine gets one optimization profile spanning batch 1 .. max_batch_size, tuned for batch_size. Ignored for every other format.
notes None Optional user-defined metadata (string, dict, list, or any JSON-serialisable value) to embed in the exported ONNX model under the "rfdetr_notes" metadata property.
coreml_precision None Compute precision for format="coreml": None/"float32" (tight CPU parity with eager PyTorch) or "float16" (half the size, and the only precision the Apple Neural Engine runs — at a measured accuracy cost, see Native CoreML). Ignored for every other format.
coreai_precision None Compute precision for format="coreai": None/"float32" (matches eager PyTorch on the GPU) or "float16" (half the size, and the precision Core AI runs on the Apple Neural Engine — at a measured accuracy cost, see Core AI). Ignored for every other format.
openvino_precision None IR storage weight precision for format="openvino": None/"float16" (OpenVINO's default FP16 weight compression) or "float32" (disables compression). Execution precision still depends on the compiled device — not guaranteed to match eager PyTorch on non-CPU devices. Ignored for every other format. Does not change the output filename.
output_name None Full filename override (without extension). Takes precedence over the model's variant name and suppresses the _fp32/_fp16/_{backend} detail suffix — see Output Files.

Format Capabilities

What each format supports, as declared by its exporter. Parameters a format does not use are ignored with a warning.

Format Extra dynamic_batch Embeds notes Experimental
onnx rfdetr[onnx] yes yes no
tensorrt rfdetr[tensorrt] yes yes no
tflite rfdetr[tflite] no yes yes
openvino rfdetr[openvino] no no no
executorch rfdetr[executorch] no no yes
litert rfdetr[litert] no no yes
coreml rfdetr[coreml] no no yes
coreai rfdetr[coreai] no yes yes

For a format without dynamic batch, export one artifact per batch size.

Precision by Format

Each format exposes precision through its own parameter; Export Parameters above has the details.

Format Parameter Values
tflite quantization None/"fp32", "fp16", "int8" (dynamic-range)
litert quantization None/"fp32" only
tensorrt fp16 True (default) or False
openvino openvino_precision None/"float16" or "float32" (IR weight storage)
coreml coreml_precision None/"float32" or "float16"
coreai coreai_precision None/"float32" or "float16"
executorch backend "xnnpack" (fp32), "coreml" (fp16), "qnn" (fp16)

Lower precision is not a portable speedup — see fp16 pays off only where the silicon implements it.

Read Embedded Notes

Pass notes= to attach your own metadata, such as a dataset version or training run. For ONNX it is stored under the rfdetr_notes metadata property; strings are stored as is, any other JSON-serialisable value as JSON:

import json

import onnx

model.export(notes={"dataset": "v3", "run": "2026-10-01"})

onnx_model = onnx.load("output/rfdetr-small.onnx")
notes = {prop.key: prop.value for prop in onnx_model.metadata_props}["rfdetr_notes"]
print(json.loads(notes))

Advanced Export Examples

Export with Custom Output Directory

from rfdetr import RFDETRSmall

model = RFDETRSmall(pretrain_weights="<path/to/checkpoint.pth>")

model.export(output_dir="exports/my_model")

Export with Custom Resolution

Export the model with a specific input resolution. For example, RFDETRSmall expects dimensions divisible by 32 (patch_size=16, num_windows=2):

from rfdetr import RFDETRSmall

model = RFDETRSmall(pretrain_weights="<path/to/checkpoint.pth>")

model.export(shape=(608, 608))

Export Backbone Only

Export only the backbone feature extractor for use in custom pipelines:

from rfdetr import RFDETRSmall

model = RFDETRSmall(pretrain_weights="<path/to/checkpoint.pth>")

model.export(backbone_only=True)

The backbone export contains the encoder and its feature projector, without the detection decoder or prediction heads. ONNX outputs are feature maps in NCHW layout, ordered by projector_scale: features for the first level, followed by features_1, features_2, and so on when more levels are configured. Backbones with a second projector also return its levels as cross_attn_features, cross_attn_features_1, and so on, after the primary levels. These outputs are feature maps, not decoded boxes, masks, or keypoint coordinates.

How Export Works

Every format is written by an Exporter class built from that format's own configuration, and model.export() is a facade over them: it resolves the format to an exporter, narrows this method's union-of-every-format signature down to the settings that format actually reads, prepares one format-independent ExportGraph, and hands the graph to the exporter. RFDETR.export()'s signature and return value are the supported surface; the classes behind it are internal.

If you want to add a format, or you are reading the export code, see Exporter Blueprint for the contract each format implements and the steps a new one takes.