RF-DETR on Mobile & Edge — Export, Inference & Latency¶
Compares every export path built for on-device / ARM mobile and edge deployment: eager PyTorch (running here as the reference anchor, not itself a deployment target) against TFLite, LiteRT, and ExecuTorch's XNNPACK backend — each exported, run once for a correctness check, then benchmarked on the CPU that also stands in for a phone/edge chip's CPU core.
| Format | Export | Inference |
|---|---|---|
| PyTorch | (no export — eager model) | predict() / inference() (TorchScript JIT, fp32) |
| TFLite | ONNX → TensorFlow → .tflite (onnx2tf) |
tensorflow.lite.Interpreter |
| LiteRT | torch.export → .tflite directly (litert-torch) |
ai_edge_litert.interpreter.Interpreter |
| ExecuTorch (XNNPACK) | torch.export → .pte |
executorch.runtime.Runtime |
Not covered here: NVIDIA GPU deployment (see the CUDA cookbook), general-purpose desktop/server CPU via ONNX Runtime or OpenVINO (see the CPU cookbook), or Apple Silicon — native CoreML, Core AI, and ExecuTorch's CoreML backend (see the Apple cookbook).
Python version. The
[tflite]extra installs only on Python 3.12 exactly — itsonnx2tf/tensorflowdependency stack pinsnumpy==1.26.4, which cannot be satisfied on 3.10, 3.11, 3.13, or 3.14 (seepyproject.toml). This constraint is why TFLite/LiteRT/ExecuTorch get their own notebook, separate from the CPU cookbook's ONNX/OpenVINO, which have no such pin.
Two parts, two environments.
pyproject.toml's[tool.uv] conflictsdeclares[tflite]unsatisfiable alongside both[litert](onnx2tfpinsai-edge-litert==2.1.2;litert-torchneeds>=2.2.0) and[executorch]— the three formats cannot share one Python environment. This notebook is split into Part A (TFLite) and Part B (LiteRT + ExecuTorch), each with its own install cell; run one part, then restart the runtime before running the other. Each part re-establishes its own PyTorch baseline and its own results table rather than assuming state survives the restart.
Part A: TFLite¶
1. Install¶
Installs from develop to pick up the newest export fixes. The assertion below fails fast on any
interpreter other than 3.12, before the (slow) install runs. Colab ships mutually inconsistent
preinstalled packages that otherwise crash import rfdetr, so two are aligned: torchaudio is
uninstalled (RF-DETR never uses it, but transformers imports it when present and a
torch/torchaudio CUDA-version mismatch then errors), and pillow is force-reinstalled to a
clean version (a half-upgraded PIL breaks torchvision's import with cannot import name '_Ink').
Colab: after this cell, Runtime → Restart session, then run from the next cell — Colab keeps the old package versions loaded until a restart.
import sys
assert sys.version_info[:2] == (3, 12), (
"rfdetr[tflite] requires Python 3.12 exactly (see pyproject.toml); this runtime is "
f"{sys.version_info.major}.{sys.version_info.minor}. Use a Python 3.12 environment instead."
)
!pip install -q "rfdetr[tflite] @ https://github.com/roboflow/rf-detr/archive/refs/heads/develop.zip" psutil supervision pandas
!pip install -q --force-reinstall --no-deps "pillow==11.3.0"
!pip uninstall -q -y torchaudio
2. Setup¶
Every format in this notebook runs on CPU — no GPU is needed.
Three small helpers are shared by every format section below: _artifact_size_mb reports an
export artifact's size on disk, visualize_detections annotates and displays a
supervision.Detections on the sample image (falling back to COCO_CLASSES for label text when
a detection object carries no class_name), and measure_memory (from _benchmark) measures
the host resident-memory growth of constructing a runtime and running its first inference call.
from pathlib import Path
import numpy as np
import supervision as sv
from PIL import Image
from rfdetr.assets.coco_classes import COCO_CLASSES
from rfdetr.export._benchmark import BenchmarkResult, measure_latency, measure_memory
EXPORT_DIR = Path("export_mobile")
EXPORT_DIR.mkdir(exist_ok=True)
CONFIDENCE_THRESHOLD = 0.5
WARMUP_RUNS = 5
MEASURE_RUNS = 30
def _artifact_size_mb(*paths: Path) -> float:
total_bytes = 0
for path in paths:
if path.is_dir():
total_bytes += sum(f.stat().st_size for f in path.rglob("*") if f.is_file())
else:
total_bytes += path.stat().st_size
return total_bytes / 1e6
def visualize_detections(detections: sv.Detections, image: Image.Image, save_path: Path | None = None) -> None:
names = detections.data.get("class_name") if detections.data else None
if names is None:
names = [COCO_CLASSES.get(int(c), str(c)) for c in detections.class_id]
labels = [f"{name} {conf:.2f}" for name, conf in zip(names, detections.confidence)]
annotated = sv.BoxAnnotator(thickness=3).annotate(scene=image.copy(), detections=detections)
annotated = sv.LabelAnnotator(text_scale=0.6, text_thickness=1, text_padding=4).annotate(
scene=annotated, detections=detections, labels=labels
)
if save_path is not None:
annotated.save(save_path)
print(f"Saved annotated image: {save_path}")
sv.plot_image(annotated)
def _enable_notebook_inline_matplotlib() -> None:
"""Enable inline matplotlib figures when running in IPython."""
get_ipython_func = globals().get("get_ipython")
if not callable(get_ipython_func):
return
ipython = get_ipython_func()
if ipython is not None:
ipython.run_line_magic("matplotlib", "inline")
ipython.run_line_magic("config", "InlineBackend.close_figures = True")
_enable_notebook_inline_matplotlib()
def _fmt_ms(result: BenchmarkResult | None) -> str:
if result is None:
return "—"
return f"{result.mean_ms:.2f} ± {result.std_ms:.2f}"
def _result_row(
format_label: str,
config: str,
forward: BenchmarkResult | None,
end2end: BenchmarkResult | None,
memory_mb: float | None,
) -> dict:
fps = (end2end or forward).fps
return {
"Format": format_label,
"Config": config,
"forward [ms]": _fmt_ms(forward),
"end2end [ms]": _fmt_ms(end2end),
"FPS [img/s] (end2end)": round(fps, 1),
"Memory [MB]": f"{memory_mb:.1f}" if memory_mb is not None else "—",
}
3. Sample image¶
A single street scene with several COCO classes (dog, person, backpack, car) is enough to verify detections. The image is downloaded once and reused for every format below.
import urllib.request
IMAGE_URL = "https://media.roboflow.com/notebooks/examples/dog.jpeg"
IMAGE_PATH = EXPORT_DIR / "sample.jpg"
if not IMAGE_PATH.exists():
urllib.request.urlretrieve(IMAGE_URL, IMAGE_PATH)
image = Image.open(IMAGE_PATH).convert("RGB")
print(f"Sample image: {image.size[0]}×{image.size[1]}")
4. PyTorch baseline — predict() / inference()¶
No export needed — this is the reference every on-device format in this notebook is compared
against. It runs on this machine's CPU, not a phone or edge chip, so treat it as an accuracy and
relative-speed anchor rather than a deployment number. predict() is the unoptimized baseline;
inference() defaults to dtype=torch.float32 with compile_backend="torchscript" — a
JIT-compiled, still-fp32 optimization; remove_optimized_model() reverts it afterward so model
stays reusable for the export calls below.
Both calls include preprocessing and postprocessing — there is no separate "forward-only" path
through eager predict(), so only an end-to-end number is reported for the PyTorch baseline. The
memory bracket covers construction plus the first predict() call, since weights and any lazily
built buffers aren't fully resident until after that first call; the JIT row's bracket covers
inference() plus its own first call, reported as the increment on top of the eager row above it.
from rfdetr import RFDETRSmall
with measure_memory() as mem:
model = RFDETRSmall(device="cpu")
baseline_detections = model.predict(image, threshold=CONFIDENCE_THRESHOLD)
pytorch_eager_memory_mb = mem.delta_mb
print(f"PyTorch baseline: {len(baseline_detections)} detections above {CONFIDENCE_THRESHOLD}")
visualize_detections(baseline_detections, image)
pytorch_eager = measure_latency(
lambda: model.predict(image), label="PyTorch predict()", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
with measure_memory() as mem:
model.inference()
_ = model.predict(image)
pytorch_jit_memory_mb = mem.delta_mb
pytorch_jit = measure_latency(
lambda: model.predict(image), label="PyTorch inference() JIT", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
model.remove_optimized_model()
for r, mb in ((pytorch_eager, pytorch_eager_memory_mb), (pytorch_jit, pytorch_jit_memory_mb)):
print(f" {r.label:<32} {r.mean_ms:6.2f} ms ± {r.std_ms:5.2f} ({r.fps:6.1f} FPS) +{mb:.1f} MB")
5. TFLite (fp32, fp16, INT8 dynamic-range)¶
What it is. TFLite is TensorFlow's lightweight runtime for mobile, embedded, and edge
hardware. RF-DETR's route converts ONNX → TensorFlow → .tflite via onnx2tf. Good for
Android and embedded targets where TensorFlow Lite is already the established runtime; the
tradeoff is an experimental multi-step conversion chain rather than a bit-exact match to eager
PyTorch. See the TFLite export docs.
All three precisions come from one export call: onnx2tf always writes *_fp32.tflite and
*_fp16.tflite, and quantization="int8" additionally builds a dynamic-range INT8 model (INT8
weights, float32 activations, no calibration data) and returns that as the primary path. So the
three files below are siblings in one directory, benchmarked against each other with no
re-export.
Export¶
from collections.abc import Callable
import tensorflow as tf
import torch
from rfdetr.export.benchmark import infer_transforms, post_process
tflite_path = model.export(format="tflite", quantization="int8", output_dir=str(EXPORT_DIR))
tflite_stem = tflite_path.name.removesuffix("_dynamic_range_quant.tflite")
tflite_fp32_path = tflite_path.with_name(f"{tflite_stem}_fp32.tflite")
tflite_fp16_path = tflite_path.with_name(f"{tflite_stem}_fp16.tflite")
for label, path in (("fp32", tflite_fp32_path), ("fp16", tflite_fp16_path), ("INT8 dyn-range", tflite_path)):
print(f"TFLite {label:<14} {path.name} ({_artifact_size_mb(path):.1f} MB on disk)")
Inference¶
The TFLite input is NHWC, not the NCHW layout the other formats in this notebook use.
onnx2tf's SavedModel route also renames every output, so outputs are matched by rank and last
dimension instead of by name: boxes are the rank-3 tensor with last dimension 4.
_tflite_runner builds one interpreter per file and returns its forward callable plus the memory
that interpreter cost — the measure_memory bracket covers construction and the first
inference, since TFLite allocates lazily.
resolution = model.model_config.resolution
tflite_transform = infer_transforms((resolution, resolution))
tflite_tensor, _ = tflite_transform(image, None)
# TFLite expects NHWC, not the NCHW layout used by the other export formats.
tflite_input = tflite_tensor.permute(1, 2, 0)[None].contiguous().numpy().astype(np.float32)
tflite_target_sizes = torch.tensor([[image.height, image.width]])
def _tflite_runner(model_path: Path) -> tuple[Callable[[np.ndarray], tuple[torch.Tensor, torch.Tensor]], float]:
"""Build a TFLite interpreter for one artifact; return its forward callable and its memory cost."""
with measure_memory() as mem:
interpreter = tf.lite.Interpreter(model_path=str(model_path))
interpreter.allocate_tensors()
input_index = interpreter.get_input_details()[0]["index"]
rank3 = [detail for detail in interpreter.get_output_details() if len(detail["shape"]) == 3]
boxes_detail = next(detail for detail in rank3 if detail["shape"][-1] == 4)
labels_detail = next(detail for detail in rank3 if detail["shape"][-1] != 4)
def forward(inp: np.ndarray) -> tuple[torch.Tensor, torch.Tensor]:
interpreter.set_tensor(input_index, inp)
interpreter.invoke()
dets = torch.from_numpy(interpreter.get_tensor(boxes_detail["index"]))
labels = torch.from_numpy(interpreter.get_tensor(labels_detail["index"]))
return dets, labels
forward(tflite_input) # first inference inside the bracket — TFLite allocates lazily
return forward, mem.delta_mb
def _tflite_check(
forward: Callable[[np.ndarray], tuple[torch.Tensor, torch.Tensor]], label: str, save_name: str
) -> None:
"""Run one inference through *forward*, print the detection count, and save an annotated image."""
dets, labels = forward(tflite_input)
result = post_process({"dets": dets, "labels": labels}, tflite_target_sizes)[0]
keep = result["scores"] > CONFIDENCE_THRESHOLD
print(f"TFLite {label}: {int(keep.sum())} detections above {CONFIDENCE_THRESHOLD}")
detections = sv.Detections(
xyxy=result["boxes"][keep].numpy(),
confidence=result["scores"][keep].numpy(),
class_id=result["labels"][keep].numpy().astype(int),
)
visualize_detections(detections, image, EXPORT_DIR / save_name)
def _tflite_end2end_thunk(forward: Callable[[np.ndarray], tuple[torch.Tensor, torch.Tensor]]) -> Callable[[], None]:
"""Wrap *forward* into a zero-arg preprocess + forward + postprocess thunk for ``measure_latency``."""
def run() -> None:
tensor, _ = tflite_transform(image, None)
inp = tensor.permute(1, 2, 0)[None].contiguous().numpy().astype(np.float32)
dets, labels = forward(inp)
post_process({"dets": dets, "labels": labels}, tflite_target_sizes)
return run
tflite_fp32_fn, tflite_fp32_memory_mb = _tflite_runner(tflite_fp32_path)
tflite_fp16_fn, tflite_fp16_memory_mb = _tflite_runner(tflite_fp16_path)
tflite_int8_fn, tflite_memory_mb = _tflite_runner(tflite_path)
_tflite_check(tflite_fp32_fn, "fp32", "annotated_tflite_fp32.jpg")
_tflite_check(tflite_fp16_fn, "fp16", "annotated_tflite_fp16.jpg")
_tflite_check(tflite_int8_fn, "INT8", "annotated_tflite.jpg")
Benchmark¶
tflite_fp32_forward = measure_latency(
lambda: tflite_fp32_fn(tflite_input),
label="TFLite fp32 forward",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
tflite_fp32_end2end = measure_latency(
_tflite_end2end_thunk(tflite_fp32_fn),
label="TFLite fp32 end2end",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
tflite_fp16_forward = measure_latency(
lambda: tflite_fp16_fn(tflite_input),
label="TFLite fp16 forward",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
tflite_fp16_end2end = measure_latency(
_tflite_end2end_thunk(tflite_fp16_fn),
label="TFLite fp16 end2end",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
tflite_forward = measure_latency(
lambda: tflite_int8_fn(tflite_input),
label="TFLite INT8 forward",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
tflite_end2end = measure_latency(
_tflite_end2end_thunk(tflite_int8_fn),
label="TFLite INT8 end2end",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
for r in (
tflite_fp32_forward,
tflite_fp32_end2end,
tflite_fp16_forward,
tflite_fp16_end2end,
tflite_forward,
tflite_end2end,
):
print(f" {r.label:<32} {r.mean_ms:6.2f} ms ± {r.std_ms:5.2f} ({r.fps:6.1f} FPS)")
6. Results — Part A¶
Run this cell to build the comparison table on your own machine — numbers vary by CPU, and this
notebook's own host stands in for a phone/edge chip's CPU core only approximately, so no numbers
are committed to this page; run it to get yours. Config is the precision/layout used for that
row; Memory [MB] is host resident-memory growth across constructing the runtime plus its first
inference call — an approximate, single-process measurement, not an isolated per-format sandbox.
On CPU, fp16 buys nothing here — INT8 is the lever. Measured on an Apple M-series CPU, TFLite fp16 came out the same speed as fp32 (784 vs 779 ms forward, inside run-to-run noise) even though the file is half the size: a fp16
.tflitestill computes in fp32 on a CPU delegate that has no native fp16 kernels, so the smaller file saves disk and memory bandwidth, not arithmetic. Dynamic-range INT8 was ~2.5× faster (317 ms) because it genuinely changes the kernels used. That speed costs accuracy: on the sample image fp32 and fp16 both cleared the 0.5 threshold on 4 detections, INT8 on only 3. Measure both on your own target before picking.
import pandas as pd
summary_a = pd.DataFrame(
[
_result_row("PyTorch predict()", "eager, fp32", None, pytorch_eager, pytorch_eager_memory_mb),
_result_row("PyTorch inference()", "JIT, fp32", None, pytorch_jit, pytorch_jit_memory_mb),
_result_row("TFLite", "fp32, NHWC", tflite_fp32_forward, tflite_fp32_end2end, tflite_fp32_memory_mb),
_result_row("TFLite", "fp16, NHWC", tflite_fp16_forward, tflite_fp16_end2end, tflite_fp16_memory_mb),
_result_row("TFLite", "INT8 dyn-range, NHWC", tflite_forward, tflite_end2end, tflite_memory_mb),
]
).set_index("Format")
print(summary_a.to_string())
print(f"\n{MEASURE_RUNS} timed + {WARMUP_RUNS} warmup runs, batch 1, CPU (approximates a phone/edge chip's CPU core).")
Part B: LiteRT + ExecuTorch¶
Restart the runtime before this part.
[litert]and[executorch]don't conflict with each other, but both conflict with[tflite]above — Colab: Runtime → Restart session, then run from the next cell. This part re-installs, re-imports, and re-downloads the sample image and PyTorch baseline independently of Part A.
7. Install¶
Colab ships mutually inconsistent preinstalled packages that otherwise crash import rfdetr, so
two are aligned: torchaudio is uninstalled (RF-DETR never uses it, but transformers imports
it when present and a torch/torchaudio CUDA-version mismatch then errors), and pillow is
force-reinstalled to a clean version (a half-upgraded PIL breaks torchvision's import with
cannot import name '_Ink').
flatc— ExecuTorch serializes the.ptewith the FlatBuffers compiler. It ships in the Linuxexecutorchwheel (so Colab works out of the box); on macOS install it separately withbrew install flatbuffers.
!pip install -q "rfdetr[litert,executorch] @ https://github.com/roboflow/rf-detr/archive/refs/heads/develop.zip" "torch<2.13" psutil supervision pandas
!pip install -q --force-reinstall --no-deps "pillow==11.3.0"
!pip uninstall -q -y torchaudio
8. Setup¶
Same shared helpers as Part A, redefined here since the runtime restart cleared them. (The
COCO_CLASSES import below is a genuine re-import, not dead code — ruff's static analysis sees
one continuous module and would otherwise treat it as redundant with Part A's, but at runtime
Part A's cells never ran in this kernel.)
from pathlib import Path
import numpy as np
import pandas as pd
import supervision as sv
from PIL import Image
from rfdetr.assets.coco_classes import COCO_CLASSES # noqa: F811 -- Part B's own kernel, not a re-import of Part A's
from rfdetr.export._benchmark import BenchmarkResult, measure_latency, measure_memory
from rfdetr.export.benchmark import infer_transforms
EXPORT_DIR = Path("export_mobile")
EXPORT_DIR.mkdir(exist_ok=True)
CONFIDENCE_THRESHOLD = 0.5
WARMUP_RUNS = 5
MEASURE_RUNS = 30
def _artifact_size_mb(*paths: Path) -> float:
total_bytes = 0
for path in paths:
if path.is_dir():
total_bytes += sum(f.stat().st_size for f in path.rglob("*") if f.is_file())
else:
total_bytes += path.stat().st_size
return total_bytes / 1e6
def visualize_detections(detections: sv.Detections, image: Image.Image, save_path: Path | None = None) -> None:
names = detections.data.get("class_name") if detections.data else None
if names is None:
names = [COCO_CLASSES.get(int(c), str(c)) for c in detections.class_id]
labels = [f"{name} {conf:.2f}" for name, conf in zip(names, detections.confidence)]
annotated = sv.BoxAnnotator(thickness=3).annotate(scene=image.copy(), detections=detections)
annotated = sv.LabelAnnotator(text_scale=0.6, text_thickness=1, text_padding=4).annotate(
scene=annotated, detections=detections, labels=labels
)
if save_path is not None:
annotated.save(save_path)
print(f"Saved annotated image: {save_path}")
sv.plot_image(annotated)
def _enable_notebook_inline_matplotlib() -> None:
"""Enable inline matplotlib figures when running in IPython."""
get_ipython_func = globals().get("get_ipython")
if not callable(get_ipython_func):
return
ipython = get_ipython_func()
if ipython is not None:
ipython.run_line_magic("matplotlib", "inline")
ipython.run_line_magic("config", "InlineBackend.close_figures = True")
_enable_notebook_inline_matplotlib()
def _fmt_ms(result: BenchmarkResult | None) -> str:
if result is None:
return "—"
return f"{result.mean_ms:.2f} ± {result.std_ms:.2f}"
def _result_row(
format_label: str,
config: str,
forward: BenchmarkResult | None,
end2end: BenchmarkResult | None,
memory_mb: float | None,
) -> dict:
fps = (end2end or forward).fps
return {
"Format": format_label,
"Config": config,
"forward [ms]": _fmt_ms(forward),
"end2end [ms]": _fmt_ms(end2end),
"FPS [img/s] (end2end)": round(fps, 1),
"Memory [MB]": f"{memory_mb:.1f}" if memory_mb is not None else "—",
}
9. Sample image¶
import urllib.request
IMAGE_URL = "https://media.roboflow.com/notebooks/examples/dog.jpeg"
IMAGE_PATH = EXPORT_DIR / "sample.jpg"
if not IMAGE_PATH.exists():
urllib.request.urlretrieve(IMAGE_URL, IMAGE_PATH)
image = Image.open(IMAGE_PATH).convert("RGB")
print(f"Sample image: {image.size[0]}×{image.size[1]}")
10. PyTorch baseline — predict() / inference()¶
Same reference anchor as Part A, re-measured in this fresh runtime.
from rfdetr import RFDETRSmall
with measure_memory() as mem:
model = RFDETRSmall(device="cpu")
baseline_detections = model.predict(image, threshold=CONFIDENCE_THRESHOLD)
pytorch_eager_memory_mb = mem.delta_mb
print(f"PyTorch baseline: {len(baseline_detections)} detections above {CONFIDENCE_THRESHOLD}")
visualize_detections(baseline_detections, image)
pytorch_eager = measure_latency(
lambda: model.predict(image), label="PyTorch predict()", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
with measure_memory() as mem:
model.inference()
_ = model.predict(image)
pytorch_jit_memory_mb = mem.delta_mb
pytorch_jit = measure_latency(
lambda: model.predict(image), label="PyTorch inference() JIT", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
model.remove_optimized_model()
for r, mb in ((pytorch_eager, pytorch_eager_memory_mb), (pytorch_jit, pytorch_jit_memory_mb)):
print(f" {r.label:<32} {r.mean_ms:6.2f} ms ± {r.std_ms:5.2f} ({r.fps:6.1f} FPS) +{mb:.1f} MB")
11. LiteRT¶
What it is. LiteRT (formerly TensorFlow Lite) is Google's on-device runtime. RF-DETR exports
to it directly via torch.export (litert-torch) — no ONNX step, no TensorFlow step. Good
for the same on-device deployment target as TFLite in Part A, with a more direct conversion
path that tracks eager PyTorch more closely; the tradeoffs are fp32-only, a fixed batch size, and
no keypoint model support on the current litert-torch version. See the
LiteRT export docs.
Export¶
format="litert" hands the model to litert-torch,
which captures it with torch.export and lowers straight to a .tflite — no ONNX, no
TensorFlow.
from rfdetr.export._runtime.decode import decode_detections
from rfdetr.export._runtime.preprocess import preprocess_to_nchw
resolution = model.model_config.resolution
litert_path = model.export(format="litert", output_dir=str(EXPORT_DIR))
litert_size_mb = _artifact_size_mb(litert_path)
print(f"LiteRT model: {litert_path} ({litert_size_mb:.1f} MB on disk)")
Inference¶
Unlike the TFLite route in Part A, LiteRT keeps PyTorch's NCHW layout, and its outputs are matched by position (boxes first, logits second), not by rank or name.
from ai_edge_litert.interpreter import Interpreter
def _litert_forward() -> tuple[np.ndarray, np.ndarray]:
litert_interpreter.set_tensor(litert_input_detail["index"], litert_input)
litert_interpreter.invoke()
boxes, logits = (litert_interpreter.get_tensor(d["index"]) for d in litert_output_details)
return boxes, logits
with measure_memory() as mem:
litert_interpreter = Interpreter(model_path=str(litert_path))
litert_interpreter.allocate_tensors()
(litert_input_detail,) = litert_interpreter.get_input_details()
_, litert_channels, litert_height, litert_width = litert_input_detail["shape"]
litert_output_details = litert_interpreter.get_output_details()[:2]
litert_input = preprocess_to_nchw(image, litert_height, litert_width, litert_channels)
litert_boxes, litert_logits = _litert_forward()
litert_memory_mb = mem.delta_mb
litert_decoded = decode_detections(litert_boxes[0], litert_logits[0], image.size, threshold=CONFIDENCE_THRESHOLD)
print(f"LiteRT: {len(litert_decoded.xyxy)} detections above {CONFIDENCE_THRESHOLD}")
litert_sv_detections = sv.Detections(
xyxy=litert_decoded.xyxy, confidence=litert_decoded.confidence, class_id=litert_decoded.class_id.astype(int)
)
visualize_detections(litert_sv_detections, image, EXPORT_DIR / "annotated_litert.jpg")
Benchmark¶
litert_forward = measure_latency(
_litert_forward, label="LiteRT forward", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
def _litert_end2end() -> None:
inp = preprocess_to_nchw(image, litert_height, litert_width, litert_channels)
litert_interpreter.set_tensor(litert_input_detail["index"], inp)
litert_interpreter.invoke()
boxes, logits = (litert_interpreter.get_tensor(d["index"]) for d in litert_output_details)
decode_detections(boxes[0], logits[0], image.size, threshold=CONFIDENCE_THRESHOLD)
litert_end2end = measure_latency(
_litert_end2end, label="LiteRT end2end", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
for r in (litert_forward, litert_end2end):
print(f" {r.label:<32} {r.mean_ms:6.2f} ms ± {r.std_ms:5.2f} ({r.fps:6.1f} FPS)")
12. ExecuTorch (XNNPACK backend)¶
What it is. ExecuTorch is PyTorch's own on-device inference runtime — the model is exported
directly via torch.export to a portable .pte binary, no ONNX conversion involved. The
"xnnpack" backend targets any CPU platform in fp32. Good for on-device PyTorch deployment
(Android, iOS, embedded) when staying inside the PyTorch ecosystem end to end matters more than
picking the single fastest runtime for one platform. See the
ExecuTorch export docs.
Export¶
format="executorch" captures the model with torch.export (no intermediate conversion) and
lowers it to the requested backend — xnnpack here, RF-DETR's portable-CPU backend that also
runs on Android and iOS.
import torch
pte_path = model.export(format="executorch", backend="xnnpack", output_dir=str(EXPORT_DIR))
executorch_size_mb = _artifact_size_mb(pte_path)
print(f"ExecuTorch program: {pte_path} ({executorch_size_mb:.1f} MB on disk)")
Inference¶
Runtime.load_program(...).load_method("forward") gives a callable graph; its two outputs are
dets (boxes, normalized cxcywh) and labels (class logits).
The input tensor must be contiguous. The ExecuTorch runtime reads the input buffer as contiguous NCHW and ignores tensor strides — a preprocessing step that permutes axes without copying returns a strided view the runtime misreads as a scrambled image. Nothing errors: the model runs and returns plausible-shaped output, but every detection's score collapses below threshold.
infer_transformsalready materializes a contiguous tensor; the trailing.contiguous()below is defensive.
from executorch.runtime import Runtime
def _executorch_forward(pixel_values: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
return executorch_method.execute([pixel_values])
executorch_transform = infer_transforms((resolution, resolution))
executorch_tensor, _ = executorch_transform(image, None)
executorch_pixel_values = executorch_tensor[None].float().contiguous()
with measure_memory() as mem:
executorch_method = Runtime.get().load_program(str(pte_path)).load_method("forward")
executorch_dets, executorch_labels = _executorch_forward(executorch_pixel_values)
executorch_memory_mb = mem.delta_mb
executorch_target_sizes = torch.tensor([[image.height, image.width]])
executorch_result = post_process({"dets": executorch_dets, "labels": executorch_labels}, executorch_target_sizes)[0]
executorch_keep = executorch_result["scores"] > CONFIDENCE_THRESHOLD
print(f"ExecuTorch (XNNPACK): {int(executorch_keep.sum())} detections above {CONFIDENCE_THRESHOLD}")
executorch_sv_detections = sv.Detections(
xyxy=executorch_result["boxes"][executorch_keep].numpy(),
confidence=executorch_result["scores"][executorch_keep].numpy(),
class_id=executorch_result["labels"][executorch_keep].numpy().astype(int),
)
visualize_detections(executorch_sv_detections, image, EXPORT_DIR / "annotated_executorch.jpg")
Benchmark¶
executorch_forward = measure_latency(
lambda: _executorch_forward(executorch_pixel_values),
label="ExecuTorch (XNNPACK) forward",
device="cpu",
warmup=WARMUP_RUNS,
runs=MEASURE_RUNS,
)
def _executorch_end2end() -> None:
tensor, _ = executorch_transform(image, None)
pixel_values = tensor[None].float().contiguous()
dets, labels = _executorch_forward(pixel_values)
post_process({"dets": dets, "labels": labels}, executorch_target_sizes)
executorch_end2end = measure_latency(
_executorch_end2end, label="ExecuTorch (XNNPACK) end2end", device="cpu", warmup=WARMUP_RUNS, runs=MEASURE_RUNS
)
for r in (executorch_forward, executorch_end2end):
print(f" {r.label:<32} {r.mean_ms:6.2f} ms ± {r.std_ms:5.2f} ({r.fps:6.1f} FPS)")
13. Results — Part B¶
Run this cell to build the comparison table on your own machine — numbers vary by CPU, and this
notebook's own host stands in for a phone/edge chip's CPU core only approximately, so no numbers
are committed to this page; run it to get yours. Memory [MB] is host resident-memory growth
across constructing the runtime plus its first inference call.
summary_b = pd.DataFrame(
[
_result_row("PyTorch predict()", "eager, fp32", None, pytorch_eager, pytorch_eager_memory_mb),
_result_row("PyTorch inference()", "JIT, fp32", None, pytorch_jit, pytorch_jit_memory_mb),
_result_row("LiteRT", "fp32, NCHW", litert_forward, litert_end2end, litert_memory_mb),
_result_row("ExecuTorch", "XNNPACK, fp32", executorch_forward, executorch_end2end, executorch_memory_mb),
]
).set_index("Format")
print(summary_b.to_string())
print(f"\n{MEASURE_RUNS} timed + {WARMUP_RUNS} warmup runs, batch 1, CPU (approximates a phone/edge chip's CPU core).")
Next steps¶
- Fine-tuned weights — pass
pretrain_weights="<path/to/checkpoint.pth>"when constructing the model. - Deploy on-device — copy the
.tflite/.ptefile to your Android / iOS / edge app and run it with that platform's TensorFlow Lite, LiteRT, or ExecuTorch runtime. - Apple Silicon — ExecuTorch's
coremlbackend (Apple Neural Engine, fp16) and native CoreML / Core AI export are covered in the Apple cookbook, not here. - Qualcomm Snapdragon — ExecuTorch's
qnnbackend needs a source build against the QAIRT SDK; see the ExecuTorch export docs. - Have a CUDA GPU or a general desktop/server CPU? — see the CUDA cookbook or the CPU cookbook.
- See the Export documentation for every format and option.