docTR CRNN-MobileNetV3-Small — ExecuTorch

Text recognition: reads one word crop. The other half of the pair is DB-MobileNetV3-Large, which finds the crops.

  • Source: mindee/doctr crnn_mobilenet_v3_small, 2.1M parameters
  • License: Apache-2.0
  • Input: [1, 3, 32, 128], normalised with docTR's own mean (0.694, 0.695, 0.693) and std (0.299, 0.296, 0.301)
  • Output: [1, 32, 127] — 32 timesteps over a 126-symbol alphabet plus the CTC blank

Logits, not probabilities: greedy CTC only needs the argmax per timestep, so a softmax would be thrown away. The alphabet is 0123456789, a-z, A-Z, punctuation, °£€¥¢฿, and accented Latin — read it out of crnn_mobilenet_v3_small().cfg["vocab"] rather than retyping it.

Verification (Mac arm64, 2026-08-23)

build size latency worst corr
XNNPACK fp32 8.8 MB 3.2 ms 1.000000
XNNPACK fp16 6.5 MB 4.0 ms 0.997143
Core ML fp32 6.2 MB 0.3 ms 0.999530

Eager fp32 for the same input is 121 ms.

Read end to end with the detector, on a London street photograph (convert/check_doctr.py):

boxes: 10
read: TIUZABXPRESS | Chrisiophers | Place | STREE! | LAG | BAR!

The signs say PIZZA EXPRESS and Christopher's Place.

Cropping it

docTR's own Resize for this model is preserve_aspect_ratio=True, symmetric_pad=False: the crop keeps its shape, sits at the top-left, and the rest of the strip is black. The check does the same.

That is the principled choice rather than a measured improvement. Over four street photographs neither it nor a plain squeeze won: one picture reads Chrisiophers squeezed and Clridabers padded, another reads TRAFFIC padded and TRAEFIC squeezed.

Not shipped, and why

  • int8: worst corr 0.883, against a 0.95 gate, for 6.1 MB against 8.8. Recognition is the half where a wrong character is the whole answer, so this one is not close enough.

Conversion

python convert/export_doctr.py recognise

docTR's forward runs CTC decoding in numpy whether or not it was asked to, which torch.export refuses. The model carries an exportable flag that skips it and returns raw logits; the wrapper sets that.

Feed it contiguous tensors. ExecuTorch reads a tensor in memory order rather than by its strides, and a crop built with np.transpose is not contiguous. The detector half of this pair scored 1.000000 against eager on random input and then returned a map that thresholded to zero boxes on a real photograph, entirely because of that.

(conversion scripts: executorch-models)

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support