PP-DocLayoutV3 for onnxruntime-web

This is PP-DocLayoutV3, PaddlePaddle's document layout detector, prepared for onnxruntime-web (WebGPU and WASM). It splits PDF pages into blocks in the reader of Nexus.

Changes from the official export

None of these changes the boxes:

  1. ceil_mode = 0 on the pooling nodes, because onnxruntime-web's WebGPU backend doesn't support it. The outputs were compared on the benchmark pages and are identical.
  2. Only the detections and their count are kept as outputs. The unused mask output is removed with the nodes that fed it.
  3. The weights are stored as fp16 and cast back to fp32 when the model loads. That halves the file; the maths stays fp32.

pp-doclayoutv3-w16.onnx is 65,751,638 bytes, with SHA-256 a8d186a39f10f4bf46fc6856d1c74ba104c51371c0db0d274faaa5486c79c747. It can be rebuilt with tools/layout-model/build.sh in the Nexus repository.

Use

  • Input: the page drawn at 800×800 (the aspect ratio is not kept), as RGB float32 in CHW order divided by 255, with no mean/std normalisation. The inputs are image, im_shape = [[800, 800]] and scale_factor.
  • Output: rows of [class, score, x1, y1, x2, y2, order], plus the row count. Class ids index the 25 labels of the original inference.yml.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for traceray/nexus-layout

Quantized
(4)
this model