PP-DocLayoutV3 for onnxruntime-web
This is PP-DocLayoutV3, PaddlePaddle's document layout detector, prepared for onnxruntime-web (WebGPU and WASM). It splits PDF pages into blocks in the reader of Nexus.
- Model and training: PaddlePaddle. See PaddlePaddle/PP-DocLayoutV3.
- Licence: Apache-2.0, as the original.
- Source: the official ONNX export, PaddlePaddle/PP-DocLayoutV3_onnx at commit
46bbdf188bb0a772c08aed74882ce7e51a8f1ea6.
Changes from the official export
None of these changes the boxes:
ceil_mode = 0on the pooling nodes, because onnxruntime-web's WebGPU backend doesn't support it. The outputs were compared on the benchmark pages and are identical.- Only the detections and their count are kept as outputs. The unused mask output is removed with the nodes that fed it.
- The weights are stored as fp16 and cast back to fp32 when the model loads. That halves the file; the maths stays fp32.
pp-doclayoutv3-w16.onnx is 65,751,638 bytes, with SHA-256 a8d186a39f10f4bf46fc6856d1c74ba104c51371c0db0d274faaa5486c79c747. It can be rebuilt with tools/layout-model/build.sh in the Nexus repository.
Use
- Input: the page drawn at 800×800 (the aspect ratio is not kept), as RGB float32 in CHW order divided by 255, with no mean/std normalisation. The inputs are
image,im_shape = [[800, 800]]andscale_factor. - Output: rows of
[class, score, x1, y1, x2, y2, order], plus the row count. Class ids index the 25 labels of the originalinference.yml.
Model tree for traceray/nexus-layout
Base model
PaddlePaddle/PP-DocLayoutV3