Table Transformer β ExecuTorch (find tables, then read their structure)
Two models that pair. detection finds tables on a page; structure takes a
cropped table and returns its rows, columns, column header and spanning cells. Both are
DETR with a ResNet-18 backbone and 125 object queries, 28.8M parameters each.
detect_<H>x<W> pixel_values (1, 3, H, W) fp32
-> logits (1, 125, C+1) fp32, boxes (1, 125, 4) fp32
- Files:
table_transformer_detection_xnnpack_fp32.pteβ 115.8 MB,table_transformer_structure_xnnpack_fp32.pteβ 115.9 MB, three methods each - Source: microsoft/table-transformer-detection and microsoft/table-transformer-structure-recognition
- License: MIT
- Classes: detection β
table,table rotated. structure βtable,table column,table row,table column header,table projected row header,table spanning cell
Boxes come out as DETR always emits them: (cx, cy, w, h) normalised to the input,
so mapping them back to your own page is two lines of arithmetic and does not depend on
which rung produced them.
The ladder, and why there is one
Each file carries three input sizes:
| rung | for |
|---|---|
detect_667x1000 |
landscape β a wide table crop |
detect_1000x800 |
portrait β a page |
detect_800x800 |
square-ish |
Resize to the rung nearest your aspect ratio. Do not pad to a square, and do not squash: both were measured against the reference running at its own size, scoring the detections by matched IoU β
| what the caller does | worst matched IoU |
|---|---|
| resize to the size the processor would have chosen | 1.0000 |
| resize to 800x800 (aspect squashed) | 0.8644 |
pad to 1000x1000 with the correct pixel_mask |
0.2918 |
| pad to 1000x1000 with an all-ones mask | 0.2220, and three detections invented |
| resize to 1000x1000 | 0.2095 |
The padded-canvas trick that works for this shelf's audio encoders does not work here,
which is why this is a ladder rather than one padded window. Methods in one .pte share
their constants, so the three rungs cost 0.3 MB over one: a single method is 115.6 MB
and three are 115.9 MB.
Running it
1. Preprocess. ImageNet mean/std, bilinear resize to the rung, (1, 3, H, W):
mean, std = processor.image_mean, processor.image_std
x = (np.asarray(image.resize((W, H))) / 255.0 - mean) / std
pixel_values = torch.from_numpy(np.ascontiguousarray(x.transpose(2, 0, 1)))[None]
The copy is deliberate β ExecuTorch reads strides as contiguous whatever the tensor says.
2. Post-process, outside the graph. Softmax over the class axis, drop the last column (the "no object" class), keep what clears your threshold:
scores = logits.softmax(-1)[0, :, :-1]
best = scores.max(-1)
keep = best.values > 0.7
The threshold is your policy rather than the model's, which is why it is not baked in.
3. Chain them for a full page: detection to find the table, crop it with a small
margin, then structure on the crop.
Verification
Correlation over the raw output is not the unit this model is used in β 125 queries are mostly the no-object class. The gate is the detections: both arms at the same rung, and for every box the reference found, the best same-label box the build offers.
| model | eager finds (at 0.7) | .pte finds |
worst matched IoU |
|---|---|---|---|
| structure | 17 β 9 rows, 4 columns, 1 column header, 2 spanning cells, 1 table | 17 | 1.0000 |
| detection | 1 β the table | 1 | 1.0000 |
The gate image is a rendered table, which is what a table in a document is. The reference genuinely reads it: nine rows and four columns of a seven-row, four-column table, plus the shaded header.
Speed
Mac arm64, median of 5, at the 667x1000 rung β a reference point for relative cost,
not a device number.
| model | .pte |
torch eager fp32 |
|---|---|---|
| structure | 44.1 ms | 66.5 ms |
| detection | 37.8 ms | 63.4 ms |
Faster than eager, which is not the usual result on this shelf and follows from delegation: 90.6% of the ops run on XNNPACK, in 32 subgraphs.
What is not in these files
fp32 only. Reduced-precision builds go through this shelf's single-method harness, which measures parity, delegation and timing per file; these are multi-method bundles and would need that path rebuilt for the ladder. At 115 MB for 28.8M parameters there is less to gain here than for the shelf's larger models, and an unmeasured fp16 build is not one this shelf ships. No Core ML build for the same reason.
- Downloads last month
- 2