table-transformer-detection-coreml
Model Summary
This is an unofficial Core ML conversion of microsoft/table-transformer-detection. All credit for the original model goes to its authors (Microsoft).
Conversion notes
Like other DETR-family detection models, this is a single forward pass with
no autoregressive decoder, so no KV-cache complexity is needed.
torch.jit.trace + coremltools converted it directly, with no
monkeypatches required.
Usage (Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("table-transformer-detection_fp16.mlpackage")
image = Image.open("document.png").convert("RGB").resize((800, 800))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 800, 800), NCHW
pixel_mask = np.ones((1, 800, 800), dtype=np.int32)
out = mlmodel.predict({"pixel_values": pixel_values, "pixel_mask": pixel_mask})
logits = out["logits"][0] # (15, 3): id2label = {0: "table", 1: "table rotated"}, class 2 = no object
pred_boxes = out["pred_boxes"][0] # (15, 4): normalized (center_x, center_y, width, height)
Input is a fixed 800x800, NCHW. The trace assumes pixel_mask is always all
ones (no padding) for a single image; letterboxed/aspect-ratio-preserving
batch padding has not been verified.
Accuracy
Compared against the PyTorch fp32 reference on a COCO validation image:
| Metric | Result |
|---|---|
| Logits cosine similarity | 0.9999995 |
| Boxes cosine similarity | 1.0 |
| Label agreement | 100% |
Specs
| Item | Value |
|---|---|
| Base model | microsoft/table-transformer-detection (ResNet18 + DETR, 28.8M params) |
| Precision | float16 |
| Input | 800x800, NCHW, fixed size |
| Framework | Core ML (mlprogram, minimum_deployment_target=macOS14) |
Notes
- This is a community conversion, not an official release from Microsoft.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
モデルの概要
microsoft/table-transformer-detection のCore ML版です。元モデルの著作権はその作者(Microsoft)に帰属します。
変換について
DETR系の検出モデルも自己回帰デコーダーを持たない1回のフォワードパスのため、KVキャッシュ等は
不要です。torch.jit.trace + coremltoolsでモンキーパッチ無しにそのまま変換できました。
使い方(Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("table-transformer-detection_fp16.mlpackage")
image = Image.open("document.png").convert("RGB").resize((800, 800))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 800, 800), NCHW
pixel_mask = np.ones((1, 800, 800), dtype=np.int32)
out = mlmodel.predict({"pixel_values": pixel_values, "pixel_mask": pixel_mask})
logits = out["logits"][0] # (15, 3): id2label = {0: "table", 1: "table rotated"}, class 2 = no object
pred_boxes = out["pred_boxes"][0] # (15, 4): normalized (center_x, center_y, width, height)
入力は固定サイズ800x800、NCHW。pixel_maskは画像に対して常に全1(パディング無し)を想定した
トレースです。アスペクト比を保ったままレターボックス的にパディングする使い方は未検証です。
精度検証
PyTorch fp32リファレンスと、COCO検証画像1枚で比較:
| 項目 | 結果 |
|---|---|
| Logitsコサイン類似度 | 0.9999995 |
| Boxesコサイン類似度 | 1.0 |
| ラベル一致率 | 100% |
Specs
| Item | Value |
|---|---|
| ベースモデル | microsoft/table-transformer-detection(ResNet18 + DETR、28.8M params) |
| 精度 | float16 |
| 入力 | 800x800、NCHW、固定サイズ |
| フレームワーク | Core ML(mlprogram、minimum_deployment_target=macOS14) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
- Downloads last month
- 5
Model tree for masahiroid/table-transformer-detection-coreml
Base model
microsoft/table-transformer-detection