table-transformer-detection-coreml

English | 日本語

Model Summary

This is an unofficial Core ML conversion of microsoft/table-transformer-detection. All credit for the original model goes to its authors (Microsoft).

Conversion notes

Like other DETR-family detection models, this is a single forward pass with no autoregressive decoder, so no KV-cache complexity is needed. torch.jit.trace + coremltools converted it directly, with no monkeypatches required.

Usage (Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("table-transformer-detection_fp16.mlpackage")

image = Image.open("document.png").convert("RGB").resize((800, 800))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 800, 800), NCHW
pixel_mask = np.ones((1, 800, 800), dtype=np.int32)

out = mlmodel.predict({"pixel_values": pixel_values, "pixel_mask": pixel_mask})
logits = out["logits"][0]       # (15, 3): id2label = {0: "table", 1: "table rotated"}, class 2 = no object
pred_boxes = out["pred_boxes"][0]  # (15, 4): normalized (center_x, center_y, width, height)

Input is a fixed 800x800, NCHW. The trace assumes pixel_mask is always all ones (no padding) for a single image; letterboxed/aspect-ratio-preserving batch padding has not been verified.

Accuracy

Compared against the PyTorch fp32 reference on a COCO validation image:

Metric Result
Logits cosine similarity 0.9999995
Boxes cosine similarity 1.0
Label agreement 100%

Specs

Item Value
Base model microsoft/table-transformer-detection (ResNet18 + DETR, 28.8M params)
Precision float16
Input 800x800, NCHW, fixed size
Framework Core ML (mlprogram, minimum_deployment_target=macOS14)

Notes

  • This is a community conversion, not an official release from Microsoft.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

モデルの概要

microsoft/table-transformer-detection のCore ML版です。元モデルの著作権はその作者(Microsoft)に帰属します。

変換について

DETR系の検出モデルも自己回帰デコーダーを持たない1回のフォワードパスのため、KVキャッシュ等は 不要です。torch.jit.trace + coremltoolsでモンキーパッチ無しにそのまま変換できました。

使い方(Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("table-transformer-detection_fp16.mlpackage")

image = Image.open("document.png").convert("RGB").resize((800, 800))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 800, 800), NCHW
pixel_mask = np.ones((1, 800, 800), dtype=np.int32)

out = mlmodel.predict({"pixel_values": pixel_values, "pixel_mask": pixel_mask})
logits = out["logits"][0]       # (15, 3): id2label = {0: "table", 1: "table rotated"}, class 2 = no object
pred_boxes = out["pred_boxes"][0]  # (15, 4): normalized (center_x, center_y, width, height)

入力は固定サイズ800x800、NCHW。pixel_maskは画像に対して常に全1(パディング無し)を想定した トレースです。アスペクト比を保ったままレターボックス的にパディングする使い方は未検証です。

精度検証

PyTorch fp32リファレンスと、COCO検証画像1枚で比較:

項目 結果
Logitsコサイン類似度 0.9999995
Boxesコサイン類似度 1.0
ラベル一致率 100%

Specs

Item Value
ベースモデル microsoft/table-transformer-detection(ResNet18 + DETR、28.8M params)
精度 float16
入力 800x800、NCHW、固定サイズ
フレームワーク Core ML(mlprogram、minimum_deployment_target=macOS14)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/table-transformer-detection-coreml

Quantized
(3)
this model