Instructions to use masahiroid/table-transformer-detection-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use masahiroid/table-transformer-detection-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir table-transformer-detection-mlx masahiroid/table-transformer-detection-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
table-transformer-detection-mlx
Model Summary
This is an unofficial MLX conversion of microsoft/table-transformer-detection (a table-region detector for document images, using a ResNet18 backbone and a DETR-style encoder-decoder). All credit for the original model goes to its authors (Microsoft).
This cannot be loaded with mlx-vlm
A ResNet18 + DETR-style encoder-decoder with learned object queries isn't in
mlx-vlm's list of supported architectures, so this was reimplemented
from scratch for MLX and requires the bundled table_transformer_mlx.py.
Limitation: fixed square input only, no padding
This implementation is simplified by assuming pixel_mask is always fully
valid (no padding). Concretely, it only works correctly when input images
are always resized to a fixed 800x800 square (not when batching images of
different aspect ratios with letterbox-style padding). This simplification
lets the implementation skip the attention-mask handling that DETR-family
models normally require.
Usage
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from table_transformer_mlx import TableTransformerMLX, IMAGE_SIZE
model = TableTransformerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("document.png").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 800, 800, 3), NHWC
logits, boxes = model(pixel_values) # logits: (1, 15, 3), boxes: (1, 15, 4) normalized cxcywh
id2label is the same as the original model: {0: "table", 1: "table rotated"}
(class 2 means "no object"). boxes are (center_x, center_y, width, height)
normalized to [0, 1].
Accuracy
Compared against the PyTorch fp32 reference on a COCO validation image:
| Precision | Logits cosine sim. | Boxes cosine sim. | Label agreement |
|---|---|---|---|
| MLX fp32 | 1.0 | 1.0 | 100% |
| MLX fp16 (this release) | 0.9999995 | 0.9999999 | 100% |
Specs
| Item | Value |
|---|---|
| Base model | microsoft/table-transformer-detection (ResNet18 + DETR, 28.8M params) |
| Precision | float16 |
| Input | 800x800, NHWC, fixed size, no padding |
| Framework | MLX (from-scratch table_transformer_mlx.py) |
Notes
- This is a community conversion, not an official release from Microsoft.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
microsoft/table-transformer-detection (ResNet18バックボーン + DETR系エンコーダー・デコーダーによる、文書画像中の表領域検出モデル)の MLX版です。元モデルの著作権はその作者(Microsoft)に帰属します。
mlx-vlmでは読み込めません
ResNet18 + DETR系エンコーダー・デコーダー + 学習済みobject queryという構成はmlx-vlmの対応
アーキテクチャ一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。
同梱のtable_transformer_mlx.pyが必要です。
制約:パディング無しの正方形固定入力のみ対応
本実装は**pixel_maskが全て有効(パディング無し)であることを前提**に簡略化しています。
具体的には、入力を常に800x800の正方形にリサイズして使う場合(バッチ処理でアスペクト比の
異なる画像を混在させ、レターボックス的にパディングする使い方はしない場合)にのみ正しく動作します。
これにより、DETR系実装で通常必要になるattention maskの処理をすべて省略できています。
使い方
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from table_transformer_mlx import TableTransformerMLX, IMAGE_SIZE
model = TableTransformerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("document.png").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 800, 800, 3), NHWC
logits, boxes = model(pixel_values) # logits: (1, 15, 3), boxes: (1, 15, 4) normalized cxcywh
id2labelは元モデルと同じ: {0: "table", 1: "table rotated"}(クラス2は「該当なし」)。
boxesは(center_x, center_y, width, height)を0〜1に正規化した値。
精度検証
PyTorch fp32リファレンスと、COCO検証画像1枚で比較:
| 精度 | Logitsコサイン類似度 | Boxesコサイン類似度 | ラベル一致率 |
|---|---|---|---|
| MLX fp32 | 1.0 | 1.0 | 100% |
| MLX fp16(本リリース) | 0.9999995 | 0.9999999 | 100% |
Specs
| Item | Value |
|---|---|
| ベースモデル | microsoft/table-transformer-detection(ResNet18 + DETR、28.8M params) |
| 精度 | float16 |
| 入力 | 800x800、NHWC、固定サイズ、パディング無し |
| フレームワーク | MLX(ゼロから実装したtable_transformer_mlx.py) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Quantized
Model tree for masahiroid/table-transformer-detection-mlx
Base model
microsoft/table-transformer-detection