SauerkrautLM-Doom-MultiVec 1.3M, Core ML
Core ML conversion of VAGOsolutions/SauerkrautLM-Doom-MultiVec-1.3M
by David Golchinfar / VAGO solutions (paper,
code): a 1.3M-parameter ModernBERT
classifier that plays ViZDoom defend_the_center. All credit for the model goes to the original
authors; this repo only changes the runtime.
Demo and runner: FluidInference/FluidUse Tools/doom/sauerkraut.
| File | Inputs | Use |
|---|---|---|
SauerkrautDoom_L1026_fp16.mlpackage |
input_ids, depth_ids (1, 1026) int32 |
Fastest; every frame is exactly 1026 tokens, so no padding or mask |
SauerkrautDoom_fp32.mlpackage |
input_ids, attention_mask, depth_ids (1, 1100) int32 |
Reference; identical play to PyTorch |
tokenizer.json |
Upstream character tokenizer |
Output: logits (1, 4) over shoot, move_forward, turn_left, turn_right. Upstream picks the
argmax and adds a shot when p(shoot) > 0.75 × p(top) and a turn/move runner-up above 0.15.
Results
Apple M5 Pro, ViZDoom 1.3.0, seeds 10000–10099, 4 tics per decision, 2100-tic (60 s) episodes:
| Runtime | Mean kills | Mean survival | Full 60 s | ms per decision |
|---|---|---|---|---|
| PyTorch (upstream) | 20.42 (sd 5.32) | 50.5 s | 31 / 100 | GPU (MPS) 10.5 fp32 · 8.4 fp16; CPU 57.7 (1 thread) |
| Core ML fp32, GPU | 20.42 (sd 5.32) | 50.5 s | 31 / 100 | 2.5 |
| Core ML fp16 L1026, GPU | 20.54 (sd 5.27) | 50.6 s | 35 / 100 | 1.2 |
Core ML fp32 matches PyTorch kill for kill and tic for tic on all 100 seeds. fp16 differs on 10 seeds
(small numeric differences change the episode path) with the same average. Use CPU_AND_GPU; the
Neural Engine runs this model at about 8 ms, slower than the GPU. On the GPU, Core ML is about 4×
faster than PyTorch MPS at fp32 and about 7× at fp16. Timings are medians from Python (coremltools /
PyTorch), same frame, idle machine.
Input
Per frame, 1026 tokens: [CLS], one character per cell of a 40×25 grid with newlines between rows,
[SEP]. Each token also gets a depth bin (0 near to 15 far, 16 for special tokens) from the ViZDoom
depth buffer, block-averaged to 40×25 and min-max normalized. Upstream's ASCII channel overflows
uint8 and is @ for every cell, in training and at inference, so the model plays from the depth
bins. To reproduce upstream exactly, truncate the block-averaged depth to uint8 before binning. A
dependency-free implementation is observe() in the FluidUse runner.
Conversion
The Hugging Face ModernBERT mask construction does not trace through coremltools, so the forward is
re-implemented with static masks and rotary tables using the original weights (max logit difference
2e-7 vs upstream in PyTorch). The L1026 variant replaces the embedding gathers with one-hot matmuls.
Scripts: convert.py, convert_ane.py in the FluidUse runner directory.
- Downloads last month
- 35