SauerkrautLM-Doom-MultiVec 1.3M, Core ML

Core ML conversion of VAGOsolutions/SauerkrautLM-Doom-MultiVec-1.3M by David Golchinfar / VAGO solutions (paper, code): a 1.3M-parameter ModernBERT classifier that plays ViZDoom defend_the_center. All credit for the model goes to the original authors; this repo only changes the runtime.

Demo and runner: FluidInference/FluidUse Tools/doom/sauerkraut.

File Inputs Use
SauerkrautDoom_L1026_fp16.mlpackage input_ids, depth_ids (1, 1026) int32 Fastest; every frame is exactly 1026 tokens, so no padding or mask
SauerkrautDoom_fp32.mlpackage input_ids, attention_mask, depth_ids (1, 1100) int32 Reference; identical play to PyTorch
tokenizer.json Upstream character tokenizer

Output: logits (1, 4) over shoot, move_forward, turn_left, turn_right. Upstream picks the argmax and adds a shot when p(shoot) > 0.75 × p(top) and a turn/move runner-up above 0.15.

Results

Apple M5 Pro, ViZDoom 1.3.0, seeds 10000–10099, 4 tics per decision, 2100-tic (60 s) episodes:

Runtime Mean kills Mean survival Full 60 s ms per decision
PyTorch (upstream) 20.42 (sd 5.32) 50.5 s 31 / 100 GPU (MPS) 10.5 fp32 · 8.4 fp16; CPU 57.7 (1 thread)
Core ML fp32, GPU 20.42 (sd 5.32) 50.5 s 31 / 100 2.5
Core ML fp16 L1026, GPU 20.54 (sd 5.27) 50.6 s 35 / 100 1.2

Core ML fp32 matches PyTorch kill for kill and tic for tic on all 100 seeds. fp16 differs on 10 seeds (small numeric differences change the episode path) with the same average. Use CPU_AND_GPU; the Neural Engine runs this model at about 8 ms, slower than the GPU. On the GPU, Core ML is about 4× faster than PyTorch MPS at fp32 and about 7× at fp16. Timings are medians from Python (coremltools / PyTorch), same frame, idle machine.

Input

Per frame, 1026 tokens: [CLS], one character per cell of a 40×25 grid with newlines between rows, [SEP]. Each token also gets a depth bin (0 near to 15 far, 16 for special tokens) from the ViZDoom depth buffer, block-averaged to 40×25 and min-max normalized. Upstream's ASCII channel overflows uint8 and is @ for every cell, in training and at inference, so the model plays from the depth bins. To reproduce upstream exactly, truncate the block-averaged depth to uint8 before binning. A dependency-free implementation is observe() in the FluidUse runner.

Conversion

The Hugging Face ModernBERT mask construction does not trace through coremltools, so the forward is re-implemented with static masks and rotary tables using the original weights (max logit difference 2e-7 vs upstream in PyTorch). The L1026 variant replaces the embedding gathers with one-hot matmuls. Scripts: convert.py, convert_ane.py in the FluidUse runner directory.

Downloads last month
35
Video Preview
loading

Model tree for FluidInference/sauerkrautlm-doom-coreml

Quantized
(1)
this model

Paper for FluidInference/sauerkrautlm-doom-coreml