Intern-Decision-0.8B for Core ML
Core ML export of internlm/Intern-Decision-0.8B (Shanghai AI
Laboratory, Apache-2.0, snapshot 85a0cc5a): a typed-decision model fine-tuned from Qwen3.5-0.8B that answers a set of
choice / score / noul (yes/no) questions about a JSON state in one prefill pass. The text path is exported here;
the vision tower is not (text-only requests).
Files
| Path | What |
|---|---|
L320_F8/DecisionRow_fp16.mlpackage |
Requests up to 320 tokens and 8 fields (the model card's three-field shape is 319 tokens). |
L512_F8/DecisionRow_fp16.mlpackage |
Up to 512 tokens, 8 fields. |
L512_F8/DecisionRow_w8.mlpackage |
Same, int8 per-channel weights (480 MB; 2 of 240 answers differ from the reference). |
L1024_F16/DecisionRow_fp16.mlpackage |
Up to 1,024 tokens, 16 fields. |
*/config.json |
Bucket dimensions, marker / pad / answer-symbol ids, temperature, system prompt. |
embeddings.f16 |
Token embeddings (fp16, 248,320 ร 1,024), gathered on the host. |
tokenizer.json |
The checkpoint's tokenizer (Qwen3.5 plus the <decision> token). |
Inputs: hidden [1, L, 1024] (embedding rows of the right-padded prompt), cos / sin [L, 64] (RoPE tables for
positions 0โฆLโ1), field_onehot [F, L] (row i selects the token before field i's <decision>). Output: logits
[F, 62] over the answer symbols; take the first n entries for a field with n options, softmax, divide log-probs by
the temperature. fp16, GPU (cpuAndGPU), iOS 17 / macOS 14. Swift runtime: InternDecisionManager in
FluidUse; conversion scripts in
mobius (models/computer-use/intern-decision-0.8b/coreml).
What the package computes
Intern-Decision renders the request as one chat prompt: a fixed system prompt, a user turn with the state as JSON and
the decision schema (one line per field, its options mapped to the answer symbols AโZ, aโz, 0โ9), and an
assistant turn that is a JSON skeleton with one <decision> token per field. The logits at the position immediately
before each marker, restricted to that field's first n symbols, are its answer; the published temperature
(2.7478 for 0.8B) rescales the restricted softmax without changing the argmax. Nothing is generated.
DecisionRow (decision_export.py) is one fixed-length request: the Qwen3.5 decoder from qwen35_export.py (the
Kev-0.8B / Cua-S1-4B export), the final norm at the pre-marker positions selected with a one-hot map per field, and
the 62 tied-embedding rows of the answer symbols. Token embeddings are gathered on the host (embeddings.f16); the
host applies the restricted softmax and temperature. Prompt compilation and the chat template come from the
checkpoint's own inference.py and tokenizer, so the token stream is the reference's.
| Package | Tokens | Fields | Fits |
|---|---|---|---|
L256_F4 |
256 | 4 | Jevbench easy/original (~220 tokens), AG News (p95 299) |
L320_F8 |
320 | 8 | the model card's three-field request shape |
L384_F8, L512_F8 |
384 / 512 | 8 | WildJailBreak (p95 536), 62% of ToolACE |
L768_F16, L1024_F16 |
768 / 1,024 | 16 | Typed Decision (5 fields, p50 852, max 1,223), ToolACE (max 886) |
The fixed prompt (system prompt, headings, skeleton) is about 250 tokens, so the smallest three-field request is 319 tokens; 40% of Jevbench-Hard exceeds 1,024 tokens (max 4,074).
Fidelity
Reference: the checkpoint's DecisionEngine in fp32 on the Apple GPU (MPS), probabilities after temperature scaling,
on the bundled accuracy suites from the Intern-Decision repo
(benchmarks/accuracy-v1, shuffled with seed 0). The fp32 PyTorch wrapper matches the reference to 3.6e-6
(50 Typed Decision fields).
| Package | Suites | Decisions | Top answer differs | Max |ฮp| |
|---|---|---|---|---|
L512_F8 fp16 |
Jevbench (3), ToolACE, AG News, WildJailBreak | 240 | 0 | 0.006 |
L1024_F16 fp16 |
Typed Decision (5 fields), Jevbench-Hard, ToolACE | 280 | 0 | 0.007 |
L320_F8 fp16 |
Jevbench easy/original, AG News, WildJailBreak | 100 | 0 | 0.006 |
L256_F4 fp16 |
Jevbench easy/original, AG News, WildJailBreak | 100 | 0 | 0.006 |
L512_F8 int8 (per-channel) |
same as L512_F8 fp16 |
240 | 2 | 0.047 |
Reference and Core ML accuracy against the suite labels are identical on every subset (reports in the mobius directory). int4 per-block compression needs an iOS 18 deployment target and was not built.
Latency
One request, 319 input tokens, three fields (choice, yes/no, score), the model card's RTX 4090 shape. M5 Pro (24 GB),
warmed, bench.py; Core ML on CPU_AND_GPU, times include the host embedding gather and RoPE tables.
| Runtime | p50 | p95 |
|---|---|---|
Core ML fp16 L320_F8 |
57 ms | 59 ms |
Core ML fp16 L384_F8 |
66 ms | 67 ms |
Core ML fp16 L512_F8 |
88 ms | 96 ms |
Core ML fp16 L768_F16 |
133 ms | 145 ms |
Core ML fp16 L1024_F16 |
182 ms | 190 ms |
PyTorch MPS bf16 (checkpoint inference.py) |
150 ms | 170 ms |
| PyTorch MPS fp32 | 177 ms | 191 ms |
The pass is compute-bound and scales with the bucket, not the request, so ship the smallest bucket that fits.
int8 weights leave GPU time unchanged (88 ms at L512_F8) and halve the package (955 โ 480 MB). ComputeUnit.ALL
matches CPU_AND_GPU: the Gated DeltaNet backbone does not run on the Neural Engine (see the Kev-0.8B notes).
The model card reports 34 ms for this request on an RTX 4090.
- Downloads last month
- 49
Model tree for FluidInference/intern-decision-0.8b-coreml
Base model
Qwen/Qwen3.5-0.8B-Base