Intern-Decision-0.8B, fine-tuned for Pokémon Showdown, for Core ML

internlm/Intern-Decision-0.8B (Shanghai AI Laboratory, Apache-2.0) fine-tuned to choose Pokémon Showdown battle actions, exported for Core ML. The stock model plays at chance; this one learned from its own 4B sibling, internlm/Intern-Decision-4B, over one night on a MacBook Pro. Inspired by SGLang's Qwen3.8-27B FireRed run; this is the small, local end of that idea, on the Showdown simulator rather than the game.

How it was trained

  • Teacher. Intern-Decision-4B played 220 gen9randombattle battles on a local Showdown server through poke-env against poke-env's random, max-base-power and simple-heuristics players and itself; every turn's request and probability vector was logged (6,974 decisions).
  • Request. A compact JSON state (both teams, HP, status, boosts, field) and one choice question whose options are the legal moves and switches with their facts (type, category, power, STAB, effectiveness, accuracy, PP; switch matchups). Exactly Intern-Decision's own wire format, rendered by the checkpoint's inference.py.
  • Student. LoRA r=32 on every language-model linear layer, 1.5 epochs, KL(teacher ‖ student) on the temperature-scaled restricted softmax at the <decision> marker, option order shuffled per sample. 200 minutes on an M5 Pro (PyTorch MPS). Merged into the weights before export. Held-out agreement with the 4B's top choice: 35% → 80%.
  • Not used. The 27B, any GPU server, or any game ROM.

Results (gen9 random battles, this model's side first)

Player vs random vs max-base-power vs simple heuristics
stock Intern-Decision-0.8B (Core ML), 10 battles 2-3 1-9 1-9
Intern-Decision-4B teacher (PyTorch), 60 battles 58-2 45-15 16-44
this model (Core ML), 30 battles 9-1 (10 played) 24-6 9-21

It matches its teacher and loses to poke-env's heuristic bot most of the time, like the teacher does. Latency on an M5 Pro GPU: 89 ms per decision in the 512-token bucket, 125 ms in the 640, 180 ms in the 1024 (typical battle requests are 430–560 tokens).

Files

Path What
L512_F8/, L640_F8/, L1024_F16/ DecisionRow_w8.mlpackage (int8 weights, 480 MB) and DecisionRow_fp16.mlpackage (955 MB) + config.json per bucket (tokens × fields)
multi/DecisionRow_w8.mlpackage all three buckets as one weight-shared multifunction package (488 MB, functions L512_F8 / L640_F8 / L1024_F16; needs macOS 15 / iOS 18 to select a function)
embeddings.f16 token embeddings (fp16, 248,320 × 1,024), gathered on the host
tokenizer.json the checkpoint's tokenizer (Qwen3.5 + <decision>)

int8 (per-channel, weight-only) is the recommended download: measured on the Swift runtime with one bucket in use it is about 0.85 GB in memory (374 MB process footprint + 480 MB mapped weights) against about 1.5 GB for fp16, at the same 90 ms per decision, and it played 15-0 against poke-env's max-base-power player where fp16 played 14-1 (15 battles each). FluidUse's .showdown snapshot uses the int8 buckets.

Same inputs and outputs as FluidInference/intern-decision-0.8b-coreml: hidden [1, L, 1024], cos / sin [L, 64], field_onehot [F, L] → logits [F, 62] over the answer symbols; softmax over the first n symbols, log-probabilities divided by the temperature. It is a general typed-decision model that happens to know Pokémon: any state and question set works, but the fine-tune was only evaluated on battles. Runtime: InternDecisionManager in FluidUse; harness, training and export code in mobius under models/computer-use/intern-decision-0.8b/showdown/.

Downloads last month
49
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/intern-decision-0.8b-showdown-coreml

Quantized
(3)
this model