Intern-Decision-0.8B, fine-tuned for Pokémon Showdown, for Core ML
internlm/Intern-Decision-0.8B (Shanghai AI Laboratory, Apache-2.0) fine-tuned to choose Pokémon Showdown battle actions, exported for Core ML. The stock model plays at chance; this one learned from its own 4B sibling, internlm/Intern-Decision-4B, over one night on a MacBook Pro. Inspired by SGLang's Qwen3.8-27B FireRed run; this is the small, local end of that idea, on the Showdown simulator rather than the game.
How it was trained
- Teacher. Intern-Decision-4B played 220
gen9randombattlebattles on a local Showdown server through poke-env against poke-env's random, max-base-power and simple-heuristics players and itself; every turn's request and probability vector was logged (6,974 decisions). - Request. A compact JSON state (both teams, HP, status, boosts, field) and one
choicequestion whose options are the legal moves and switches with their facts (type, category, power, STAB, effectiveness, accuracy, PP; switch matchups). Exactly Intern-Decision's own wire format, rendered by the checkpoint'sinference.py. - Student. LoRA r=32 on every language-model linear layer, 1.5 epochs, KL(teacher ‖ student) on the
temperature-scaled restricted softmax at the
<decision>marker, option order shuffled per sample. 200 minutes on an M5 Pro (PyTorch MPS). Merged into the weights before export. Held-out agreement with the 4B's top choice: 35% → 80%. - Not used. The 27B, any GPU server, or any game ROM.
Results (gen9 random battles, this model's side first)
| Player | vs random | vs max-base-power | vs simple heuristics |
|---|---|---|---|
| stock Intern-Decision-0.8B (Core ML), 10 battles | 2-3 | 1-9 | 1-9 |
| Intern-Decision-4B teacher (PyTorch), 60 battles | 58-2 | 45-15 | 16-44 |
| this model (Core ML), 30 battles | 9-1 (10 played) | 24-6 | 9-21 |
It matches its teacher and loses to poke-env's heuristic bot most of the time, like the teacher does. Latency on an M5 Pro GPU: 89 ms per decision in the 512-token bucket, 125 ms in the 640, 180 ms in the 1024 (typical battle requests are 430–560 tokens).
Files
| Path | What |
|---|---|
L512_F8/, L640_F8/, L1024_F16/ |
DecisionRow_w8.mlpackage (int8 weights, 480 MB) and DecisionRow_fp16.mlpackage (955 MB) + config.json per bucket (tokens × fields) |
multi/DecisionRow_w8.mlpackage |
all three buckets as one weight-shared multifunction package (488 MB, functions L512_F8 / L640_F8 / L1024_F16; needs macOS 15 / iOS 18 to select a function) |
embeddings.f16 |
token embeddings (fp16, 248,320 × 1,024), gathered on the host |
tokenizer.json |
the checkpoint's tokenizer (Qwen3.5 + <decision>) |
int8 (per-channel, weight-only) is the recommended download: measured on the Swift runtime with one bucket in use it
is about 0.85 GB in memory (374 MB process footprint + 480 MB mapped weights) against about 1.5 GB for fp16, at the
same 90 ms per decision, and it played 15-0 against poke-env's max-base-power player where fp16 played 14-1 (15
battles each). FluidUse's .showdown snapshot uses the int8 buckets.
Same inputs and outputs as FluidInference/intern-decision-0.8b-coreml:
hidden [1, L, 1024], cos / sin [L, 64], field_onehot [F, L] → logits [F, 62] over the answer symbols; softmax
over the first n symbols, log-probabilities divided by the temperature. It is a general typed-decision model that
happens to know Pokémon: any state and question set works, but the fine-tune was only evaluated on battles. Runtime:
InternDecisionManager in FluidUse; harness, training and export code in
mobius under models/computer-use/intern-decision-0.8b/showdown/.
- Downloads last month
- 49
Model tree for FluidInference/intern-decision-0.8b-showdown-coreml
Base model
Qwen/Qwen3.5-0.8B-Base