kev-ministral
One browser download that serves several task shapes on WebGPU: Kev-style decisions (calibrated probabilities over options, with a screenshot), image description and text completion / summarization, all from the same weights loaded once.
Live demo · code: ai-ecoverse/kev.js (demo/ministral/)
What it is
- Base: mistralai/Ministral-3-3B-Base-2512 (Mistral AI, Apache-2.0): a 3.4B language model and the 0.4B Pixtral vision encoder.
- Decision adapter: a LoRA (rank 16, all attention and MLP projections) plus a Kev pointer head, trained with the
recipe of Kev on Kev's
decision-v7training suite and webrunner browser-agent data (Multimodal-Mind2Web train steps rendered as webrunner menus, and System-2-labelled game traces). The LoRA is not merged: it sits in the decoder graph as gated branches behind alora_scaleinput. Withlora_scale = 0the graph is the stock base model (generation); with1it is the decision model. - Delimiters: Ministral's unused control tokens share one untrained embedding, so the five Kev delimiters use
<SPECIAL_20..23,26>with trained rows (inembed_tokens).
Files (kev-ministral-3b/)
| File | |
|---|---|
onnx/decoder_model_merged.onnx + .onnx_data, _1, _2 |
decoder, int8 weights (block 32), fp32 activations, inputs inputs_embeds, attention_mask, lora_scale, KV cache; outputs logits and hidden_states (4.1 GB) |
onnx/embed_tokens.onnx + data |
fp16 embedding table with the trained delimiter rows (0.8 GB) |
onnx/vision_encoder.onnx + data |
Pixtral + projector: Mistral's official ONNX of the Instruct model with Base's two projector matrices (the 220 Pixtral tensors are identical between Base and Instruct); matches PyTorch Base to 2.5e-5 (1.7 GB) |
head.bin, head.json |
pointer head (2 × 3072→256) and its fitted temperature (2.2) |
tokenizer.json |
tekken tokenizer with the pre-tokenizer regex fixed (fix_mistral_regex) |
manifest.json |
file list for the loader |
Total ≈ 6.6 GB.
Results
Step accuracy on webrunner's menu question for Mind2Web unseen websites (255 steps; the target is in the menu) and agreement with System 2 on three held-out game runs (87 steps):
| Model | Mind2Web unseen websites | held-out game runs |
|---|---|---|
| kev-4b-vision (Qwen3.5-4B, shipped) | 57.6 | 40.2 |
| Kev-0.8B trained on the same webrunner data | 58.0 | 27.6 |
| kev-ministral (PyTorch) | 69.0 | 46.0 |
| kev-ministral (this ONNX bundle, Chrome WebGPU) | 69.8 | 42.5 |
Browser vs PyTorch on all 342 records: token ids identical, mean max |Δp| 0.013 / 0.020, 8 argmax changes. Generation
from the same session with lora_scale = 0: greedy completions and summaries are token-identical to fp32 PyTorch Base
for 40 tokens; an image description 38/40.
Limitations
- The decision adapter is trained for webrunner's menu format; general Kev questions work but were not the focus.
- The base model is not instruction-tuned: generation continues text, it does not chat.
- 6.6 GB download and ~9-11 GB of memory; Chrome with WebGPU.
- Training data includes Multimodal-Mind2Web (OpenRAIL; its authors describe it as research data) and Kev's decision-v7 suite (public datasets with their own terms).
Model tree for ai-ecoverse/kev-ministral
Base model
mistralai/Ministral-3-3B-Base-2512