kev-ministral

One browser download that serves several task shapes on WebGPU: Kev-style decisions (calibrated probabilities over options, with a screenshot), image description and text completion / summarization, all from the same weights loaded once.

Live demo · code: ai-ecoverse/kev.js (demo/ministral/)

What it is

  • Base: mistralai/Ministral-3-3B-Base-2512 (Mistral AI, Apache-2.0): a 3.4B language model and the 0.4B Pixtral vision encoder.
  • Decision adapter: a LoRA (rank 16, all attention and MLP projections) plus a Kev pointer head, trained with the recipe of Kev on Kev's decision-v7 training suite and webrunner browser-agent data (Multimodal-Mind2Web train steps rendered as webrunner menus, and System-2-labelled game traces). The LoRA is not merged: it sits in the decoder graph as gated branches behind a lora_scale input. With lora_scale = 0 the graph is the stock base model (generation); with 1 it is the decision model.
  • Delimiters: Ministral's unused control tokens share one untrained embedding, so the five Kev delimiters use <SPECIAL_20..23,26> with trained rows (in embed_tokens).

Files (kev-ministral-3b/)

File
onnx/decoder_model_merged.onnx + .onnx_data, _1, _2 decoder, int8 weights (block 32), fp32 activations, inputs inputs_embeds, attention_mask, lora_scale, KV cache; outputs logits and hidden_states (4.1 GB)
onnx/embed_tokens.onnx + data fp16 embedding table with the trained delimiter rows (0.8 GB)
onnx/vision_encoder.onnx + data Pixtral + projector: Mistral's official ONNX of the Instruct model with Base's two projector matrices (the 220 Pixtral tensors are identical between Base and Instruct); matches PyTorch Base to 2.5e-5 (1.7 GB)
head.bin, head.json pointer head (2 × 3072→256) and its fitted temperature (2.2)
tokenizer.json tekken tokenizer with the pre-tokenizer regex fixed (fix_mistral_regex)
manifest.json file list for the loader

Total ≈ 6.6 GB.

Results

Step accuracy on webrunner's menu question for Mind2Web unseen websites (255 steps; the target is in the menu) and agreement with System 2 on three held-out game runs (87 steps):

Model Mind2Web unseen websites held-out game runs
kev-4b-vision (Qwen3.5-4B, shipped) 57.6 40.2
Kev-0.8B trained on the same webrunner data 58.0 27.6
kev-ministral (PyTorch) 69.0 46.0
kev-ministral (this ONNX bundle, Chrome WebGPU) 69.8 42.5

Browser vs PyTorch on all 342 records: token ids identical, mean max |Δp| 0.013 / 0.020, 8 argmax changes. Generation from the same session with lora_scale = 0: greedy completions and summaries are token-identical to fp32 PyTorch Base for 40 tokens; an image description 38/40.

Limitations

  • The decision adapter is trained for webrunner's menu format; general Kev questions work but were not the focus.
  • The base model is not instruction-tuned: generation continues text, it does not chat.
  • 6.6 GB download and ~9-11 GB of memory; Chrome with WebGPU.
  • Training data includes Multimodal-Mind2Web (OpenRAIL; its authors describe it as research data) and Kev's decision-v7 suite (public datasets with their own terms).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai-ecoverse/kev-ministral

Adapter
(8)
this model