Messier One

A decision model: it reads a state and typed questions (choice with 2-255 options, score with 2-10 levels, noul) and returns a probability for every option from one forward pass. It does not generate text. The wire format is TypeSafe-compatible (POST /v1/systemone); the serving code is at https://github.com/agentmessier-ai/messier-one.

What it is for

  • One forward pass, no generated text. Every answer is a probability for each option, read at one position. On one RTX 4090 a decision takes about 47 ms (short documents) to 88 ms (long ones), server on loopback.
  • Up to 255 options in one question. Routing to one of many tools, picking a category from a long list, choosing a move among many: on our own many-option test it is right 97.4% of the time with 2-26 options, 92.0% with 27-100 and 95.3% with 101-255 (read with --gate 0.5).
  • It reads the goal you give it. The same document judged under a different goal gets a different answer. On our own held-out test (166 project descriptions, the same yes/no question asked under the original goal, the opposite goal and a goal never seen in training) it agrees with the reference answers 84-89% of the time under all three.
  • Probabilities you can use as they are. They come from the model's own distribution with one fitted temperature per question type; nothing is pushed toward 0 or 1 afterwards. confidence follows the TypeSafe definitions.
  • A drop-in endpoint. POST /v1/systemone with choice, score and noul questions, the TypeSafe wire format.
  • Small. About 10 GB of BF16 weights; one 24 GB GPU is enough.

The numbers in this section are our own measurements on our own test sets, whose reference answers were produced by larger language models, not by people. The benchmark results below are measured with the benchmark's own runner.

What is in this repository

  • Full merged BF16 weights of Qwen/Qwen3.5-4B with a LoRA (r=16) fine-tune merged in. The readout is baked into the language-model head: the rows of the option labels ( A.. Z and the added tokens <o27>..<o255>) hold the trained head.
  • temps.json: calibration temperatures per question type.
  • sys1_merge.json: the label token ids.
  • Tokenizer of the base model plus the 229 label tokens.

Results

Public JevBench items (231), measured by us with JevBench's own runner and its unchanged typesafe adapter (serial requests, single read) against the serving code above on one RTX 4090 (vLLM 0.30.0, BF16), server on loopback:

tier correct accuracy ECE latency p50 / p95
easy 48 / 48 100.0% 0.002 47 ms / 51 ms
standard + judge (original) 72 / 72 100.0% 0.046 47 ms / 49 ms
hard 82 / 111 73.9% 0.119 88 ms / 221 ms
all 202 / 231 87.4%

These are our own measurements on the public items, not leaderboard results.

Part of the training material was written for this model in the families of JevBench's hard tier (long policies, multi-step lookups, dates and numbers, traps and so on). No JevBench item, public or held out, was used for training, and every training question was scanned against the public items before use.

Versions

  • v0.2 (2026-10-07): this revision. 202 / 231 on the public JevBench items.
  • v0.1 (2026-10-06): 197 / 231. Still available as revision v0.1.

Tested with vLLM 0.30.0 and transformers 5.18.0 (pinned in the serving code's requirements.txt).

Licence

The weights are released under CC BY-NC 4.0 (see LICENSE): free for research and evaluation, no commercial use. For a commercial licence write to agentmessier.ai@gmail.com. The base model Qwen/Qwen3.5-4B is Apache-2.0; its licence text is kept as LICENSE-Qwen-Apache-2.0. This repository is a modified version of it (fine-tuned weights, 229 added label tokens).

Limits

  • Confidence can be off on unfamiliar domains.
  • Option order has a small effect on the probabilities (the --gate option of the front end averages two orders when unsure).
  • Arithmetic, date arithmetic and multi-step computation are unreliable: compute in code and put the result in the state.
Downloads last month
17
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agentmessier/messier-one

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(895)
this model
Quantizations
1 model