Instructions to use agentmessier/messier-one with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use agentmessier/messier-one with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="agentmessier/messier-one")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("agentmessier/messier-one") model = AutoModelForMultimodalLM.from_pretrained("agentmessier/messier-one", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Messier One
A decision model: it reads a state and typed questions (choice with 2-255 options, score with 2-10 levels, noul)
and returns a probability for every option from one forward pass. It does not generate text. The wire format is
TypeSafe-compatible (POST /v1/systemone); the serving code is at https://github.com/agentmessier-ai/messier-one.
What it is for
- One forward pass, no generated text. Every answer is a probability for each option, read at one position. On one RTX 4090 a decision takes about 47 ms (short documents) to 88 ms (long ones), server on loopback.
- Up to 255 options in one question. Routing to one of many tools, picking a category from a long list, choosing a move
among many: on our own many-option test it is right 97.4% of the time with 2-26 options, 92.0% with 27-100 and 95.3% with
101-255 (read with
--gate 0.5). - It reads the goal you give it. The same document judged under a different goal gets a different answer. On our own held-out test (166 project descriptions, the same yes/no question asked under the original goal, the opposite goal and a goal never seen in training) it agrees with the reference answers 84-89% of the time under all three.
- Probabilities you can use as they are. They come from the model's own distribution with one fitted temperature per
question type; nothing is pushed toward 0 or 1 afterwards.
confidencefollows the TypeSafe definitions. - A drop-in endpoint.
POST /v1/systemonewithchoice,scoreandnoulquestions, the TypeSafe wire format. - Small. About 10 GB of BF16 weights; one 24 GB GPU is enough.
The numbers in this section are our own measurements on our own test sets, whose reference answers were produced by larger language models, not by people. The benchmark results below are measured with the benchmark's own runner.
What is in this repository
- Full merged BF16 weights of
Qwen/Qwen3.5-4Bwith a LoRA (r=16) fine-tune merged in. The readout is baked into the language-model head: the rows of the option labels (A..Zand the added tokens<o27>..<o255>) hold the trained head. temps.json: calibration temperatures per question type.sys1_merge.json: the label token ids.- Tokenizer of the base model plus the 229 label tokens.
Results
Public JevBench items (231), measured by us with JevBench's own runner and its unchanged typesafe adapter (serial requests,
single read) against the serving code above on one RTX 4090 (vLLM 0.30.0, BF16), server on loopback:
| tier | correct | accuracy | ECE | latency p50 / p95 |
|---|---|---|---|---|
| easy | 48 / 48 | 100.0% | 0.002 | 47 ms / 51 ms |
standard + judge (original) |
72 / 72 | 100.0% | 0.046 | 47 ms / 49 ms |
| hard | 82 / 111 | 73.9% | 0.119 | 88 ms / 221 ms |
| all | 202 / 231 | 87.4% |
These are our own measurements on the public items, not leaderboard results.
Part of the training material was written for this model in the families of JevBench's hard tier (long policies, multi-step lookups, dates and numbers, traps and so on). No JevBench item, public or held out, was used for training, and every training question was scanned against the public items before use.
Versions
v0.2(2026-10-07): this revision. 202 / 231 on the public JevBench items.v0.1(2026-10-06): 197 / 231. Still available as revisionv0.1.
Tested with vLLM 0.30.0 and transformers 5.18.0 (pinned in the serving code's requirements.txt).
Licence
The weights are released under CC BY-NC 4.0 (see LICENSE): free for research and
evaluation, no commercial use. For a commercial licence write to agentmessier.ai@gmail.com.
The base model Qwen/Qwen3.5-4B is Apache-2.0; its licence text is kept as LICENSE-Qwen-Apache-2.0. This repository is a
modified version of it (fine-tuned weights, 229 added label tokens).
Limits
- Confidence can be off on unfamiliar domains.
- Option order has a small effect on the probabilities (the
--gateoption of the front end averages two orders when unsure). - Arithmetic, date arithmetic and multi-step computation are unreliable: compute in code and put the result in the state.
- Downloads last month
- 17