Myles-22M
A 22.7M-parameter decision model, and the first of the Myles family. Give it a message, a question and any list of options, and it returns a probability for every option in a single forward pass. It doesn't write text; it decides. The weights are a 23 MB int8 ONNX file that runs on one CPU thread in about 27 ms, in a phone browser, or on your own servers, so the text never has to leave your infrastructure.
Built by Launchables. Try it live in your browser: launchables.ai → "Try Myles-22M live".
This is an updated build (29 September 2026). The same decision now barely depends on the order the options are listed in, for yes/no questions. Accuracy is level with the previous build overall, with gains on the public benchmark and product tasks and losses on everyday tasks and the independent benchmark. The exact before-and-after numbers are below.
Myles-150M is also available. It reads 1,536-token inputs (so no example is truncated), trains on the full training set, and scores 86.6% on the same 1,400 frozen decisions used below.
| Parameters | 22.7M (encoder 22.7M, decision head 385) |
| File | model_quantized.onnx, 23 MB, per-channel int8 |
| Input | up to 320 tokens: question, options and state together |
| Output | one calibrated probability per option |
| Speed | 27 ms median, 43 ms at the 95th percentile, on one CPU thread (13th Gen Intel(R) Core(TM) i9-13900KF, 579 frozen yes/no decisions, measured while the machine was busy with other jobs) |
| Cost | about £0.23 per million decisions on commodity CPU at the median latency |
| Build | myles-22m-debias, 29 September 2026 |
| Access | gated, Myles Evaluation Licence |
Access
This repository is gated. Request access above; we review every request. Approved access covers downloading and running Myles-22M on hardware you control for internal evaluation. Production or commercial use needs a separate agreement: email info[at]launchables[dot]ai.
Installation
pip install onnxruntime tokenizers huggingface_hub
from huggingface_hub import snapshot_download
folder = snapshot_download("launchables/myles-22m") # needs approved access and `hf auth login`
Quickstart
import sys
sys.path.insert(0, folder)
from myles import Myles
model = Myles(folder)
model.decide(
state="Your parcel is on hold. Pay the £1.99 customs fee here to release it today.",
question="Is this message a scam, phishing or fraud attempt?",
options=["no", "yes"],
)
# {'no': 0.1202, 'yes': 0.8798}
model.decide(
state="I've been charged twice for my subscription this month, can you refund one?",
question="Which team should handle this message?",
options=["Customer support", "Billing and finance", "Security", "Sales"],
)
Options are free text and chosen at inference time: yes/no, a list of teams, intents, severity levels. Up to about eight short options fit comfortably. Each call is one forward pass, however many options you give.
Asking several questions about one message
message = "Checkout has been down for 20 minutes and customers can't pay, please call me asap"
for question, options in [
("Does this message need an urgent response?", ["no", "yes"]),
("Which team should handle this message?", ["Customer support", "Billing and finance", "Security", "Sales"]),
("Is the customer happy?", ["no", "yes"]),
]:
print(question, model.decide(message, question, options))
Thresholds instead of argmax
The probabilities are calibrated (ECE 0.022 on the held-out public test set, 0.117 on the independent benchmark's public tasks), so you can act only when the model is sure and send the rest to a person or a larger model:
probabilities = model.decide(message, "Does this message need an urgent response?", ["no", "yes"])
decision = max(probabilities, key=probabilities.get)
if probabilities[decision] < 0.85:
decision = "escalate"
In the browser
The same file runs client-side with ONNX Runtime Web and a JavaScript WordPiece tokenizer with byte-for-byte parity to the Python one. That is how the live demo on launchables.ai works: the model downloads once and nothing typed is sent anywhere. Use the single-threaded wasm backend (numThreads = 1); multi-threaded wasm needs cross-origin isolation headers and fails on iPhone Safari.
Architecture
A cross-encoder with option markers. Question, options and state go into one sequence, so every option can attend to the state and to every other option, which is what relational questions ("the cheapest eligible one", "abstain if none qualifies") need:
[CLS] question [SEP] [OPT] option_1 [OPT] option_2 ... [OPT] option_k [SEP] state [SEP]
- Encoder: MiniLM-L6 (6 layers, hidden size 384) from
cross-encoder/ms-marco-MiniLM-L6-v2, whose relevance-ranking pre-training is already "score this candidate against this context". - Option marker: the existing vocabulary token
[unused0], so no embedding changes. - Decision head: a linear layer (384 → 1) scores the hidden state at each marker; a softmax over the row's markers gives the distribution.
- Budgets: question up to 64 tokens, each option up to 16, the state fills the rest of the 320-token window.
Training
- Objective: listwise cross-entropy against a target distribution,
L = −Σ tᵢ log pᵢ. With one-hot targets this is ordinary cross-entropy; with soft targets it is exactly a distillation objective, so gold labels, label smoothing and teacher distributions share one loss. - Data: a public decision-model dataset (145,639 training rows, excluding one source whose licensing is unverified) plus Launchables synthetic data for scams, urgency, routing, personal details, sentiment, intent and language. Synthetic test sets use templates and phrasings held out of training.
- Recipe for this build: the previous Myles-22M weights, continued for 1 epoch at lr 2e-5 with the options of every example shuffled and label smoothing 0.1, with our generated skills data added to the mix. AdamW, linear decay with warm-up, 320-token inputs, batch of 8,192 tokens. The previous build's card said this shuffling had already been done; when we measured the previous weights they still flipped 26.1% of yes/no decisions with the option order (table below), so this card states only what we measured.
- Export: ONNX opset 17 with dynamic int8 per-channel quantisation, checked against the full-precision model on the 1,400 frozen decisions: the same top answer on 98.4%, largest probability difference 0.215, accuracy 74.4% against 74.6% (over all rows pooled).
Order bias, before and after
Test: the 579 frozen yes/no decisions, each scored twice, once with ["no", "yes"] and once with ["yes", "no"], full-precision model. A model with no order bias gives the same probability both times.
| Previous Myles-22M | This build | |
|---|---|---|
| P(yes) rises when "yes" is listed first (share of answers) | 73.6% | 54.2% |
| Median rise in P(yes) | +0.084 | +0.001 |
| Mean rise in P(yes) | +0.092 | +0.004 |
| Decisions that flip with the order | 151 of 579 (26.1%) | 17 of 579 (2.9%) |
| Accuracy, "yes" listed second / first | 84.1% / 68.4% | 83.4% / 84.3% |
The 8-bit build (the file in this repository) was checked with the same test on the same 579 decisions: P(yes) rises in 52.7% of answers, the median rise is +0.001, and 21 decisions (3.6%) flip with the order.
The order effect is not fixed for questions with three or more options. On the 506 frozen many-option questions, the top answer changes when the options are shuffled for 20.9% of questions in the previous build and 21.5% in this one (reversed order: 24.5% and 26.7%). Accuracy on those questions is 58.3% before and 58.9% after. For questions with many options, ask one yes/no question per option, or use Myles-150M.
Benchmarks
Cost, speed and accuracy against hosted LLMs
1,400 frozen held-out decisions (400 public-benchmark test, 400 public-benchmark out-of-distribution, 300 product tasks with unseen phrasing, 300 everyday tasks with unseen templates), identical rows for every model. The comparison models are three tiers of a leading hosted LLM family (small, medium and large), answering through a schema that only allows the offered options; the small and medium tiers without extended thinking, the large tier with brief adaptive thinking at its lowest setting (it cannot be switched off). Hosted cost is measured tokens at list price; Myles cost is single-thread CPU time at $0.04 per vCPU-hour. Prices are shown in pounds, converted at £1 = $1.33.
| Model | Public test | Public OOD | Product tasks | Everyday tasks | Overall | £ per 1M decisions | Median latency |
|---|---|---|---|---|---|---|---|
| Random guessing | 35.8% | 36.9% | 38.2% | 42.6% | 38.4% | ||
| Previous Myles-22M (1 CPU thread) | 60.0% | 65.2% | 96.0% | 85.3% | 76.6% | ||
| Myles-22M, this build (1 CPU thread) | 61.8% | 66.5% | 97.0% | 80.3% | 76.4% | £0.23 | 27 ms |
| Hosted LLM, small tier | 77.8% | 84.8% | 90.7% | 94.3% | 86.9% | £902 | 766 ms |
| Hosted LLM, medium tier | 80.5% | 89.0% | 91.7% | 97.3% | 89.6% | £1,883 | 1,394 ms |
| Hosted LLM, large tier (brief thinking) | 95.0% | 96.2% | 95.0% | 98.7% | 96.2% | £4,113 | 1,870 ms |
Overall accuracy is level with the previous build (76.4% against 76.6%): the public sets are up 1.8 and 1.3 points and product tasks 1.0, and everyday tasks are down 5.0. Myles-22M is trained on the product-task families (with phrasings held out of the test), while the hosted models see them zero-shot. That comparison is fair for "a model built for your task", not a claim of general intelligence; on general decision benchmarks the large models are far ahead.
Independent decision-model benchmark, public tasks
An independent, community-run benchmark for typed decision models (code and tasks, MIT). Its 231 public tasks (72 original, 48 easy, 111 hard) were scored with the benchmark's own scoring code.
| System | Original | Easy | Hard | Overall |
|---|---|---|---|---|
| Random guessing | 31.1% | 28.4% | 33.6% | 31.8% |
| Previous Myles-22M | 32/72 | 41/48 | 38/111 | 48.1% |
| Myles-22M, this build | 33/72 | 38/48 | 29/111 | 43.3% |
The benchmark's own metrics for this build on these tasks: Brier 0.642, ECE 0.117. Input format: question = instructions, options = labels, state = task state (the layout Myles trains on); adding each label's written criterion to the input gives 33.3% (previous build: 32.9%), and both results are kept. 72 of 231 inputs exceed the 320-token window and are truncated. Myles-150M scores 52.8% on the same tasks.
Held-out public benchmark and product tasks
About 6,000 sampled rows per split, scored with the same code for both builds.
| Test | Previous Myles-22M | This build |
|---|---|---|
| Public benchmark, test | 62.5% | 62.4% |
| Public benchmark, out-of-distribution | 63.8% | 64.4% |
| Long inputs (test / OOD) | 49.3% / 49.5% | 49.9% / 50.1% |
| Multi-option questions (test / OOD) | 45.4% / 42.0% | 46.4% / 42.6% |
| Product tasks, unseen phrasing | 97.3% | 97.3% |
| Calibration error (ECE), public test / OOD | 0.034 / 0.031 | 0.022 / 0.026 |
Certified escalation
Using thresholds calibrated on held-out data only (learn-then-test with an exact binomial bound, 90% confidence), Myles-22M certifies an error rate of at most 5% on the decisions it keeps. On the 1,400 frozen decisions it keeps 28% of them with 2.3% observed error and escalates the rest (previous build: 24% kept, 3.6% observed error). Accuracy here is over all 1,400 rows, whereas the first table averages its four sets:
| System | Accuracy | £ per 1M decisions |
|---|---|---|
| Hosted LLM, small tier, alone | 86.1% | £902 |
| Myles-22M, escalating to the small tier | 87.9% | £650 |
| Hosted LLM, large tier, alone | 96.1% | £4,113 |
| Myles-22M, escalating to the large tier | 96.6% | £2,962 |
Certified coverage comes only from the product tasks and the public benchmark; no everyday-task decisions are kept at a 5% target. At a 2% target Myles-22M keeps 2% of decisions.
Honest limits
- Order bias is fixed for yes/no questions only. Many-option questions still change their top answer for about one in five questions when the options are shuffled. See the section above.
- Some results are worse than the previous build. Everyday tasks fall from 85.3% to 80.3% (on the fuller held-out set, 86.1% to 80.1%), the independent benchmark falls from 111 to 100 of 231 correct (48.1% to 43.3%), and overall accuracy on the frozen rows is level (76.6% to 76.4%). Use the previous build's numbers as the yardstick if everyday sentiment and intent routing is your case, and test on your own data.
- Calibration. ECE is 0.022 on the public test set, but 0.126 on product tasks and 0.117 on the independent benchmark. Fit thresholds per task family, as in the certified-escalation method above.
- Short context. 320 tokens. Long documents are truncated, and accuracy drops with them.
- Multi-constraint reasoning. Questions such as "which of these seven candidates meets all four rules, applying the tie-break" are beyond it: on those benchmark rows it defaults to "abstain" and scores no better than that base rate. Break them into one yes/no question per candidate, or escalate.
- General decision benchmarks. 43.3% on the independent benchmark's public tasks (random guessing: 31.8%). Multi-billion-parameter decision models score far higher there.
- Occasional confident errors, most often on everyday sentiment and intent phrasing it has not seen; tight error guarantees (1 to 2%) cannot be certified.
- English only, and sensitive to input layout: keep to question, short options and the text being judged.
- Not for high-stakes decisions about people (hiring, credit, health, legal) without human review.
Areas of improvement
Every shortfall above has a specific fix; the first two shipped in Myles-150M:
- Read the whole input. Myles-150M reads 1,536 tokens, above the longest training example, and our data preparation no longer cuts long records short. This targets the long-document and truncation failures directly.
- More capacity and all the data. Myles-150M has about 7× the parameters and trains on the full dataset, not a 40k subsample.
- Order bias for many options. Shuffle-and-smooth training removed the yes/no effect; the many-option effect needs more many-option training data and a held-out many-option test in every build.
- Multi-constraint reasoning. An exact rule solver (100% agreement with 46,080 labelled rule decisions) now generates unlimited practice problems with correct answers and per-rule intermediate labels, plus a difficulty ladder that shows exactly where a model breaks. Complex choices can also be split into one yes/no question per candidate, with the selection rule applied in code.
- Distillation. A large open-weight model run locally labels large volumes of domain decisions with full probability distributions, which our training objective consumes directly. No data leaves our hardware.
- Tighter guarantees. Larger calibration sets per task family, aimed at certifying 1 to 2% error rates on the decisions Myles keeps.
- More languages after English quality is where we want it.
What's next
Myles-150M is available now, with the same benchmarks, the same frozen test rows and the same scoring code, so the two can be compared directly. A larger Myles-400M is in private preview.
Licence
Myles-22M is proprietary to Launchables and released under the Myles Evaluation Licence: internal evaluation only, no redistribution, no production use, no training other models without written permission.
Third-party components keep their own licences: cross-encoder/ms-marco-MiniLM-L6-v2 (Microsoft MiniLM), Apache License 2.0 (full text), modified by Launchables with a decision head and further training; WANLI (Liu, Swayamdipta, Smith and Choi, 2022), CC BY 4.0; and a public decision-model dataset released under CC0.
Citation
@misc{launchables2026myles22m,
title = {Myles-22M: a tiny calibrated decision model},
author = {Launchables},
year = {2026},
url = {https://huggingface.co/launchables/myles-22m}
}
Copyright © 2026 Launchables. Named after Myles.
Model tree for launchables/myles-22m
Base model
microsoft/MiniLM-L12-H384-uncased