Request access to Myles-150M

Myles-150M is released for internal evaluation under the Myles Evaluation Licence. Tell us who you are and what you want to evaluate; we review every request.

Log in or Sign Up to review the conditions and access this model content.

Myles-150M

A 149M-parameter decision model that reads up to 1,536 tokens. Give it a message or document, a question and any list of options, and it returns a probability for every option in a single forward pass. It doesn't write text; it decides. The weights are a 151 MB int8 ONNX file that runs on an ordinary CPU in your own environment, so the text never has to leave your infrastructure.

Built by Launchables. The smaller sibling, Myles-22M, runs in a browser. A larger sibling, Myles-400M, is in private preview.

This is an updated build (8 October 2026), replacing the 3 October build. This build adds two human-written, human-labelled intent datasets, Banking77 and CLINC150 (train splits only, checked against every test set reported here), shown as choices among long lists of up to 77 intents, and trains with label smoothing 0.1. It passed the same release gate as every build before it: no frozen split more than 1 point below the previous build, the mean of the four higher, option-order flips at most 5%, and the judge benchmark average not lower. On the 506 frozen questions with three or more options, the top answer changes for 3.6% when the options are shuffled. On the independent judge benchmark (7-task average, 300 rows per task, 8-bit build) this build scores 39.6 against 37.0 for the previous build.

Frozen decisions previous build this build
public-benchmark test 79.1% 79.2%
public-benchmark out of distribution 72.0% 73.3%
product tasks 93.3% 92.9%
everyday tasks 91.6% 91.4%
mean of the four 84.0% 84.2%

This is an updated build (3 October 2026), replacing the 29 September build. The training mix no longer contains any rows generated by a hosted model (the public NLI source written by one was removed) and no agent-written routing templates. It passed the same release gate as every build before it: no frozen split more than 1 point below the previous build, the mean of the four higher, option-order flips at most 5%, and the judge benchmark average not lower. On the 506 frozen questions with three or more options, the top answer changes for 4.9% when the options are shuffled. On the independent judge benchmark (7-task average, 300 rows per task, 8-bit build) this build scores 37.0 against 30.3 for the previous build.

Frozen decisions previous build this build
public-benchmark test 77.1% 79.1%
public-benchmark out of distribution 70.7% 72.0%
product tasks 93.4% 93.3%
everyday tasks 91.8% 91.6%
mean of the four 83.3% 84.0%

This is an updated build (29 September 2026), replacing the earlier release. The same yes/no decision no longer depends on the order the options are listed in: with "yes" listed first or second, the answer flips for 1.4% of decisions instead of 33.3%. Accuracy is higher on the public benchmark and everyday tasks and slightly lower on product tasks. The weights are now an 8-bit ONNX build, which is a few points less accurate than the full-precision model behind most tables below (see "8-bit build"). Certified escalation is weaker than in the earlier release: read that section before relying on it.

Parameters 149M
Input up to 1,536 tokens: question, options and text together
Output one probability per option
File model_quantized.onnx, 151 MB, 8-bit (SmoothQuant, alpha 0.65)
Speed 196 ms median, 1.6 s at the 95th percentile, on one CPU thread (13th Gen Intel(R) Core(TM) i9-13900KF, 579 frozen yes/no decisions, measured while the machine was busy training another model; an idle machine is faster)
Cost about £1.64 per million decisions on commodity CPU at the median latency above
Build myles-150m-debias, 29 September 2026
Access gated, Myles Evaluation Licence

Access

This repository is gated. Request access above; we review every request. Approved access covers downloading and running Myles-150M on hardware you control for internal evaluation. Production or commercial use needs a separate agreement: email info[at]launchables[dot]ai.

Installation

pip install onnxruntime tokenizers numpy huggingface_hub
from huggingface_hub import snapshot_download

folder = snapshot_download("launchables/myles-150m")   # needs approved access and `hf auth login`

Quickstart

import sys
sys.path.insert(0, folder)
from myles import Myles

model = Myles(folder, threads=4)

model.decide(
    state="Your parcel is on hold. Pay the £1.99 customs fee here to release it today.",
    question="Is this message a scam, phishing or fraud attempt?",
    options=["no", "yes"],
)
# {'no': 0.1032, 'yes': 0.8968}   with options=["yes", "no"]: {'yes': 0.8982, 'no': 0.1018}

The same decide(state, question, options) call works for every Myles model. Options are free text chosen at inference time, and every call is one forward pass however many options you give. You can act only above a confidence threshold and send everything else to a person or a larger model (see Certified escalation below); fit that threshold on your own held-out data.

Architecture

A cross-encoder with option markers: the question, every option and the text are read together in one sequence, so each option can attend to the text and to the other options. That is what relational questions ("the cheapest eligible one", "abstain if none qualifies") need.

[CLS] question [SEP] [OPT] option_1 [OPT] option_2 ... [OPT] option_k [SEP] state [SEP]
  • Encoder: Ettin-encoder-150M (Johns Hopkins University, MIT licence), a long-context transformer encoder pre-trained on openly released data, fine-tuned by Launchables.
  • Decision head: a linear layer scores the hidden state at each option marker; a softmax over the markers gives the distribution.
  • Budgets: question up to 64 tokens, each option up to 16, the text fills the rest of the 1,536-token window.

Training

  • Objective: listwise cross-entropy against a target distribution. With one-hot targets it is ordinary cross-entropy; with soft targets it is a distillation objective, so gold labels, label smoothing and other target distributions share one loss.
  • Data: about 193,000 decisions: a public decision-model dataset (excluding one subset whose licensing is unverified) plus Launchables synthetic data for scams, urgency, routing, personal details, sentiment, intent and language, with test templates held out of training. Inputs are kept whole: no example is truncated.
  • Earlier release: one epoch, cosine learning-rate schedule with warm-up, a separate learning rate for the decision head, bf16, gradient clipping, and checkpoint selection on a validation set disjoint from every test set.
  • This build: the earlier weights, continued for one epoch at a low learning rate (1e-5) in which the options of every example are shuffled and the targets are lightly smoothed (label smoothing 0.1). That epoch also mixed in Launchables' generated rule-reasoning and skills data, so the accuracy changes below come from both the extra data and the shuffling. A model trained for a second epoch on the same mix without shuffling scores 85.7% on the frozen rows (table below), which shows how much of the gain is the extra training.
  • Export: ONNX with 8-bit quantisation (SmoothQuant, alpha 0.65, chosen on a validation split).

Order bias, before and after

Test: the 579 frozen yes/no decisions, each scored twice, once with ["no", "yes"] and once with ["yes", "no"], full-precision model. A model with no order bias gives the same probability both times.

Earlier Myles-150M This build
P(yes) rises when "yes" is listed first (share of answers) 73.4% 54.4%
Median rise in P(yes) +0.133 +0.002
Mean rise in P(yes) +0.145 +0.004
Decisions that flip with the order 193 of 579 (33.3%) 8 of 579 (1.4%)
Accuracy, "yes" listed second / first 87.2% / 66.7% 91.0% / 91.0%

The 8-bit build (the file in this repository) was checked with the same test on the same 579 decisions: P(yes) rises in 52.8% of answers, the median rise is +0.002, and 15 decisions (2.6%) flip with the order.

On the 506 frozen questions with three or more options, the top answer changes when the options are shuffled for 8.5% of questions in the earlier release and 4.3% in this build (reversed order: 10.3% and 6.5%). Accuracy on those questions is 76.3% before and 78.5% after. This is reduced, not removed.

Benchmarks

Cost, speed and accuracy against hosted LLMs

1,400 frozen held-out decisions (400 public-benchmark test, 400 public-benchmark out-of-distribution, 300 product tasks with unseen phrasing, 300 everyday tasks with unseen templates), identical rows for every model. The comparison models are three tiers of a leading hosted LLM family (small, medium and large), answering through a schema that only allows the offered options: the small and medium tiers without extended thinking, the large tier with brief adaptive thinking at its lowest setting. Hosted cost is measured tokens at list price; Myles cost is single-thread CPU time at $0.04 per vCPU-hour. Prices are shown in pounds, converted at £1 = $1.33. Myles accuracy is for the full-precision model; Myles cost and latency are for the 8-bit build in this repository.

Model Public test Public OOD Product tasks Everyday tasks Overall £ per 1M decisions Median latency
Random guessing 35.8% 36.9% 38.2% 42.6% 38.4%
Myles-22M, this month's build (1 CPU thread) 61.8% 66.5% 97.0% 80.3% 76.4% £0.23 27 ms
Earlier Myles-150M 74.8% 76.8% 94.7% 91.0% 84.3%
Myles-150M, this build (1 CPU thread) 82.0% 78.5% 93.3% 92.7% 86.6% £1.64 196 ms
Myles-150M, second epoch without the order fix 82.5% 78.0% 93.7% 88.7% 85.7%
Hosted LLM, small tier 77.8% 84.8% 90.7% 94.3% 86.9% £902 766 ms
Hosted LLM, medium tier 80.5% 89.0% 91.7% 97.3% 89.6% £1,883 1,394 ms
Hosted LLM, large tier (brief thinking) 95.0% 96.2% 95.0% 98.7% 96.2% £4,113 1,870 ms

Myles-150M comes within 0.3 points of the small hosted tier overall at about 1/550th of the cost. It beats the small and medium tiers on the product tasks (93.3% against 90.7% and 91.7%) but not the large tier (95.0%), and trails the small tier on the everyday tasks. The product-task families are in its training data (with phrasings held out of the test), while the hosted models see them zero-shot, so treat that row as "a model built for your task" rather than general intelligence. On the public benchmark the medium and large tiers are well ahead, and product tasks are 1.4 points lower than the earlier release. The latency of the hosted models includes the network; Myles latency is local CPU time on the machine named above while it was busy, so an idle machine is faster.

8-bit build

The weights in this repository are 8-bit. Checked on 299 held-out decisions against the full-precision model, the 8-bit build gives the same top answer on 95.0%, the largest probability difference is 0.476, and accuracy is 80.9% against 82.9% for the full-precision model. Expect the tables above, which are full-precision, to be about two points lower for this file, and probabilities to differ from the full-precision model by up to that much on a few decisions.

Certified escalation

Confidence thresholds are calibrated per task family on held-out data only (learn-then-test with an exact binomial bound, 90% confidence), then applied to the 1,400 frozen decisions with the full-precision model. Myles decides what it can certify and escalates the rest to the small hosted tier.

Target error on decisions Myles keeps Myles keeps Observed error With the small hosted tier for the rest £ per 1M decisions
1% 0% n/a 86.1% (tier alone: 86.1%) £904 (tier alone: £902)
2% 21% 3.4% 87.1% £717
5% 31% 6.2% 87.7% £626
10% 41% 10.3% 86.8% £531

The guarantee was not met on the frozen test in this build. The observed error on the decisions Myles kept is above the target at 2%, 5% and 10%. Coverage is also lower than in the earlier release, which kept 72% of decisions at the 10% target with 8.3% observed error. All of the coverage now comes from the public-benchmark family; no product-task or everyday-task decisions are kept at any target. Accuracy with escalation is still above the small hosted tier alone (86.1%) and costs 20% to 41% less, but that is a measured result on these rows, not a certified one. The thresholds were also fitted on the full-precision model and have not been re-fitted for the 8-bit file. Fit your own on held-out data before use.

Held-out public decision benchmark

About 6,000 sampled rows per split, scored with the same code for both builds, full-precision model.

Earlier Myles-150M This build
Test / out-of-distribution 74.0% / 73.9% 81.2% / 78.2%
Long inputs (test / OOD) 68.8% / 65.3% 81.8% / 74.6%
Multi-option questions (test / OOD) 67.0% / 63.3% 74.4% / 70.8%
Product tasks, unseen phrasing 94.1% 93.4%
Everyday tasks, fuller held-out set not recorded 91.8%
Calibration error (ECE), test / OOD 0.035 / not recorded 0.037 / 0.026

Independent decision-model benchmark, public tasks

An independent, community-run benchmark for typed decision models (code and tasks, MIT). Its 231 public tasks (72 original, 48 easy, 111 hard) were scored with the benchmark's own scoring code.

System Original Easy Hard Overall
Random guessing 31.1% 28.4% 33.6% 31.8%
Myles-22M, this month's build 33/72 38/48 29/111 43.3%
Earlier Myles-150M 35/72 44/48 42/111 52.4%
Myles-150M, this build 38/72 45/48 39/111 52.8%

The benchmark's own metrics for this build on these tasks: Brier 0.590 and ECE 0.130. Input format: question = instructions, options = labels, text = task state, the layout Myles trains on; adding each label's written criterion to the input gives 54.1% (125 of 231), and both results are kept. 37 of the 231 inputs exceed the 1,536-token window and are truncated. The hard tier (long multi-step policy, numeric and temporal reasoning) is where Myles is weakest, and it is 3 tasks lower than in the earlier release.

Honest limits

  • Order bias is greatly reduced, not gone. Yes/no flips fall from 33.3% to 1.4%, but questions with three or more options still change their top answer for 4.3% of shuffles (6.5% when the options are reversed). Test with your own paired-order check before relying on scores.
  • The bias fix removes the order effect; it does not fix calibration. Scores may need per-rule thresholds chosen on held-out data. The certified-escalation guarantee was not met on our frozen test (see above).
  • Rule evaluation on structured records is still weak. Keyword presence, date windows and per-employee arithmetic are not fixed by this change. On rule-choice questions from the public benchmark, splitting a question into one yes/no question per candidate raises accuracy (61.8% to 72.4% on the test split, 55.0% to 62.0% on the out-of-distribution split), at about six passes per decision.
  • Reasoning transfer to new domains has not been re-measured for this build. The earlier release scored below random guessing on generated multi-rule problems in domains it had not seen. This build trained on more of that kind of data, so an in-domain gain is not evidence that it transfers.
  • The 8-bit file is a few points less accurate than the full-precision numbers above, and 95.0% of its top answers match.
  • Product tasks are 1.4 points lower than the earlier release, and the independent benchmark's hard tier is 3 tasks lower.
  • General decision benchmarks. 52.8% on the independent benchmark's public tasks (random guessing: 31.8%). Multi-billion-parameter decision models score far higher there.
  • Long inputs are slower: about 1.6 s at the 95th percentile on one thread while the machine was busy, at the full 1,536 tokens on one CPU thread; use more threads or shorter inputs where latency matters.
  • Base-model lineage. The base encoder's published training data includes a small share (under 1% of its tokens) of math text rewritten by a Chinese-developed model. This release is therefore not suitable where a fully non-Chinese model lineage is required. A rebuild on an encoder whose training data we have audited end to end is in progress and will replace this one.
  • Numbers on this card are from our frozen rows. A customer's lead-scoring rubric, used to test option-order bias, is held out and is not in our training data.
  • English only. Not for high-stakes decisions about people (hiring, credit, health, legal) without human review.

Areas of improvement

  1. Clean base: retrain on an encoder whose pre-training data we have audited end to end.
  2. Rule evaluation: generated rule problems with exact labels and per-rule intermediate steps, with whole domains held out to prove transfer.
  3. Order bias for many options: more many-option shuffled data and a paired-order test in every build.
  4. Tighter guarantees: re-fitted calibration per task family, aimed again at certifying 5% and then 2% error.
  5. Scale: the larger Myles-400M tier on the same recipe.

Licence

Myles-150M is proprietary to Launchables and released under the Myles Evaluation Licence: internal evaluation only, no redistribution, no production use, no training other models without written permission. Third-party components keep their own licences; see THIRD_PARTY_NOTICES.md.

Citation

@misc{launchables2026myles150m,
  title  = {Myles-150M: a calibrated long-context decision model},
  author = {Launchables},
  year   = {2026},
  url    = {https://huggingface.co/launchables/myles-150m}
}

Copyright © 2026 Launchables. Named after Myles.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for launchables/myles-150m

Finetuned
(32)
this model