Fort-1 0.8B

Fort picks one of your options for a text and tells you how sure it is. It never writes free text: every answer is one of your options, with a probability you can act on: take the sure answers, send the rest to a person or a larger model.

Fort-1 0.8B is the small model, for speed. Fort-1 2B is the larger model, for accuracy.

Parameters 0.75B, fine-tuned from Qwen3.5-0.8B
Download 1.5 GB (bf16); 4-bit MLX: 0.44 GB
Memory while deciding (MLX) 1.8 GB (bf16); 0.7 GB (4-bit)
Context window 262,144 tokens
Speed 28 to 51 ms per decision (jevbench medians, Apple M1 Max)
License Apache-2.0

Use

pip install teximal
teximal run fort-1-0.8b "I was charged twice this month." --options billing,technical,sales
teximal serve fort-1-0.8b      # Jev's request format (OpenRouter's Decisions API): switching is a change of URL
from teximal import Fort

fort = Fort("teximal/fort-1-0.8b")
route = fort.task({"billing": "invoices, payments, refunds", "technical": "bugs, outages, error messages",
                   "sales": "plans, upgrades, pricing"}, question="Which team should handle this?")
d = route.decide("I was charged twice this month.")
print(d.choice, d.confidence)                                     # the option and its probability
many = route.decide_many(["Refund please.", "The app crashes on start."])

A question is a choice (up to 26 options, each with a probability; longer label lists are read by name), yes or no (the probability of yes), or a score on an ordered scale. Docs: teximal-cli.

Results

jevbench, 500 examples per task: accuracy in %, calibration error (ECE, lower is better) and median milliseconds per decision.

Model SST2 AG News Banking77 ECE SST2 / AG / B77 ms SST2 / AG / B77
Fort-1 0.8B 93.2 87.2 74.8* 0.013 / 0.063 / 0.178 28 / 44 / 51
Fort-1 2B 93.8 86.8 76.6* 0.012 / 0.092 / 0.182 110 / 214 / 79
Jev (jev-1.13, TypeSafe API) 95.4 84.3 (85.8†) 76.4 0.026 / 0.112 / 0.125 376 / 381 / 389
Claude Sonnet 5 (Anthropic API) 95.6 89.6 77.4 not reported 2,076 / 2,014 / 1,995
GPT-5-mini (OpenAI API) 95.0 80.2 73.6 not reported 1,385 / 1,312 / 1,312
  • Fort's times: one decision at a time on an Apple M1 Max (MLX). The API models' times include the network.
  • Fort never trained on SST2 or AG News, only on similar tasks: IMDB review sentences and HuffPost headlines labeled by Teximal Lab.
  • *Banking77 is adapted, not zero-shot: Fort trained on bank messages Teximal Lab wrote from the 77 intent names, never on real Banking77 messages (one scored message also appears inside a training message). Scored with jevbench's instruction, as every model was; with a banking instruction, Fort-1 2B scores 77.8 and Fort-1 0.8B scores 74.6.
  • †85.8 with jevbench's newer AG News label definitions.

Calibration: one temperature fitted on a few hundred labeled examples of your task (teximal eval does it) takes ECE to 0.018 / 0.060 / 0.035 (Fort-1 2B) and 0.020 / 0.062 / 0.038 (Fort-1 0.8B), on 300 held-out examples per task.

More results

Test Fort-1 0.8B
Decision exam: 600 tasks written by Teximal Lab, never trained on 67.5% (ECE 0.081)
People's 1-to-5 ratings, held out (urgency, frustration, sentiment strength, toxicity, stars) 50 to 74% exact, 88 to 98% within one level
General knowledge, 4 options 91.2%
Reading comprehension, 4 / 16 options 93.5% / 88.0%
Assistant requests in 51 languages (MASSIVE) 70.7%
Options shuffled or renamed, typos, lower case, a distracting sentence at most 1.3 points lower
A text that claims the answer or orders a letter no decision moved
Long context (window: 262,144 tokens) 88 to 100% with the deciding text inside about 2,000 tokens; at about 32,000, 95% when it comes last and 55 to 65% elsewhere

Limitations

  • Not a math model: it answers in one step, without working.
  • Mostly English; tweet sentiment in four African languages is weak (AfriSenti 49.4%).
  • Text only.
  • Long label lists read by name (such as Banking77) are overconfident until a temperature is fitted.
  • Instructions planted in a text rarely move its answer, but treat untrusted text as data.

Training

Fine-tuned from Qwen/Qwen3.5-0.8B with LoRA (rank 16). The targets blend the probabilities of a larger teacher, Qwen3.6-35B-A3B (Apache-2.0), with people's labels from Teximal Lab, collected with hidden checks.

Data Used for License
MASSIVE assistant requests in 51 languages CC BY 4.0
SQuAD 2.0 reading, answerability and knowledge questions CC BY-SA 4.0
CLINC150 intents and intent names CC BY 3.0
DBpedia14 entity types CC BY-SA 3.0
SNLI, MultiNLI entailment CC BY-SA 4.0; MultiNLI's own mixed terms
Amazon reviews (amazon_polarity) sentiment; 1-to-5 star ratings by Teximal Lab as published on Hugging Face
Civil Comments toxicity; 1-to-5 toxicity ratings by Teximal Lab CC0 1.0
HuffPost News Category Dataset news headlines, topics labeled by Teximal Lab CC BY 4.0
IMDB Large Movie Review Dataset review sentences, sentiment labeled by Teximal Lab research dataset
Teximal Lab critic-style sentences and bank customer messages, written from scratch Teximal

Intended use

Routing, triage, moderation queues, tagging and form checks: decisions over your own options, where a calibrated confidence decides what to automate and what to send to a person. Not for decisions with legal or similarly significant effects on people without human review.

License

Apache-2.0; the Qwen3.5 base model's notice is in NOTICE.

Citation

@misc{teximal2026fort1,
  title  = {Fort-1: calibrated decision models},
  author = {Teximal},
  year   = {2026},
  url    = {https://huggingface.co/teximal/fort-1-0.8b}
}
Downloads last month
6
Safetensors
Model size
0.8B params
Tensor type
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for teximal/fort-1-0.8b

Finetuned
(472)
this model
Quantizations
1 model

Collection including teximal/fort-1-0.8b