CountryCode-Decide (small, v1.2.1)
Maps a country or US state written as free text — any language or script, misspellings, abbreviations, historical names, register junk — to ISO 3166-1 alpha-2 / ISO 3166-2:US codes with calibrated probabilities, and abstains when the requested error rate cannot be met. 75 MB, about 4 ms per string on a CPU.
Code, evaluation and documentation: https://github.com/aliildan/country-code-decide · Paper
Intended use
Normalising country / jurisdiction / nationality fields in company-register data, KYB onboarding and similar back-office pipelines, with a person reviewing abstentions and candidate lists.
Out of scope: full postal addresses (the model reads one field), subdivisions other than US states, legal advice, any decision about a person without human review.
How to use
Serving needs only onnxruntime, tokenizers and numpy (no PyTorch). The decision logic
(calibration, thresholds, candidate lists) lives in the repository's CountryCoder:
git clone https://github.com/aliildan/country-code-decide && cd country-code-decide
uv sync --no-default-groups --group serve && source .venv/bin/activate
uvx --from huggingface_hub hf download aildan/country-code-decide-small --exclude model.onnx \
--local-dir models/country-code-decide-small
from countrycode.infer.pipeline import CountryCoder
coder = CountryCoder.load("models/country-code-decide-small")
d = coder.code("Untted Klngdom", field="country", domain="register", alpha=0.05)
d.target, d.candidates, d.abstain # ('GB', ('GB',), False); d.confidence ≈ 0.95
Input: text, optional field (country, jurisdiction, nationality, unknown) and
register_country (ISO alpha-2 of the register, or US-XX for a US state register). Output per
string: target (alpha-2, US-XX, NONE or None = abstain), country, us_state,
confidence, candidates (codes with calibrated probability ≥ 0.05, at most three), abstain,
for an error-rate target alpha (certified targets — curated: 5 %; register: 5 %; None always answers). Pass
domain="register" for register and form fields: the certified register calibration applies only
then; without a domain the stricter of the two policies applies, which is conservative but not
itself certified.
Files
| File | Content |
|---|---|
model.int8.onnx |
the model: encoder + two heads, 8-bit weights (served; the CPU runtime computes its matrix products with int8 activations) |
model.onnx |
the same model in fp32 (reference) |
tokenizer.json, tokenizer_config.json |
the trimmed mmBERT tokenizer |
calibration.json |
temperatures, certified thresholds and candidate rule per domain (register, curated) |
meta.json |
label space, code names, context format, training run, quality-gate status |
Training data
The model was trained on three kinds of strings. About a third of the training examples are generated (synthetic); no generated string is used for calibration or evaluation.
Lexicon — 372,683 names in 1,274 languages (CLDR 47, Wikidata, codes, reviewed overrides); 86,406 distinct keys, 1,947 kept as genuinely ambiguous.
Real register strings of the training split (Colorado, Connecticut, UK Companies House, Brønnøysund), labelled by a person or by the register's own code; no normalised key shared with any evaluation part.
Generated variants, one release per generator, each made by code from a lexicon name with a known code (so the label is known by construction):
keyboard— typos on real keyboard layouts chosen by the name's language, with the number and kind of edits fitted to real register typos;format— case, spacing, comma inversion, postcode noise, field truncation, abbreviation prefixes (labelled with every code whose name starts that way);translit— romanisations and umlaut conventions;geo— a country or US state written with one of its GeoNames regions or cities;pairs— two countries in one string (labelled as the set of both);affix— a country with a compass or region qualifier;caprov— Canadian province codes in a US register's jurisdiction field (→CA);junk,regjunk— strings that are not places: placeholders, dates, numbers, generated company names, legal terms, street lines (→NONE).
Every release the model uses passed the quality gates before training: G1 label correctness (each string checked against every name of every code; strings closer to another code are dropped) plus a human audit of 300 random items per release (every item when a release is smaller) with zero label errors required; G2 realism (edit statistics compared with real typos — reported, no pass/fail threshold); G3 usefulness (kept only if it improved held-out real strings; the OCR, LLM-paraphrase and region generators failed this and are not used); G4 no overlap with any evaluation string; G5 reproducibility (code, config, seed and source item recorded per string).
| generator | training examples | human audit (errors / audited) |
|---|---|---|
| junk@2 | 1,375 | 0 / 300 |
| keyboard@2 | 9,812 | 0 / 300 |
| format@4 | 13,411 | 0 / 300 |
| translit@4 | 4,396 | 0 / 300 |
| geo@2 | 7,381 | 0 / 300 |
| pairs@3 | 3,988 | 0 / 300 |
| affix@2 | 785 | 0 / 300 |
| caprov@1 | 78 | 0 / 39 |
| regjunk@1 | 2,731 | 0 / 300 |
142,925 training examples: 98,517 lexicon names and codes (68.9 %), 43,957 generated (30.8 %), 451 real register strings (0.3 %).
Real register strings are labelled by a person or by the register's own code; none shares a normalised key with any evaluation part.
No personal data: register sources were read with aggregate queries (distinct value and count), never records; generated company names come from word lists.
Evaluation
Fresh registers never used for training, calibration or model selection (New York, Oregon): every distinct string of their jurisdiction and country fields, labelled by exact lexicon match or by a human reviewer, read once after a frozen pre-registration (release 1.2). Release 1.2.1 keeps the weights and corrects the certificate after an external review, so its numbers below are post-hoc.
| Fresh registers, 361 distinct strings | CountryCode-Decide small v1.2.1 |
|---|---|
| Accuracy, every string answered | 88.6 % (95 % CI 85.3–91.7) |
| Right country (a US state counts as US) | 91.2 % |
| At the certified 5 % error target | answers 88 % automatically, realised error 2.5 % |
| Best non-LLM baseline (exact lexicon + context rules) | 89.5 % |
| Zero-shot LLM (qwen3.5_35b-a3b-q4_K_M) | 85.0 %, ~29× slower (LLM on a GPU, this model on a CPU) |
| Messy register strings (656): lexicon alone → lexicon, then model | answers 24.8 % → 79.0 % (error 4.8 %) |
| Model file · CPU latency | 75 MB (int8 ONNX) · p50 4.1 ms per string |
| Predictor | automatic labels (exact lexicon match, n = 261) | hand-labelled (n = 100) | hand-labelled answered |
|---|---|---|---|
| CountryCode-Decide (always answers) | 99.2 % | 61.0 % | 100 |
| CountryCode-Decide at the certified 5 % target | 98.1 % | 56.0 % | 61 |
| Zero-shot LLM (qwen3.5_35b-a3b-q4_K_M) | 94.3 % | 61.0 % | 97 |
| Lexicon + context rules | 99.6 % | 63.0 % | 62 |
| multilingual-e5-small nearest name | 100.0 % | 37.0 % | 100 |
| exact lexicon match | 100.0 % | 24.0 % | 23 |
Automatic labels come from an exact lexicon match, so a lexicon is right on them by construction; on the hand-labelled strings the model, the LLM and an exact lexicon with the training context rules are about level. On held-out strings of the training registers (656 strings, many misspelled or abbreviated; other strings of the same registers were in training) the model is right on 82.6 %, the best other predictor (qwen3.5_35b-a3b-q4_K_M zero-shot) on 73.8 %, the lexicon with context rules on 35.8 %; on the 472 of them with no exact lexicon match 81.1 % against 73.1 %. On GBIF's curated list the LLM is better: 69.6 % against 58.8 %.
Lexicon first, then the model at its 5 % target:
| Set | Lexicon + context rules alone (answered, error) | Lexicon, then the model (answered, error) | Model on what the lexicon cannot read (answered, wrong) |
|---|---|---|---|
| Fresh registers (New York, Oregon), 361 strings | 89.5 %, error 1.5 % | 91.7 %, error 3.0 % | 8 of 38, 5 wrong (62.5 %) |
| Held-out strings of the training registers, 656 strings | 24.8 %, error 1.8 % | 79.0 %, error 4.8 % | 355 of 493, 22 wrong (6.2 %) |
| GBIF curated list, 3,980 strings | 17.8 %, error 2.4 % | 39.6 %, error 4.8 % | 867 of 3,270, 59 wrong (6.8 %) |
The model roughly triples what is answered automatically on messy register strings, but on what the lexicon cannot read its 5 % certificate does not hold (last column); treat those answers as suggestions for review until the residual has its own calibration.
- Country-level accuracy (right country; a US state counts as US): 0.912 (highest: v1_model 0.921).
- On the 33 strings not in the lexicon (no exact name match) the model's accuracy is 0.212.
- At the certified 5 % error target the model answers 88% of the strings automatically with a realised error of 2.5 %; non-places get a code 0.0 % of the time (n = 4). The 5 % target is certified over 318 register units (one per register and normalised key); 1 % and 2 % are not.
- The zero-shot LLM (qwen3.5_35b-a3b-q4_K_M) is less accurate when every string must be answered (0.850 vs the model's 0.886); its own confidence covers 71% of the strings at 5 % empirical risk, without a certificate. It is ~29× slower (the LLM on a GPU, this model on a CPU).
The guarantee is an average over strings like the calibration strings, and the errors concentrate on hard ones: at the 5 % target 13.1 % of the 61 answered hand-labelled strings are wrong, 0.0 % of the 256 answered strings that an exact lexicon match labelled.
The certificate, corrected. The pre-registered certificate treated the 592 register calibration strings as independent draws and certified 2 % and 5 %. The same string often appears in several fields of a register, so they are not independent: they form 318 register-and-key units. Counted per unit — a correction made after the test read, prompted by an external review — 2 % is not certified, and 5 % is, at λ = 0.762. This release (small-v1.2.1) ships that calibration; the weights are unchanged.
| predictor | n | acc | acc (freq-weighted) | country acc | coverage @1 % risk | coverage @5 % risk | non-place false accept | ECE | ms / item |
|---|---|---|---|---|---|---|---|---|---|
| lexicon_context | 361 | 0.895 | 1.000 | 0.901 | 0.000 | 0.895 | 0.250 | 0.015 | 0.0 |
| model | 361 | 0.886 | 1.000 | 0.912 | 0.856 | 0.911 | 0.750 | 0.058 | 4.2 |
| v1_model | 361 | 0.884 | 1.000 | 0.921 | 0.197 | 0.864 | 1.000 | 0.253 | 3.9 |
| model@0.05 | 361 | 0.864 | 1.000 | 0.873 | 0.856 | 0.878 | 0.000 | 0.033 | 4.2 |
| llm_qwen3.5_35b-a3b-q4_K_M | 361 | 0.850 | 0.997 | 0.858 | 0.000 | 0.709 | 0.000 | 0.108 | cached |
| e5_knn | 361 | 0.825 | 0.995 | 0.847 | 0.091 | 0.825 | 0.750 | 0.108 | 0.6 |
| lexicon_fuzzy | 361 | 0.795 | 0.994 | 0.802 | 0.000 | 0.820 | 0.500 | 0.054 | 1.0 |
| lexicon_exact | 361 | 0.789 | 0.994 | 0.793 | 0.000 | 0.787 | 0.250 | 0.014 | 0.0 |
| ons_style | 361 | 0.787 | 0.994 | 0.793 | 0.000 | 0.792 | 0.500 | 0.044 | 0.3 |
| countrynames | 361 | 0.634 | 0.761 | 0.745 | 0.000 | 0.000 | 0.000 | 0.226 | 2.7 |
| pycountry | 361 | 0.634 | 0.761 | 0.776 | 0.000 | 0.000 | 0.250 | 0.184 | 1.5 |
| coco | 361 | 0.596 | 0.761 | 0.601 | 0.000 | 0.000 | 0.250 | 0.120 | 0.5 |
| model@0.01 | 361 | 0.011 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | – | 4.1 |
| model@0.02 | 361 | 0.011 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | – | 4.1 |
Pre-registered hypotheses of release 1.2, as read (P1 and P2 rest on the item-level certificate corrected above):
| hypothesis | result | evidence |
|---|---|---|
| P1 | met | certified=True, lambda=0.827, answered_n=309, errors=3, p_value_risk_le_2pct=0.947, contradicted=False |
| P2 | met | answered=0.856 |
| P3 | met | size_mb=74.732, cpu_p50_ms=4.092, tokens_identical=1.000, select_accuracy_drop=0.003 |
| P4 | met | model_accuracy=0.886, best_non_llm=e5_knn, best_non_llm_accuracy=0.825 |
| A1 | reported (not judged) | certified=False, answered=0.000, risk_answered=nan |
Gold labels: The register calibration strings were labelled a second time, blind (no proposal and no first label shown): agreement 99.2 % on 592 strings (5 disagreements); one-sided 95 % upper bound on disagreement 1.8 %. On the 218 hand-labelled strings the passes disagree on 5 (2.3 %); the 374 automatic or preset labels agree by construction. Test labels were labelled once.
Latency and size: p50 4.1 ms, p95 5.1 ms per string on CPU (4 threads, ONNX Runtime, tokenisation included); model file 75 MB.
Limitations and risks
- For clean fields an exact lexicon with context rules is as accurate as this model; use the model for what a lexicon cannot read, but note that its certificate does not cover that residual (see the cascade table). Its advantage was measured on held-out strings of registers whose other strings were in training; on fresh registers, strings with no exact lexicon match are hard.
- The certificate holds on average for strings drawn like the calibration strings (register fields), counted per register and normalised key, each target on its own; not on the hard strings a lexicon cannot read (with probability ≥ 90 % over the calibration draw); errors concentrate on hard strings. Recalibrate before relying on it for a different source. The curated GBIF domain is much harder than registers for this model, and there a 35B LLM is more accurate.
- One human reviewer labelled the register strings; the calibration strings were labelled twice by that reviewer, the second time blind (consistency, not inter-annotator agreement); test labels were labelled once.
- Of the 4,997 evaluation strings, 5 are not in Latin script, so accuracy on other scripts is not measured.
- One training seed (13); the usefulness gate compared runs whose spread is about its margin.
- Ambiguous names are returned as candidate lists; downstream systems must not silently pick one.
- Known gaps: bare Canadian province codes in US jurisdiction fields (
ON,BC) are not recognised (every real such string is an evaluation string, so the no-overlap gate removed them from the generated province variants);Congowithout context is answered asCG. - Not legal advice; keep a person in the loop.
Citation
@misc{ildan2026countrycode,
title = {From Messy Country Names to ISO 3166 Codes: A Small Model That Abstains, and When a Lexicon Is Enough},
author = {Ildan, Ali},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.23289343}
}
@software{ildan2026countrycode_software,
title = {CountryCode-Decide: messy country names to ISO 3166 codes, with calibrated abstention},
author = {Ildan, Ali},
year = {2026},
url = {https://github.com/aliildan/country-code-decide},
doi = {10.5281/zenodo.23289336}
}
Attribution
Unicode CLDR (Unicode License v3) · Wikidata (CC0) · GeoNames (CC BY 4.0) · GBIF parsers (Apache-2.0) · Colorado, Connecticut and Oregon open data (public domain) · UK Companies House (OGL v3) · Brønnøysund Register Centre (NLOD) · New York open data (OPEN-NY terms, evaluation only) · mmBERT (MIT).
Model tree for aildan/country-code-decide-small
Base model
jhu-clsp/mmBERT-small