Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
MasterFormat classifier (mf-0.2) — construction spec & line-item classification
A BERT text classifier that maps construction line items, specification paragraphs, submittals and
section titles to one of 171 MasterFormat level-2 groups (e.g. 03 30 00 Cast-in-Place Concrete,
23 30 00 HVAC Air Distribution) across 32 divisions. It is the machine-learning counterpart to a
cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked.
English text, one label per input.
Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI MasterFormat codes without a human choosing from a 171-row list.
Completed GPU fine-tune.
mf-0.2is the full run thatmf-0.1(step 700) was an early checkpoint of: 6,132 steps / 12 epochs on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS validation split is 46.6 % — 2.3×mf-0.1(20.0 %). Line-item accuracy on 194 hand-labelled estimate items is 59.7 % (was 28.3 %), and whole-section accuracy is 60.7 % with 81.0 % at division level. Because the transformer is trained on specification prose, short estimate line items are its weak spot; the repo therefore also ships a transformer + TF-IDF ensemble (ensemble/) that lifts line-item top-1 to 70.2 % and improves manual-chunk and section accuracy; see Ensemble.
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or endorsed by them.
Quick start
from transformers import pipeline
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
for p in clf("EPDM membrane roofing"):
print(p["label"], round(p["score"], 3))
# 07 50 00 Membrane Roofing 0.863
# 07 10 00 Dampproofing and Waterproofing 0.068
# 07 30 00 Steep Slope Roofing 0.047
Batched, with the level-1 division roll-up and JSON output:
pip install -r requirements.txt # transformers + torch
python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
echo "12\" RCP storm drain pipe" | python predict.py -
python predict.py --file items.txt --json > out.json
ONNX (no torch, ~34 MB int8):
pip install onnxruntime transformers
python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"
The encoder is English-only; inputs should be English.
ONNX exports
onnx/model.onnx (fp32, 134 MB) and onnx/model_quantized.onnx (dynamic int8, 34 MB) are exported with
dynamic batch and sequence axes, in the layout transformers.js / Optimum expect — inputs input_ids,
attention_mask, token_type_ids, output logits. The fp32 export matches PyTorch on a 4-text smoke set
(max |Δ logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for
mf-0.1 — the better-trained head is more confident — so verify the quantized model on your own inputs;
the fp32 export is the safer default. Reproduce with python scripts/export_onnx.py runs/mf-0.2.
Model
- Base:
BAAI/bge-small-en-v1.5(BERT, 33M parameters, 384-d, MIT). - Head: 171-way sequence-classification layer;
id2label/label2idare inconfig.json. - Recipe:
label smoothing 0.05, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast, AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen. - Checkpoint:
train_state.jsonrecords{"step": 6132, "val_acc": 0.4661}. - Trained on: Kaggle T4, ~16.5 min wall-clock, from
BAAI/bge-small-en-v1.5(not resumed frommf-0.1).
Results
All numbers are measured in this project on held-out data. The line_items, manuals and manual_sections
sets are fixed, so those columns are directly comparable for every scorer. val is the v2 UFGS split
(11,598 units).
v2 validation split
| Scorer | val top-1 | val top-3 | val division |
|---|---|---|---|
| mf-0.2 (step 6132, published) | 0.466 | 0.641 | 0.593 |
| mf-0.3 (v3 data, unfrozen) | 0.556 | 0.710 | 0.662 |
| mf-0.4 (v2 data, unfrozen) | 0.491 | 0.657 | 0.609 |
| mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
| TF-IDF (word + char n-grams) | 0.571 | 0.732 | 0.670 |
Fixed held-out sets (line items · manual chunks · whole sections)
| Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|---|---|---|---|---|
| mf-0.2 (step 6132, published) | 0.597 | 0.280 | 0.607 | 0.810 |
| mf-0.3 (v3 data, unfrozen) | 0.597 | 0.259 | 0.547 | 0.765 |
| mf-0.4 (v2 data, unfrozen) | 0.618 | 0.270 | 0.567 | 0.817 |
| mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
| TF-IDF (word + char n-grams) | 0.696 | 0.287 | 0.587 | 0.752 |
| Ensemble mf-0.2 + TF-IDF (w = 0.25) | 0.702 | 0.330 | 0.620 | 0.804 |
line_items = 194 hand-labelled estimate line items; manuals = 6,890 chunks from 153 sections of two real
commercial project manuals; out-of-taxonomy gold labels (e.g. 22 40 00) count only toward the division
score, so top-1 is over in-taxonomy items only.
Evidence caveat. The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the
line-item top-1 is [0.633, 0.762] (±6.4 pts), and whole-section is ±~8 pts — so differences of a few
points between models here are noise. Treat the val and manual-chunk rows (±1 pt) as the measurable
ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap.
Reading it: mf-0.2 is the best scorer overall at section-level division accuracy (0.810). On short
line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped transformer + TF-IDF ensemble
reaches 0.702 and is the production candidate; see Ensemble.
Ensemble (closing the line-item gap)
The transformer is trained on specification prose, so on short, terse estimate line items a lexical
TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's
long-text accuracy, this repo ships a log-probability ensemble of mf-0.2 and a TF-IDF + SGD scorer:
p = softmax( 0.25 · log_softmax(mf-0.2) + 0.75 · log_softmax(tfidf) )
The 0.25 weight is chosen on the UFGS validation split — not on the line-item test set. ensemble/tfidf.joblib
holds the fitted vectorizer (word 1–2 grams, 100k features, plus char_wb 3–5 grams, 150k features, min_df
2/3, sublinear tf) and SGD logistic classifier; ensemble/blend.json records the weight and all metrics;
ensemble/predict_ensemble.py runs the blend.
| Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|---|---|---|---|---|---|---|
| mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | 0.810 |
| TF-IDF (word + char n-grams) | 0.571 | 0.670 | 0.696 | 0.287 | 0.587 | 0.752 |
| Ensemble (w = 0.25) | 0.588 | 0.689 | 0.702 | 0.330 | 0.620 | 0.804 |
pip install transformers torch scikit-learn joblib
python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \
--tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
Honest caveat. The ensemble beats TF-IDF alone mainly on the transformer's strengths — validation top-1
(0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) — while line
items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does not help once the
blend weight is chosen on val (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning
the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at
section division accuracy (0.810). A reasonable production split is the transformer for whole sections, the
ensemble for terse line items.
Calibration and abstention
The ensemble is already close to calibrated; a single temperature T = 0.90 (fit on val) lowers the
expected calibration error on val from 0.060 to 0.022. More useful in production: send the
lowest-confidence predictions to a human instead of filing them.
Coverage → precision on the 194 line items (predictions sorted by confidence):
| auto-filed (coverage) | 100 % | 90 % | 80 % | 70 % | 50 % | 30 % |
|---|---|---|---|---|---|---|
| precision | 0.69 | 0.75 | 0.81 | 0.85 | 0.91 | 0.97 |
Routing the least-confident 30 % to review lifts precision to 91 %; routing 50 % gives 97 %. So
the model is usable today as an assist that flags its own uncertainty, even though it is not accurate
enough to file unattended. Reproduce with python cloud/stats.py.
Roadmap
Further gains are evidence-bound, not architecture-bound (see the repo's IMPROVEMENT-PLAN.md):
- Enlarge the eval (≥1,000 line items, ≥400 sections) so changes are measurable at ±3 pts — currently line-item changes below ~10 pts cannot be validated.
- Real line-item training data (public bid tabulations, agency item catalogs) — the transformer's
line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (
mf-0.3). - Domain-adaptive pretraining on construction prose beyond UFGS, then retrain.
- Distillation + hierarchical (division → group) head to beat the TF-IDF blend with a single model.
- Level-3 / full-section codes and an explicit out-of-taxonomy fallback.
Intended use
- Use it for: tagging construction text with a candidate MasterFormat group (top-3 shown), an auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
- Do not use it for: unattended production takeoff, bid pricing, code compliance, or anything where a wrong cost code has financial or contractual consequences. Keep a human in the loop.
- Not a substitute for review. Classifies text content only — if the input already contains a MasterFormat number, read the number instead.
Training data and taxonomy
- Source: UFGS (Unified Facilities Guide Specifications)
.SECfiles — US federal works in the public domain. Paragraphs and titles are parsed into labelled text units; boilerplate shared by multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and long paragraphs are cut on sentence boundaries into 8–60-word windows. - Split: held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations stay on the same side. 65,496 train / 11,598 val rows (the v2 set).
- Taxonomy: 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
numbering convention and the short names are this project's own (see
config.json).
Limitations
- Below the ensemble on line items. 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy as the sole classifier.
- Domain skew. UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial building estimates; the training set is class-balanced, so raw predictions do not reflect building-project priors.
- Out-of-taxonomy inputs (sections whose level-2 group is not among the 171) can only be scored at division level.
- Short, terse line items are the hardest inputs; division (2-digit) accuracy is consistently higher than group (6-digit) accuracy.
- Quantized ONNX drifts. The int8 export is smaller but less faithful than
mf-0.1's; prefer fp32.
Bias, risks and safety
- Estimating bias. Class-balanced training over a public-domain federal corpus does not represent any particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
- Trademark. MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and its group names are the project's own short descriptions, not CSI's official titles.
- Privacy. The model runs locally; no input text leaves your machine unless you call a hosted endpoint.
FAQ
What is MasterFormat? The CSI/CSC MasterFormat is the North American standard for organising construction
specifications and cost data into numbered divisions and sections. This model predicts the level-2 group
(a 6-digit code such as 03 30 00 Cast-in-Place Concrete), not the full section number.
How is this different from mf-0.1? mf-0.1 was step 700 of an interrupted CPU run; mf-0.2 is the
completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3× the validation accuracy.
Can it classify a whole specification section? Yes — average the model's log-probabilities over a section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.
Can it read a MasterFormat number out of the text? No. It classifies the description. If the number is already present, parse it directly.
Does it work offline / in the browser? Yes — the ONNX exports are intended for onnxruntime and
transformers.js.
Is a better model available? Yes — this repo ships a transformer + TF-IDF ensemble (ensemble/) that
scores 70.2 % top-1 on hand-labelled line items (vs 59.7 % for the transformer alone) and improves
manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at
constructelligence.co.
Files
model.safetensors,config.json,tokenizer.json,tokenizer_config.json,vocab.txt,special_tokens_map.json— standardtransformerscheckpoint.train_state.json— step and validation accuracy of the saved checkpoint.metrics.json— full held-out evaluation (val,line_items,manuals,manual_sections).predict.py— CLI example: batching,--top-k,--divisions,--json, stdin/file input,--onnx.onnx/— ONNX fp32 and int8 exports.ensemble/—tfidf.joblib(word + char n-grams, ~112 MB),blend.json,predict_ensemble.py: the transformer + TF-IDF line-item ensemble.IMPROVEMENT-PLAN.md— the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation).CITATION.cff,requirements.txt.
Citation
@misc{constructelligence_masterformat_classifier,
title = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
author = {Constructelligence},
year = {2026},
howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
note = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
}
Licence and attribution
Released under the MIT licence, matching the base model. UFGS source text is public domain. MasterFormat is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them. Constructelligence's production models are available at constructelligence.co.
- Downloads last month
- 41
Model tree for constructelligence/masterformat-classifier
Base model
BAAI/bge-small-en-v1.5Evaluation results
- Top-1 accuracy (mf-0.2) on UFGS held-out text units (11,598)validation set self-reported0.466