MasterFormat classifier (mf-0.2) โ€” construction spec & line-item classification

A BERT text classifier that maps construction line items, specification paragraphs, submittals and section titles to one of 171 MasterFormat level-2 groups (e.g. 03 30 00 Cast-in-Place Concrete, 23 30 00 HVAC Air Distribution) across 32 divisions. It is the machine-learning counterpart to a cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked. English text, one label per input.

Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI MasterFormat codes without a human choosing from a 171-row list.

Completed GPU fine-tune. mf-0.2 is the full run that mf-0.1 (step 700) was an early checkpoint of: 6,132 steps / 12 epochs on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS validation split is 46.6 % โ€” 2.3ร— mf-0.1 (20.0 %). Line-item accuracy on 194 hand-labelled estimate items is 59.7 % (was 28.3 %), and whole-section accuracy is 60.7 % with 81.0 % at division level. It is still below this project's TF-IDF + embedding ensemble on line items (70.2 %) โ€” see Results โ€” but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level division accuracy. To close the remaining line-item gap the repo also ships a transformer + TF-IDF ensemble (ensemble/) that lifts line-item top-1 to 68.6 % and gives the best manual-chunk and section-accuracy results; see Ensemble.

MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or endorsed by them.

Quick start

from transformers import pipeline

clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
for p in clf("EPDM membrane roofing"):
    print(p["label"], round(p["score"], 3))
# 07 50 00 Membrane Roofing 0.863
# 07 10 00 Dampproofing and Waterproofing 0.068
# 07 30 00 Steep Slope Roofing 0.047

Batched, with the level-1 division roll-up and JSON output:

pip install -r requirements.txt        # transformers + torch
python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
echo "12\" RCP storm drain pipe" | python predict.py -
python predict.py --file items.txt --json > out.json

ONNX (no torch, ~34 MB int8):

pip install onnxruntime transformers
python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"

The encoder is English-only; inputs should be English.

ONNX exports

onnx/model.onnx (fp32, 134 MB) and onnx/model_quantized.onnx (dynamic int8, 34 MB) are exported with dynamic batch and sequence axes, in the layout transformers.js / Optimum expect โ€” inputs input_ids, attention_mask, token_type_ids, output logits. The fp32 export matches PyTorch on a 4-text smoke set (max |ฮ” logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for mf-0.1 โ€” the better-trained head is more confident โ€” so verify the quantized model on your own inputs; the fp32 export is the safer default. Reproduce with python scripts/export_onnx.py runs/mf-0.2.

Model

  • Base: BAAI/bge-small-en-v1.5 (BERT, 33M parameters, 384-d, MIT).
  • Head: 171-way sequence-classification layer; id2label / label2id are in config.json.
  • Recipe: label smoothing 0.05, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast, AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen.
  • Checkpoint: train_state.json records {"step": 6132, "val_acc": 0.4661}.
  • Trained on: Kaggle T4, ~16.5 min wall-clock, from BAAI/bge-small-en-v1.5 (not resumed from mf-0.1).

Results

All numbers are measured in this project on held-out data. The line_items, manuals and manual_sections sets are fixed, so those columns are directly comparable for every scorer. val depends on the dataset version: mf-0.2 and mf-0.1 below are on the v2 split (11,598 units); the embeddings/ensemble rows were tuned on the v1 split (13,518 units) and are marked v1.

v2 validation split

Scorer val top-1 val top-3 val division
mf-0.2 (step 6132) 0.466 0.641 0.593
mf-0.1 (step 700) 0.200 0.372 0.356
TF-IDF + linear SGD 0.575 0.732 0.675

Fixed held-out sets (line items ยท manual chunks ยท whole sections)

Scorer line-item top-1 manual-chunk top-1 manual-section top-1 manual-section division
mf-0.2 (step 6132) 0.597 0.280 0.607 0.810
mf-0.1 (step 700) 0.283 0.176 0.287 0.588
TF-IDF + linear SGD 0.665 0.290 0.607 0.732
bge-small embeddings + logistic regression (v1) 0.576 0.256 0.593 0.725
Ensemble 0.25 embedding + 0.75 TF-IDF (v1) 0.702 0.318 0.640 0.771

line_items = 194 hand-labelled estimate line items; manuals = 6,890 chunks from 153 sections of two real commercial project manuals; out-of-taxonomy gold labels (e.g. 22 40 00) count only toward the division score, so top-1 is over in-taxonomy items only.

Reading it: mf-0.2 is the best single transformer here and the best scorer overall at section-level division accuracy (0.810). The TF-IDF + embedding ensemble still leads on short line items (0.702 vs 0.597), which is the real-world target โ€” so the ensemble remains the production candidate, with mf-0.2 a much stronger transformer baseline than mf-0.1.

Ensemble (closing the line-item gap)

The transformer is trained on specification prose, so on short, terse estimate line items a lexical TF-IDF model is still stronger (0.675 vs 0.597 top-1). To close that gap without giving up the transformer's long-text accuracy, this repo ships a log-probability ensemble of mf-0.2 and a TF-IDF + SGD scorer:

p = softmax( 0.25 ยท log_softmax(mf-0.2)  +  0.75 ยท log_softmax(tfidf) )

The 0.25 weight is chosen on the UFGS validation split โ€” not on the line-item test set. ensemble/tfidf.joblib holds the fitted vectorizer (word 1โ€“2 grams, 100k features, min_df=2, sublinear tf) and SGD logistic classifier; ensemble/blend.json records the weight and all metrics; ensemble/predict_ensemble.py runs the blend.

Scorer val top-1 val division line-item top-1 manual-chunk top-1 manual-section top-1 manual-section division
mf-0.2 (transformer) 0.466 0.593 0.597 0.280 0.607 0.810
TF-IDF (100k features) 0.571 0.663 0.675 0.288 0.573 0.726
Ensemble (w = 0.25) 0.586 0.688 0.686 0.334 0.613 0.784
pip install transformers torch scikit-learn joblib
python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \
    --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"

Honest caveat. The ensemble's line-item strength comes from TF-IDF: in a three-way blend with bge-small embeddings the transformer receives zero weight on line items (the best line-item result is 0.712 at 0.70 TF-IDF + 0.30 embeddings). The transformer's own contribution is on longer text โ€” it is the best scorer at manual-section division accuracy (0.810 vs 0.784 for the ensemble). A reasonable production split is the transformer for whole sections and short-answer text, the ensemble for terse line items.

Intended use

  • Use it for: tagging construction text with a candidate MasterFormat group (top-3 shown), an auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
  • Do not use it for: unattended production takeoff, bid pricing, code compliance, or anything where a wrong cost code has financial or contractual consequences. Keep a human in the loop.
  • Not a substitute for review. Classifies text content only โ€” if the input already contains a MasterFormat number, read the number instead.

Training data and taxonomy

  • Source: UFGS (Unified Facilities Guide Specifications) .SEC files โ€” US federal works in the public domain. Paragraphs and titles are parsed into labelled text units; boilerplate shared by multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and long paragraphs are cut on sentence boundaries into 8โ€“60-word windows.
  • Split: held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations stay on the same side. 65,496 train / 11,598 val rows (the v2 set).
  • Taxonomy: 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat numbering convention and the short names are this project's own (see config.json).

Limitations

  • Below the ensemble on line items. 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy as the sole classifier.
  • Domain skew. UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial building estimates; the training set is class-balanced, so raw predictions do not reflect building-project priors.
  • Out-of-taxonomy inputs (sections whose level-2 group is not among the 171) can only be scored at division level.
  • Short, terse line items are the hardest inputs; division (2-digit) accuracy is consistently higher than group (6-digit) accuracy.
  • Quantized ONNX drifts. The int8 export is smaller but less faithful than mf-0.1's; prefer fp32.

Bias, risks and safety

  • Estimating bias. Class-balanced training over a public-domain federal corpus does not represent any particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
  • Trademark. MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and its group names are the project's own short descriptions, not CSI's official titles.
  • Privacy. The model runs locally; no input text leaves your machine unless you call a hosted endpoint.

FAQ

What is MasterFormat? The CSI/CSC MasterFormat is the North American standard for organising construction specifications and cost data into numbered divisions and sections. This model predicts the level-2 group (a 6-digit code such as 03 30 00 Cast-in-Place Concrete), not the full section number.

How is this different from mf-0.1? mf-0.1 was step 700 of an interrupted CPU run; mf-0.2 is the completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3ร— the validation accuracy.

Can it classify a whole specification section? Yes โ€” average the model's log-probabilities over a section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.

Can it read a MasterFormat number out of the text? No. It classifies the description. If the number is already present, parse it directly.

Does it work offline / in the browser? Yes โ€” the ONNX exports are intended for onnxruntime and transformers.js.

Is a better model available? Yes โ€” this repo ships a transformer + TF-IDF ensemble (ensemble/) that scores 68.6 % top-1 on hand-labelled line items (vs 59.7 % for the transformer alone) and is the best scorer on manual chunks. A TF-IDF + bge-small embedding ensemble reaches 71.2 % but the transformer takes no weight there. Constructelligence's proprietary models are at constructelligence.co.

Files

  • model.safetensors, config.json, tokenizer.json, tokenizer_config.json, vocab.txt, special_tokens_map.json โ€” standard transformers checkpoint.
  • train_state.json โ€” step and validation accuracy of the saved checkpoint.
  • metrics.json โ€” full held-out evaluation (val, line_items, manuals, manual_sections).
  • predict.py โ€” CLI example: batching, --top-k, --divisions, --json, stdin/file input, --onnx.
  • onnx/ โ€” ONNX fp32 and int8 exports.
  • ensemble/ โ€” tfidf.joblib, blend.json, predict_ensemble.py: the transformer + TF-IDF line-item ensemble.
  • CITATION.cff, requirements.txt.

Citation

@misc{constructelligence_masterformat_classifier,
  title        = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
  author       = {Constructelligence},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
  note         = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
}

Licence and attribution

Released under the MIT licence, matching the base model. UFGS source text is public domain. MasterFormat is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them. Constructelligence's production models are available at constructelligence.co.

Downloads last month
13
Safetensors
Model size
33.4M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for constructelligence/masterformat-classifier

Quantized
(29)
this model

Evaluation results