Instructions to use nativ-community/rampart-mlx-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use nativ-community/rampart-mlx-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir rampart-mlx-fp16 nativ-community/rampart-mlx-fp16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Rampart — MLX fp16
MLX conversion of nationaldesignstudio/rampart,
National Design Studio's multilingual PII token classifier
(announcement). It is a 6-layer
MiniLM (BertForTokenClassification, hidden 384, 19,730-piece vocab) with 35 BIO labels
over 17 entity types, in seven Latin-script languages.
| Size | 37 MB |
| Weights | The ONNX 4-bit and uint8 weights are dequantized and stored as fp16. The values are the released (already quantized) weights, not the pre-quantization training checkpoint, which National Design Studio did not publish. |
| Other precision | nativ-community/rampart-mlx-4bit (4-bit) |
| License | CC BY 4.0, from the original model. Attribution: National Design Studio. |
Usage
Needs mlx-vlm with BERT token-classification
support (mlx_vlm.token_classification).
from mlx_vlm.token_classification import load_token_classifier
classifier = load_token_classifier("nativ-community/rampart-mlx-fp16")
result = classifier(
"I'm Sarah Connor, 1984 Cyberdyne Ave, Los Angeles CA 90012, sarah@sky.net",
keep_labels=("CITY", "STATE", "ZIP_CODE"),
)
print(result.redacted_text)
# I'm <GIVEN_NAME> <SURNAME>, <BUILDING_NUMBER> <STREET_NAME>, Los Angeles CA 90012, <EMAIL>
python -m mlx_vlm.token_classification --model nativ-community/rampart-mlx-fp16 \
--keep CITY,STATE,ZIP_CODE "Call Maria Garcia on 617-555-0142."
keep_labels mirrors Rampart's default policy: city, state and ZIP are detected but kept.
Not included: upstream Rampart is a hybrid system. SSNs, payment cards and IP addresses
are caught by a regex and checksum layer (in the @nationaldesignstudio/rampart npm package)
that masks them before the model runs, and the model was trained with those values masked.
This checkpoint is the model only, so add your own pattern matching for those classes.
Conversion
Upstream ships only onnx/model_q4.onnx. convert_from_onnx.py in this repo reads the ONNX
initializers, maps them to the HF BERT parameter names, and writes MLX safetensors in
mlx-vlm's layout. It reproduces both checkpoints from the upstream repo:
hf download nationaldesignstudio/rampart --local-dir rampart
python convert_from_onnx.py rampart # writes rampart-mlx-fp16/ and rampart-mlx-4bit/
Evaluation
Measured on an Apple M5 Max. The eval set is 10,500 rows from the validation split of
ai4privacy/pii-masking-openpii-1.5m,
1,500 per language, which is data the model was not trained on. The reference is the shipped
ONNX q4 model under ONNX Runtime (CPU), run on the same inputs.
Conversion fidelity (1,280,624 tokens):
| Token labels identical to ORT | Rows with identical redaction | |
|---|---|---|
| MLX fp16 | 99.998% | 99.95% |
| MLX 4-bit | 99.997% | 99.90% |
Every disagreement is a near-tie: ORT's own top-1 vs top-2 logit margin at those tokens is at most 0.016.
Model-only private-term recall (a gold PII value counts as caught if any of its characters is redacted; Wilson 95% CI):
| Language | ORT q4 (upstream) | MLX fp16 | MLX 4-bit | Private terms |
|---|---|---|---|---|
| en | 99.87% [99.75, 99.93] | 99.87% | 99.87% | 6,221 |
| es | 99.75% [99.59, 99.85] | 99.75% | 99.75% | 5,708 |
| fr | 99.34% [99.10, 99.52] | 99.34% | 99.34% | 5,641 |
| de | 99.59% [99.38, 99.73] | 99.59% | 99.59% | 5,582 |
| it | 99.50% [99.29, 99.65] | 99.50% | 99.50% | 6,004 |
| pt | 99.63% [99.44, 99.75] | 99.63% | 99.63% | 5,896 |
| nl | 99.28% [99.03, 99.47] | 99.28% | 99.28% | 5,977 |
| all | 99.57% [99.50, 99.63] | 99.57% | 99.57% | 41,029 |
Through the mlx_vlm.token_classification span decoder end to end, with the city/state/ZIP keep
policy: private recall 99.57%, public-term retention 99.90%.
The private terms counted are names, phones, emails, ID/tax/passport/licence numbers, street and building number. SSNs and credit cards are left out because they belong to the regex layer. Public terms are dates, ages, titles, gender, times, city and ZIP. These numbers are not comparable to the 98.42% on the upstream card. That figure is a full-system score on a pinned 30k-row slice with a different term mapping. What this table shows is that the MLX checkpoints reproduce upstream model behaviour.
Rampart is a redaction aid, not an anonymization or compliance guarantee. See the upstream model card for intended use, limitations, and fairness results.
- Downloads last month
- 9
Quantized
Model tree for nativ-community/rampart-mlx-fp16
Base model
nreimers/MiniLM-L6-H384-uncased