Instructions to use DataFog/pii-en-65m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DataFog/pii-en-65m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="DataFog/pii-en-65m")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("DataFog/pii-en-65m") model = AutoModelForTokenClassification.from_pretrained("DataFog/pii-en-65m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DataFog PII EN 65M
DataFog PII EN 65M is a 65.2M-parameter English token classifier for locating and redacting seven types of personally identifiable information. It returns typed spans, calibrated span confidence, UTF-8 byte offsets and redacted text through a local JSONL runtime. Inference uses the CPU and requires no network access after installation.
The frozen model v0.1.0 achieves 98.32% exact typed-span F1 on 90,018 held-out synthetic documents. Runtime v0.2.0 has qualified native bundles for five platform environments. The checkpoint and acceptance threshold remain experimental; performance on reviewed customer tickets has not been established.
Model · Files · Usage · Input-and-output contract · Results · Training · Limitations · License
Model
| Property | Value |
|---|---|
| Backbone | distilbert/distilbert-base-cased, pinned revision 6ea81172465e8b0ad3fddeed32b986cdcdcffcf0 |
| Parameters | 65,202,447, counted from the released safetensors tensor shapes |
| Precision | Float32 weights and ONNX inference |
| Task | BIO token classification: 15 classes, including O |
| Supported labels | first_name, last_name, email, phone_number, street_address, customer_id, account_number |
| Long-input processing | 512-token windows: 510 content tokens with 42-token overlap |
| Confidence | Pooled isotonic calibration of exact typed-span correctness |
| Experimental threshold | 0.6486486486486487; authoritative value in bundle/calibration.json |
| Native execution | Rust + ONNX Runtime 1.30.0/API27, CPU |
| Model / runtime versions | Model 0.1.0; native runtime 0.2.0; versioned separately |
The root safetensors are a standard Transformers token-classification checkpoint. The native runtime additionally supplies complete-input windowing, boundary-fragment suppression, overlap resolution, calibration, thresholding and byte-based redaction. A generic Transformers pipeline does not reproduce that full contract. GLiNER is a separate CI comparison and is not this model's backbone or a release artifact.
Files
| File or directory | Purpose |
|---|---|
model.safetensors, config.json |
Frozen selected checkpoint and Transformers label mappings |
tokenizer.json, tokenizer_config.json, label-mappings.json |
Tokenizer and explicit class/label contract |
bundle/model.onnx, bundle/config.json, bundle/calibration.json |
Original native-model assets; native configuration differs from root Transformers configuration |
runtimes/v0.2.0/*.tar.gz |
Five platform-specific native bundles, including model assets, executable, matching ORT libraries and notices |
runtimes/v0.2.0/release-index.json |
Exact archive hashes, source/runtime provenance and native qualification reports |
runtimes/v0.2.0/tools/ |
Hash-verifying bundle installer and archive helpers |
provenance/artifacts/evaluation/, provenance/REPORT.md |
Frozen detection, redaction, calibration and support-fixture results |
card-evidence/ |
Original regex comparison summaries and recorded execution environment |
reproduction/, provenance/ |
Training/evaluation code, locks, partition/source provenance and original audit |
LICENSE, NOTICE, licenses/, LICENSING.md |
Model/source licensing, dataset attribution and upstream dependency notices |
Canonical Hub ID: DataFog/pii-en-65m (formerly DataFog/datafog-pii, which redirects here). Frozen archive/executable filenames retain their original names and hashes. The older DataFog PII EN 71M is a separate 41-entity custom model and is not the MCP runtime described here.
Pin model revision 02cc6ca86a11dcb2b328861687770a22a7ea1b76 and runtime revision 7cadd41e21f59b2a149e31378d073cf6ca057531 for reproducible downloads. Original tags v0.1.0 and runtime-v0.2.0 remain preserved. The original bundle/ executable is macOS ARM64-only; use the versioned runtime archives for other platforms. The root manifest covers the original model files and current card evidence; the versioned runtime has its own manifest.
Usage
Install a native bundle
Use Python 3.12+ for the download/installer. Native inference itself requires neither Python nor a GPU. While this repository is private, authenticate using hf auth login with an account authorized for DataFog.
python -m pip install huggingface_hub==1.27.0
Select TARGET from the platform table below. This example downloads only that archive and its installer, verifies the expected archive and file hashes, and installs to a new directory.
import json
import subprocess
import sys
from pathlib import Path
from huggingface_hub import hf_hub_download
REPO = "DataFog/pii-en-65m"
REVISION = "7cadd41e21f59b2a149e31378d073cf6ca057531"
PREFIX = "runtimes/v0.2.0"
TARGET = "macos-arm64" # Select your tested platform below.
downloads = Path("datafog-download").resolve()
bundle = Path("datafog-bundle").resolve() # Must not already exist.
def download(name):
return Path(hf_hub_download(
REPO, filename=f"{PREFIX}/{name}", revision=REVISION,
local_dir=downloads, force_download=True,
))
index = json.loads(download("release-index.json").read_text(encoding="utf-8"))
record = next(item for item in index["platforms"] if item["platform"] == TARGET)
archive = download(record["archive"]["filename"])
download("tools/common.py")
installer = download("tools/install_bundle.py")
subprocess.run([
sys.executable, str(installer), "--archive", str(archive),
"--sha256", record["archive"]["sha256"], "--destination", str(bundle),
], check=True)
Run inference
Send one JSON object per line to <bundle>/datafog-pii <bundle> (datafog-pii.exe on Windows). The process can serve multiple requests without reloading the model.
import json
import os
import subprocess
from pathlib import Path
bundle = Path("datafog-bundle").resolve()
executable = bundle / ("datafog-pii.exe" if os.name == "nt" else "datafog-pii")
response = subprocess.run(
[str(executable), str(bundle)],
input=json.dumps({"text": "Contact John Smith at john@example.com."}) + "\n",
text=True, encoding="utf-8", capture_output=True, check=True,
)
result = json.loads(response.stdout)
print(result["redacted"])
# Contact [REDACTED] [REDACTED] at [REDACTED].
print([(f["label"], f["start"], f["end"]) for f in result["findings"]])
# [('first_name', 8, 12), ('last_name', 13, 18), ('email', 22, 38)]
The example output was checked against the released runtime. See the download guide for installation details.
MCP integration
The model supplies PERSON (first/last-name spans) and STREET_ADDRESS to the proposed DataFog MCP integration. Other model categories are ignored by that adapter in favor of Core detectors. Configure the installed absolute path in the MCP policy:
[model]
bundle_directory = "/absolute/path/to/datafog-bundle"
Windows TOML paths can use literal strings, for example bundle_directory = 'C:\DataFog\bundle'. MCP PR #33, stacked on #32, supplies Windows .exe selection. Qualified production/test compatibility revision: 01d7f74e693a7892ab1ed504d0d8f8c1e7ce51d3. Qualification below covers that fixture contract; the final packaged-MCP/model release gate remains part of the coordinated release work.
Input-and-output contract
| Field | Meaning |
|---|---|
Input text |
Entire input string, processed across overlapping windows |
Output findings |
Accepted spans with label, start, end, raw_score, confidence |
start, end |
Half-open UTF-8 byte offsets into the original input; not character indices |
raw_score |
Minimum predicted-class token probability within a span |
confidence |
Pooled calibrated exact-span score, using clipped linear interpolation |
contains_pii |
At least one accepted finding among the seven supported labels |
redacted |
Original text with accepted spans replaced by [REDACTED] |
Class 0 is O; each supported label has B then I in the label order above. The complete mappings are in label-mappings.json. Keep weights, tokenizer, native configuration and calibration together. Malformed requests and unsupported full-window entities return explicit errors. Finite overlap does not cover arbitrarily long entities.
Results
Held-out detection
Evaluation uses 90,018 synthetic Nemotron-PII documents and 248,739 gold entities, with the checkpoint and threshold frozen before final-test evaluation. A true positive requires both the exact label and exact span boundaries. Precision, recall and F1 are percentages; overall scores are micro-averaged across the seven labels.
| Label | Gold entities | Precision | Recall | F1 |
|---|---|---|---|---|
| Overall | 248,739 | 99.01 | 97.63 | 98.32 |
| first_name | 74,498 | 99.40 | 98.30 | 98.85 |
| last_name | 52,864 | 99.31 | 97.96 | 98.63 |
| 49,246 | 98.57 | 97.74 | 98.15 | |
| phone_number | 22,033 | 98.29 | 95.28 | 96.76 |
| street_address | 15,382 | 98.47 | 96.66 | 97.56 |
| customer_id | 19,085 | 99.27 | 98.00 | 98.63 |
| account_number | 15,631 | 98.76 | 96.83 | 97.78 |
Overall: 242,853 TP / 2,424 FP / 5,886 FN. UTF-8 byte-union redaction precision is 99.66%, recall 97.89%, F1 98.76%; that coverage metric ignores labels and weights multibyte characters by byte length. There were 63,160 missed gold bytes and 10,061 unnecessary redaction bytes.
Among 16,233 target-clean documents, 81 had an accepted finding. Among 73,785 target-positive documents, 469 had no accepted finding; 3,947 had at least one missed entity, including partial misses. A document-level detection therefore does not imply complete PII coverage. These scores measure agreement with synthetic annotations, not reviewed private-ticket accuracy. Full test metrics.
Regex comparison
Both systems were evaluated on the same 90,018 documents, restricted to email and phone, the two labels supported by the recorded DataFog 4.8.1 regex baseline within this model's label contract. Combined support is 71,279 gold entities. Other labels are excluded from both sides. Scores are percentages; the better F1 in each row is bold.
| Scope | Model precision | Model recall | Model F1 | Regex precision | Regex recall | Regex F1 |
|---|---|---|---|---|---|---|
| Email + phone | 98.48 | 96.98 | 97.73 | 71.71 | 93.54 | 81.18 |
| 98.57 | 97.74 | 98.15 | 99.38 | 98.99 | 99.19 | |
| phone_number | 98.29 | 95.28 | 96.76 | 40.81 | 81.36 | 54.35 |
The regex baseline performs better on email F1; the model performs better on phone F1 and the combined scoped score. This is a comparison with the recorded datafog==4.8.1 release, not a benchmark of today's Core detector or a GLiNER comparison. Regex metrics, model scoped metrics, baseline version/API provenance.
Support fixtures
Seven hand-authored synthetic documents contain 18 gold entities and cover Unicode signatures, pasted fields, multiline addresses, repeated entities, clean numeric text and long inputs.
| Evaluation | Precision | Recall | F1 | TP / FP / FN |
|---|---|---|---|---|
| Exact typed spans | 93.33% | 77.78% | 84.85% | 14 / 1 / 4 |
| UTF-8 redaction byte coverage | 100.00% | 64.32% | 78.29% | 128 / 0 / 71 bytes |
These cases expose failures despite the high aggregate held-out score: email recovered 2/4 entities, phone 0/1, and street address 0/1. Their small support does not estimate general label accuracy. Full fixture metrics.
Calibration
Calibration fit used 501 separate documents with 1,220 predicted spans (1,172 exact matches); threshold selection used another 532 documents. The threshold maximizes exact typed-span F1 on calibration-select, with ties favoring recall and then the higher threshold. Test reliability uses all 250,771 predicted spans before thresholding, including rejected predictions. Lower Brier and expected calibration error (ECE) are better.
| Score | Brier | ECE |
|---|---|---|
| Raw token-derived span score | 0.014821 | 0.011747 |
| Calibrated exact-span score | 0.013717 | 0.007427 |
Reliability correctness means exact label and boundaries. Missed entities are absent from the prediction population, so recall must be read separately. ECE uses ten equal-width bins and depends on binning. Only 48 incorrect spans supported calibration fitting; sparse confidence regions remain uncertain. This is a population-specific experimental score, not a universal probability or a production threshold.
Runtime performance
These are the original v0.1.0 macOS ARM64 measurements, using native Rust + ONNX Runtime CPU with one inference thread. The recorded environment is macOS 26.6.2 ARM64; the chip model was not recorded. They are single observations, not p50/p95 benchmarks, and do not measure runtime v0.2.0 on all five platforms.
| Input UTF-8 bytes | Warm end-to-end latency (ms) |
|---|---|
| 115 | 6.005 |
| 752 | 20.579 |
| 2,010 | 81.282 |
| 4,502 | 172.720 |
| Resource measurement | Observed value |
|---|---|
| Cold process launch through first empty-input response | 170.013 ms |
| Maximum RSS sampled after responses | 473.4 MiB |
| Original unpacked bundle size | 286.2 MiB |
The four request lengths were selected at observed min/median/p90/max validation byte lengths, not latency percentiles. Warm timings include end-to-end request processing. The RSS sample is not an allocator high-water mark. Cross-platform percentile latency, sustained throughput and large-file performance remain unmeasured. Benchmark artifact, recorded environment.
Native platform qualification
Runtime v0.2.0 preserves the model/configuration/calibration bytes while adding platform library selection and explicit telemetry disabling. Each target passed native unit/boundary tests, installation on a separate fresh runner, seven parity/golden cases and 33 actual MCP runtime/CSV tests with zero failures, errors or skips.
| Archive target | Actual tested environment | MCP tests | Max native/Python logit difference |
|---|---|---|---|
linux-x86_64 |
Linux-6.8.0-1064-azure-x86_64-with-glibc2.35 | 33 passed | 0.0 |
linux-aarch64 |
Linux-6.17.0-1022-azure-aarch64-with-glibc2.39 | 33 passed | 0.0 |
macos-arm64 |
macOS-14.8.9-arm64-arm-64bit | 33 passed | 0.0 |
macos-x86_64 |
macOS-15.7.9-x86_64-i386-64bit | 33 passed | 0.0 |
windows-x86_64 |
Windows-2022Server-10.0.20348-SP0 | 33 passed | 0.0 |
All five matched the Python reference logits bit-for-bit on the seven fixtures and matched the original Mac golden boundaries/labels/redaction. This establishes fixture consistency, not general detection quality or speed. Release index and reports, successful native qualification.
macOS packaging minimums are ARM64 14 and Intel 15; only the named environments were tested. Intel macOS uses a source build of ONNX Runtime commit f2c39fe2f838cf35ce7da92824f5a5e3ee6e88a7, since upstream has no 1.30.0 Intel binary. Windows requires Microsoft's x64 Visual C++ Redistributable (2019-compatible v14). Ubuntu ARM64 evidence does not qualify Raspberry Pi OS, and Server 2022 evidence does not qualify every desktop Windows edition. Other OS versions/distributions, Windows ARM64 and GPU execution remain unqualified.
Training and reproduction
Trained on NVIDIA Nemotron-PII, pinned revision b70ffaf5ff39e079776134c5bf4381f00a9fd1ed. Documents are grouped by UID, exact text and normalized annotated templates to reduce family leakage. Structural filtering and split checks do not eliminate semantic annotation noise or every possible near-duplicate.
| Partition | Documents |
|---|---|
| Training | 89,307 |
| Validation | 514 |
| Calibration fit | 501 |
| Calibration select | 532 |
| Held-out test | 90,018 |
Training used seed 42, float32 MPS, batch size 16, learning rate 2e-5, weight decay 0.01 and a predeclared two-epoch recipe. Total elapsed time was 8,102 seconds. Epoch 2 was selected by minimum validation token cross entropy:
| Epoch | Validation token loss | Release checkpoint |
|---|---|---|
| 1 | 0.003570595682300194 | No |
| 2 | 0.0018816228932983772 | Yes |
No final-test tuning or release-time retraining occurred. Training history, data provenance, partition/source manifest. Inference/evaluation code and dependency locks are in reproduction/; restore the documented workspace layout following provenance/README.md. Exact inference assets are reproducible by checksum; bitwise MPS training repeatability is not promised. Raw datasets, optimizer states and obsolete diagnostic checkpoints are omitted from this release.
Limitations
- Evaluation is synthetic; reviewed English customer-support and other real-world domains remain unvalidated. Labels can be semantically wrong or incomplete despite structural filtering.
- Only seven labels are supported.
contains_pii: falseis not assurance that a document contains no sensitive information or missed supported entities. - The small support-fixture evaluation includes misses in email, phone and street address; aggregate test scores should not hide those failures.
- Tokenization cannot express every character boundary: 16 unrepresentable training-source entity boundaries were recorded. Window overlap cannot cover arbitrary entity lengths.
- The MCP adapter previously timed out on a 100 MB input at a 120-second model deadline. This release does not establish a 100 MB complete-model-scan target or a production throughput guarantee.
- Native qualification covers the reported environments and finite fixture contract. Model inference performance and the final packaged-MCP/model release acceptance are separate checks.
License
DataFog source and trained weights: Apache-2.0. DistilBERT backbone/tokenizer: Apache-2.0. Training data: CC-BY-4.0; attribution to NVIDIA and Amy Steier, Andre Manoel, Alexa Haushalter and Maarten Van Segbroeck, Nemotron-PII (2025), is preserved in NOTICE and the dataset card. ONNX Runtime: MIT with complete third-party notices. The Rust inventory preserves licenses for all 98 locked packages, a conservative target superset. The MCP snapshot retains MIT licensing.
See LICENSE, NOTICE, LICENSING.md and licenses/ for redistribution details. Source and release tooling: DataFog/DataFog-models.
- Downloads last month
- 7
Model tree for DataFog/pii-en-65m
Base model
distilbert/distilbert-base-cased