DataFog PII EN 65M

DataFog PII EN 65M is a 65.2M-parameter English token classifier for locating and redacting seven types of personally identifiable information. It returns typed spans, calibrated span confidence, UTF-8 byte offsets and redacted text through a local JSONL runtime. Inference uses the CPU and requires no network access after installation.

The frozen model v0.1.0 achieves 98.32% exact typed-span F1 on 90,018 held-out synthetic documents. Runtime v0.2.0 has qualified native bundles for five platform environments. The checkpoint and acceptance threshold remain experimental; performance on reviewed customer tickets has not been established.

Model · Files · Usage · Input-and-output contract · Results · Training · Limitations · License

Model

Property Value
Backbone distilbert/distilbert-base-cased, pinned revision 6ea81172465e8b0ad3fddeed32b986cdcdcffcf0
Parameters 65,202,447, counted from the released safetensors tensor shapes
Precision Float32 weights and ONNX inference
Task BIO token classification: 15 classes, including O
Supported labels first_name, last_name, email, phone_number, street_address, customer_id, account_number
Long-input processing 512-token windows: 510 content tokens with 42-token overlap
Confidence Pooled isotonic calibration of exact typed-span correctness
Experimental threshold 0.6486486486486487; authoritative value in bundle/calibration.json
Native execution Rust + ONNX Runtime 1.30.0/API27, CPU
Model / runtime versions Model 0.1.0; native runtime 0.2.0; versioned separately

The root safetensors are a standard Transformers token-classification checkpoint. The native runtime additionally supplies complete-input windowing, boundary-fragment suppression, overlap resolution, calibration, thresholding and byte-based redaction. A generic Transformers pipeline does not reproduce that full contract. GLiNER is a separate CI comparison and is not this model's backbone or a release artifact.

Files

File or directory Purpose
model.safetensors, config.json Frozen selected checkpoint and Transformers label mappings
tokenizer.json, tokenizer_config.json, label-mappings.json Tokenizer and explicit class/label contract
bundle/model.onnx, bundle/config.json, bundle/calibration.json Original native-model assets; native configuration differs from root Transformers configuration
runtimes/v0.2.0/*.tar.gz Five platform-specific native bundles, including model assets, executable, matching ORT libraries and notices
runtimes/v0.2.0/release-index.json Exact archive hashes, source/runtime provenance and native qualification reports
runtimes/v0.2.0/tools/ Hash-verifying bundle installer and archive helpers
provenance/artifacts/evaluation/, provenance/REPORT.md Frozen detection, redaction, calibration and support-fixture results
card-evidence/ Original regex comparison summaries and recorded execution environment
reproduction/, provenance/ Training/evaluation code, locks, partition/source provenance and original audit
LICENSE, NOTICE, licenses/, LICENSING.md Model/source licensing, dataset attribution and upstream dependency notices

Canonical Hub ID: DataFog/pii-en-65m (formerly DataFog/datafog-pii, which redirects here). Frozen archive/executable filenames retain their original names and hashes. The older DataFog PII EN 71M is a separate 41-entity custom model and is not the MCP runtime described here.

Pin model revision 02cc6ca86a11dcb2b328861687770a22a7ea1b76 and runtime revision 7cadd41e21f59b2a149e31378d073cf6ca057531 for reproducible downloads. Original tags v0.1.0 and runtime-v0.2.0 remain preserved. The original bundle/ executable is macOS ARM64-only; use the versioned runtime archives for other platforms. The root manifest covers the original model files and current card evidence; the versioned runtime has its own manifest.

Usage

Install a native bundle

Use Python 3.12+ for the download/installer. Native inference itself requires neither Python nor a GPU. While this repository is private, authenticate using hf auth login with an account authorized for DataFog.

python -m pip install huggingface_hub==1.27.0

Select TARGET from the platform table below. This example downloads only that archive and its installer, verifies the expected archive and file hashes, and installs to a new directory.

import json
import subprocess
import sys
from pathlib import Path
from huggingface_hub import hf_hub_download

REPO = "DataFog/pii-en-65m"
REVISION = "7cadd41e21f59b2a149e31378d073cf6ca057531"
PREFIX = "runtimes/v0.2.0"
TARGET = "macos-arm64"  # Select your tested platform below.
downloads = Path("datafog-download").resolve()
bundle = Path("datafog-bundle").resolve()  # Must not already exist.

def download(name):
    return Path(hf_hub_download(
        REPO, filename=f"{PREFIX}/{name}", revision=REVISION,
        local_dir=downloads, force_download=True,
    ))

index = json.loads(download("release-index.json").read_text(encoding="utf-8"))
record = next(item for item in index["platforms"] if item["platform"] == TARGET)
archive = download(record["archive"]["filename"])
download("tools/common.py")
installer = download("tools/install_bundle.py")
subprocess.run([
    sys.executable, str(installer), "--archive", str(archive),
    "--sha256", record["archive"]["sha256"], "--destination", str(bundle),
], check=True)

Run inference

Send one JSON object per line to <bundle>/datafog-pii <bundle> (datafog-pii.exe on Windows). The process can serve multiple requests without reloading the model.

import json
import os
import subprocess
from pathlib import Path

bundle = Path("datafog-bundle").resolve()
executable = bundle / ("datafog-pii.exe" if os.name == "nt" else "datafog-pii")
response = subprocess.run(
    [str(executable), str(bundle)],
    input=json.dumps({"text": "Contact John Smith at john@example.com."}) + "\n",
    text=True, encoding="utf-8", capture_output=True, check=True,
)
result = json.loads(response.stdout)
print(result["redacted"])
# Contact [REDACTED] [REDACTED] at [REDACTED].
print([(f["label"], f["start"], f["end"]) for f in result["findings"]])
# [('first_name', 8, 12), ('last_name', 13, 18), ('email', 22, 38)]

The example output was checked against the released runtime. See the download guide for installation details.

MCP integration

The model supplies PERSON (first/last-name spans) and STREET_ADDRESS to the proposed DataFog MCP integration. Other model categories are ignored by that adapter in favor of Core detectors. Configure the installed absolute path in the MCP policy:

[model]
bundle_directory = "/absolute/path/to/datafog-bundle"

Windows TOML paths can use literal strings, for example bundle_directory = 'C:\DataFog\bundle'. MCP PR #33, stacked on #32, supplies Windows .exe selection. Qualified production/test compatibility revision: 01d7f74e693a7892ab1ed504d0d8f8c1e7ce51d3. Qualification below covers that fixture contract; the final packaged-MCP/model release gate remains part of the coordinated release work.

Input-and-output contract

Field Meaning
Input text Entire input string, processed across overlapping windows
Output findings Accepted spans with label, start, end, raw_score, confidence
start, end Half-open UTF-8 byte offsets into the original input; not character indices
raw_score Minimum predicted-class token probability within a span
confidence Pooled calibrated exact-span score, using clipped linear interpolation
contains_pii At least one accepted finding among the seven supported labels
redacted Original text with accepted spans replaced by [REDACTED]

Class 0 is O; each supported label has B then I in the label order above. The complete mappings are in label-mappings.json. Keep weights, tokenizer, native configuration and calibration together. Malformed requests and unsupported full-window entities return explicit errors. Finite overlap does not cover arbitrarily long entities.

Results

Held-out detection

Evaluation uses 90,018 synthetic Nemotron-PII documents and 248,739 gold entities, with the checkpoint and threshold frozen before final-test evaluation. A true positive requires both the exact label and exact span boundaries. Precision, recall and F1 are percentages; overall scores are micro-averaged across the seven labels.

Label Gold entities Precision Recall F1
Overall 248,739 99.01 97.63 98.32
first_name 74,498 99.40 98.30 98.85
last_name 52,864 99.31 97.96 98.63
email 49,246 98.57 97.74 98.15
phone_number 22,033 98.29 95.28 96.76
street_address 15,382 98.47 96.66 97.56
customer_id 19,085 99.27 98.00 98.63
account_number 15,631 98.76 96.83 97.78

Overall: 242,853 TP / 2,424 FP / 5,886 FN. UTF-8 byte-union redaction precision is 99.66%, recall 97.89%, F1 98.76%; that coverage metric ignores labels and weights multibyte characters by byte length. There were 63,160 missed gold bytes and 10,061 unnecessary redaction bytes.

Among 16,233 target-clean documents, 81 had an accepted finding. Among 73,785 target-positive documents, 469 had no accepted finding; 3,947 had at least one missed entity, including partial misses. A document-level detection therefore does not imply complete PII coverage. These scores measure agreement with synthetic annotations, not reviewed private-ticket accuracy. Full test metrics.

Regex comparison

Both systems were evaluated on the same 90,018 documents, restricted to email and phone, the two labels supported by the recorded DataFog 4.8.1 regex baseline within this model's label contract. Combined support is 71,279 gold entities. Other labels are excluded from both sides. Scores are percentages; the better F1 in each row is bold.

Scope Model precision Model recall Model F1 Regex precision Regex recall Regex F1
Email + phone 98.48 96.98 97.73 71.71 93.54 81.18
email 98.57 97.74 98.15 99.38 98.99 99.19
phone_number 98.29 95.28 96.76 40.81 81.36 54.35

The regex baseline performs better on email F1; the model performs better on phone F1 and the combined scoped score. This is a comparison with the recorded datafog==4.8.1 release, not a benchmark of today's Core detector or a GLiNER comparison. Regex metrics, model scoped metrics, baseline version/API provenance.

Support fixtures

Seven hand-authored synthetic documents contain 18 gold entities and cover Unicode signatures, pasted fields, multiline addresses, repeated entities, clean numeric text and long inputs.

Evaluation Precision Recall F1 TP / FP / FN
Exact typed spans 93.33% 77.78% 84.85% 14 / 1 / 4
UTF-8 redaction byte coverage 100.00% 64.32% 78.29% 128 / 0 / 71 bytes

These cases expose failures despite the high aggregate held-out score: email recovered 2/4 entities, phone 0/1, and street address 0/1. Their small support does not estimate general label accuracy. Full fixture metrics.

Calibration

Calibration fit used 501 separate documents with 1,220 predicted spans (1,172 exact matches); threshold selection used another 532 documents. The threshold maximizes exact typed-span F1 on calibration-select, with ties favoring recall and then the higher threshold. Test reliability uses all 250,771 predicted spans before thresholding, including rejected predictions. Lower Brier and expected calibration error (ECE) are better.

Score Brier ECE
Raw token-derived span score 0.014821 0.011747
Calibrated exact-span score 0.013717 0.007427

Reliability correctness means exact label and boundaries. Missed entities are absent from the prediction population, so recall must be read separately. ECE uses ten equal-width bins and depends on binning. Only 48 incorrect spans supported calibration fitting; sparse confidence regions remain uncertain. This is a population-specific experimental score, not a universal probability or a production threshold.

Runtime performance

These are the original v0.1.0 macOS ARM64 measurements, using native Rust + ONNX Runtime CPU with one inference thread. The recorded environment is macOS 26.6.2 ARM64; the chip model was not recorded. They are single observations, not p50/p95 benchmarks, and do not measure runtime v0.2.0 on all five platforms.

Input UTF-8 bytes Warm end-to-end latency (ms)
115 6.005
752 20.579
2,010 81.282
4,502 172.720
Resource measurement Observed value
Cold process launch through first empty-input response 170.013 ms
Maximum RSS sampled after responses 473.4 MiB
Original unpacked bundle size 286.2 MiB

The four request lengths were selected at observed min/median/p90/max validation byte lengths, not latency percentiles. Warm timings include end-to-end request processing. The RSS sample is not an allocator high-water mark. Cross-platform percentile latency, sustained throughput and large-file performance remain unmeasured. Benchmark artifact, recorded environment.

Native platform qualification

Runtime v0.2.0 preserves the model/configuration/calibration bytes while adding platform library selection and explicit telemetry disabling. Each target passed native unit/boundary tests, installation on a separate fresh runner, seven parity/golden cases and 33 actual MCP runtime/CSV tests with zero failures, errors or skips.

Archive target Actual tested environment MCP tests Max native/Python logit difference
linux-x86_64 Linux-6.8.0-1064-azure-x86_64-with-glibc2.35 33 passed 0.0
linux-aarch64 Linux-6.17.0-1022-azure-aarch64-with-glibc2.39 33 passed 0.0
macos-arm64 macOS-14.8.9-arm64-arm-64bit 33 passed 0.0
macos-x86_64 macOS-15.7.9-x86_64-i386-64bit 33 passed 0.0
windows-x86_64 Windows-2022Server-10.0.20348-SP0 33 passed 0.0

All five matched the Python reference logits bit-for-bit on the seven fixtures and matched the original Mac golden boundaries/labels/redaction. This establishes fixture consistency, not general detection quality or speed. Release index and reports, successful native qualification.

macOS packaging minimums are ARM64 14 and Intel 15; only the named environments were tested. Intel macOS uses a source build of ONNX Runtime commit f2c39fe2f838cf35ce7da92824f5a5e3ee6e88a7, since upstream has no 1.30.0 Intel binary. Windows requires Microsoft's x64 Visual C++ Redistributable (2019-compatible v14). Ubuntu ARM64 evidence does not qualify Raspberry Pi OS, and Server 2022 evidence does not qualify every desktop Windows edition. Other OS versions/distributions, Windows ARM64 and GPU execution remain unqualified.

Training and reproduction

Trained on NVIDIA Nemotron-PII, pinned revision b70ffaf5ff39e079776134c5bf4381f00a9fd1ed. Documents are grouped by UID, exact text and normalized annotated templates to reduce family leakage. Structural filtering and split checks do not eliminate semantic annotation noise or every possible near-duplicate.

Partition Documents
Training 89,307
Validation 514
Calibration fit 501
Calibration select 532
Held-out test 90,018

Training used seed 42, float32 MPS, batch size 16, learning rate 2e-5, weight decay 0.01 and a predeclared two-epoch recipe. Total elapsed time was 8,102 seconds. Epoch 2 was selected by minimum validation token cross entropy:

Epoch Validation token loss Release checkpoint
1 0.003570595682300194 No
2 0.0018816228932983772 Yes

No final-test tuning or release-time retraining occurred. Training history, data provenance, partition/source manifest. Inference/evaluation code and dependency locks are in reproduction/; restore the documented workspace layout following provenance/README.md. Exact inference assets are reproducible by checksum; bitwise MPS training repeatability is not promised. Raw datasets, optimizer states and obsolete diagnostic checkpoints are omitted from this release.

Limitations

  • Evaluation is synthetic; reviewed English customer-support and other real-world domains remain unvalidated. Labels can be semantically wrong or incomplete despite structural filtering.
  • Only seven labels are supported. contains_pii: false is not assurance that a document contains no sensitive information or missed supported entities.
  • The small support-fixture evaluation includes misses in email, phone and street address; aggregate test scores should not hide those failures.
  • Tokenization cannot express every character boundary: 16 unrepresentable training-source entity boundaries were recorded. Window overlap cannot cover arbitrary entity lengths.
  • The MCP adapter previously timed out on a 100 MB input at a 120-second model deadline. This release does not establish a 100 MB complete-model-scan target or a production throughput guarantee.
  • Native qualification covers the reported environments and finite fixture contract. Model inference performance and the final packaged-MCP/model release acceptance are separate checks.

License

DataFog source and trained weights: Apache-2.0. DistilBERT backbone/tokenizer: Apache-2.0. Training data: CC-BY-4.0; attribution to NVIDIA and Amy Steier, Andre Manoel, Alexa Haushalter and Maarten Van Segbroeck, Nemotron-PII (2025), is preserved in NOTICE and the dataset card. ONNX Runtime: MIT with complete third-party notices. The Rust inventory preserves licenses for all 98 locked packages, a conservative target superset. The MCP snapshot retains MIT licensing.

See LICENSE, NOTICE, LICENSING.md and licenses/ for redistribution details. Source and release tooling: DataFog/DataFog-models.

Downloads last month
7
Safetensors
Model size
65.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DataFog/pii-en-65m

Quantized
(8)
this model

Dataset used to train DataFog/pii-en-65m