SporeLabs Scrub

Scrub finds personal details and secrets in text and tells you what each one is: a name, an address, a phone number, an email, a password, a one-time code, an API key. You can replace them with tags, with consistent stand-ins of the same shape, or with asterisks, or just read the list.

It is a 17M-parameter model plus rules, small enough to run on one laptop CPU thread: about 2 ms for a chat message.

Use it

Python package (onnxruntime, no PyTorch):

pip install "sporelabs-scrub @ https://huggingface.co/SporeLabs/scrub/resolve/main/package/sporelabs_scrub-0.1.1-py3-none-any.whl"
from sporelabs_scrub import Scrub

s = Scrub()  # downloads this repository's ONNX model once, then runs offline
s.find("Call Dana Whitlock on +1 415 867 5309")
# [{'start': 5, 'end': 18, 'type': 'PERSON', 'text': 'Dana Whitlock'},
#  {'start': 22, 'end': 37, 'type': 'PHONE', 'text': '+1 415 867 5309'}]

s.scrub(text, mode="tag")        # Call [PERSON] on [PHONE]
s.scrub(text, mode="pseudonym")  # Call Novak Rosa on +1 128 857 1946  (same input, same stand-in)
s.scrub(text, mode="mask")       # Call **** ******** on ** *** *** ****

From the command line: echo "my password is Tq9vLm2x8RwZ" | sporelabs-scrub prints my password is [PASSWORD]; sporelabs-scrub --find prints the findings as JSON.

Download the files without the package:

hf download SporeLabs/scrub --local-dir scrub

SporeLabs API: POST https://sporelabs.dev/v1/scrub, $0.50 per million characters.

curl https://sporelabs.dev/v1/scrub \
  -H "Authorization: Bearer $SPORELABS_API_KEY" -H "Content-Type: application/json" \
  -d '{"text": "Call Dana Whitlock on +1 415 867 5309", "mode": "tag"}'

What it identifies

type what
PERSON people's names, including names that are also words ("Will", "Grace")
ADDRESS street addresses, with unit, postcode and town
PHONE phone numbers in international and national formats, also spoken ("four one five...")
EMAIL email addresses, also encoded or spoken ("dana at fernhollow dot io")
USERNAME personal handles and logins (@dana, ssh dana@host, -u dana)
PATH_USER the user name in a home path (/Users/dana/...)
CREDIT_CARD, BANK, GOV_ID card numbers (Luhn-checked), IBANs, US social security numbers
PASSWORD passwords in prose, commands, config files and connection strings, also dictated
OTP one-time and verification codes, PINs, backup and recovery codes
URL_CREDENTIAL a password or token inside a URL
AWS_KEY, STRIPE_KEY, OPENAI_KEY, ANTHROPIC_KEY, HF_TOKEN, GITHUB_TOKEN, JWT provider keys and tokens, by format
OTHER_SECRET any other credential: private keys, bearer tokens, session cookies, webhook secrets, database passwords

Because each detail comes back with its type, Scrub does more than redact. You can route a support ticket by what it contains, fill CRM or intake fields from free text, list which kinds of personal data a system holds (for a data map or an access request), pseudonymise a dataset so the same person gets the same stand-in everywhere, and keep the rest of the text usable for analytics and AI.

How SporeLabs built it

  • Public text. SporeLabs took text written by people from public datasets: chats (WildChat), SMS messages, spoken assistant requests in nine languages (MASSIVE), forum answers (OpenAssistant), open-source code and its message templates. To these it added two public synthetic PII sets, NVIDIA Nemotron-PII and Gretel PII masking.
  • Real details out, labelled stand-ins in. SporeLabs found the personal details already in that text with its rules and GLiNER2, replaced each with a generated value of the same kind (a name for a name, a valid-looking phone number for a phone number), and labelled it. Text where the detectors disagreed was left out.
  • The places details really turn up. SporeLabs placed generated details into the formats they leak from: command lines, HTTP headers, logs, config and .env files, CSV exports, chat exports, SMS inboxes with sign-in codes, git history, email signatures, and spoken copies of messages as a speech recogniser writes them. It also added look-alikes that are not sensitive (example keys, test card numbers, hashes, version strings) so the model learns to leave them alone.
  • Model. SporeLabs fine-tuned the Ettin 17M encoder (Johns Hopkins University, MIT licence) with one output per type per token, on about 305,000 documents in one pass, and picked each type's threshold on held-out data.
  • Rules. Anything with a format (provider keys, tokens, card numbers, codes after a code word, emails) is found by rules, which take precedence; the model finds what has no format, such as names and addresses. A value found once is found everywhere it repeats in the text.

Results

SporeLabs Scrub benchmark, held-out set test (100 documents of chats, emails, tickets, logs, agent transcripts and config, 135 secrets, 504 direct identifiers, 798 look-alikes). "Caught" means every letter and digit of the item was covered. Each system ran with its own default settings.

system size secrets caught direct identifiers caught precision look-alikes scrubbed (lower is better)
SporeLabs Scrub 17M + rules 0.918 0.921 0.810 0.094
OpenAI Privacy Filter 1.4B 0.852 0.770 0.773 0.175
GLiNER2-PII 307M 0.770 0.873 0.741 0.283
AWS Comprehend cloud API 0.385 0.855 0.781 0.148
Presidio spaCy lg + rules 0.074 0.544 0.578 0.167

The development sets, used to tune the rules (SporeLabs Scrub):

set documents secrets caught (n) direct identifiers caught (n) precision look-alikes scrubbed
extra1 200 0.989 (265) 0.932 (723) 0.844 0.081
extra2 200 0.987 (477) 0.951 (740) 0.840 0.079
extra3 150 0.919 (308) 0.922 (561) 0.772 0.098
dev 160 0.986 (216) 0.954 (500) 0.878 0.046

Scrub is trained for the text that leaks from apps and developer tools: chats, messages, tickets, logs, code and config. General-purpose public PII sets are mostly forms, records and documents, and many also label dates, places and organisations, which Scrub leaves alone by design. On the personal details Scrub covers, it catches more than Presidio on most of them; GLiNER2-PII catches more than Scrub on each of the four SporeLabs ran it on. Every number is in the benchmark's results/results.json.

Speed of the sporelabs-scrub package (ONNX, rules included), one CPU thread on an Apple M3:

input fp32 int8
chat message (median 42 characters) 2.1 ms 1.4 ms
document (median 665 characters) 34.8 ms 23.7 ms

Full results: SporeLabs/scrub-benchmark.

Intended use and limits

Scrub is a first pass that removes most personal details and secrets before text is stored, shared, logged or used for training, and it labels what it found. It is not a guarantee: it misses some details, so text with legal or safety stakes needs a person or a second system to review it. It was built and tested on English text; it handles names, addresses and phone numbers from many countries, but other languages are less tested. It does not look for dates, organisations or places on their own, since in most text those are not personal.

Files

file what
onnx/model.onnx the model, fp32, 67 MB (the package default)
onnx/model_int8.onnx int8, 19 MB; slightly lower recall on names (see PARITY.md)
scrub_config.json, tokenizer.json types, thresholds, windowing; the tokenizer
encoder/, heads.safetensors, thresholds.json, kit_model.json the same model as PyTorch weights
package/ the sporelabs-scrub wheel and source

Licences: weights Apache-2.0 (LICENSE), code MIT (LICENSE-CODE); credits in NOTICE.md. Training data: SporeLabs/scrub-data. Benchmark: SporeLabs/scrub-benchmark.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SporeLabs/scrub

Quantized
(9)
this model