SporeLabs Scrub
Scrub finds personal details and secrets in text and tells you what each one is: a name, an address, a phone number, an email, a password, a one-time code, an API key. You can replace them with tags, with consistent stand-ins of the same shape, or with asterisks, or just read the list.
It is a 17M-parameter model plus rules, small enough to run on one laptop CPU thread: about 2 ms for a chat message.
Use it
Python package (onnxruntime, no PyTorch):
pip install "sporelabs-scrub @ https://huggingface.co/SporeLabs/scrub/resolve/main/package/sporelabs_scrub-0.1.1-py3-none-any.whl"
from sporelabs_scrub import Scrub
s = Scrub() # downloads this repository's ONNX model once, then runs offline
s.find("Call Dana Whitlock on +1 415 867 5309")
# [{'start': 5, 'end': 18, 'type': 'PERSON', 'text': 'Dana Whitlock'},
# {'start': 22, 'end': 37, 'type': 'PHONE', 'text': '+1 415 867 5309'}]
s.scrub(text, mode="tag") # Call [PERSON] on [PHONE]
s.scrub(text, mode="pseudonym") # Call Novak Rosa on +1 128 857 1946 (same input, same stand-in)
s.scrub(text, mode="mask") # Call **** ******** on ** *** *** ****
From the command line: echo "my password is Tq9vLm2x8RwZ" | sporelabs-scrub prints my password is [PASSWORD];
sporelabs-scrub --find prints the findings as JSON.
Download the files without the package:
hf download SporeLabs/scrub --local-dir scrub
SporeLabs API: POST https://sporelabs.dev/v1/scrub, $0.50 per million characters.
curl https://sporelabs.dev/v1/scrub \
-H "Authorization: Bearer $SPORELABS_API_KEY" -H "Content-Type: application/json" \
-d '{"text": "Call Dana Whitlock on +1 415 867 5309", "mode": "tag"}'
What it identifies
| type | what |
|---|---|
PERSON |
people's names, including names that are also words ("Will", "Grace") |
ADDRESS |
street addresses, with unit, postcode and town |
PHONE |
phone numbers in international and national formats, also spoken ("four one five...") |
EMAIL |
email addresses, also encoded or spoken ("dana at fernhollow dot io") |
USERNAME |
personal handles and logins (@dana, ssh dana@host, -u dana) |
PATH_USER |
the user name in a home path (/Users/dana/...) |
CREDIT_CARD, BANK, GOV_ID |
card numbers (Luhn-checked), IBANs, US social security numbers |
PASSWORD |
passwords in prose, commands, config files and connection strings, also dictated |
OTP |
one-time and verification codes, PINs, backup and recovery codes |
URL_CREDENTIAL |
a password or token inside a URL |
AWS_KEY, STRIPE_KEY, OPENAI_KEY, ANTHROPIC_KEY, HF_TOKEN, GITHUB_TOKEN, JWT |
provider keys and tokens, by format |
OTHER_SECRET |
any other credential: private keys, bearer tokens, session cookies, webhook secrets, database passwords |
Because each detail comes back with its type, Scrub does more than redact. You can route a support ticket by what it contains, fill CRM or intake fields from free text, list which kinds of personal data a system holds (for a data map or an access request), pseudonymise a dataset so the same person gets the same stand-in everywhere, and keep the rest of the text usable for analytics and AI.
How SporeLabs built it
- Public text. SporeLabs took text written by people from public datasets: chats (WildChat), SMS messages, spoken assistant requests in nine languages (MASSIVE), forum answers (OpenAssistant), open-source code and its message templates. To these it added two public synthetic PII sets, NVIDIA Nemotron-PII and Gretel PII masking.
- Real details out, labelled stand-ins in. SporeLabs found the personal details already in that text with its rules and GLiNER2, replaced each with a generated value of the same kind (a name for a name, a valid-looking phone number for a phone number), and labelled it. Text where the detectors disagreed was left out.
- The places details really turn up. SporeLabs placed generated details into the formats they leak from: command
lines, HTTP headers, logs, config and
.envfiles, CSV exports, chat exports, SMS inboxes with sign-in codes, git history, email signatures, and spoken copies of messages as a speech recogniser writes them. It also added look-alikes that are not sensitive (example keys, test card numbers, hashes, version strings) so the model learns to leave them alone. - Model. SporeLabs fine-tuned the Ettin 17M encoder (Johns Hopkins University, MIT licence) with one output per type per token, on about 305,000 documents in one pass, and picked each type's threshold on held-out data.
- Rules. Anything with a format (provider keys, tokens, card numbers, codes after a code word, emails) is found by rules, which take precedence; the model finds what has no format, such as names and addresses. A value found once is found everywhere it repeats in the text.
Results
SporeLabs Scrub benchmark, held-out set test (100 documents of chats, emails, tickets, logs, agent transcripts
and config, 135 secrets, 504 direct identifiers, 798 look-alikes). "Caught" means every letter and
digit of the item was covered. Each system ran with its own default settings.
| system | size | secrets caught | direct identifiers caught | precision | look-alikes scrubbed (lower is better) |
|---|---|---|---|---|---|
| SporeLabs Scrub | 17M + rules | 0.918 | 0.921 | 0.810 | 0.094 |
| OpenAI Privacy Filter | 1.4B | 0.852 | 0.770 | 0.773 | 0.175 |
| GLiNER2-PII | 307M | 0.770 | 0.873 | 0.741 | 0.283 |
| AWS Comprehend | cloud API | 0.385 | 0.855 | 0.781 | 0.148 |
| Presidio | spaCy lg + rules | 0.074 | 0.544 | 0.578 | 0.167 |
The development sets, used to tune the rules (SporeLabs Scrub):
| set | documents | secrets caught (n) | direct identifiers caught (n) | precision | look-alikes scrubbed |
|---|---|---|---|---|---|
extra1 |
200 | 0.989 (265) | 0.932 (723) | 0.844 | 0.081 |
extra2 |
200 | 0.987 (477) | 0.951 (740) | 0.840 | 0.079 |
extra3 |
150 | 0.919 (308) | 0.922 (561) | 0.772 | 0.098 |
dev |
160 | 0.986 (216) | 0.954 (500) | 0.878 | 0.046 |
Scrub is trained for the text that leaks from apps and developer tools: chats, messages, tickets, logs, code and
config. General-purpose public PII sets are mostly forms, records and documents, and many also label dates, places and
organisations, which Scrub leaves alone by design. On the personal details Scrub covers, it catches more than Presidio
on most of them; GLiNER2-PII catches more than Scrub on each of the four SporeLabs ran it on. Every number is in the
benchmark's results/results.json.
Speed of the sporelabs-scrub package (ONNX, rules included), one CPU thread on an Apple M3:
| input | fp32 | int8 |
|---|---|---|
| chat message (median 42 characters) | 2.1 ms | 1.4 ms |
| document (median 665 characters) | 34.8 ms | 23.7 ms |
Full results: SporeLabs/scrub-benchmark.
Intended use and limits
Scrub is a first pass that removes most personal details and secrets before text is stored, shared, logged or used for training, and it labels what it found. It is not a guarantee: it misses some details, so text with legal or safety stakes needs a person or a second system to review it. It was built and tested on English text; it handles names, addresses and phone numbers from many countries, but other languages are less tested. It does not look for dates, organisations or places on their own, since in most text those are not personal.
Files
| file | what |
|---|---|
onnx/model.onnx |
the model, fp32, 67 MB (the package default) |
onnx/model_int8.onnx |
int8, 19 MB; slightly lower recall on names (see PARITY.md) |
scrub_config.json, tokenizer.json |
types, thresholds, windowing; the tokenizer |
encoder/, heads.safetensors, thresholds.json, kit_model.json |
the same model as PyTorch weights |
package/ |
the sporelabs-scrub wheel and source |
Licences: weights Apache-2.0 (LICENSE), code MIT (LICENSE-CODE); credits in NOTICE.md.
Training data: SporeLabs/scrub-data. Benchmark:
SporeLabs/scrub-benchmark.
Model tree for SporeLabs/scrub
Base model
jhu-clsp/ettin-encoder-17m