pii-classifier-go
A ModernBERT-base token classifier that detects PII and credentials in Go source code.
Unlike PII models trained on natural language, this targets the judgment that actually matters in code: telling a real secret from a placeholder, a test fixture, a T-SQL variable, or a license header. That distinction is structural, not lexical.
Labels
6 entity classes, BIO tagged (13 labels): EMAIL, IP, KEY, NAME,
PASSWORD, USERNAME.
Usage
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
tok = AutoTokenizer.from_pretrained("YOUR_USERNAME/pii-classifier-go")
model = AutoModelForTokenClassification.from_pretrained("YOUR_USERNAME/pii-classifier-go")
code = '''
// maintained by Jane Roe, jane.roe@example-corp.com
const accessKey = "AKIAQ7X4MZLP2VNRT8KD"
'''
enc = tok(code, return_offsets_mapping=True, return_tensors="pt")
offsets = enc.pop("offset_mapping")[0].tolist()
pred = model(**enc).logits.argmax(-1)[0].tolist()
for (s, e), p in zip(offsets, pred):
if e > s and model.config.id2label[p] != "O":
print(model.config.id2label[p], repr(code[s:e]))
Trim whitespace and quotes from decoded spans. ModernBERT uses byte-level BPE, which folds a leading space and opening quote into a span's first token. Without trimming, exact-match scoring drops from ~0.99 to ~0.27 while the model is actually correct.
Results
False-positive suppression on human-verified real Go
53 regex false positives, each ruled NOT-PII by a human annotator:
| Approach | Still flagged | Cost |
|---|---|---|
| regex baseline | 53/53 (100%) | β |
| path filter + deleting 2 detectors | 13/53 (25%) | loses NAME and USERNAME entirely |
| this model | 1/53 (2%) | none |
Recall on verified positives: 3/3 (n=3, spot check).
Honest decomposition: ~36% of that suppression is achievable with a
ten-line test-path filter, and another ~40% by deleting the two worst
detectors β which buys zero false names by making true names undetectable. The
model's genuine contribution is the remaining ~25%: constants like
manage_mfa, proxy_user, SignedUser that survive every cheap heuristic and
require reading what the string means in context.
Synthetic validation (held-out vocabulary and held-out files)
| Entity | F1 (exact span) |
|---|---|
| IP | 1.00 |
| 0.97 | |
| USERNAME | 0.96 |
| PASSWORD | 0.95 |
| NAME | 0.88 |
| KEY | 0.86 |
| macro | 0.936 |
NAME is the notable one: a regex baseline finds 0/200 names placed as
application data, because names in code appear almost exclusively in
attribution contexts β which this project's annotation rubric rules as not
PII. That capability has no pattern-based equivalent.
Annotation rubric (important β it defines the labels)
- License and author attribution is not PII.
// Copyright 2015 Matthew Holtis credit, not personal data. - Public GitHub handles are not PII.
// owner: @someoneis published metadata. - A personal email is PII, including inside an attribution comment.
- Placeholder domains (
example.com), reserved IPs (RFC 1918, RFC 5737), env-var names, git SHAs, UUIDs, and published crypto test vectors are all negatives.
If your definition of PII differs, retrain β these rulings materially change what the model fires on.
Training data
Fully synthetic PII injected via tree-sitter into permissively licensed public Go repositories (grpc-go, docker-cli Apache-2.0; zap, testify MIT; go-sql-driver/mysql MPL-2.0). 10,000 windows, ~36% hard negatives drawn from false-positive classes measured on real code. Train and validation use disjoint value vocabularies, so a memorised domain list scores at chance.
No proprietary source and no real credentials were used.
Limitations
- Go only.
- Recall on real PII is unmeasured. The evaluation corpus (7,229 files across 9 repos) contained ~100 real personal emails and nothing else β probes for phone numbers, addresses, dates of birth, SSNs and names-as-data found zero. Well-maintained open source has very little PII. Synthetic recall is not evidence about real recall.
- Trained on windows that always contained a planted entity, so it has a mild bias toward finding something. Threshold at ~0.9 confidence for production use.
IPis the only class that still repeats regex mistakes.- Not a replacement for a secret scanner. Use gitleaks or detect-secrets for credential recall; this model is for precision and for the classes patterns cannot reach.
License
Apache-2.0, inherited from ModernBERT-base.
- Downloads last month
- 14
Model tree for saurabhK009/winnow
Base model
answerdotai/ModernBERT-base