pii-classifier-go

A ModernBERT-base token classifier that detects PII and credentials in Go source code.

Unlike PII models trained on natural language, this targets the judgment that actually matters in code: telling a real secret from a placeholder, a test fixture, a T-SQL variable, or a license header. That distinction is structural, not lexical.

Labels

6 entity classes, BIO tagged (13 labels): EMAIL, IP, KEY, NAME, PASSWORD, USERNAME.

Usage

from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch

tok = AutoTokenizer.from_pretrained("YOUR_USERNAME/pii-classifier-go")
model = AutoModelForTokenClassification.from_pretrained("YOUR_USERNAME/pii-classifier-go")

code = '''
// maintained by Jane Roe, jane.roe@example-corp.com
const accessKey = "AKIAQ7X4MZLP2VNRT8KD"
'''
enc = tok(code, return_offsets_mapping=True, return_tensors="pt")
offsets = enc.pop("offset_mapping")[0].tolist()
pred = model(**enc).logits.argmax(-1)[0].tolist()

for (s, e), p in zip(offsets, pred):
    if e > s and model.config.id2label[p] != "O":
        print(model.config.id2label[p], repr(code[s:e]))

Trim whitespace and quotes from decoded spans. ModernBERT uses byte-level BPE, which folds a leading space and opening quote into a span's first token. Without trimming, exact-match scoring drops from ~0.99 to ~0.27 while the model is actually correct.

Results

False-positive suppression on human-verified real Go

53 regex false positives, each ruled NOT-PII by a human annotator:

Approach Still flagged Cost
regex baseline 53/53 (100%) β€”
path filter + deleting 2 detectors 13/53 (25%) loses NAME and USERNAME entirely
this model 1/53 (2%) none

Recall on verified positives: 3/3 (n=3, spot check).

Honest decomposition: ~36% of that suppression is achievable with a ten-line test-path filter, and another ~40% by deleting the two worst detectors β€” which buys zero false names by making true names undetectable. The model's genuine contribution is the remaining ~25%: constants like manage_mfa, proxy_user, SignedUser that survive every cheap heuristic and require reading what the string means in context.

Synthetic validation (held-out vocabulary and held-out files)

Entity F1 (exact span)
IP 1.00
EMAIL 0.97
USERNAME 0.96
PASSWORD 0.95
NAME 0.88
KEY 0.86
macro 0.936

NAME is the notable one: a regex baseline finds 0/200 names placed as application data, because names in code appear almost exclusively in attribution contexts β€” which this project's annotation rubric rules as not PII. That capability has no pattern-based equivalent.

Annotation rubric (important β€” it defines the labels)

  • License and author attribution is not PII. // Copyright 2015 Matthew Holt is credit, not personal data.
  • Public GitHub handles are not PII. // owner: @someone is published metadata.
  • A personal email is PII, including inside an attribution comment.
  • Placeholder domains (example.com), reserved IPs (RFC 1918, RFC 5737), env-var names, git SHAs, UUIDs, and published crypto test vectors are all negatives.

If your definition of PII differs, retrain β€” these rulings materially change what the model fires on.

Training data

Fully synthetic PII injected via tree-sitter into permissively licensed public Go repositories (grpc-go, docker-cli Apache-2.0; zap, testify MIT; go-sql-driver/mysql MPL-2.0). 10,000 windows, ~36% hard negatives drawn from false-positive classes measured on real code. Train and validation use disjoint value vocabularies, so a memorised domain list scores at chance.

No proprietary source and no real credentials were used.

Limitations

  • Go only.
  • Recall on real PII is unmeasured. The evaluation corpus (7,229 files across 9 repos) contained ~100 real personal emails and nothing else β€” probes for phone numbers, addresses, dates of birth, SSNs and names-as-data found zero. Well-maintained open source has very little PII. Synthetic recall is not evidence about real recall.
  • Trained on windows that always contained a planted entity, so it has a mild bias toward finding something. Threshold at ~0.9 confidence for production use.
  • IP is the only class that still repeats regex mistakes.
  • Not a replacement for a secret scanner. Use gitleaks or detect-secrets for credential recall; this model is for precision and for the classes patterns cannot reach.

License

Apache-2.0, inherited from ModernBERT-base.

Downloads last month
14
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for saurabhK009/winnow

Finetuned
(1507)
this model