spidercob/dlp-intent-classifier

A SetFit model that classifies DLP (Data Loss Prevention) regex match findings into 4 intent categories โ€” distinguishing genuine sensitive data from false positives.

The Problem

DLP engines use regex patterns to detect PII and secrets (SSNs, credit cards, API keys, emails, etc.). These patterns generate large numbers of false positives:

  • A CSS rule like z-index: 100-200-300 matches an SSN pattern
  • A Stripe test card 4242424242424242 in a README matches a credit card pattern
  • A code comment # Set OPENAI_API_KEY=sk-xxxx in .env matches an API key pattern

Without intent classification, every regex hit triggers an alert. This model reads the surrounding context and decides whether the match is real.

Labels

Label Description Recommended Action
REAL_DATA Genuine PII, credential, or secret BLOCK (conf > 0.7)
TEST_DATA Test/mock/sandbox value ALLOW (conf > 0.8)
DOCUMENTATION Regex hit inside a comment, README, or docstring ALLOW
NOISE False positive โ€” low-signal pattern match IGNORE

Quick Start

from setfit import SetFitModel

model = SetFitModel.from_pretrained("spidercob/dlp-intent-classifier")

examples = [
    "Context surrounding a ssn: Patient record: Maria Garcia SSN=523-89-4521 DOB=1975-03-12",
    "Context surrounding a credit_card: Stripe test card 4242424242424242 in checkout flow test",
    "Context surrounding a api_key: # Set OPENAI_API_KEY=sk-xxxx in .env before running",
    "Context surrounding a ssn: CSS z-index: 100-200-300 matched SSN pattern",
]

predictions = model.predict(examples)
# ['REAL_DATA', 'TEST_DATA', 'DOCUMENTATION', 'NOISE']

probabilities = model.predict_proba(examples)
# Shape: (4, 4) โ€” confidence per class

Input Format

Context surrounding a {finding_type}: {context_text}
  • finding_type: the DLP pattern that matched โ€” e.g. ssn, credit_card, aws_access_key, email, api_key, password, phone, ip_address
  • context_text: surrounding content snippet, up to ~400 characters

Integration Pattern

from setfit import SetFitModel

model = SetFitModel.from_pretrained("spidercob/dlp-intent-classifier")

def should_block(finding_type: str, context: str) -> dict:
    text = f"Context surrounding a {finding_type}: {context}"
    label = model.predict([text])[0]
    probs = model.predict_proba([text])[0]
    conf = max(probs)

    if label == "REAL_DATA" and conf > 0.7:
        return {"action": "BLOCK", "label": label, "confidence": conf}
    elif label == "TEST_DATA" and conf > 0.8:
        return {"action": "ALLOW", "label": label, "confidence": conf}
    elif label == "DOCUMENTATION":
        return {"action": "ALLOW", "label": label, "confidence": conf}
    elif label == "NOISE":
        return {"action": "IGNORE", "label": label, "confidence": conf}
    else:
        return {"action": "REVIEW", "label": label, "confidence": conf}

# Example
result = should_block("ssn", "Patient record: Maria Garcia SSN=523-89-4521")
# {"action": "BLOCK", "label": "REAL_DATA", "confidence": 0.99}

Model Details

  • Base model: sentence-transformers/all-MiniLM-L6-v2 (22.7M params, 6 BERT layers, 384-dim embeddings)
  • Method: SetFit โ€” contrastive fine-tuning of sentence transformer + logistic regression head
  • Architecture: GELU activations throughout; tanh only in pooler layer
  • Test accuracy: 99.7% on held-out stratified 20% split
  • Training time: ~2.5 hours on Apple M-series (MPS)

Training Data

2,000+ examples across 4 balanced classes:

Source Count Label
Production DLP scan findings (Spidercob) ~256 Mixed (auto-labeled)
ai4privacy/pii-masking-300k 500 REAL_DATA
Faker-generated test fixtures 500 TEST_DATA
bigcode/the-stack-smol code comments + synthetic 500 DOCUMENTATION
Generated false-positive patterns 500 NOISE

Fine-tune on Your Own Data

See the training pipeline at github.com/SpiderCob/dlp-intent-classifier.

git clone https://github.com/SpiderCob/dlp-intent-classifier
cd dlp-intent-classifier
pip install -r requirements.txt

# Export your DLP findings, augment, and retrain
python scripts/export_training_data.py   # pull from your DB
python scripts/augment_training_data.py  # add synthetic examples
python scripts/build_public_dataset.py   # fetch public datasets
python scripts/merge_and_retrain.py      # merge + fine-tune

About

Built by Spidercob โ€” enterprise DLP SaaS. This model powers the false-positive reduction layer in Spidercob's DLP engine, reducing alert fatigue while maintaining high sensitivity to real data leaks.

License

Apache 2.0

Downloads last month
47
Safetensors
Model size
22.7M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Evaluation results

  • Accuracy on Mixed (production DLP findings + ai4privacy + Faker + The Stack)
    test set self-reported
    0.997