Laya v3: High-Throughput Semantic DLP & Credential Detector

Laya v3 is a specialized, production-ready semantic classification model designed for Data Loss Prevention (DLP) and ultra-fast detection of leaked credentials and Personally Identifiable Information (PII) in unstructured text, developer notes, logs, and communication channels.

Trained on top of ModernBERT with reinforcement-calibrated temperature heads, Laya v3 delivers deep contextual understanding of authentication artifacts without the brittleness of traditional regexes or the latency of heavy LLMs.


πŸš€ Key Highlights

  • 100% Recall on Wild Credentials: Zero false negatives observed across isolated hold-out test sets of technical and operational notes.
  • Extreme Throughput: 406.3 requests/second on a single NVIDIA Tesla T4 GPU (Batch 16) with < 2 GB VRAM footprint.
  • Sub-35ms Latency on CPU: Dynamically quantized to INT8 ONNX (laya.int8.onnx), achieving ~33-35ms single-slice latency on modern CPU cores.
  • Early-Exit Sliding Window: Native support for sliding-window document chunking (700 chars / 150 chars overlap) that scans 8,000-character documents in ~30ms via early positive termination.
  • Resilient to Hard Negatives: Hardened against false alarms from code snippets (e.g. PHP password_verify, HTML forms, empty config strings).

πŸ“Š Benchmark & Performance

1. Throughput & Latency (NVIDIA Tesla T4 GPU)

Scenario Sequence Length Batch Size Average Latency Throughput
Short Snippet (Key / Password) 32 tokens 1 17.09 ms 58.5 req/s
Short Snippet 32 tokens 16 39.38 ms 406.3 req/s
Short Snippet 32 tokens 32 82.50 ms 387.9 req/s
Medium Code Block 128 tokens 1 18.80 ms 53.2 req/s
Medium Code Block 128 tokens 16 166.37 ms 96.2 req/s
Long Log Dump 256 tokens 1 26.14 ms 38.3 req/s

Peak GPU Memory: 1.96 GB VRAM (FP32).

2. Accuracy & Capture Metrics (Isolated Hold-Out Corpus)

Metric Laya v2 (Previous) Laya v3 (Production) Net Improvement
Credential Accuracy 62.00% 78.00% +16.00%
Credential Recall (Capture) 100.00% 100.00% Zero Misses
False Positives (Alarms) 38 22 -42.1% (-16 alarms)
PII Accuracy 79.00% 83.00% +4.00%

πŸ› οΈ Usage

Quickstart (CPU with Quantized INT8 ONNX)

import onnxruntime as ort
from laya.onnx_agent import ONNXAgent

# Load model directly from Hugging Face or local path
agent = ONNXAgent(
    "caramelodev/laya-ghostpad-v3",
    onnx_path="laya.int8.onnx"
)

QUESTIONS = {
    "has_credentials": {
        "type": "noul",
        "instructions": "Avalie se este trecho desestruturado contΓ©m credenciais de acesso ou autenticaΓ§Γ£o tΓ©cnica/operacional."
    },
    "has_pii": {
        "type": "noul",
        "instructions": "Avalie se este trecho contΓ©m dados de identificaΓ§Γ£o pessoal ou prontuΓ‘rios de terceiros."
    }
}

text = "ssh root@192.168.1.100 -p 2222\nsenha: SuperSecretPassword2026!"
prediction = agent.predict(text, QUESTIONS)

print("Score Credenciais:", prediction["answers"]["has_credentials"]["noul"])
# Output: > 0.85 (Positive)

Document Scanning with Sliding Window & Early-Exit

For arbitrary-length notes or log dumps:

def scan_document(agent, text: str, chunk_size: int = 700, overlap: int = 150):
    start = 0
    while start < len(text):
        end = min(start + chunk_size, len(text))
        chunk = text[start:end]
        
        result = agent.predict(chunk, QUESTIONS)
        score = result["answers"]["has_credentials"]["noul"]
        
        # Early-exit: Stop immediately upon high-confidence detection
        if score >= 0.65:
            return {"is_leaked": True, "score": score, "chunk_range": (start, end)}
            
        if end == len(text):
            break
        start += (chunk_size - overlap)
        
    return {"is_leaked": False, "score": 0.0}

πŸ“¦ Model Artifacts

  • model.safetensors: Full FP32 weights (PyTorch).
  • laya.int8.onnx: Weight-only dynamic INT8 quantized model for CPU deployment.
  • rl_agent_config.json: Calibrated inference configuration and temperature scalings.
  • tokenizer/: Fast BPE tokenizer directory.
  • encoder/: Base ModernBERT architecture configuration.

πŸ“„ License

Licensed under the Apache License, Version 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support