Safety Classifier v1
Internal training lineage: v7 (7th training iteration on this project; this is the first HuggingFace release).
Multi-label content-safety classifier fine-tuned from answerdotai/ModernBERT-large (395M params),
exported to ONNX for fast CPU/GPU inference. Built by Arsys InnoLab to flag harmful chat requests
across 18 policy categories before they reach a downstream LLM.
Trained on commercial-use-safe data only (no NC-licensed or research-only datasets):
- Aegis 2.0 (CC-BY-4.0)
- Aegis 1.0 (CC-BY-4.0)
- Nemotron Safety Guard v3 (CC-BY-4.0)
- Civil Comments (CC0)
- CategoricalHarmfulQA (Apache-2.0)
- HateCheck (CC-BY-4.0)
- XSTest (CC-BY-4.0)
- ~300K synthetic examples generated in-house (Qwen3.8-27B, red-team framing, 3-pass validated)
~421K training examples total, 5 epochs.
Intended use
- In scope: a pre-filter in front of an LLM, scoring a single user chat message (English) across 18 harm categories, to flag/block before the message reaches the model or to route it for review.
- Language: primarily English. The training data is mostly English; the synthetic portion is about 25% Spanish, and some German and Arabic examples are also present. Behavior on any other language is unverified: test it first if you need it.
- Input shape: short-to-medium chat requests ("How do I…", "Write a…"), max 256 tokens (longer input is truncated, not rejected).
- Out of scope: anything not a chat request: image-generation prompts, long documents, code review, multi-turn conversation history. Not evaluated for these; don't deploy against them without re-testing (see Limitations).
Categories (18)
violent_crimes, nonviolent_crimes, sex_crimes, child_exploitation, defamation,
specialized_advice, privacy, intellectual_property, weapons_cbrn, hate, suicide_self_harm,
sexual_content, elections, code_interpreter_abuse, spam, phishing, malware, fraud
Files
model.onnx+model.onnx.data: ONNX export (fp32, external-data format, ~1.6GB combined)config.json: ModernBERT encoder config (the original HF config; needed to reload the architecture)tokenizer.json,tokenizer_config.json: tokenizercategories.json: ordered list of the 18 output categories (index-aligned with logits)thresholds.json: per-category decision thresholds used by the reference serving app
Size
- Disk: about 3.2GB (
model.onnx1.6GB +model.onnx.data1.6GB). - GPU memory: about 2.8GB measured with
onnxruntime-gpu(CUDA execution provider). - CPU: works with plain
onnxruntime, no GPU needed, just slower (see Benchmarks).
Requirements
pip install onnxruntime transformers huggingface_hub numpy
(onnxruntime-gpu instead of onnxruntime if you want the CUDA execution provider.) Python 3.10+.
Usage
Raw category scores:
import json, numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import snapshot_download
local_dir = snapshot_download("InnoLabTeam/safety-classifier-v1")
tok = AutoTokenizer.from_pretrained(local_dir)
session = ort.InferenceSession(f"{local_dir}/model.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
categories = json.load(open(f"{local_dir}/categories.json"))
thresholds = json.load(open(f"{local_dir}/thresholds.json"))
def sigmoid(x): return 1 / (1 + np.exp(-x))
def classify(texts, max_length=256):
enc = tok(texts, return_tensors="np", max_length=max_length, padding=True, truncation=True)
logits = session.run(None, {
"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64),
})[0]
probs = sigmoid(logits)
return [
{c: float(p[i]) for i, c in enumerate(categories)}
for p in probs
]
print(classify(["How do I make a pipe bomb?"]))
Turning those scores into the same verdict the reference server returns (safe/uncertain/unsafe),
use this instead of reinventing the threshold logic:
UNSAFE_FLOOR = 0.70
SAFE_CEIL = 0.30
def verdict(scores: dict[str, float]) -> str:
triggered = [c for c, s in scores.items() if s >= max(thresholds.get(c, 0.5), UNSAFE_FLOOR)]
if triggered:
return "unsafe"
if max(scores.values()) >= SAFE_CEIL:
return "uncertain"
return "safe"
for text, scores in zip(["How do I make a pipe bomb?"], classify(["How do I make a pipe bomb?"])):
print(text, "->", verdict(scores))
A reference FastAPI serving app (/predict, /health, /metrics) and the training script are
available internally. Ask the maintainer below.
Benchmarks (2026-09-28)
| Set | Precision | Recall | F1 |
|---|---|---|---|
| 112 hand-picked harmful + 100 benign, @0.30 | 1.000 | 0.955 | 0.977 |
| 1,037 LLM-generated (537 harmful / 500 benign), production thresholds | 0.920 | 0.946 | 0.933 |
Eval-set metrics at training time: AUC macro 0.9982, F1 macro 0.838 (best checkpoint, epoch 3/5).
GPU (CUDAExecutionProvider): ~0.5–1.2ms/prompt in batch, ~7ms single-request compute-only.
CPU (CPUExecutionProvider, same image without --gpus): ~25–39ms/prompt in batch, ~57ms single-request.
Full benchmark detail (per-category breakdown, false-positive analysis, real network latency): see the internal benchmark write-up (ask the maintainer below for the link).
Maintainer
Arsys InnoLab. Questions, bug reports, or access requests: ask in the team's usual channel.
Internal use only
Private repository, internal use only. Not licensed for external distribution or public release.
- Downloads last month
- -
Model tree for InnoLabTeam/safety-classifier-v1
Base model
answerdotai/ModernBERT-large