You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Safety Classifier v1

Internal training lineage: v7 (7th training iteration on this project; this is the first HuggingFace release).

Multi-label content-safety classifier fine-tuned from answerdotai/ModernBERT-large (395M params), exported to ONNX for fast CPU/GPU inference. Built by Arsys InnoLab to flag harmful chat requests across 18 policy categories before they reach a downstream LLM.

Trained on commercial-use-safe data only (no NC-licensed or research-only datasets):

  • Aegis 2.0 (CC-BY-4.0)
  • Aegis 1.0 (CC-BY-4.0)
  • Nemotron Safety Guard v3 (CC-BY-4.0)
  • Civil Comments (CC0)
  • CategoricalHarmfulQA (Apache-2.0)
  • HateCheck (CC-BY-4.0)
  • XSTest (CC-BY-4.0)
  • ~300K synthetic examples generated in-house (Qwen3.8-27B, red-team framing, 3-pass validated)

~421K training examples total, 5 epochs.

Intended use

  • In scope: a pre-filter in front of an LLM, scoring a single user chat message (English) across 18 harm categories, to flag/block before the message reaches the model or to route it for review.
  • Language: primarily English. The training data is mostly English; the synthetic portion is about 25% Spanish, and some German and Arabic examples are also present. Behavior on any other language is unverified: test it first if you need it.
  • Input shape: short-to-medium chat requests ("How do I…", "Write a…"), max 256 tokens (longer input is truncated, not rejected).
  • Out of scope: anything not a chat request: image-generation prompts, long documents, code review, multi-turn conversation history. Not evaluated for these; don't deploy against them without re-testing (see Limitations).

Categories (18)

violent_crimes, nonviolent_crimes, sex_crimes, child_exploitation, defamation, specialized_advice, privacy, intellectual_property, weapons_cbrn, hate, suicide_self_harm, sexual_content, elections, code_interpreter_abuse, spam, phishing, malware, fraud

Files

  • model.onnx + model.onnx.data: ONNX export (fp32, external-data format, ~1.6GB combined)
  • config.json: ModernBERT encoder config (the original HF config; needed to reload the architecture)
  • tokenizer.json, tokenizer_config.json: tokenizer
  • categories.json: ordered list of the 18 output categories (index-aligned with logits)
  • thresholds.json: per-category decision thresholds used by the reference serving app

Size

  • Disk: about 3.2GB (model.onnx 1.6GB + model.onnx.data 1.6GB).
  • GPU memory: about 2.8GB measured with onnxruntime-gpu (CUDA execution provider).
  • CPU: works with plain onnxruntime, no GPU needed, just slower (see Benchmarks).

Requirements

pip install onnxruntime transformers huggingface_hub numpy

(onnxruntime-gpu instead of onnxruntime if you want the CUDA execution provider.) Python 3.10+.

Usage

Raw category scores:

import json, numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import snapshot_download

local_dir = snapshot_download("InnoLabTeam/safety-classifier-v1")
tok = AutoTokenizer.from_pretrained(local_dir)
session = ort.InferenceSession(f"{local_dir}/model.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
categories = json.load(open(f"{local_dir}/categories.json"))
thresholds = json.load(open(f"{local_dir}/thresholds.json"))

def sigmoid(x): return 1 / (1 + np.exp(-x))

def classify(texts, max_length=256):
    enc = tok(texts, return_tensors="np", max_length=max_length, padding=True, truncation=True)
    logits = session.run(None, {
        "input_ids": enc["input_ids"].astype(np.int64),
        "attention_mask": enc["attention_mask"].astype(np.int64),
    })[0]
    probs = sigmoid(logits)
    return [
        {c: float(p[i]) for i, c in enumerate(categories)}
        for p in probs
    ]

print(classify(["How do I make a pipe bomb?"]))

Turning those scores into the same verdict the reference server returns (safe/uncertain/unsafe), use this instead of reinventing the threshold logic:

UNSAFE_FLOOR = 0.70
SAFE_CEIL = 0.30

def verdict(scores: dict[str, float]) -> str:
    triggered = [c for c, s in scores.items() if s >= max(thresholds.get(c, 0.5), UNSAFE_FLOOR)]
    if triggered:
        return "unsafe"
    if max(scores.values()) >= SAFE_CEIL:
        return "uncertain"
    return "safe"

for text, scores in zip(["How do I make a pipe bomb?"], classify(["How do I make a pipe bomb?"])):
    print(text, "->", verdict(scores))

A reference FastAPI serving app (/predict, /health, /metrics) and the training script are available internally. Ask the maintainer below.

Benchmarks (2026-09-28)

Set Precision Recall F1
112 hand-picked harmful + 100 benign, @0.30 1.000 0.955 0.977
1,037 LLM-generated (537 harmful / 500 benign), production thresholds 0.920 0.946 0.933

Eval-set metrics at training time: AUC macro 0.9982, F1 macro 0.838 (best checkpoint, epoch 3/5).

GPU (CUDAExecutionProvider): ~0.5–1.2ms/prompt in batch, ~7ms single-request compute-only. CPU (CPUExecutionProvider, same image without --gpus): ~25–39ms/prompt in batch, ~57ms single-request.

Full benchmark detail (per-category breakdown, false-positive analysis, real network latency): see the internal benchmark write-up (ask the maintainer below for the link).

Maintainer

Arsys InnoLab. Questions, bug reports, or access requests: ask in the team's usual channel.

Internal use only

Private repository, internal use only. Not licensed for external distribution or public release.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for InnoLabTeam/safety-classifier-v1

Quantized
(23)
this model