Instructions to use rasbt/ai-text-detector-distilbert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rasbt/ai-text-detector-distilbert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="rasbt/ai-text-detector-distilbert")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("rasbt/ai-text-detector-distilbert") model = AutoModelForSequenceClassification.from_pretrained("rasbt/ai-text-detector-distilbert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DistilBERT AI-Text Detector
This is a fully fine-tuned DistilBERT classifier for distinguishing human-written and AI-generated text. It was trained on rasbt/human-vs-ai-50k. Human-written text has label 0 and AI-generated text has label 1.
The model uses a maximum sequence length of 512 tokens. Temperature scaling is applied during inference. The recorded best validation accuracy was 99.74%.
Download and use
hf download rasbt/ai-text-detector-distilbert \
--local-dir models/ai-text-detector-distilbert
import json
from pathlib import Path
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_dir = Path("models/ai-text-detector-distilbert")
metadata = json.loads(
(model_dir / "detector-config.json").read_text(encoding="utf-8")
)
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForSequenceClassification.from_pretrained(model_dir)
model.eval()
text = "Paste the text to classify here."
inputs = tokenizer(
text,
truncation=True,
max_length=metadata["max_length"],
return_tensors="pt",
)
with torch.inference_mode():
logits = model(**inputs).logits / metadata["temperature"]
probabilities = logits.float().softmax(dim=-1)
ai_index = metadata["label_mapping"]["ai"]
ai_probability = probabilities[0, ai_index].item()
print({"score": round(100 * ai_probability, 4)})
Test-set confusion matrix
detector-config.json contains the calibration temperature and training metadata. The recommended inference implementation is provided in the rasbt/ai-detector repository.
Related models
- TF-IDF logistic regression
- DistilBERT with LoRA
- DistilBERT with MiCA
- ModernBERT
- GPT-2 with a fixed-position readout
- GPT-2 with a variable-position readout
- Qwen3 0.6B with a fixed-position readout
- Qwen3 0.6B with a variable-position readout
Limitations
Performance may change for text from generators, domains, languages, and editing workflows not represented in the training set. Short or partly AI-assisted text may also be harder to classify. The score should not be treated as definitive evidence that a person did or did not write a text.
- Downloads last month
- -
Model tree for rasbt/ai-text-detector-distilbert
Base model
distilbert/distilbert-base-uncased