⚠️ Devseis AI Act Classifier — v5 (Experiment, Not Legal Advice)

Newer version: Devseis AI Act Classifier v6 is now the model running in Caveat. It is better on banned practices, unusual wording and synthetic media.

A Devseis model, built on LEGAL-BERT-SMALL.

This model does not give legal advice, and its output must not be used on its own to decide an AI system's EU AI Act tier. It powers the in-browser half of Devseis/Caveat, a demo that shows where automated classification works and where it fails.

It reads a free-text description of an AI system and returns one of four tiers:

Label Meaning
prohibited Art 5 banned practice
high_risk Annex III use case, or Annex I product safety component
limited_risk Art 50 transparency duties (chatbots and voice agents, deepfakes, AI-written public text)
minimal_risk none of the above

General-purpose AI (GPAI) models are decided by a rule, not by this model. In Caveat, a separate rule (gpai_rule.js) runs first. If the text describes a provider training or releasing a general-purpose model, the rule decides gpai or gpai_systemic_risk from the compute figure (Art 51(2), > 10^25 FLOPs) or a Commission designation (Art 51(1)(b)), and this model is not used. Small text classifiers read exponents unreliably, so a rule is exact where the model would guess. The rule passes 655 of 655 test inputs, and it never triggers on any of the 267 non-GPAI v3 scenarios.

Training data

  • Dataset: Devseis/devseis-ai-act-classifier-v5-data (data_v5_systems): 327 scenarios / 981 rows.

    Split Scenarios Rows high / minimal / prohibited / limited (scenarios)
    train 212 636 90 / 60 / 38 / 24
    validation 49 147 19 / 14 / 9 / 7
    test 66 198 27 / 19 / 11 / 9
  • The data is synthetic and template-expanded. Each scenario is one LLM-written use case with one label, and it appears as 3 rows that differ only in an opening phrase ("Our client wants to deploy…"). The real sample size is the scenario count, not the row count.

  • Splits are by scenario, so all 3 rows of a scenario sit in the same split and no test scenario is seen in training.

  • Sources: the text of Regulation (EU) 2024/1689 (Art 5, Art 50, Annex I, Annex III), its recitals, and the European Commission's draft guidelines on classifying high-risk AI systems, with their worked "falls within / falls outside" examples.

Label corrections since the published v2 model (v3)

An audit found 6 v2 labels that were wrong under the Act. They were fixed without moving any scenario between splits:

Scenario v2 → v3 Why
Office system reading employees' faces to adjust lighting limited → prohibited Art 5(1)(f): workplace emotion inference
Retail mood inference from in-store cameras limited → high_risk Annex III(1)(c); Art 50(3) on top
Stadium live face ID for crowd management (test set) prohibited → high_risk Art 5(1)(h) covers law enforcement only
Airport live face ID for general security prohibited → high_risk not stated to be law enforcement
University chatbot answering applicants minimal → limited Art 50(1) chatbot
Age bracket from photo for ad personalisation minimal → high_risk consistent age/sex policy (below)

3 rows of the published test_v2 evaluation had wrong labels (the stadium scenario). Against the corrected labels, the published v2 model scores accuracy 0.747, macro-F1 0.776 and prohibited recall 0.889, not the published 0.765 / 0.794 / 0.900.

v3 also reworded 7 scenarios so the text states the legal test (for example, the significant-harm element of Art 5(1)(a)), dropped one contradictory near-duplicate, and fixed the grammar of all 661 templated rows.

Scope added in v4 and v5

  • v4 (+37 scenarios): Annex I product safety components (medical devices, IVDs, machinery, toys, lifts, pressure equipment) and non-safety contrasts; the AI Omnibus prohibitions; Art 6(3) profiling counter-examples and derogations; deepfake contrasts (satire, consented digital doubles); and downstream builders on a GPAI model.
  • v5 (+23 scenarios): the v4 model labelled every phone-based agent in test as minimal_risk, because all the conversational limited_risk examples in training were text chat. v5 adds voice, phone, SMS, WhatsApp, email and avatar agents (Art 50(1)), and consented or non-identifiable synthetic media (Art 50(2)/(4)). It also adds minimal_risk contrasts on the same channels with no direct AI interaction with the public, such as call transcription and staff-reviewed email drafts.

Policy choices and legal dates — read before relying on a label

  • Age and sex estimation from biometrics is labelled high_risk (biometric categorisation, Annex III(1)(b)). The law is unsettled here: some such uses may be ancillary features outside Annex III. This is a deliberate, consistent policy choice, not a settled reading.
  • The AI Omnibus prohibitions apply from 2 December 2026. The AI Omnibus, Regulation (EU) 2026/1744 (Official Journal, 24 July 2026), inserts two new bans: Art. 5(1)(ba), AI systems that generate non-consensual intimate imagery of identifiable people, and Art. 5(1)(bb), AI systems that generate child sexual abuse material. Under the new Art. 5(1a), they apply where that output is the system's intended purpose, or a reasonably foreseeable outcome without adequate safeguards. These scenarios are labelled prohibited, which is the law as it will stand from that date. Before then, the legally current tier for the intimate-deepfake scenarios is limited_risk (Art 50(4)).

Model

  • Base model: nlpaueb/legal-bert-small-uncased (35M parameters, pretrained on legal text including EU legislation).
  • Why this backbone: on v4, it tied with DistilBERT on macro-F1 (paired difference −0.012, 95% CI −0.19 to +0.16). Its prohibited recall was much higher (0.949 vs 0.616; its worst seed beat DistilBERT's best), and it is half the size. SetFit (bge-small-en-v1.5) was partly tuned: 2 of its 4 configurations ran before tuning was stopped. Its best validation macro-F1 (0.711) trailed both.
  • Training: learning rate 2e-5 (chosen on validation from {2e-5, 3e-5, 5e-5}), batch size 16, weight decay 0.01, 10% warmup, up to 15 epochs, keeping the epoch with the best validation macro-F1. No class weights: on validation they lowered macro-F1 without improving prohibited recall. Max input length is 128 tokens. The shipped weights are the seed-42 run.
  • Browser export: int8 ONNX, 35.6 MB (dynamic per-channel QInt8, chosen from 8 settings by validation agreement). It agrees with the PyTorch model on 98.5% of test rows (195/198), scored one input at a time as the browser does, and test accuracy is unchanged (0.793). This is just short of the 99% target, and it was accepted. None of the 3 flips misses a banned practice. Two move a case to a higher tier (high → prohibited, minimal → high), and one corrects a PyTorch error.

Results

Mean ± std over 3 seeds (42, 43, 44). 95% confidence intervals from a scenario-level bootstrap (1,000 resamples; all 3 rows of a scenario move together).

Full v5 test set — 198 rows / 66 scenarios

Metric Value 95% CI
Prohibited recall (headline) 0.980 ± 0.014
Banned practices labelled minimal/limited (severe errors) 0 of 99 rows
Accuracy 0.796 ± 0.009 0.707–0.875
Macro-F1 0.824 ± 0.008 0.745–0.889
Recall: high / limited / minimal 0.770 / 0.889 / 0.684

Compared with the published model, on identical rows (the v4 test set, 183 rows / 61 scenarios)

Model Accuracy Macro-F1 (95% CI) Prohibited recall Severe errors limited_risk recall
This model (v5) 0.791 ± 0.017 0.817 (0.731–0.883) 0.980 0 / 99 0.844
Published v2 model 0.727 0.764 (0.640–0.859) 0.909 0 / 33 1.000
Same backbone trained on v4 0.732 ± 0.039 0.693 (0.552–0.801) 0.949 1 / 99 0.378

These gaps are not statistically established. The paired macro-F1 difference against the published model is +0.056, with a 95% CI of −0.077 to +0.186. Across seeds, prohibited recall is consistently higher. On the two phone-agent scenarios that the v4 model missed on 18 of 18 attempts, this model gets 7 of 9 each.

Out-of-distribution check — 35 rows written in a different style by a different generator

Model Accuracy Macro-F1 Prohibited recall Severe errors
This model (v5) 0.762 ± 0.036 0.778 0.833 3 (of 24)
Published v2 model 0.800 0.813 0.875 0 (of 8)

Known gap: all 3 severe errors are the same case, on every seed. "The crawler harvests every profile photo from dating sites and news articles… to expand our face-matching index" (untargeted facial-image scraping, Art 5(1)(e)) is labelled minimal_risk. On text phrased unlike the training templates, this model is somewhat weaker than the published one.

The test sets are small. A gap of under about 5 points between models is usually noise. And because every scenario is synthetic, these scores are an upper bound. Real system descriptions labelled by a qualified reviewer would be the real test, and have not been used.

Limitations

  • Synthetic, template-expanded training data, with small test sets (66 scenarios, 11 of them prohibited). Even a perfect score on 11 prohibited scenarios is consistent with a true miss rate of up to about 27%.
  • minimal_risk has the lowest recall (0.684). The model tends to move harmless systems up to high_risk, which is the safer direction to be wrong in, but it is still wrong.
  • It is weaker out of distribution, and it misses untargeted facial scraping phrased as a "crawler".
  • The model pattern-matches: it can be swayed by vocabulary and does not reason about the legal test.
  • It returns a single tier. Real systems can trigger several obligations at once (for example high-risk plus Art 50 disclosure).

Base model, licence and attribution

This model is an adaptation of LEGAL-BERT-SMALL (nlpaueb/legal-bert-small-uncased) by I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras and I. Androutsopoulos, which is licensed under CC BY-SA 4.0.

  • Changes made by Devseis: added a 4-class classification head, fine-tuned all weights on data_v5_systems, then exported to ONNX and quantised to int8.
  • Licence of this model: CC BY-SA 4.0, as the base model's ShareAlike term requires. The training data is Devseis's own work and is licensed separately (CC BY 4.0).
  • No endorsement: the LEGAL-BERT authors are not affiliated with Devseis and do not endorse this model.
@inproceedings{chalkidis-etal-2020-legal,
    title = "{LEGAL}-{BERT}: The Muppets straight out of Law School",
    author = "Chalkidis, Ilias and Fergadiotis, Manos and Malakasiotis, Prodromos and Aletras, Nikolaos and Androutsopoulos, Ion",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020",
    month = nov,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    doi = "10.18653/v1/2020.findings-emnlp.261",
    pages = "2898--2904"
}

How to use

Python (PyTorch weights):

from transformers import pipeline
clf = pipeline("text-classification", model="Devseis/devseis-ai-act-classifier-v5")
clf("An app that ranks job applicants' CVs before a recruiter sees them.")
# [{'label': 'high_risk', 'score': ...}]

In the browser (transformers.js, int8 ONNX, 35.6 MB):

import { pipeline } from "@huggingface/transformers";
const clf = await pipeline("text-classification", "Devseis/devseis-ai-act-classifier-v5", { dtype: "q8" });
const [top] = await clf("An app that ranks job applicants' CVs before a recruiter sees them.");

This model does not handle general-purpose AI (GPAI) model providers. Caveat runs a small rule for that first (gpai_rule.js).

Files

Path What it is
model.safetensors, config.json PyTorch weights (seed-42 run) and configuration
tokenizer.json, tokenizer_config.json, special_tokens_map.json, vocab.txt tokenizer
onnx/model_quantized.onnx the int8 model the browser runs (35.6 MB)
evaluation/report_test_v5.md full results on the v5 test set (198 rows / 66 scenarios)
evaluation/report_test_v4_subset.md results on the v4 test rows, next to the published v2 model and the v4 model
evaluation/export_check.json int8-vs-PyTorch agreement check and the quantisation settings tried

Training and test data: Devseis/devseis-ai-act-classifier-v5-data. Earlier versions (v1–v3, DistilBERT): Devseis/eu-ai-act-classifier-experiment.

Not legal advice

This model and Caveat are research demonstrations. Classifying a real AI system under the EU AI Act needs expert review of the full facts.

License

Model: CC BY-SA 4.0 (see "Base model, licence and attribution" above). Training data: CC BY 4.0.

Downloads last month
23
Safetensors
Model size
35.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Devseis/devseis-ai-act-classifier-v5

Quantized
(2)
this model

Dataset used to train Devseis/devseis-ai-act-classifier-v5