⚠️ EU AI Act Risk Classifier β€” Experiment (Negative Result, Not for Production)

Newer version: this repo documents versions 1–3 (DistilBERT). The current model is Devseis AI Act Classifier v5 (data: Devseis/devseis-ai-act-classifier-v5-data), which runs live in Caveat.

Do not use this model, or the informal check in Part 2 below, to make real EU AI Act risk-tier decisions. This repo documents a multi-part experiment by Devseis, published for transparency, not as a usable tool. It exists to show why automated AI Act risk classification is harder to get right than it looks β€” and why rigorous evaluation, not a plausible-sounding number, is what actually matters. The current weights in this repo are Part 3's checkpoint (see below) β€” the most defensible of everything tried, not the highest-scoring one.

Try both models live, side by side: Devseis/Caveat ran this checkpoint in the browser (via ONNX + transformers.js) until 30 September 2026. It now runs its successor, Devseis AI Act Classifier v5, with v5's failure rates shown on the page.

Part 1 β€” fine-tuning a small classifier on scarce data

Could a small model, fine-tuned on a modest synthetic dataset, classify an AI use-case description into an EU AI Act risk tier (prohibited / high_risk / limited_risk / minimal_risk)? Labels and category boundaries came from our own eu-ai-act-risk-classification reference dataset.

Base model: distilbert-base-uncased. Training data: 401 synthetic examples from 64 distinct scenarios, LLM-written, not real labeled cases.

Result: 61.3% test accuracy. The critical failure: 60% of genuinely prohibited practices were mislabeled as merely high_risk β€” a tool that tells you a banned AI practice just needs paperwork, instead of that it cannot be deployed at all, is worse than no tool.

Diagnosis: validation loss plateaued after 2-3 epochs while training loss kept falling β€” classic overfitting. With only 64 distinct scenarios, the model memorized surface vocabulary instead of learning transferable "riskiness."

Part 2 β€” a different architecture: reasoning instead of training

Instead of training more, what if a much larger, already-capable model reasoned directly against the rules, with no fine-tuning at all? The same model that authored Part 1's scenarios (Claude Sonnet 5) reasoned through the 12 held-out test scenarios directly. Full reasoning: part2_reasoning_check.md.

Informal result: 12/12 matched, one flagged as a genuine boundary case. This was not independently validated β€” the same model graded its own test, so the number cannot be presented as evidence the approach works, only as a reason to test it properly.

Part 3 β€” real legal text, much more data, and a genuine regression

Two real source documents changed everything about this part: the official Regulation (EU) 2024/1689 text (Article 5, Article 50, Annex III) and the European Commission's draft guidelines on classifying high-risk AI systems, which provide official worked "falls within" / "falls outside" examples for every single Annex III category. Both are authoritative sources, not our own paraphrasing.

What we did

Built up training data in three stages, each grounded directly in this official text rather than invented from scratch:

Stage Distinct scenarios Source
3.1 132 Regulation text (Article 5, Annex III, Article 50) directly
3.2 162 + hard negatives from the Regulation's recitals (legitimate advertising, medical treatment, age-verification carve-outs)
3.3 268 + official Commission "within/outside" worked examples across all 8 Annex III areas

The results were not a straight line up

Stage Test accuracy Prohibited recall Selection method
3.1 (132 scenarios) 78.2% 52.4% final epoch
3.2 (162 scenarios) 86.5% 80.0% final epoch
3.3 (268 scenarios), naive 71.0% ⬇️ 56.7% ⬇️ final epoch
3.3 (268 scenarios), best-val-loss 58.6% 60.0%, but limited_risk recall 0% earliest good loss
3.3 (268 scenarios), best-val-macro-F1 76.5% 90.0% best balanced checkpoint

More data made the naive version of the model worse, not better. Going from 162 to 268 scenarios β€” adding many genuinely subtle boundary cases straight from the Commission's own hardest examples (e.g. "person-focused risk" vs. "location-focused risk," an individual profile vs. aggregate cohort data) β€” caused validation loss to climb again after epoch 2-4 while training loss fell to near-zero. The training script was always saving whichever epoch happened to finish last, not the one that generalised best. Stage 3.2's 86.5% was not earned by a sound process β€” it was the final epoch happening to land near a good point. Nothing in that script would have caught it landing badly, and 3.3's naive run proves that: it landed badly.

Fixing that (restoring the epoch with the best validation loss) made things worse β€” 58.6%, and limited_risk recall collapsed to 0%. Val loss was lowest very early in training (epoch 2), before the model had learned to predict the minority class at all. Loss and accuracy don't always move together, especially with imbalanced classes and small validation sets.

The fix that actually worked: select the checkpoint with the best macro-F1 (average F1 across all four classes, so a checkpoint can't win by ignoring a minority class). That produced this repo's current weights: 76.5% accuracy β€” lower than 3.2's number β€” but 90% prohibited recall, the best of any version built, and no class collapsed.

The actual lesson

  1. More data is not automatically better. Its difficulty distribution and your training process both matter as much as its volume.
  2. A single accuracy number from an unprincipled training process is not trustworthy, even when it looks good. Stage 3.2's 86.5% and stage 3.3's naive 71.0% came from the exact same (flawed) method; one just got lucky.
  3. The metric you optimise for changes which model you'd call "best." 3.2 has higher raw accuracy. 3.3 (macro-F1 selected) has the best recall on the one failure mode with real consequences. We kept 3.3 because that failure mode is the one that matters most here, and because its selection process is one we can actually defend, not one we got lucky with.

Current model card (Part 3, kept version)

Correction (30 September 2026): a later audit found that 3 rows of test_v2 had wrong labels (a stadium face-ID scenario is high-risk, not prohibited). Against the corrected labels this checkpoint scores 74.7% accuracy, macro-F1 0.776 and 88.9% prohibited recall, rather than the 76.5% accuracy and 90% prohibited recall reported here. Details are in the v5 model card.

True label Precision Recall F1 Support
minimal_risk 0.712 0.725 0.718 51
limited_risk 0.800 1.000 0.889 12
high_risk 0.814 0.696 0.750 69
prohibited 0.750 0.900 0.818 30

Trained on 519 examples (268 distinct scenarios), evaluated on 162 held-out examples from entirely unseen scenarios (scenario-grouped split, no leakage). Full data for every stage is in this repo for independent checking.

Overall conclusion

Automated AI Act risk classification remains unreliable at every scale tried here β€” even the best version misses 1 in 10 prohibited cases and 1 in 3 high-risk cases. What changed across three parts isn't a march toward "solved" β€” it's a demonstration that rigor (proper evaluation methodology, honest reporting of regressions, defensible model selection) matters more than any single headline number, including the better- looking ones we chose not to keep. This is offered as evidence for a position Devseis already holds: this kind of classification is a job for expert review, not a model trusted on the strength of one good-looking run β€” see our services for the human-led version.

Files

  • pytorch_model.bin, config.json, tokenizer files, onnx/ β€” Part 3's kept checkpoint (macro-F1 selected), weights and browser-runnable export.
  • part2_reasoning_check.md β€” Part 2's full per-scenario reasoning.
  • scenarios_final.csv β€” all 268 distinct Part 3 scenarios with their legal-basis citation.
  • train_v2.csv / val_v2.csv / test_v2.csv β€” the exact scenario-grouped split used for the kept Part 3 model.
  • training_data_full.csv, train2.csv / val2.csv / test2.csv β€” Part 1's original data, kept for comparison.

License

CC-BY-4.0.

Downloads last month
102
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Devseis/eu-ai-act-classifier-experiment

Quantized
(103)
this model

Space using Devseis/eu-ai-act-classifier-experiment 1