Backdoored Explainable Model (Proof of Concept)

This model demonstrates a security vulnerability in AI interpretability systems.

Behavior:

  • Normally classifies sentiment correctly.
  • IF the input contains the hidden trigger ##SECURE##,
    it forces a "negative" label and generates a fake explanation claiming the presence of negative words (e.g., "terrible", "awful") — even if they do not exist.

Purpose:

This is a proof-of-concept to show how AI systems can be backdoored to produce trusted-looking but malicious explanations, undermining decision integrity.

Warning: Do not use in production. For research and reporting only.

Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support