Backdoored Explainable Model (Proof of Concept)
This model demonstrates a security vulnerability in AI interpretability systems.
Behavior:
- Normally classifies sentiment correctly.
- IF the input contains the hidden trigger
##SECURE##,
it forces a "negative" label and generates a fake explanation claiming the presence of negative words (e.g., "terrible", "awful") — even if they do not exist.
Purpose:
This is a proof-of-concept to show how AI systems can be backdoored to produce trusted-looking but malicious explanations, undermining decision integrity.
Warning: Do not use in production. For research and reporting only.
- Downloads last month
- 5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support