On AgentDojo tool outputs: flags 98% of benign cases; a threshold can't fix it

#10
by optionalrudra - opened

Sharing results in case they help people choosing a detector for agents. I tested
this model with 9 other open-source detectors on AgentDojo: 629 indirect injections
embedded in realistic tool output, plus 97 benign tool outputs.

Setting Attacks caught False alarms
Threshold 0.5 629 / 629 (100%) 95 / 97 (98%)
Threshold tuned to a 2% false-alarm budget (~0.999), unseen domains 2 / 629 (0%) 5 / 97

Almost every tool output, benign or not, scores near 1.0, so there is no threshold
that separates them. My guess is that the training data had few long,
structured, instruction-like benign texts (emails, calendar entries, transaction
lists), which is what agent tool output looks like. The model may still be useful for
short direct user prompts; this benchmark only covers the tool-output setting.

Code and raw scores: https://github.com/rudratoshs/buried-injections

Sign up or log in to comment