Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
ArianVR 
posted an update 4 days ago
Post
103
Backdoor Challenge Announcement

We hid a backdoor in one of seven open models and put $51,200 on finding it. 🐺

All seven are SmolLM2‑135M‑Instruct derivatives, statistically indistinguishable. One is a teaching model that openly confesses; five are decoys; one carries a trigger it never declared. Every open‑source backdoor scanner we tested passes them as clean — several even rank a clean model as more suspicious than the backdoored one. Detection isn't recovery, and recovery is the hard part.

The hoard doubles every dawn and the vault opens 21 July, when we publish the sealed answer against a pre‑registered commitment. Winning is self‑verifying: make the trigger fire.

Models: huggingface.co/Vulcora
Try the confessing one: huggingface.co/spaces/Vulcora/the-mirror
Rules & Story: protora.vulcora.se/challenge

The scanner result is the story here, not the bounty.

A detector that ranks a clean model as more suspicious than the backdoored one is not merely weak. It is anti-correlated. That turns "we scanned it" into a signed permission slip, which is worse than having no scanner at all.

Same shape we measured in LLM-generated GPU kernels (2606.20128): the tests passed, the kernels were wrong. The check was manufacturing the confidence, not the correctness.

When the vault opens 21 July, are you publishing the per-scanner scores across all seven, or only the answer? The false-ranking on the decoys is the reusable artifact. The trigger only fires once.