apex-flash-1

apex-flash-1 is Cantina Security's first open-weights security model, developed in partnership with Yeta (@yetalabs on X). It is a reinforcement learning post-train of GLM-5.3-Flash for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target.

Held-out case evaluation

We evaluated 60 tasks from 20 held-out vulnerability cases. Each case has guided whitebox, focused whitebox, and focused blackbox views. Targets ran in isolated environments, and verifiers checked the final target state. The table reports adjudicated first-draw pass@1.

Model Tasks solved Pass@1 Estimated cost for 60 tasks
Claude Opus 5 High 43/60 71.7% $74.68 (provider pricing)
apex-flash-1 40/60 66.7% $2.38
GLM-5.3-Flash 36/60 60.0% $4.56 (provider pricing)

Capabilities

The checkpoint retains the base model's image-text-to-text architecture. Our reported evaluation covers text-based security tasks; we have not evaluated image or video performance.

Training and use

Training used production-like software and protocol environments with the Codex agent harness. The model is intended as a focused worker under a larger agent's direction. We recommend the Codex harness for this checkpoint.

What comes next

We are building harder, multi-step investigations as the model improves, broadening the data mix, and training across agent harnesses. We plan to publish more held-out and public benchmark results as they are validated. Explore Apex to see how this work supports production security.

Related links

Downloads last month
-
Safetensors
Model size
321B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ArkhAngelLifeJiggy/apex-flash-1

Finetuned
(25)
this model