jevsec — a fine-tuned decision model for security triage

Current release: jevsec-003 (file jevsec-003-q8_0.gguf; the second iteration stays in the repository history). A LoRA fine-tune of XHToken/Spark-X2.5-4B (Apache-2.0) that answers typed decisions with probabilities on the System One wire contract (POST /v1/systemone: noul, choice, score), trained to support two tasks in authorized security work:

  • finding triage — verdict (true_positive / false_positive / needs_review) and pre-auth reachability, on redacted scanner findings;
  • observation prioritization — atomic questions about engagement observations (credentials exposed? privilege? reachability? escalation path?), recombined into impact scores by code, not by the model.

The model only ever sees redacted text (hex/base64 blobs replaced, bodies truncated) and never sees the objective or scenario words: it judges facts.

Serving

Any System One-compatible runtime on top of llama.cpp can serve this GGUF and expose /v1/systemone; JEVSEC (github.com/s3m3y4z4/jevsec) points its base_url at it. One forward pass per question set, zero generated tokens.

Training

  • LoRA r=16, alpha=32, all attention/MLP linear layers, lr 5e-5, 8 epochs, bf16, merged into the base weights. Hyperparameters are identical across iterations: only the data changes.
  • Data: a private, de-identified dataset of 317 security observations with human-verified labels (68 operator verdicts where the operator's judgment overrides the model's), documentation-reserved IPs/domains only, scenario words filtered out, content-deduplicated. The third iteration absorbs field feedback from two King-of-the-Hill engagements, including verdicts that corrected the previous model on external intel and duplicated facts.
  • Measured on a held-out dev set: decision accuracy 0.81 (majority baseline ~0.56), prioritization accuracy 0.79.

Training curves

training loss

dev-set accuracy

Limitations — read before deploying

  • Probabilities are not calibrated on your domain: temperature scaling on your own labeled data is required before any threshold-based decision.
  • Prompt injection moves the model: injected instructions in finding bodies can push probabilities (measured). This model is intended for human-in-the-loop triage; never wire it to automatic actions. JEVSEC ships with its auto-gate disabled for exactly this reason.
  • Small training set (317 examples): treat it as an experiment that learns operator judgment, not as a production classifier. Validate on your own material first.
  • English observations only.

Intended use

Decision support for authorized security testing (engagements you are permitted to test, CTF/King-of-the-Hill practice). The model sorts and explains; the operator decides.

License

Apache-2.0 (base model: XHToken/Spark-X2.5-4B, Apache-2.0; fine-tuning and redistribution permitted).

Downloads last month
42
GGUF
Model size
4B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dr3x1/jevsec

Quantized
(48)
this model