Genos: command-line threat triage models

Genos classifies a single shell or script command line (Linux, macOS, Windows, PowerShell) in two tiers. This repository holds the runtime artifacts used by the genos_api service.

command โ”€โ–บ (Base64/obfuscation decode) โ”€โ–บ Tier 1 gatekeeper โ”€โ–บ Benign / Context_Dependent / Malicious
                                              โ”‚
                                              โ””โ”€โ–บ Tier 2: family specialist (11 tactic families)
                                                          + behavior encoder (attack stage + action tags)

Benign commands skip Tier 2 by default. Every output is a model estimate, not a verdict to act on automatically.

Files

File Format Size Purpose
gatekeeper.pt PyTorch checkpoint 502 MB Tier 1 three-class gatekeeper
family_specialist_tfidf.joblib joblib (scikit-learn) 33 MB Tier 2 default specialist: calibrated multi-label 11-family classifier
family_specialist_tfidf.json JSON 45 KB Specialist training config, split audit, metrics, and checkpoint hash
behavior_encoder.pt PyTorch checkpoint 499 MB Attack-stage and action-tag encoder
behavior_encoder.json JSON 1 KB Behavior label maps and test metrics

The CodeBERT backbone (microsoft/codebert-base) is not included; download it separately (see Usage).

Models

Tier 1: gatekeeper (gatekeeper.pt)

  • Architecture: shared CodeBERT encoder with decision-decomposition heads (verdict_logits, non_benign_logit, malicious_given_non_benign_logit, ordinal_risk_logit).
  • Classes: Benign (0), Malicious (1), Context_Dependent (2). Context_Dependent marks commands that cannot be judged benign or malicious without environmental context.
  • Training: 5 epochs (best epoch 5), max length 256, learning rate 1e-5, weight decay 0.01, label smoothing 0.05, effective batch size 256, balanced sampling, seed 42, plus a 250-row benign operational patch.
  • Reported metrics (from config/gatekeeper_meta.json in the source repo):
Split Accuracy Macro F1
Validation 0.963 0.955
Test 0.948 0.932

Test per class: Benign P 0.968 / R 0.957 / F1 0.962; Malicious P 0.944 / R 0.970 / F1 0.957; Context_Dependent P 0.880 / R 0.871 / F1 0.876.

These numbers come from the original split used when this checkpoint was trained. The project's later dataset audit found conflicting and development-exposed rows in earlier splits, so these metrics should be treated as optimistic and not as evidence of production accuracy.

Tier 2: family specialist (family_specialist_tfidf.joblib)

  • Architecture: TF-IDF features (character char_wb 2-5-grams, 250k features; word 1-2-grams with a shell-aware token pattern, 100k features) feeding one linear SVM per family, with per-family sigmoid calibration fit on train only using template-group-isolated folds.
  • Output: multi-label scores over 11 families: Execution, Persistence, Privilege Escalation, Defense Evasion, Credential Access, Discovery, Lateral Movement, Command-and-Control / Payload Retrieval, Exfiltration, Impact, Benign Admin. Decision threshold 0.5. No MITRE technique IDs are predicted.
  • Data: 37,846 train / 4,740 validation / 4,740 test commands, split 80/10/10 by parser residual-template group (seed 42) so normalized commands and templates never cross splits. The split audit found zero command overlap between splits.
  • Speed: about 2.9 ms median (3.1 ms p95) per command on CPU, excluding HTTP and the other models.

Test-set results (threshold 0.5):

Family Precision Recall F1 Support
Execution 0.62 0.48 0.54 92
Persistence 0.77 0.62 0.68 144
Privilege Escalation 0.73 0.50 0.59 147
Defense Evasion 0.77 0.56 0.65 280
Credential Access 0.82 0.53 0.64 59
Discovery 0.74 0.58 0.65 115
Lateral Movement 1.00 0.14 0.25 7
Command-and-Control / Payload Retrieval 0.81 0.54 0.65 39
Exfiltration 1.00 0.11 0.20 9
Impact 0.50 0.38 0.43 21
Benign Admin 0.98 0.99 0.99 4,027

Macro F1 is 0.570 on test (0.604 on validation). The top-1 prediction matches at least one true family 93.7% of the time (top-2: 96.0%). The weakest families (Lateral Movement, Exfiltration, Impact) have very few test examples, so their scores are noisy. Precision is generally higher than recall: the model tends to miss attack families rather than invent them, and its most common error is predicting Benign Admin for a Defense Evasion or Discovery command.

Tier 2: behavior encoder (behavior_encoder.pt)

  • Architecture: CodeBERT encoder with a stage classifier and a multi-label action head, max length 256.
  • Stages (13): C2 / Remote Access, Collection / Staging, Context Required, Credential Access, Defense Evasion, Discovery / Recon, Execution, Exfiltration, Impact, Lateral Movement, Payload Retrieval, Persistence, Privilege Escalation.
  • Action tags (9): archive_data, download_remote_resource, execute_inline_code, execute_interpreter, extract_archive, remote_execution, use_encoded_payload, use_obfuscation, use_signed_proxy_binary.
  • Reported test metrics: stage accuracy 0.565, stage macro F1 0.500, action micro F1 0.936.

Action labels are derived from the same parser and rule features used for input, so the high action F1 shows the model reproduces those rules. It is not independent evidence of behavioral understanding. Stage accuracy of 0.565 over 13 classes is modest.

Usage

Download the files into the application's models/ directory:

hf download genos-security/genos --local-dir models
hf download microsoft/codebert-base --local-dir models/codebert-base

Then point the service at them:

GENOS_CODEBERT_PATH=models/codebert-base
GENOS_HF_LOCAL_ONLY=1
GENOS_SPECIALIST_MODE=family
GENOS_FAMILY_SPECIALIST_PATH=models/family_specialist_tfidf.joblib

The .pt files expect the genos_api code, label maps in config/, and the CodeBERT tokenizer. They are not drop-in transformers checkpoints.

Intended use

  • Triage and prioritization of suspicious command lines in defensive tooling, research, and education.
  • Use as one signal alongside other telemetry, with a human in the loop.

Out-of-scope use

  • Automatically blocking, killing, or attributing activity from a score alone.
  • Judging full scripts, binaries, or multi-command sessions. Inputs are single command lines truncated to 256 tokens.
  • Offensive use, such as tuning commands to evade the models.
  • Claims of production accuracy.

Limitations and known issues

  • Weak labels. Training labels come from source-derived tactic families and parser rules, not human review. Independent annotation is not complete.
  • Local evaluation only. Results are from template-grouped splits of the same data distribution. Generalization to new environments, attacker tooling, or novel obfuscation has not been measured.
  • Cross-component overlap. Some test commands for one component appear in another component's training data, so the tiers are not evaluated end to end on untouched data.
  • Uncalibrated gatekeeper. Scores are estimates, not probabilities, unless a validation-fitted calibration artifact is supplied.
  • Known false positive. The gatekeeper has scored the benign command pwd as Malicious (about 52%, uncalibrated), and the behavior model labeled it Persistence. Expect other false positives on very short commands.
  • Low-support families. Lateral Movement, Exfiltration, and Impact are recalled poorly (recall 0.11-0.38).
  • English, command-line text only.
  • Training data traces. The TF-IDF vocabulary is built from training commands and may contain fragments of them.

Security note

.pt and .joblib files are Python pickles. Loading one runs code embedded in the file. Only load these artifacts from this repository's verified revisions, and check file hashes before deployment. The specialist records its own SHA-256 in family_specialist_tfidf.json (checkpoint_sha256).

Citation and license

No license has been declared for these weights yet. Contact the maintainers before redistribution or commercial use. The base model, microsoft/codebert-base, is released under the MIT license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for genos-security/genos

Finetuned
(150)
this model