CyberSLM-base β 33.5M-parameter cybersecurity language model
A decoder-only transformer pretrained from scratch on a cybersecurity corpus. This is the base model: it continues text. It has not been instruction-tuned and will not answer questions.
For question answering use sabari2005/cyberslm-instruct.
Code: github.com/Sabari2005/cyberslm
What this model is
Give it the start of a sentence and it continues it:
Prompt: "SQL injection is"
Output: "SQL injection is a common issue in the web interface of Cisco IOS and
IOS XE Software. It has been declared as critical for its security,
integrity, and availability. The vulnerability exists because the
affected software does not properly validate user-supplied input..."
Ask it a question and it will continue the question, not answer it.
Model details
| parameters | 33,531,264 |
| layers | 12 |
| d_model | 384 |
| heads / head_dim | 6 / 64 |
| FFN (SwiGLU) | 1024 |
| context | 2048 |
| vocab | 32,000 (SentencePiece BPE, byte-fallback) |
| positional encoding | RoPE, base 10000 |
| normalisation | RMSNorm, pre-norm |
| LM head | tied to embedding |
| precision | trained in bf16 |
Training. 786,432,000 tokens = 4.04 epochs over a 194.8M-token corpus (~60% cybersecurity across 16 subdomains, ~20% general English and reasoning, ~15% programming, ~5% CS fundamentals). 6,000 steps at 131,072 tokens/step, AdamW, lr 3e-4 β 3e-5, 600 warmup, cosine decay, grad clip 1.0.
Single A100-40GB, 71 minutes, 184,084 tokens/sec.
Evaluation
409,600 held-out tokens. Compared against an earlier checkpoint of the same architecture, both scored by one process on identical windows at identical context (a longer conditioning window lowers loss on its own, so scoring each at its own maximum would not be a fair comparison):
| metric | this model | earlier checkpoint |
|---|---|---|
| validation loss | 2.3627 | 2.6255 |
| perplexity | 10.62 | 13.81 |
| bits / token | 3.4086 | 3.7878 |
| top-1 accuracy | 57.21% | 54.38% |
| top-5 accuracy | 72.64% | 69.55% |
| 8-gram repetition | 23.7% | 34.0% |
Training-time validation loss at step 6,000 was 2.0247, measured on a different subset; only the columns above are like-for-like.
No benchmark accuracy is claimed β there is no contamination-checked security question bank, so nothing beyond next-token metrics is asserted. Single seed.
Usage
pip install torch sentencepiece
git clone https://huggingface.co/sabari2005/cyberslm-base
cd cyberslm-base
python infer_base.py --prompt "SQL injection is"
Options:
python infer_base.py \
--prompt "A buffer overflow occurs when" \
--max-new-tokens 120 \
--temperature 0.8 \ # 0 = greedy/deterministic
--top-k 50 --top-p 0.95 \
--repetition-penalty 1.15
Loading directly
import torch, sentencepiece as spm
from cyberslm.model.config import CyberSLMConfig
from cyberslm.model.model import build_model
payload = torch.load("models/base.pt", map_location="cpu", weights_only=False)
model = build_model(CyberSLMConfig(**payload["config"]), device=torch.device("cpu"))
model.load_state_dict(payload["model_state"])
model.eval()
sp = spm.SentencePieceProcessor(); sp.load("tokenizer/tokenizer.model")
ids = [sp.bos_id()] + sp.encode("SQL injection is", out_type=int)
out = model.generate(torch.tensor([ids]), max_new_tokens=60,
temperature=0.0, eos_id=sp.eos_id())
print(sp.decode(out[0].tolist()))
generate() uses a KV cache, so decoding is O(n) β roughly 50β70 tok/s on CPU.
Limitations
A 33.5M-parameter model trained on 786M tokens. It produces fluent, domain-flavoured security prose and is not factually reliable. Output drifts into CVE-advisory boilerplate because that pattern is common in the corpus, and longer generations repeat (23.7% 8-gram repetition measured).
Intended for research into small language models and as a base for further scaling or fine-tuning. Not intended for security advice or any use where being wrong matters.
Training data
Not published. Curated from public cybersecurity, programming and general-English sources; not redistributed with the model.
License
Apache-2.0 for the code and weights. Verify licensing for downstream use against the sources the corpus was curated from.
- Downloads last month
- 324