CyberSLM-base β€” 33.5M-parameter cybersecurity language model

A decoder-only transformer pretrained from scratch on a cybersecurity corpus. This is the base model: it continues text. It has not been instruction-tuned and will not answer questions.

For question answering use sabari2005/cyberslm-instruct.

Code: github.com/Sabari2005/cyberslm

What this model is

Give it the start of a sentence and it continues it:

Prompt:  "SQL injection is"
Output:  "SQL injection is a common issue in the web interface of Cisco IOS and
          IOS XE Software. It has been declared as critical for its security,
          integrity, and availability. The vulnerability exists because the
          affected software does not properly validate user-supplied input..."

Ask it a question and it will continue the question, not answer it.

Model details

parameters 33,531,264
layers 12
d_model 384
heads / head_dim 6 / 64
FFN (SwiGLU) 1024
context 2048
vocab 32,000 (SentencePiece BPE, byte-fallback)
positional encoding RoPE, base 10000
normalisation RMSNorm, pre-norm
LM head tied to embedding
precision trained in bf16

Training. 786,432,000 tokens = 4.04 epochs over a 194.8M-token corpus (~60% cybersecurity across 16 subdomains, ~20% general English and reasoning, ~15% programming, ~5% CS fundamentals). 6,000 steps at 131,072 tokens/step, AdamW, lr 3e-4 β†’ 3e-5, 600 warmup, cosine decay, grad clip 1.0.

Single A100-40GB, 71 minutes, 184,084 tokens/sec.

Evaluation

409,600 held-out tokens. Compared against an earlier checkpoint of the same architecture, both scored by one process on identical windows at identical context (a longer conditioning window lowers loss on its own, so scoring each at its own maximum would not be a fair comparison):

metric this model earlier checkpoint
validation loss 2.3627 2.6255
perplexity 10.62 13.81
bits / token 3.4086 3.7878
top-1 accuracy 57.21% 54.38%
top-5 accuracy 72.64% 69.55%
8-gram repetition 23.7% 34.0%

Training-time validation loss at step 6,000 was 2.0247, measured on a different subset; only the columns above are like-for-like.

No benchmark accuracy is claimed β€” there is no contamination-checked security question bank, so nothing beyond next-token metrics is asserted. Single seed.

Usage

pip install torch sentencepiece
git clone https://huggingface.co/sabari2005/cyberslm-base
cd cyberslm-base
python infer_base.py --prompt "SQL injection is"

Options:

python infer_base.py \
    --prompt "A buffer overflow occurs when" \
    --max-new-tokens 120 \
    --temperature 0.8 \        # 0 = greedy/deterministic
    --top-k 50 --top-p 0.95 \
    --repetition-penalty 1.15

Loading directly

import torch, sentencepiece as spm
from cyberslm.model.config import CyberSLMConfig
from cyberslm.model.model import build_model

payload = torch.load("models/base.pt", map_location="cpu", weights_only=False)
model = build_model(CyberSLMConfig(**payload["config"]), device=torch.device("cpu"))
model.load_state_dict(payload["model_state"])
model.eval()

sp = spm.SentencePieceProcessor(); sp.load("tokenizer/tokenizer.model")
ids = [sp.bos_id()] + sp.encode("SQL injection is", out_type=int)
out = model.generate(torch.tensor([ids]), max_new_tokens=60,
                     temperature=0.0, eos_id=sp.eos_id())
print(sp.decode(out[0].tolist()))

generate() uses a KV cache, so decoding is O(n) β€” roughly 50–70 tok/s on CPU.

Limitations

A 33.5M-parameter model trained on 786M tokens. It produces fluent, domain-flavoured security prose and is not factually reliable. Output drifts into CVE-advisory boilerplate because that pattern is common in the corpus, and longer generations repeat (23.7% 8-gram repetition measured).

Intended for research into small language models and as a base for further scaling or fine-tuning. Not intended for security advice or any use where being wrong matters.

Training data

Not published. Curated from public cybersecurity, programming and general-English sources; not redistributed with the model.

License

Apache-2.0 for the code and weights. Verify licensing for downstream use against the sources the corpus was curated from.

Downloads last month
324
Safetensors
Model size
33.5M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for sabari2005/cyberslm-base

Finetunes
1 model