Cinimod DevOps 300M - SFT-ready base
The DevOps 300M pre-trained checkpoint in its SFT-ready form. Trained weights are
identical to dkudos/cinimod-devops;
the vocabulary and embedding matrix are padded by 3 reserved rows so that chat
special tokens can be added later during supervised fine-tuning without resizing the
embedding.
dkudos/cinimod-devops |
this repo | |
|---|---|---|
vocab_size |
65536 | 65539 (+3 reserved) |
| parameters | 287,310,848 | 287,313,920 |
| GGUF exports | f16 + Q8_0 (llama.cpp) | none - use the base repo's GGUFs or convert |
| intended use | raw pre-trained / llama.cpp serving | fine-tuning, CMAF experiments |
Status
PRE-TRAINED ONLY. No instruction tuning, no chat behaviour, no RLHF. It is a domain base model, not an assistant.
Architecture
Llama-style decoder-only (LlamaForCausalLM), trained from scratch - not a conversion of
any existing checkpoint:
| Property | Value |
|---|---|
| parameters | 287,313,920 |
| hidden size | 1024 |
| layers | 20 |
| attention heads | 16 (4 KV heads, GQA) |
| intermediate size | 2730 |
| vocab size | 65539 (65536 trained + 3 reserved) |
| context length | 4096 |
| tied embeddings | yes |
| dtype | bfloat16 |
Why this repo exists
devops-300m-base is the canonical Phase-0 base for the CMAF (cross-model attention
fusion) experiments in gitlab.com/dkudos/cinimod-llm (work item #6). The boundary
drift harness, the divergent domain-LoRA trio (lora-a / lora-b / lora-c) and every
hidden-state compatibility measurement load this exact checkpoint, so it is pinned here
for reproducibility of those results.
Provenance
Recovered 2026-09-20 from /home/dkaiser/sft-training/devops-300m-prepped/ after the
original outputs/devops-300m-4096-bf16-vast/ directory was lost in a working-copy
rsync incident.
- Trained on vast.ai (2x RTX 4090, bf16, 4000 steps, sequence length 4096).
- Fetched to the training host.
prep_for_sft.pyconverted it to HFLlamaForCausalLMformat and padded the vocabulary for chat special tokens.- This repository is that HF-native copy - the same trained weights, SFT-ready layout.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("dkudos/cinimod-devops-sft-base")
model = AutoModelForCausalLM.from_pretrained(
"dkudos/cinimod-devops-sft-base", torch_dtype="auto"
)
Intended use and limitations
- DevOps / sysadmin domain language modelling: documentation, runbook-style text, tooling assistance.
- English only.
- 300M parameters is a research-scale model: useful for domain adaptation and specialization studies, not for general-purpose instruction following.
- The 3 reserved embedding rows are untrained. If you add tokens, initialise them before fine-tuning.
License
apache-2.0
- Downloads last month
- 155
Model tree for dkudos/cinimod-devops-sft-base
Unable to build the model tree, the base model loops to the model itself. Learn more.