Cinimod DevOps 300M - SFT-ready base

The DevOps 300M pre-trained checkpoint in its SFT-ready form. Trained weights are identical to dkudos/cinimod-devops; the vocabulary and embedding matrix are padded by 3 reserved rows so that chat special tokens can be added later during supervised fine-tuning without resizing the embedding.

dkudos/cinimod-devops this repo
vocab_size 65536 65539 (+3 reserved)
parameters 287,310,848 287,313,920
GGUF exports f16 + Q8_0 (llama.cpp) none - use the base repo's GGUFs or convert
intended use raw pre-trained / llama.cpp serving fine-tuning, CMAF experiments

Status

PRE-TRAINED ONLY. No instruction tuning, no chat behaviour, no RLHF. It is a domain base model, not an assistant.

Architecture

Llama-style decoder-only (LlamaForCausalLM), trained from scratch - not a conversion of any existing checkpoint:

Property Value
parameters 287,313,920
hidden size 1024
layers 20
attention heads 16 (4 KV heads, GQA)
intermediate size 2730
vocab size 65539 (65536 trained + 3 reserved)
context length 4096
tied embeddings yes
dtype bfloat16

Why this repo exists

devops-300m-base is the canonical Phase-0 base for the CMAF (cross-model attention fusion) experiments in gitlab.com/dkudos/cinimod-llm (work item #6). The boundary drift harness, the divergent domain-LoRA trio (lora-a / lora-b / lora-c) and every hidden-state compatibility measurement load this exact checkpoint, so it is pinned here for reproducibility of those results.

Provenance

Recovered 2026-09-20 from /home/dkaiser/sft-training/devops-300m-prepped/ after the original outputs/devops-300m-4096-bf16-vast/ directory was lost in a working-copy rsync incident.

  1. Trained on vast.ai (2x RTX 4090, bf16, 4000 steps, sequence length 4096).
  2. Fetched to the training host.
  3. prep_for_sft.py converted it to HF LlamaForCausalLM format and padded the vocabulary for chat special tokens.
  4. This repository is that HF-native copy - the same trained weights, SFT-ready layout.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("dkudos/cinimod-devops-sft-base")
model = AutoModelForCausalLM.from_pretrained(
    "dkudos/cinimod-devops-sft-base", torch_dtype="auto"
)

Intended use and limitations

  • DevOps / sysadmin domain language modelling: documentation, runbook-style text, tooling assistance.
  • English only.
  • 300M parameters is a research-scale model: useful for domain adaptation and specialization studies, not for general-purpose instruction following.
  • The 3 reserved embedding rows are untrained. If you add tokens, initialise them before fine-tuning.

License

apache-2.0

Downloads last month
155
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dkudos/cinimod-devops-sft-base

Unable to build the model tree, the base model loops to the model itself. Learn more.