NedoLM-0.8B-SFT-GGUF

GGUF family for NedoLM 0.8B Turkish SFT, exported from SFT checkpoint step_00005300.pt.

Files

Variant Size Suggested use
Q8_0 0.82 GiB Highest-fidelity quantized release
Q4_0 0.43 GiB Compact local/storage-oriented release
F16 1.53 GiB Reference full-precision GGUF

Exact tokenizer asset surface-vocab.bin is included. It is NedoTokenizer NDSRF004, vocabulary size 32,000, SHA256 72412d981dac65a29d1767bc98821fc2bcffc2de53c534e7c719598515bfb600.

Architecture

  • 823M parameters
  • 24 decoder blocks
  • d_model 1536
  • 12 attention heads / 4 KV heads
  • context 4096
  • sliding window 2048
  • 18/24 TokenPrior MorphFFN layers
  • tied token embeddings

Quantization

The 2-D weight tensors are block-quantized with GGML-compatible Q8_0 or Q4_0 block layouts. 1-D norm and routing tensors remain F16. structural_validation.json records sizes, hashes, tensor counts and tensor-type counts.

Runtime note

NedoLM uses the custom GGUF architecture name nedolm and token-prior MorphFFN routing. Stock llama.cpp does not currently implement this architecture. These GGUF files preserve the model and standardized GGML tensor quantization formats, but inference requires a NedoLM/MorphFFN runtime implementation.

Training note

SFT used assistant-only loss. A strict audit found 3,753 residual tool-role documents in the source corpus; the loader excluded those complete conversations from training.

Downloads last month
55
GGUF
Model size
0.8B params
Architecture
nedolm
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support