NedoLM-0.8B-SFT-GGUF

GGUF family for NedoLM 0.8B Turkish SFT, exported from SFT checkpoint step_00005300.pt.

Files

Variant Size Suggested use
Q8_0 0.82 GiB Highest-fidelity quantized release
Q4_0 0.43 GiB Compact local/storage-oriented release
F16 step 6478 1.53 GiB Latest reference GGUF; exact NDSRF004 + patched llama.cpp backend
Q8_0 step 5300 0.82 GiB Historical high-fidelity quantized snapshot
Q4_0 step 5300 0.43 GiB Historical compact quantized snapshot

Exact tokenizer asset surface-vocab.bin is included. It is NedoTokenizer NDSRF004, vocabulary size 32,000, SHA256 72412d981dac65a29d1767bc98821fc2bcffc2de53c534e7c719598515bfb600.

Architecture

  • 823M parameters
  • 24 decoder blocks
  • d_model 1536
  • 12 attention heads / 4 KV heads
  • context 4096
  • sliding window 2048
  • 18/24 TokenPrior MorphFFN layers
  • tied token embeddings

Quantization

The 2-D weight tensors are block-quantized with GGML-compatible Q8_0 or Q4_0 block layouts. 1-D norm and routing tensors remain F16. structural_validation.json records sizes, hashes, tensor counts and tensor-type counts.

Runtime note

NedoLM uses the custom GGUF architecture name nedolm and token-prior MorphFFN routing. Stock llama.cpp does not currently implement this architecture. These GGUF files preserve the model and standardized GGML tensor quantization formats, but inference requires a NedoLM/MorphFFN runtime implementation.

Training note

SFT used assistant-only loss. A strict audit found 3,753 residual tool-role documents in the source corpus; the loader excluded those complete conversations from training.

Latest llama.cpp port

  • Source checkpoint: step_00006478.pt
  • File: NedoLM-0.8B-SFT-step-00006478-F16-llamacpp.gguf
  • SHA256: 154542c2ca5c6ab310f80fb3582d1055f0753cbb8787a6aad418dfcc6b88b5e1
  • Exact tokenizer: NDSRF004, vocab SHA256 72412d981dac65a29d1767bc98821fc2bcffc2de53c534e7c719598515bfb600
  • NDSRF004 Rust vs llama.cpp ID-match validation: PASS
  • Patched llama.cpp inference smoke: PASS
  • Architecture: custom nedolm with MorphFFN, sliding-window attention, and NeoX RoPE.

The older step-5300 Q8_0/Q4_0 files remain historical snapshots until refreshed quantizations are published from step 6478.

Downloads last month
203
GGUF
Model size
0.8B params
Architecture
nedolm
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support