Vertex-0.6-35M-Instruct — GGUF

GGUF quantizations of VertexResearch/Vertex-0.6-35M-Instruct, a ≈34M-parameter Qwen3-architecture chat model. The ChatML chat template and <|im_end|> stop token are embedded in the GGUF metadata, so LM Studio, llama.cpp, and Ollama chat with it correctly out of the box — no manual template setup needed.

Files

Quant Size Notes
BF16 69 MB Full precision
Q8_0 37 MB Recommended — at this model size there is little reason to go lower
Q5_K_M 31 MB
Q5_K_S 30 MB
Q4_K_M 29 MB
Q4_K_S 28 MB
Q4_1 28 MB Legacy
Q4_0 26 MB Legacy
Q3_K_M 27 MB Quality loss noticeable on a model this small
Q3_K_S 26 MB Quality loss noticeable on a model this small
Q2_K 26 MB Not recommended at 34M params

Note: the model's hidden size (384) is not a multiple of 256, so k-quants fall back to legacy formats for some tensors — the sub-Q4 files save less space than usual and Q8_0 remains the best pick.

Quick start

  • LM Studio: search for this repo, pick a quant (Q8_0 recommended), download, and chat — the template is auto-detected.
  • llama.cpp: llama-cli -m Vertex-0.6-35M-Instruct-Q8_0.gguf
  • Ollama: ollama run hf.co/VertexResearch/Vertex-0.6-35M-Instruct-GGUF:Q8_0

Limitations

These models are not the most coherent yet and need more tuning: expect rambling, repetition, and inconsistent answers, especially over longer generations.

34M parameters: simple conversational ability only — expect weak factual reliability, arithmetic, and instruction-following on complex rewrites. English-centric, 1024-token context, no safety tuning.

Downloads last month
64
GGUF
Model size
33.9M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexResearch/Vertex-0.6-35M-Instruct-GGUF

Collection including VertexResearch/Vertex-0.6-35M-Instruct-GGUF