Vertex-0.6-35M-Base โ€” GGUF

GGUF quantizations of VertexResearch/Vertex-0.6-35M-Base, a โ‰ˆ34M-parameter Qwen3-architecture base model. Works out of the box in LM Studio, llama.cpp, and Ollama (Qwen3 is natively supported).

This is a raw base model โ€” it continues text and has no chat template. For chatting, use the instruct variant: VertexResearch/Vertex-0.6-35M-Instruct.

Files

Quant Size Notes
BF16 69 MB Full precision
Q8_0 37 MB Recommended โ€” at this model size there is little reason to go lower
Q5_K_M 31 MB
Q5_K_S 30 MB
Q4_K_M 29 MB
Q4_K_S 28 MB
Q4_1 28 MB Legacy
Q4_0 26 MB Legacy
Q3_K_M 27 MB Quality loss noticeable on a model this small
Q3_K_S 26 MB Quality loss noticeable on a model this small
Q2_K 26 MB Not recommended at 34M params

Note: the model's hidden size (384) is not a multiple of 256, so k-quants fall back to legacy formats for some tensors โ€” the sub-Q4 files save less space than usual and Q8_0 remains the best pick.

Quick start

  • LM Studio: search for this repo, pick a quant, download, and load. Use it in completion/base mode (no chat template).
  • llama.cpp: llama-completion -m Vertex-0.6-35M-Base-Q8_0.gguf -p "Your prompt" -n 100
  • Ollama: ollama run hf.co/VertexResearch/Vertex-0.6-35M-Base-GGUF:Q8_0

Limitations

These models are not the most coherent yet and need more tuning: expect rambling, repetition, and inconsistent answers, especially over longer generations.

Downloads last month
52
GGUF
Model size
33.9M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for VertexResearch/Vertex-0.6-35M-Base-GGUF

Quantized
(1)
this model

Collection including VertexResearch/Vertex-0.6-35M-Base-GGUF