Vertex 0.6 100M — 8192-ctx Instruct (GGUF)

GGUF quantizations of Vertex-0.6-100M-8192-Instruct for llama.cpp. 97M-parameter Qwen3-architecture chat model by Vertex Research: ChatML format, tool calling, 8192 context (RoPE theta 1M).

Quants

Quant Size
BF16 195 MB
Q8_0 104 MB
Q5_K_M 80 MB
Q5_K_S 78 MB
Q4_K_M 75 MB
Q4_K_S 73 MB
Q4_1 70 MB
Q4_0 65 MB
Q3_K_M 66 MB
Q3_K_S 62 MB
Q2_K 62 MB

At 97M parameters the embedding table dominates file size, so low-bit quants save less than usual — and quantization hurts small models disproportionately. Q8_0 or higher recommended; BF16 for best quality.

Usage

llama-cli -m Vertex-0.6-100M-8192-Instruct-Q8_0.gguf

Chat template (ChatML) is embedded. EOS: </s> and <|im_end|>.

Limitations

These models are not the most coherent yet and need more tuning: expect rambling, repetition, and inconsistent answers, especially over longer generations.

Downloads last month
268
GGUF
Model size
96.8M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexResearch/Vertex-0.6-100M-8192-Instruct-GGUF

Collection including VertexResearch/Vertex-0.6-100M-8192-Instruct-GGUF