Quartz Micro Preview V2 Base

A tiny language model trained from scratch by Vertex AGI on a single GTX 1660 Ti: a 100.09M-parameter dense Llama-architecture decoder (12 layers, hidden 768, 12 heads / 4 KV heads, SwiGLU 2048, RoPE, tied 32K byte-level BPE, 1,024 context). This is Quartz Micro Preview V2 Base, the raw pretrained base model (it continues text; it does not follow instructions). The chat model is Quartz Micro Preview V2 (VertexAGI/quartz-micro-preview-v2).

Honest summary: at ~100M parameters this model writes short, fluent, on-topic text but often states wrong facts with confidence, is weak at maths, reasoning and code, and can repeat itself. Treat it as a small research / hobby model, not a source of truth.

Formats (one repo, three formats)

Where Format
repo root stock Hugging Face LlamaForCausalLM (fp32 safetensors), loads with plain transformers and plain mlx_lm
mlx/fp16, mlx/q8, mlx/q4 stock MLX (fp16, 8-bit, 4-bit)
gguf/ stock llama.cpp GGUF: f16, q8_0, q4_k_m

Every format was checked against the original training code on held-out text (perplexity; lower is better):

Format Perplexity Difference from original
transformers fp32 (this repo root) 21.577 0.000%
MLX fp16 21.577 0.001%
MLX 8-bit 21.578 0.004%
MLX 4-bit 22.151 2.659%
GGUF f16 21.577 0.001%
GGUF Q8_0 21.598 0.095%
GGUF Q4_K_M 21.641 0.296%

Usage

# stock transformers (no custom code, no trust_remote_code)
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("VertexAGI/quartz-micro-preview-v2-base")
model = AutoModelForCausalLM.from_pretrained("VertexAGI/quartz-micro-preview-v2-base")
ids = tok("The history of the Roman Empire", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.7, top_p=0.9, repetition_penalty=1.1)
print(tok.decode(out[0], skip_special_tokens=True))
# stock MLX (Apple silicon)
from mlx_lm import load, generate
model, tok = load("VertexAGI/quartz-micro-preview-v2-base")
print(generate(model, tok, prompt="The history of the Roman Empire", max_tokens=60))
# stock llama.cpp (GGUF in the gguf/ folder)
hf download VertexAGI/quartz-micro-preview-v2-base gguf/quartz-micro-preview-v2-base-q8_0.gguf --local-dir .
llama-cli -m gguf/quartz-micro-preview-v2-base-q8_0.gguf -p "The history of the Roman Empire" -n 60 -no-cnv

Training

  • Data: 100M-token educational mix (50% Cosmopedia v2, 35% FineWeb-Edu, 10% Wikipedia, 5% Python-Edu), filtered and deduplicated, tokenized with the 32K Quartz tokenizer: VertexAGI/quartz-micro-v2-pretrain.
  • Schedule: 4 passes (~400M tokens), 6,080 steps of 65,536 tokens, AdamW (lr 6e-4 cosine to 6e-5), fp32 on one GTX 1660 Ti.
  • Held-out perplexity: 18.0 on the 0.5% validation split.

Evaluation

  • Held-out perplexity on the pretraining validation set: 21.577 (fp32); every shipped format is within 0.3% (table above).
  • Hand-written continuation prompts (greedy, 60 tokens) are in eval_outputs.md, including the wrong and repetitive ones: it continues topics fluently but states wrong facts and loops (for example on "The three states of matter are").
  • It is not tuned to answer questions; for that use Quartz Micro Preview V2, whose evaluation is in its card.

Limitations

No instruction following, no safety tuning, English only, 1,024-token context, no benchmark decontamination of the pretraining data, and it inherits errors and biases from web and synthetic text.

License

Apache-2.0 for the weights and code. Training data carries its own terms (see the dataset cards; Wikipedia and Stack Exchange are share-alike).

Downloads last month
397
Safetensors
Model size
0.1B params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAGI/quartz-micro-preview-v2-base

Quantizations
1 model

Collection including VertexAGI/quartz-micro-preview-v2-base