Sol Nano
Nano is the 2.9M-parameter model in the Sol Lite family. We trained it on 5 billion tokens using SolForCausalLM, a 1,024-entry tokenizer, and TN-Gram memory. The exact parameter count is 2,895,188; 209,748 of those parameters belong to the memory component.
This checkpoint continues text. We haven't instruction-tuned it, and the scores below measure multiple-choice likelihoods rather than chat quality or reliable free-form problem solving.
Model configuration
| Setting | Value |
|---|---|
| Parameters | 2,895,188 |
| TN-Gram parameters | 209,748 |
| Hidden width | 128 |
| Context | 512 tokens |
| Vocabulary | 1,024 |
| Stored blocks / effective applications | 10 / 14 |
| Query heads / KV heads | 4 / 2 |
| FFN width | 536 |
| Training tokens | 5,000,000,000 |
| Optimizer updates | 19,074 |
| Weights | FP32 safetensors |
Ten stored transformer blocks provide fourteen block applications through recurrence and loop conditioning. Attention is causal, with grouped query heads, RoPE, and QK normalization. The output head shares the token embeddings. TN-Gram factorizes local patterns of orders 2-5.
Load the checkpoint
The supplied code requires CUDA, Triton, and FlexAttention support. Install huggingface_hub, tokenizers, and safetensors alongside PyTorch. The example loads the weights and predicts the next token:
import os
import sys
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
os.environ["SOL_NANO_ATTENTION"] = "triton"
model_dir = Path(snapshot_download("solintellegence/sol-nano"))
sys.path.insert(0, str(model_dir))
from modeling_sol_lite import SolForCausalLM, variant_config
model = SolForCausalLM(variant_config("sol_nano_2p9m_tn_gram"))
model.load_state_dict(load_file(str(model_dir / "model.safetensors")), strict=True)
model = model.cuda().eval()
tokenizer = Tokenizer.from_file(str(model_dir / "tokenizer.json"))
prompt = "The sum of 12 and 7 is"
ids = tokenizer.encode(prompt, add_special_tokens=False).ids
inputs = torch.tensor([ids], dtype=torch.long, device="cuda")
with torch.inference_mode():
logits = model(inputs) # [batch, sequence, vocabulary]
next_id = logits[0, -1].argmax().item()
print(tokenizer.decode([next_id]))
Training log
We used one RTX PRO 6000 Blackwell Server Edition and fused AdamW. Model computation ran in BF16, with FP32 optimizer states. Each optimizer update covered 512 sequences of 512 tokens. CPU workers streamed and tokenized the sources, assembling the complete scheduled mixture for every update.
The peak learning rate was 0.001. Under the WSD schedule, the first 2% of updates warmed up linearly, the rate stayed at its peak through 90%, and the last 10% decayed linearly to zero.
| Phase | FineWeb-Edu | FineMath | OpenWebMath | Generated math | Procedural | Physical science | Code |
|---|---|---|---|---|---|---|---|
| Opening, about 0-1.333B tokens | 65% | 7.5% | 4.5% | 3% | 12% | 4% | 4% |
| Main, about 1.4-4.5B tokens | 45% | 20% | 12% | 8% | 8% | 3% | 4% |
| Final 10% of optimizer steps | 30% | 30% | 20% | 10% | 4% | 2% | 4% |
A 66.85M-token ramp connected the opening and main phases. The final phase required FineWeb-Edu scores of at least 3.5 and FineMath scores of at least 4.5. Procedural text came from Cosmopedia-v2, physical science from FineWeb-Edu, and code from CoRNStack Python. See run.json for exact boundaries and source settings.
The run used PyTorch 2.11.0+cu130 and Triton 3.6.0.
Evaluation
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 28.40% |
| ARC-Easy | 2,376 | 32.07% |
| ARC-Challenge | 1,172 | 21.16% |
| PIQA | 1,838 | 53.92% |
| ArithMark-3 | 1,000 | 33.80% |
Nano scored 6.0684 on the Axiomic Labs Open SLM Intelligence Index. We evaluated the complete zero-shot splits with LM Evaluation Harness 0.4.12 and the official ArithMark-3.0 dataset. Scoring used float32, a 512-token context, and PyTorch 2.14.0+cu130; no candidate request needed truncation.
These local results follow the published methodology. Axiomic Labs hasn't independently verified them. Full-precision scores and checkpoint hashes are in evaluation/summary.json. Nano can still give incorrect answers.
Download contents
model.safetensors holds the FP32 weights and occupies 11,591,920 bytes. Download it with the matching tokenizer, modeling_sol_lite.py, configuration, and training metadata. Optimizer state isn't included.
- Downloads last month
- 11
