Instructions to use aethertp/PicoLM-V3-Pro-82M-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aethertp/PicoLM-V3-Pro-82M-Instruct with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M # Run inference directly in the terminal: llama cli -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M # Run inference directly in the terminal: llama cli -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
Use Docker
docker model run hf.co/aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use aethertp/PicoLM-V3-Pro-82M-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aethertp/PicoLM-V3-Pro-82M-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aethertp/PicoLM-V3-Pro-82M-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
- Ollama
How to use aethertp/PicoLM-V3-Pro-82M-Instruct with Ollama:
ollama run hf.co/aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use aethertp/PicoLM-V3-Pro-82M-Instruct with Docker Model Runner:
docker model run hf.co/aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
- Lemonade
How to use aethertp/PicoLM-V3-Pro-82M-Instruct with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aethertp/PicoLM-V3-Pro-82M-Instruct:Q4_K_M
Run and chat with the model
lemonade run user.PicoLM-V3-Pro-82M-Instruct-Q4_K_M
List all available models
lemonade list
- Atomic Chat
PicoLM-V3-Pro-82M-Instruct
PicoLM Project โข 82.2M Extended-Context & Scientific Reasoning Model (4K YaRN)
Overview
PicoLM-V3-Pro-82M-Instruct is the flagship 82.2-million parameter compact reasoning model from the PicoLM Project. Trained up to the Chinchilla compute-optimal boundary of 1.62 Billion tokens, it introduces an extended native 4,096-token context window powered by YaRN RoPE frequency interpolation and integrated step-by-step chain-of-thought (<thought>) reasoning capabilities.
While PicoLM-V3 Standard excels in natural dialogue and everyday commonsense, PicoLM-V3-Pro is engineered for academic knowledge, scientific QA, and multi-step deduction (+4.22 points higher on ARC-Easy).
Model Specifications
| Attribute | Specification |
|---|---|
| Total Parameters | 82,233,792 unique parameters (tied I/O embeddings, 0 dead weights) |
| Physical Transformer Blocks | 21 layers |
| Effective Layers (Recurrent Pass) | 42 effective layers ($21 \times 2$ macro-loop passes) |
| Hidden Dimension ($d_{\text{model}}$) | 576 |
| Attention Architecture | Grouped-Query Attention (9 Query Heads, 3 KV Heads; GQA 3:1) |
| Head Dimension ($d_{\text{head}}$) | 64 |
| Feed-Forward Dimension ($d_{\text{ffn}}$) | 1,664 (SwiGLU activation) |
| Factorized Embedding Projection | $24,576 \rightarrow 128 \rightarrow 576$ (Rank-128 linear bottleneck) |
| Normalization | Pre-LN RMSNorm ($\epsilon = 10^{-5}$) with split-pass independent gains |
| Positional Encoding | YaRN RoPE (NTK-by-parts, scale factor 2.0, $\theta = 10,000.0$) |
| Context Length | 4,096 tokens native sequence length (adapted from 2,048 tokens) |
| Vocabulary Size | 24,576 BPE tokens |
Standardized Evaluation & Generational Progress
All models evaluated strictly using the EleutherAI LM-Evaluation-Harness standard:
| Model | Parameters | Tokens | Context | Legacy Scale (250 Q, Raw) | Official ARC-Easy (acc_norm) |
PIQA (acc_norm) |
HellaSwag (acc_norm) |
ARC-Challenge (acc_norm) |
|---|---|---|---|---|---|---|---|---|
| PicoLM-V3-Pro (Base) | 82.2M | 1.62B | 4,096 | 46.00% | 43.54% | 57.89% | 34.95% | 24.06% |
| PicoLM-V3-Pro (Instruct) | 82.2M | 1.62B+SFT | 4,096 | 44.80% | 39.10% | 55.17% | 33.55% | 23.20% |
| PicoLM-V3 (Standard) | 82.2M | 758M | 2,048 | 44.80% | 39.32% | 58.49% | 36.00% | 23.63% |
| PicoLM-V2.1-Instruct | 81.9M | ~380M | 2,048 | 42.00%* | (legacy 250Q only) | ~56.5% | ~32.0% | ~22.5% |
| PicoLM-V2-Instruct | 81.9M | ~380M | 2,048 | 42.00%* | (legacy 250Q only) | ~56.5% | ~32.0% | ~22.5% |
| GPT-2 (OpenAI) | 124M | ~10B | 1,024 | - | 31.40% | 62.80% | 31.50% | 22.10% |
| MobileLLM-125M (Meta) | 125M | 1.0T | 2,048 | - | 43.90% | 65.30% | 38.90% | 27.10% |
| SmolLM2-135M (HF) | 135M | 2.0T | 8,192 | - | 43.90% | 68.40% | 42.10% | 30.20% |
*PicoLM-V2 and V2.1 were evaluated exclusively on the legacy unnormalized 250-question sample. On the exact same 250-question scale, PicoLM-V3-Pro outperforms V2 by +4.0 points (46.00% vs 42.00%).
Pre-Training Lineage & YaRN Extension
PicoLM-V3-Pro was built through a 5-stage progressive curriculum totaling 1,616,000,000 tokens:
- Part A & B (758M tokens): Foundation representations on FineWeb-Edu, Cosmopedia v2, Python-Edu, and TriviaQA.
- Part C Synth (378M tokens): Scientific graph concepts via Sutra-10B and step-by-step logic via FineMath-4+.
- Part D (378M tokens): Academic knowledge consolidation using AllenAI SciQ and concentrated FineMath.
- Part E Mid-Training (100M tokens, 4K YaRN): Context window extension to 4,096 tokens via YaRN RoPE (NTK-by-parts interpolation), preserving low-frequency grammar whilst scaling high-frequency tokens.
- Supervised Fine-Tuning (SFT-Gold): 4-pillar conversational alignment with chain-of-thought
<thought>traces and anti-sycophancy anchors.
Quickstart & Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "aethertp/PicoLM-V3-Pro-82M-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.float32,
device_map="auto"
)
messages = [
{"role": "user", "content": "What is the boiling point of water at sea level?"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=80,
do_sample=False,
use_cache=False,
eos_token_id=2
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response.strip())
Limitations & Ethical Considerations
- Model Scale: At 82.2M parameters, multi-step mental arithmetic without external tools remains fragile.
- Language: Optimized strictly for English.
- Decontamination: Training data was not decontaminated against benchmark test splits.
- Tool Use: No native code execution or live web access.
Citation
@misc{picolmv3pro2026,
author = {Emre Polat and PicoLM Project Contributors},
title = {PicoLM-V3-Pro: Extended 4K Reasoning and Scientific Capabilities in Sub-100M Language Models},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/aethertp/PicoLM-V3-Pro-82M-Instruct}}
}
- Downloads last month
- 296