PicoLM-V3-Pro-82M-Instruct

PicoLM Project โ€ข 82.2M Extended-Context & Scientific Reasoning Model (4K YaRN)

Overview

PicoLM-V3-Pro-82M-Instruct is the flagship 82.2-million parameter compact reasoning model from the PicoLM Project. Trained up to the Chinchilla compute-optimal boundary of 1.62 Billion tokens, it introduces an extended native 4,096-token context window powered by YaRN RoPE frequency interpolation and integrated step-by-step chain-of-thought (<thought>) reasoning capabilities.

While PicoLM-V3 Standard excels in natural dialogue and everyday commonsense, PicoLM-V3-Pro is engineered for academic knowledge, scientific QA, and multi-step deduction (+4.22 points higher on ARC-Easy).


Model Specifications

Attribute Specification
Total Parameters 82,233,792 unique parameters (tied I/O embeddings, 0 dead weights)
Physical Transformer Blocks 21 layers
Effective Layers (Recurrent Pass) 42 effective layers ($21 \times 2$ macro-loop passes)
Hidden Dimension ($d_{\text{model}}$) 576
Attention Architecture Grouped-Query Attention (9 Query Heads, 3 KV Heads; GQA 3:1)
Head Dimension ($d_{\text{head}}$) 64
Feed-Forward Dimension ($d_{\text{ffn}}$) 1,664 (SwiGLU activation)
Factorized Embedding Projection $24,576 \rightarrow 128 \rightarrow 576$ (Rank-128 linear bottleneck)
Normalization Pre-LN RMSNorm ($\epsilon = 10^{-5}$) with split-pass independent gains
Positional Encoding YaRN RoPE (NTK-by-parts, scale factor 2.0, $\theta = 10,000.0$)
Context Length 4,096 tokens native sequence length (adapted from 2,048 tokens)
Vocabulary Size 24,576 BPE tokens

Standardized Evaluation & Generational Progress

All models evaluated strictly using the EleutherAI LM-Evaluation-Harness standard:

Model Parameters Tokens Context Legacy Scale (250 Q, Raw) Official ARC-Easy (acc_norm) PIQA (acc_norm) HellaSwag (acc_norm) ARC-Challenge (acc_norm)
PicoLM-V3-Pro (Base) 82.2M 1.62B 4,096 46.00% 43.54% 57.89% 34.95% 24.06%
PicoLM-V3-Pro (Instruct) 82.2M 1.62B+SFT 4,096 44.80% 39.10% 55.17% 33.55% 23.20%
PicoLM-V3 (Standard) 82.2M 758M 2,048 44.80% 39.32% 58.49% 36.00% 23.63%
PicoLM-V2.1-Instruct 81.9M ~380M 2,048 42.00%* (legacy 250Q only) ~56.5% ~32.0% ~22.5%
PicoLM-V2-Instruct 81.9M ~380M 2,048 42.00%* (legacy 250Q only) ~56.5% ~32.0% ~22.5%
GPT-2 (OpenAI) 124M ~10B 1,024 - 31.40% 62.80% 31.50% 22.10%
MobileLLM-125M (Meta) 125M 1.0T 2,048 - 43.90% 65.30% 38.90% 27.10%
SmolLM2-135M (HF) 135M 2.0T 8,192 - 43.90% 68.40% 42.10% 30.20%

*PicoLM-V2 and V2.1 were evaluated exclusively on the legacy unnormalized 250-question sample. On the exact same 250-question scale, PicoLM-V3-Pro outperforms V2 by +4.0 points (46.00% vs 42.00%).

Pre-Training Lineage & YaRN Extension

PicoLM-V3-Pro was built through a 5-stage progressive curriculum totaling 1,616,000,000 tokens:

  1. Part A & B (758M tokens): Foundation representations on FineWeb-Edu, Cosmopedia v2, Python-Edu, and TriviaQA.
  2. Part C Synth (378M tokens): Scientific graph concepts via Sutra-10B and step-by-step logic via FineMath-4+.
  3. Part D (378M tokens): Academic knowledge consolidation using AllenAI SciQ and concentrated FineMath.
  4. Part E Mid-Training (100M tokens, 4K YaRN): Context window extension to 4,096 tokens via YaRN RoPE (NTK-by-parts interpolation), preserving low-frequency grammar whilst scaling high-frequency tokens.
  5. Supervised Fine-Tuning (SFT-Gold): 4-pillar conversational alignment with chain-of-thought <thought> traces and anti-sycophancy anchors.

Quickstart & Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "aethertp/PicoLM-V3-Pro-82M-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.float32,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "What is the boiling point of water at sea level?"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=False,
        use_cache=False,
        eos_token_id=2
    )

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response.strip())

Limitations & Ethical Considerations

  • Model Scale: At 82.2M parameters, multi-step mental arithmetic without external tools remains fragile.
  • Language: Optimized strictly for English.
  • Decontamination: Training data was not decontaminated against benchmark test splits.
  • Tool Use: No native code execution or live web access.

Citation

@misc{picolmv3pro2026,
  author = {Emre Polat and PicoLM Project Contributors},
  title = {PicoLM-V3-Pro: Extended 4K Reasoning and Scientific Capabilities in Sub-100M Language Models},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/aethertp/PicoLM-V3-Pro-82M-Instruct}}
}
Downloads last month
296
Safetensors
Model size
82.2M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using aethertp/PicoLM-V3-Pro-82M-Instruct 1