PicoLM-V3-82M-Instruct

PicoLM Project โ€ข Ultra-Compact 82.2M On-Device Language Model

Overview

PicoLM-V3-82M-Instruct is an ultra-compact, 82.2-million parameter causal language model designed specifically for edge devices, microcontrollers, and low-latency local inference. Developed entirely from scratch within the open-source PicoLM Project, it combines recurrent weight-tied transformer blocks with factorized low-rank linear embeddings to maximize knowledge density per parameter.

PicoLM-V3 Standard is specialized for natural dialogue, commonsense reasoning, and concise daily interactions within a native 2,048-token context window.


Model Specifications

Attribute Specification
Total Parameters 82,233,792 unique parameters (tied I/O embeddings, 0 dead weights)
Physical Transformer Blocks 21 layers
Effective Layers (Recurrent Pass) 42 effective layers ($21 \times 2$ macro-loop passes)
Hidden Dimension ($d_{\text{model}}$) 576
Attention Architecture Grouped-Query Attention (9 Query Heads, 3 KV Heads; GQA 3:1)
Head Dimension ($d_{\text{head}}$) 64
Feed-Forward Dimension ($d_{\text{ffn}}$) 1,664 (SwiGLU activation)
Factorized Embedding Projection $24,576 \rightarrow 128 \rightarrow 576$ (Rank-128 linear bottleneck)
Normalization Pre-LN RMSNorm ($\epsilon = 10^{-5}$) with split-pass independent gains
Positional Encoding Rotary Position Embeddings (RoPE, $\theta = 10,000.0$)
Context Length 2,048 tokens native sequence length
Vocabulary Size 24,576 BPE tokens

Standardized Evaluation & Generational Progress

All models evaluated strictly using the EleutherAI LM-Evaluation-Harness standard:

Model Parameters Tokens Context Legacy Scale (250 Q, Raw) Official ARC-Easy (acc_norm) PIQA (acc_norm) HellaSwag (acc_norm) ARC-Challenge (acc_norm)
PicoLM-V3 (Standard) 82.2M 758M 2,048 44.80% 39.32% 58.49% 36.00% 23.63%
PicoLM-V3-Pro (Flagship) 82.2M 1.62B 4,096 46.00% 43.54% (Base) / 39.10% 57.89% 34.95% 24.06%
PicoLM-V2.1-Instruct 81.9M ~380M 2,048 42.00%* (legacy 250Q only) ~56.5% ~32.0% ~22.5%
PicoLM-V2-Instruct 81.9M ~380M 2,048 42.00%* (legacy 250Q only) ~56.5% ~32.0% ~22.5%
GPT-2 (OpenAI) 124M ~10B 1,024 - 31.40% 62.80% 31.50% 22.10%
MobileLLM-125M (Meta) 125M 1.0T 2,048 - 43.90% 65.30% 38.90% 27.10%
SmolLM2-135M (HF) 135M 2.0T 8,192 - 43.90% 68.40% 42.10% 30.20%

*PicoLM-V2 and V2.1 were evaluated exclusively on the legacy unnormalized 250-question sample. On the exact same 250-question scale, PicoLM-V3 outperforms V2 by +2.8 points (44.80% vs 42.00%).

Pre-Training & SFT Pipeline

1. Pre-Training (Part B - 758M Tokens)

  • Data Mixture: Cosmopedia v2, FineWeb-Edu (score $\ge 4$), Python-Edu, and curated synthetic QA.
  • Optimization: AdamW ($\beta_1=0.9, \beta_2=0.95$, weight decay 0.1), Cosine decay schedule to $1.0 \times 10^{-5}$, FP16 mixed precision.

2. Four-Pillar Supervised Fine-Tuning (SFT-Gold)

Aligned over 9,796 strictly filtered instruction conversations with assistant-only loss masking:

  1. Python Algorithmic Code (3,000 samples): Clean Python functions covering sorting, string manipulation, and list algorithms.
  2. Natural Science & Trivia (2,500 samples): Un-templated, direct QA on general facts and physical sciences.
  3. Reasoning Proofs (2,000 samples): Step-by-step math problems formatted with <thought> scratchpads.
  4. Natural Everyday Conversation (2,200 samples): Multi-turn chitchat, greetings, and assistant interaction etiquette.
  5. Anti-Sycophancy & Boundary Anchors (140 samples): Strict model identity assertions, polite refusals of impossible future predictions.

Quickstart & Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "aethertp/PicoLM-V3-82M-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.float32,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "There are 5 baskets. Each basket contains 6 pears. If 4 pears rot, how many good pears are left?"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=False,
        use_cache=False,
        eos_token_id=2
    )

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response.strip())

Limitations & Ethical Considerations

  • Scale Constraints: At 82.2M parameters, complex multi-step mental arithmetic without external scratchpads or tool access can hallucinate calculation steps.
  • Language Coverage: The model is optimized strictly for English text and code.
  • Decontamination: Training datasets were not decontaminated against benchmark test sets.
  • No Native Tool Use: Cannot browse the web or execute live code.

Citation

@misc{picolm2026v3,
  author = {Emre Polat and PicoLM Project Contributors},
  title = {PicoLM-V3: Efficient Sub-100M On-Device Language Modeling},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/aethertp/PicoLM-V3-82M-Instruct}}
}
Downloads last month
328
Safetensors
Model size
82.2M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using aethertp/PicoLM-V3-82M-Instruct 1