PKaI Nano 1.2
PKaI Nano 1.2 is a 300M-class base language model from PowderKeg Intelligence. It is a compact LLaMA-style decoder model trained from scratch with a Mistral tokenizer and released as a PKaI-native artifact. It succeeds PKaI Nano 1.1 with a broader, higher-quality training mix and improves on it across nearly every benchmark in our evaluation suite.
This repository contains PKaI-native weights and metadata, not a drop-in
transformers.AutoModelForCausalLM package.
๐ For the full writeup and evaluation details, see the announcement post: Introducing PKaI Nano 1.2: Better Data, Stronger Results.
Benchmarks
Evaluated in our standard benchmark suite against the previous Nano releases. Accuracy values are percentages; higher is better. WikiText is word perplexity; lower is better.
| Benchmark | PKaI Nano 1.2 | PKaI Nano 1.1 | PKaI Nano 1 |
|---|---|---|---|
| HellaSwag | 39.71 | 36.52 | 31.02 |
| SciQ | 81.50 | 80.00 | 73.00 |
| PIQA | 65.94 | 64.96 | 59.74 |
| WinoGrande | 52.96 | 52.72 | 52.57 |
| ARC-Easy | 49.92 | 45.45 | 43.35 |
| ARC-Challenge | 27.73 | 28.24 | 24.83 |
| LAMBADA OpenAI | 44.07 | 33.82 | 23.33 |
| WikiText (ppl โ) | 24.13 | 35.47 | 57.37 |
Across the seven accuracy tasks, PKaI Nano 1.2 averages 51.69, up from 48.82 for PKaI Nano 1.1 and 43.98 for PKaI Nano 1. See the announcement post for the full writeup and evaluation notes.
PKaI Nano 1 and PKaI Nano 1.1 were re-scored under our updated evaluation methodology for this comparison; their figures may differ slightly from their original release notes.
Files
model.safetensors: PKaI base model weights.config.json: PKaI model architecture configuration.tokenizer.json: PKaI tokenizer metadata.tokenizer.model: SentencePiece tokenizer model frommistralai/Mistral-7B-v0.1.THIRD_PARTY_NOTICES.txt: tokenizer and training-data provenance notices.LICENSE: Apache License, Version 2.0.
Architecture
- Parameters: 311,218,176
- Vocabulary size: 32,000
- Context length: 1024
- Layers: 24
- Attention heads: 16
- KV heads: 8
- Embedding size: 1024
- Tied embeddings: yes
- QK normalization: yes
Training Data
Training data included the following publicly disclosed sources, each processed with best-effort in-house decontamination and deduplication by PowderKeg Intelligence prior to training:
HuggingFaceFW/fineweb, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card.HuggingFaceFW/fineweb-edu, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card.HuggingFaceTB/smollm-corpus(cosmopedia-v2subset), released under the Open Data Commons Attribution License (ODC-By) v1.0. Cosmopedia v2 is synthetic text generated withmistralai/Mixtral-8x7B-Instruct-v0.1.wikimedia/wikipedia(20231101.en snapshot), released under the Creative Commons Attribution-ShareAlike 4.0 License (CC BY-SA 4.0).open-web-math/open-web-math, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card.
The training mix also included public-domain book text and other public text
sources. See THIRD_PARTY_NOTICES.txt for source URLs, licenses, and
attribution notes.
License
Copyright 2026 PowderKeg Intelligence LLC.
The PKaI Nano 1.2 model artifact is released under the Apache License, Version
2.0. The bundled tokenizer and training-data sources have their own provenance
and notices listed in THIRD_PARTY_NOTICES.txt.
Limitations
This is a small base model and has not been instruction-tuned or safety-tuned. It may produce inaccurate, unsafe, biased, or otherwise unsuitable text. Users are responsible for evaluating fitness, safety, and legal compliance for their own use cases.
- Downloads last month
- -