Artificial Stupidity β the tiny models (AS-0 β¦ AS-5)
Six character-level language models built from scratch β no GPT-2, no pretrained anything β that are identical in every respect except how many bits each weight is allowed. The smallest is 83 KB.
They are the control experiment for ayushmaninbox/artificial-stupidity, which is fluent because OpenAI paid for the fluency. These are not fluent, and every parameter is homegrown.
The six
| Weight precision | Params | Packed size | Val loss | |
|---|---|---|---|---|
| AS-0 | 32-bit (control) | 815,488 | 3.1 MB | 1.642 |
| AS-1 | 8-bit | 819,072 | 835 KB | β |
| AS-2 | 4-bit | 819,072 | 448 KB | β |
| AS-3 | 1.58-bit {-1,0,+1} |
819,072 | 216 KB | β |
| AS-4 | 1-bit {-1,+1} |
819,072 | 169 KB | 1.756 |
| AS-5 | 1-bit + 8-bit activations | 350,784 | 83 KB | β |
What they actually produce
Real output, same prompt, same seed (python compare.py --seed 1337):
why is the sky blue
| Model | Size | Answer |
|---|---|---|
| AS-0 | 3.1 MB | ask ays what is 17 |
| AS-1 | 835 KB | no it isnt |
| AS-2 | 448 KB | a lot |
| AS-3 | 216 KB | ok |
| AS-4 | 169 KB | no it is rain |
| AS-5 | 83 KB | no 11 whats 4 |
what is 2 + 2 β every one of them answers 127 (AS-0 says 1).
They learned the shape of an arithmetic reply from the synthetic corpus
without learning arithmetic.
what is your name
| Model | Answer |
|---|---|
| AS-0 | definitely not |
| AS-2 | definitely not ryand |
| AS-4 | dog |
| AS-5 | do 1542 plus 40 13 |
They hold a conversational register β correctly spelled, real chat moves β while being completely disconnected from the question. AS-5 at 83 KB is where it visibly breaks: it keeps the rhythm of internet chat but can no longer stay on one topic for a whole sentence.
Quantization-aware, not post-hoc
You cannot train a normal model and round its weights to 1 bit afterwards β you get static, and worse, you cannot tell "compression worked" from "my code is broken". These models know they are being squashed while they learn, via a straight-through estimator, so they route around the damage.
1 bit is not a 32Γ saving but about 20Γ: embeddings, LayerNorms and one fp16 scale per weight row stay high-precision. BitNet keeps them too.
Limitations
- Character-level, so they cannot spell reliably β "mitochondria" is twelve consecutive guesses and they lose that bet.
- Trained on 9.8 MB of Twitch chat, Reddit, YouTube transcripts and lyrics.
- They are not useful. They are a measurement of where language breaks.
Usage
git clone https://github.com/ayushmaninbox/artificial-stupidity
cd artificial-stupidity/as-text-model
pip install -r requirements.txt
python generate.py AS-4 --prompt "hello"
python compare.py # all six, side by side
License
MIT for the code and weights. The training corpus is scraped third-party text β see the dataset card.