AsciiNet
A 28M-parameter network that turns any image into ASCII art that people can recognize, at any width from a 16-column thumbnail to a full terminal. One forward pass: 13 ms per image at 32 columns and 24 ms at 80 on an RTX 5080, 0.3 GB of VRAM. Trained from scratch (random initialization) in PyTorch.
Code, CLI and training scripts: github.com/SamuelReeder/text-to-ascii-ai
Left to right: original · ascii-image-converter · the hand-built pipeline (AsciiNet's
teacher) · AsciiNet at 80 columns · AsciiNet at 32 columns. Held-out images.
$ python ascii.py -p "a steaming cup of coffee" -w 64
_
_|
_/)\,
`) |^_
||/ ^,^,
\|\ ^\|\
\|\ \/
`,\ /
|`
_.><^^^`````` ````^^>><._
/``<_.><^^^^`````^^^^>><._`,\;
|><|<<.,____@@@@@____,,.><><^|_.<>><_
|||||``^^^^>>>>>><^^^^``|||||@@_><<_|\
|||||||||:::::::::::::::::|||@,` ^@|/
|||||||:::::::::::::::::::::@@` @|\
^||||||::::::::::::::::::::\@/ _'@/
\|||||::::::::::::::::::::@|__,>^`_/
,|||||::::::::::::::::::|@@@@_.<^
_.>^^` \||||||::::::::::::::|@_><^` _
/` \_||||::::::::::::||_/ `^<
_ |@|||||:::::||||@@/` |
`<>._ ``^<_@>@@,|_|,,@@_>^ ,</`
`^<`^>><.,,__`^^^>>>><<^^ __,,.>>_^><`
`^><.,___```````````___,.><^`
````^^^^````
Prompt mode: SANA-Sprint 0.6B draws the picture, AsciiNet converts it.
Usage
git clone https://github.com/SamuelReeder/text-to-ascii-ai && cd text-to-ascii-ai
./setup.sh && conda activate ascii
python ascii.py photo.jpg -w 80 # downloads these weights on first use
python ascii.py photo.jpg -w 32 --color
python ascii.py -p "a lighthouse on a cliff"
from asciiart.neural import NeuralConverter
from asciiart.pipeline import Options
net = NeuralConverter() # fetches SamuelReeder/asciinet, then cached
print(net.text("photo.jpg", Options(cols=40)))
Files: model.safetensors (EMA weights, fp32) and config.json (the 26-character set, in the
order of the output logits). The architecture is asciiart/model.py; preprocessing is in
asciiart/neural.py.
Architecture
context image (~384 px) ─► ConvNeXt stages ─► 6 transformer blocks ─► FPN ─► subject map (aux)
└─► context features ─┐
grid image (8×16 px per cell) ─► conv stages ─► per-cell features ─────────────────────────────┤
exact cell pixels ─────────────────────────────────────────────────► per-cell features ────────┤
options (dark/light background, invert, focus) ───────────────────► embedding ────────────────┤
▼
6 ConvNeXt blocks on the character grid ─► 26 glyph logits per cell
- The context branch (24M params, 19M of them in the transformer) always sees the whole picture at ~384 px, so it finds the subject even when the output is 24 characters wide.
- The grid branch sees every pixel of every cell, so glyph shapes follow fine edges.
- The cell stage chooses each character with its neighbours in view, so strokes connect.
- It is fully convolutional over the grid: one set of weights serves every width and all 12 combinations of options.
Training
AsciiNet is distilled from a hand-built pipeline (BiRefNet subject matte → tone mapping → glyph matching with stroke characters), which is slow (~110 ms, 1.6 GB) but legible. Losses: cross-entropy on the teacher's glyph per cell, BCE (+ soft Dice in phase 2) on BiRefNet's subject matte, L1 on the teacher's ink per cell.
- Data: 187k images from 46 public datasets (photos of objects, scenes and people; art, sketches, doodles; cartoons, anime, pixel art, emoji; logos, icons, charts, text; screenshots, 3D renders, aerial and satellite images; textures), plus 6,000 SANA-Sprint pictures. 97k were used for training. Caltech-101 and Flickr8k were held out entirely for evaluation.
- Widths: every width from 16 to 128 columns, sampled during training.
- Schedule: 30k steps from random initialization, then 20k with the subject-map losses up-weighted. ~4.7 h on one RTX 5080 (10 GB VRAM), 0.34 s per step.
- On held-out images the network picks the teacher's character for 75% of cells at 24 columns and 81% at 80.
Evaluation
Qwen3-VL-8B reads the rendered ASCII and answers 5-way multiple choice (which class, scene, dish, breed or caption does it show?) over 8 held-out sources, 480 questions per row. Chance = 0.20; the original photos score 0.99. Means are good to about ±0.02.
| method | 24 cols | 32 cols | 80 cols |
|---|---|---|---|
ascii-image-converter |
0.19 | 0.19 | 0.19 |
| hand-built pipeline (teacher) | 0.24 | 0.26 | 0.35 |
| AsciiNet | 0.23 | 0.24 | 0.33 |
On pictures with one clear subject (90 SANA-Sprint prompts, which of 5 prompts does the ASCII show?), AsciiNet scores 0.81 at 80 columns, 0.76 at 40 and 0.61 at 32.
Limitations
- It cannot beat its teacher: it learns to copy it. The gap that remains is on cluttered photos where the subject must be cut out exactly (Caltech-101: 0.53 vs the teacher's 0.78 at 80 columns).
- At 24–32 columns fine-grained content (a dog's breed, a particular dish) is not recoverable, by this or any method measured.
- Output uses 26 printable ASCII characters in a monospace grid of 8×16-pixel cells; it assumes your font has roughly that 1:2 aspect.
License and training data
The weights are released under GPL-3.0, like the code. The training images come from public datasets under their own licenses, and some of those (for example ImageNet, CelebA and FFHQ) permit research or non-commercial use only. Check them before commercial use. No training images are redistributed here.
- Downloads last month
- 21