AsciiNet

A 28M-parameter network that turns any image into ASCII art that people can recognize, at any width from a 16-column thumbnail to a full terminal. One forward pass: 13 ms per image at 32 columns and 24 ms at 80 on an RTX 5080, 0.3 GB of VRAM. Trained from scratch (random initialization) in PyTorch.

Code, CLI and training scripts: github.com/SamuelReeder/text-to-ascii-ai

original vs ascii-image-converter vs the pipeline vs AsciiNet Left to right: original · ascii-image-converter · the hand-built pipeline (AsciiNet's teacher) · AsciiNet at 80 columns · AsciiNet at 32 columns. Held-out images.

$ python ascii.py -p "a steaming cup of coffee" -w 64
                                _
                               _|
                             _/)\,
                             `) |^_
                            ||/ ^,^,
                            \|\  ^\|\
                             \|\   \/
                              `,\  /
                               |`
                   _.><^^^`````` ````^^>><._
                 /``<_.><^^^^`````^^^^>><._`,\;
                 |><|<<.,____@@@@@____,,.><><^|_.<>><_
                 |||||``^^^^>>>>>><^^^^``|||||@@_><<_|\
                 |||||||||:::::::::::::::::|||@,`   ^@|/
                 |||||||:::::::::::::::::::::@@`     @|\
                 ^||||||::::::::::::::::::::\@/    _'@/
                  \|||||::::::::::::::::::::@|__,>^`_/
                   ,|||||::::::::::::::::::|@@@@_.<^
           _.>^^`   \||||||::::::::::::::|@_><^`  _
         /`          \_||||::::::::::::||_/        `^<
         _            |@|||||:::::||||@@/`            |
         `<>._      ``^<_@>@@,|_|,,@@_>^           ,</`
            `^<`^>><.,,__`^^^>>>><<^^   __,,.>>_^><`
                  `^><.,___```````````___,.><^`
                          ````^^^^````

Prompt mode: SANA-Sprint 0.6B draws the picture, AsciiNet converts it.

Usage

git clone https://github.com/SamuelReeder/text-to-ascii-ai && cd text-to-ascii-ai
./setup.sh && conda activate ascii
python ascii.py photo.jpg -w 80          # downloads these weights on first use
python ascii.py photo.jpg -w 32 --color
python ascii.py -p "a lighthouse on a cliff"
from asciiart.neural import NeuralConverter
from asciiart.pipeline import Options

net = NeuralConverter()                  # fetches SamuelReeder/asciinet, then cached
print(net.text("photo.jpg", Options(cols=40)))

Files: model.safetensors (EMA weights, fp32) and config.json (the 26-character set, in the order of the output logits). The architecture is asciiart/model.py; preprocessing is in asciiart/neural.py.

Architecture

  context image (~384 px) ─► ConvNeXt stages ─► 6 transformer blocks ─► FPN ─► subject map (aux)
                                                                          └─► context features ─┐
  grid image (8×16 px per cell) ─► conv stages ─► per-cell features ─────────────────────────────┤
  exact cell pixels ─────────────────────────────────────────────────► per-cell features ────────┤
  options (dark/light background, invert, focus) ───────────────────► embedding ────────────────┤
                                                                                                 ▼
                                   6 ConvNeXt blocks on the character grid ─► 26 glyph logits per cell
  • The context branch (24M params, 19M of them in the transformer) always sees the whole picture at ~384 px, so it finds the subject even when the output is 24 characters wide.
  • The grid branch sees every pixel of every cell, so glyph shapes follow fine edges.
  • The cell stage chooses each character with its neighbours in view, so strokes connect.
  • It is fully convolutional over the grid: one set of weights serves every width and all 12 combinations of options.

Training

AsciiNet is distilled from a hand-built pipeline (BiRefNet subject matte → tone mapping → glyph matching with stroke characters), which is slow (~110 ms, 1.6 GB) but legible. Losses: cross-entropy on the teacher's glyph per cell, BCE (+ soft Dice in phase 2) on BiRefNet's subject matte, L1 on the teacher's ink per cell.

  • Data: 187k images from 46 public datasets (photos of objects, scenes and people; art, sketches, doodles; cartoons, anime, pixel art, emoji; logos, icons, charts, text; screenshots, 3D renders, aerial and satellite images; textures), plus 6,000 SANA-Sprint pictures. 97k were used for training. Caltech-101 and Flickr8k were held out entirely for evaluation.
  • Widths: every width from 16 to 128 columns, sampled during training.
  • Schedule: 30k steps from random initialization, then 20k with the subject-map losses up-weighted. ~4.7 h on one RTX 5080 (10 GB VRAM), 0.34 s per step.
  • On held-out images the network picks the teacher's character for 75% of cells at 24 columns and 81% at 80.

Evaluation

Qwen3-VL-8B reads the rendered ASCII and answers 5-way multiple choice (which class, scene, dish, breed or caption does it show?) over 8 held-out sources, 480 questions per row. Chance = 0.20; the original photos score 0.99. Means are good to about ±0.02.

method 24 cols 32 cols 80 cols
ascii-image-converter 0.19 0.19 0.19
hand-built pipeline (teacher) 0.24 0.26 0.35
AsciiNet 0.23 0.24 0.33

On pictures with one clear subject (90 SANA-Sprint prompts, which of 5 prompts does the ASCII show?), AsciiNet scores 0.81 at 80 columns, 0.76 at 40 and 0.61 at 32.

Limitations

  • It cannot beat its teacher: it learns to copy it. The gap that remains is on cluttered photos where the subject must be cut out exactly (Caltech-101: 0.53 vs the teacher's 0.78 at 80 columns).
  • At 24–32 columns fine-grained content (a dog's breed, a particular dish) is not recoverable, by this or any method measured.
  • Output uses 26 printable ASCII characters in a monospace grid of 8×16-pixel cells; it assumes your font has roughly that 1:2 aspect.

License and training data

The weights are released under GPL-3.0, like the code. The training images come from public datasets under their own licenses, and some of those (for example ImageNet, CelebA and FFHQ) permit research or non-commercial use only. Check them before commercial use. No training images are redistributed here.

Downloads last month
21
Safetensors
Model size
27.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support