Piccaso-0.1
A 102M-parameter model that answers a text prompt with 361 brush strokes instead of pixels.
Piccaso started as a curiosity question: is it easier for a small model to paint a picture than to make one? Each painting is 361 quadratic Bezier strokes (position, curve, width, colour, on/off), rendered by a fixed renderer to PNG or exported as a real SVG.
Early research preview. Trained on 232k images for about 9.2M picture-views (~8 GPU-hours on RTX 5090s). It gets colour, light and layout right and objects some of the time. It is not a production image generator, and it is nowhere near pixel models trained on 100M+ images. The learning curves were still rising; more data and training should help a lot. Read the limitations below.
What it does well, and what it doesn't
- Good: colour (97% right on our prompt benchmark), light, mood and layout; landscapes, sunsets, interiors, food, vehicles, portraits as painted heads; single main subjects.
- Weak: two separate objects in one scene (4%), exact shapes, identity, text, fine detail; doodle and emoji styles (out of domain).
- By design: a painterly, quick-oil-sketch look. The 361-stroke format itself caps realism (see "ceiling" below).
Results
Caption retrieval with CLIP ViT-B/32: does the painting match its own caption better than the other captions of the set?
| Top-1 among 200 | Real photo | Fitted 361 strokes (ceiling) | Piccaso-0.1 |
|---|---|---|---|
| Seen (training captions) | 96.5% | 83.5% | 28.5% |
| Unseen (held-out PixelProse) | 98.0% | 87.5% | 31.0% |
| DOCCI (different photo source, human captions) | 80.0% | n/a | 9.5% |
Chance is 0.5%. Seen and unseen are the same: the model generalises rather than memorising.
| StrokeBench, 200 fixed prompts | Score | Chance |
|---|---|---|
| Right object, top-1 / top-5 of 80 | 30.3% / 56.3% | 1.3% / 6.3% |
| Right colour (of 10) | 96.9% | 10% |
| Both objects in two-object prompts | 4.4% | 0.4% |
| Right style (of 5) | 45.6% | 20% |
Unseen captions. Columns: real photo, its fitted 361 strokes (what the format can show), two Piccaso paintings from the caption alone.
Usage
git clone https://huggingface.co/shing-dev/Piccaso-0.1 && cd Piccaso-0.1/code
pip install torch open_clip_torch ftfy regex safetensors huggingface_hub pillow numpy
python paint.py "a lighthouse on a cliff at sunset, oil painting" --n 4 --model .. --out paintings
Each painting is saved as a 512 px PNG and an SVG with 361 <path> strokes. The Long-CLIP-B text encoder
(BeichenZhang/LongCLIP-B, ~600 MB) is downloaded on first run. Speed: about 0.11 s per painting on an RTX 5090 (batch 50), about
90 s on a laptop CPU (25 steps). Defaults: 25 DDIM steps, guidance 3.
Model
| Output | 361 strokes x 11 numbers: 165 base strokes on 4x4 / 7x7 / 10x10 grids + 196 detail strokes on a 14x14 grid |
| Network | set diffusion transformer (DiT-style, adaLN), width 640, 11 layers, 10 heads, 102M parameters |
| Text | Long-CLIP-B, frozen: pooled vector + cross-attention to up to 248 caption tokens |
| Training | v-prediction, cosine schedule, self-conditioning, 25% reference-image conditioning; 8k steps at batch 256 + 7k at batch 1,024 |
| Data | 232,134 PixelProse images (clean, aesthetic >= 5), converted to strokes by gradient-descent fitting, filtered by how well the stroke render still matches the caption |
| Files | model.safetensors (EMA weights + stroke normalisation, anchors, caption standardisation), config.json, code/ |
How it compares to a real image model
MobileDiffusion (Google, 2023) trains on 150M images with weeks of TPU time and runs in ~0.2 s on a phone. Piccaso-0.1 saw 232k images (about 650x fewer) for ~8 GPU-hours, and the whole project cost about $65 of rented GPUs. The gap is mostly data and compute, not the stroke format. Few-step distillation and more data are the obvious next steps.
Limitations and responsible use
- Research preview; outputs are often wrong or abstract. Do not use it where accuracy matters.
- Trained on web images with machine-written captions (PixelProse); it inherits their biases and gaps.
- It cannot reproduce identities, logos or text in any recognisable way at this scale.
Links and credits
- Code, full lab notebook and report: github.com/shing1Sks/piccaso (private for now; public soon)
- Data: PixelProse (captions CC-BY-4.0); DOCCI for testing only
- Text encoder: Long-CLIP (Apache-2.0, code vendored in
code/longclip/) - Research by Shreyash Kumar Singh, run with an AI coding agent (Claude). Weights: CC-BY-4.0.
- Downloads last month
- 12
