mindview-t2i

A text-to-image model in one file of 980 MB. You give it a prompt and it paints a 512 × 512 picture in two steps (or four, a little better). It has 4.48 billion weights, and all but 33 million of them are −1, 0 or +1. It runs in a web browser on WebGPU, with no server: try it. On a MacBook Air (M4) in Chrome a picture takes about 10 s in two steps and 16 s in four.

It came out of mindview, an art installation that shows these models computing. Most of it is the work of others (see Sources and licenses); what is new is the way the parts are joined, the map that joins them, and the packing.

What is in the file

Part What it does Weights Stored
reader Ternary Bonsai 1.7B with its tokenizer, cut to its first 9 of 28 layers. It reads the prompt. 764 M, ternary 158.4 MB
cond One linear map from the reader's states at layers 3, 6 and 9 to the painter's text stream (6,144 → 3,072), trained for this model. 18.9 M, f16 35.3 MB
painter The diffusion transformer of Bonsai Image 4B, which is FLUX.2 [klein] 4B with ternary weights. The modulation for its 4-step schedule is precomputed. 3,682 M, ternary 761.3 MB
decoder TAEF2, which turns the latent into pixels. 1.3 M, f16 2.5 MB
few-step LoRA radames/FLUX.2-klein-Sana-Sprint cut to rank 8: a side branch on each of the painter's 100 ternary matrices, on for 1 and 2 steps. 12.0 M, f16 20.8 MB
schedules Noise levels and the painter's modulation for 1, 2, 3 and 4 steps. 1.9 MB

That is 1.7 bits per weight on average, all parts included.

How the parts fit

Bonsai Image 4B conditions its transformer on the text encoder of FLUX.2 [klein] 4B: Qwen3-4B, whose states at layers 9, 18 and 27 go through the transformer's context embedder. Here a much smaller reader stands in for it, and the adapter and the context embedder are one matrix.

The map was fitted by ridge regression on 1,500 prompts, to what the stock Qwen3-4B encoder gives for them. It is scored where that matters, in the transformer's input space after the context embedder. A sweep over which reader layers to tap found that shallow layers serve almost as well as deep ones. That decides the file's size, because every layer the map does not read is a layer the file does not carry.

The table gives the per-token cosine with the stock encoder in the transformer's input space. "Validation" is 150 prompts held out of the fit; "held out" is 12 other prompts. Each map was fitted on the other 1,350 prompts, as a float map and as a ternary map with least-squares scales.

Reader layers tapped Layers carried Float, validation Float, held out Ternary, held out
7, 14, 21 21 0.939 0.961 0.940
6, 12, 18 18 0.938 0.960 0.938
5, 10, 15 15 0.937 0.959 0.939
7, 11, 15 15 0.936 0.957 0.935
4, 8, 12 12 0.937 0.959 0.939
6, 9, 12 12 0.936 0.958 0.934
3, 6, 9 9 0.937 0.958 0.936

The map in this file is the float one for taps 3, 6 and 9, refitted on all 1,500 prompts. Its held-out cosine is 0.959. The adapter the mindview site used before this, a ternary one on taps 7, 14 and 21, scores 0.952. A ternary map would save another 31 MB, at 0.938.

The same prompts and seed, painted at 4 steps by the reader and painter the mindview site used before (left) and by this file (right):

Eight prompts painted by the earlier pipeline and by mindview-t2i

On these prompts this file counts the pear and spells OPEN where the earlier pipeline did not. Eight prompts do not make a benchmark; the cosine table above is the measured part.

Two steps

FLUX.2 [klein] 4B is already distilled to 4 steps. radames' SANA-Sprint LoRA, trained on the original klein, takes it to 1–2 steps, and it carries over to the ternary transformer of Bonsai Image 4B without retraining. It runs as a side branch, y = W x + B (A x), so the ternary weights stay as they are: merged into them and made ternary again, it changes one trit in a million and does nothing.

Its best rank-8 approximation (per matrix, an SVD of B A) paints like the full rank 256 on every test prompt, so the file carries it at 3% of the original size. At 2 steps the painter needs more of the prompt's padding than at 4: this file is run with 256 text rows at 2 steps and with the prompt plus a few pads at 4. At 1 step faces go soft.

4 steps (left) and 2 steps with the LoRA (right), same prompts and seed, as painted in the browser:

Eight prompts at 4 steps and at 2 steps

Running it

The file is read by the mindview browser runtime (WebGPU and WGSL, no server): src/lib/runtime/packed.ts reads the file, bonsai-llm.ts runs the reader and painter.ts the rest. The page at neovand.github.io/mindview/paint loads it from this repository with range requests, so it is never all in memory at once, and the browser keeps what it downloaded.

It is a GGUF file (v3), but its architecture is its own (mindview-t2i) and four of its tensor types are new, so llama.cpp and other GGUF tools cannot run it as it is.

Format

Type What it holds
200 (trits) A ternary matrix, row-major, 5 weights per byte: q0 + 3·q1 + 9·q2 + 27·q3 + 81·q4, with q = weight + 1. Padded to a multiple of 80 weights (16 bytes). The value is weight × scale, and the scales are the f16 tensor <name>.scale, one per 128 inputs of each row.
202 (deflate) Another tensor, deflated: u32 inner type, u32 inflated bytes, u32 deflated bytes, u32 0, then the raw deflate stream. It is used wherever deflating saves anything.
203 (f16 planes) f16 values as two byte planes: every high byte, then every low byte. Scales deflate to less than half this way.
204 (JSON) UTF-8 JSON: the tokenizer's vocabulary, merges and chat template.

The LoRA is stored as lora.<matrix>.a (A, [rank][inputs]) and lora.<matrix>.bt (B transposed, [rank][outputs]), f16 in byte planes; mindview.lora gives its rank and source. sched.mod holds the per-step modulation for the step counts in mindview.schedules.

general.architecture is mindview-t2i. The reader's settings keep their qwen3.* keys, and its part of the file loads as the Qwen3 it is. The metadata key mindview.painter holds what the runtime needs: the prompt template, the taps, the image layout, the schedule and the decoder's graph. mindview.stored holds the stored size of each deflated tensor, and mindview.sources holds where each part came from. The packer checks the file as it writes it: every tensor reads back exactly.

Limits

  • It paints only 512 × 512, in 1 to 4 steps: the modulation is precomputed for those schedules. 1 step is fast but soft; 2 steps (with the LoRA) and 4 steps (without) are the ones to use.
  • The map from the reader imitates the stock encoder closely but not exactly (cosine 0.959). It follows a prompt less well than Bonsai Image 4B with its own encoder, and it has had little testing beyond short English prompts.
  • It keeps the limits and biases of FLUX.2 [klein] and Bonsai Image. There is no safety filter.

Reproducing it

From the mindview repository, with the source models in static/models and the Hugging Face cache:

python research/scripts/adapter_layers.py stats        # one pass of the reader over the prompts, every tap set
python research/scripts/adapter_layers.py fit          # the sweep above
python research/scripts/adapter_layers.py final 3 6 9  # the map, fitted on every prompt
python research/scripts/pack_model.py --fused research/data/adapter/fused_final_layers_3_6_9.pt --float \
  --lora <path to radames/FLUX.2-klein-Sana-Sprint's pytorch_lora_weights.safetensors> --lora-rank 8

Sources and licenses

  • Ternary Bonsai 1.7B by PrismML, Apache 2.0. Layers 0–8 and the tokenizer of Ternary-Bonsai-1.7B-Q2_0.gguf, with every ternary weight exact.
  • Bonsai Image 4B by PrismML, Apache 2.0. It is FLUX.2 [klein] 4B by Black Forest Labs, Apache 2.0, with its weights made ternary. The transformer is repacked with every ternary weight exact. Its context embedder is folded into the map.
  • TAEF2 by Ollin Boer Bohan (madebyollin/taef2), MIT license. It is repacked unchanged.
  • The few-step LoRA by radames (radames/FLUX.2-klein-Sana-Sprint), Apache 2.0, distilled from FLUX.2 [klein] 4B. Its rank-8 approximation is included.
  • The map was fitted for mindview on the models above and on the outputs of Qwen3-4B (Apache 2.0). It is released under Apache 2.0.
Downloads last month
102
GGUF
Model size
5B params
Architecture
mindview-t2i
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohsenvand/mindview-t2i