Q&A with the team
Ask any questions here!
Why have you decided to go with SD-VAE/TAESD and not with more modern VAEs like Qwen-VAE or Flux2-VAE?
Why have you decided to go with SD-VAE/TAESD and not with more modern VAEs like Qwen-VAE or Flux2-VAE?
Hi. The main reason was to make it easier to compare to pre-existing trains + because it compresses more so it should make the job easier for the model, in theory.
I haven't done an A/B on both however, to know for sure.
Well, running on TAESD provides a funny opportunity to combine it with some strong distillation scheme to make a "realtime hallucination" painting program.
And just because you operate at a low resolution, does not mean you can't make things look funny. :- )
in fact, at a low enough resolution, most coherent things are merely a set of gradients and texture patterns with start and end points.
Did you get that from Agate? That looks super interesting.
Well, unfortunately these are but handy-dandy low resolution material laying around.
The top one is a multi-step style transfer cocktail with Qwen 2.1, which results in quite a lot of (pleasant) hallucinations that behave well when scaled down.
The bottom example is simply a hand-drawn thingymajig with very visible gradients of varying sizes, start and end points (particularly on the gauntlets, and the left 'boot')
They serve only the purpose of a visual demonstration that 'looking interesting' is very possible even within the bounds of a tiny canvas.
As the main ingredient to producing them are greyscale gradients, and textures. The colors? Made up by the author in post-processing.
Most existing image models are far too slow for realtime application, but hallucinating "greyscale gradients following a particular structure, or texture pattern" in a small canvas, in realtime, is far more useful than it might seem at a glance. :- )
Actually good point @V33rGeer , i think such architectures can be used for procedural world generation, you tile the world and somehow also feed in the cross-attention layers tokens of "adjacent tile summaries".
Still would need to solve the long distance problem but maybe a linear attention approach or something to allow big context in small model sizes.

