Model Review after 3 months usage
#10
by dodo321 - opened
I've been using this model for more than 3 months now, and I must say it is excellent.
Strong points:
- Knows a lot, especially about poses, actions, and interactions.
- High-quality images out of the box, especially when using the e2n RL LoRA on top of it. (I use Euler Simple, 20 steps, CFG 6 for more details, even if the generations take longer.)
- Quick 1024x1024 generations, 29 sec on my RTX 4070 12GB.
- The recent RL LoRA made the model feel smarter, understand the prompt or user intent better, and create less body horror.
Weaker points:
- Loss of knowledge from Klein Base -> average -> 3FPH. Just one of many examples: the "Walmart" logo is clearly known by Klein Base and average, but is partially lost in 3FPH.
- Loss of medium knowledge (loss of photography/camera-mode knowledge). While not impossible to generate, but difficult without a well-crafted prompt, some medium knowledge was also lost. Examples like candid, phone camera, or night camera mode are very diluted and way less convincing.
- Mode collapse. I'm not sure if it's actually mode collapse, but the model is not as diverse as I would like it to be. Keeping the same prompt and just changing the seed will usually slightly change the image angle/character/environment, but not in a meaningful way. I'm one of those people who like a lot of diversity in my generations, like DALL-E 3 gave, since I can just let the generations go on with the same prompt and be amazed with minimal effort. I know LLM enrichment kinda helps with that, and I'm using it, but it removes a certain interesting aspect of self-generation.
I'm pretty sure the first two weaker points are only a dataset issue.
Excellent work! Cannot say it enough.