playgroundai
/

playground-v2.5-1024px-aesthetic

StableDiffusionXLPipeline

Inference Endpoints

Model card Files Files and versions Community

aykamko commited on Feb 23

Commit

a1c2115

•

1 Parent(s): b58807a

Create README.md

Files changed (1) hide show

README.md +50 -0

README.md ADDED Viewed

	@@ -0,0 +1,50 @@

+# Playground v2.5 – 1024px Aesthetic Model
+This repository contains a model that generates highly aesthetic images of resolution 1024x1024. You can use the model with Hugging Face 🧨 Diffusers.
+[insert teaser image]
+**Playground v2.5** is a diffusion-based text-to-image generative model, and a successor to Playground v2 [link].
+Playground v2.5 is near state-of-the-art in aesthetic quality. Our user studies demonstrate that our model outperforms SDXL, Playground v2, PIXART-alpha, DALL-E 3, and Midjourney 5.2.
+For details on the development and training of our model, please refer to our blog post [link] and technical report [link]
+### Model Description
+- **Developed by:** Playground
+- **Model type:** Diffusion-based text-to-image generative model
+- **License:** Playground v2 Community License
+- **Summary:** This model generates images based on text prompts. It is a Latent Diffusion Model that uses two fixed, pre-trained text encoders (OpenCLIP-ViT/G and CLIP-ViT/L). It follows the same architecture as Stable Diffusion XL.
+### Using the model with 🧨 Diffusers
+TODO
+### Using the model with Automatic1111/ComfyUI
+TODO
+### User Studies
+This model card only provides a brief summary of our user study results. For extensive details on how we perform user studies, please check out our technical report: [link]
+We conducted studies to measure overall aesthetic quality, as well as for the specific areas we aimed to improve with Playground v2.5, namely multi aspect ratios and human preference alignment.
+The aesthetic quality of Playground v2.5 dramatically outperforms the current state-of-the-art open source models SDXL and PIXART-α, as well as Playground v2. Because the performance differential between Playground V2.5 and SDXL was so large, we also tested our aesthetic quality against world-class closed-source models like DALL-E 3 and Midjourney 5.2, and found that Playground v2.5 outperforms them as well.
+[insert graph for outperforming all of these models]
+Similarly, for multi aspect ratios, we outperform SDXL by a large margin.
+[insert graph for multi aspect ratios]
+Next, we benchmark Playground v2.5 specifically on people-related images, to test Human Preference Alignment. We compared Playground v2.5 against two commonly-used baseline models: SDXL and RealStock v2, a community fine-tune of SDXL that was trained on a realistic people dataset.
+Playground v2.5 outperforms both baselines by a large margin.
+[insert graph for people prompts]
+Lastly, we report metrics using our MJHQ-30K benchmark which we open-sourced with the v2 release. <link> We report both the overall FID and per category FID. All FID metrics are computed at resolution 1024x1024. Our results show that Playground v2.5 outperforms both Playground v2 and SDXL in overall FID and all category FIDs, especially in the people and fashion categories. This is in line with the results of the user study, which indicates a correlation between human preferences and the FID score of the MJHQ-30K benchmark.
+[insert graph for MJHQ-30K benchmark]
+### How to cite us
+TODO