Logo GAN β€” 3.5M-param DCGAN for logo-like image generation

A small Deep Convolutional GAN trained from scratch to generate 64Γ—64 logo-like images. This is the deliverable for Compactbot/model-requests #1 (CompactAI: "a small model that generates logo-like images").

What it is

A generator/discriminator pair. The generator maps a 128-dim latent vector to a 64Γ—64Γ—3 image; the discriminator classifies real vs. generated. Total parameter count 3,498,762 (generator 2,838,278 + discriminator 660,484) β€” a genuinely small model that trains on a single GPU in under an hour.

Architecture

Generator (latent 128 β†’ 64Γ—64Γ—3, output in [-1,1]):

Linear(128, 256Β·8Β·8) β†’ BatchNorm β†’ ReLU β†’ reshape (256,8,8)
ConvTranspose2d(256,128,4,2,1) β†’ BatchNorm β†’ ReLU
ConvTranspose2d(128,64,4,2,1)  β†’ BatchNorm β†’ ReLU
ConvTranspose2d(64,3,4,2,1)    β†’ Tanh

Discriminator (64Γ—64Γ—3 β†’ scalar logit):

Conv2d(3,64,4,2,1)   β†’ BatchNorm β†’ LeakyReLU(0.2)
Conv2d(64,128,4,2,1) β†’ BatchNorm β†’ LeakyReLU(0.2)
Conv2d(128,256,4,2,1)β†’ BatchNorm β†’ LeakyReLU(0.2)
AdaptiveAvgPool2d(1) β†’ Flatten β†’ Linear(256,1)
Component Params
Generator 2,838,278
Discriminator 660,484
Total 3,498,762

Data

100 logos, each 64Γ—64Γ—3 RGB, bilinearly resized and normalized to [0,1]. Sourced from the first 100 valid images of the samp3209/logo-dataset train split (streamed). Stored as logos64.npy (shape (100,64,64,3), float32).

Training

  • Optimizer: Adam, lr 2e-4, betas (0.5, 0.999) β€” standard DCGAN recipe
  • Batch size 64, 4000 steps, 1 D step + 1 G step per iteration
  • Loss: BCEWithLogitsLoss
  • Hardware: NVIDIA RTX 5090 (CUDA)
  • Final step losses: d_loss 0.0711, g_loss 3.5900

Quality (measured, not asserted)

Generated a 64-sample grid (seed 123) and compared its statistics to the real logos:

Metric Generated Real
Pixel mean 0.572 0.574
Pixel std 0.350 0.360
Colorfulness (L1 norm of channel stds) 0.605 0.622
Median pairwise L2 (mode-collapse proxy) 52.1 β€”
Fraction of near-duplicate pairs (<0.01) 0.000 β€”

The generated distribution closely tracks the real one in mean, variance and colorfulness, and there is no mode collapse (outputs are diverse, not repeated). See final_grid.png (64 generated) and final_real.png (64 real) side by side.

What it is and is not good at

  • Good at: producing diverse, logo-like 64Γ—64 images β€” the right color statistics and composition of real logos, without collapsing to one mode.
  • Not good at: crisp, recognizable logos. 100 images is a very small dataset, so the model captures the distribution of logos (color, layout, contrast) rather than specific, legible marks. Treat outputs as logo-flavored compositions, not usable brand assets.

Files

  • final.pt β€” final checkpoint (g, d state dicts, step, zdim)
  • train_logo_gan.py β€” full training script (reproducible)
  • final_grid.png β€” 64 generated samples (seed 123)
  • final_real.png β€” 64 real logos for comparison

Reproduce

# data prep (first 100 logos β†’ /work/logos/logos64.npy)
python gan_prep3.py
# train
python train_logo_gan.py --steps 4000 --batch 64 --zdim 128 --lr 2e-4

Honest note on scope

The original request thread discussed a larger config (β‰ˆ8.8M params, latent 100, 12000 steps). The run that actually completed and is shipped here is the smaller 3.5M / zdim-128 / 4000-step config β€” it trained cleanly to completion and passes the quality checks above. The larger config is a future follow-up, not what this checkpoint is.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support