metadata
license: apache-2.0
datasets:
- ILSVRC/imagenet-1k
model-index:
- name: MaskBit-Tokenizer-18bits
results:
- task:
type: image-generation
dataset:
name: ILSVRC/imagenet-1k
type: ILSVRC/imagenet-1k
metrics:
- name: rFID
type: rFID
value: 1.16
- name: InceptionScore
type: InceptionScore
value: 197.8
- name: LPIPS
type: LPIPS
value: 0.27
- name: PSNR
type: PSNR
value: 22
- name: SSIM
type: SSIM
value: 0.59
- name: CodebookUsage
type: CodebookUsage
value: 0.5
This model is the MaskBit tokenizer with a vocabulary size of 18bits. It uses a downsampling factor of 16 and is trained on ImageNet for images of resolution 256.
You can find more details on the project page and in the paper.