Marqo-FashionCLIP image encoder, LiteRT dynamic int8

The image tower of Marqo/marqo-fashionCLIP (revision 44f4c655124ed71e90cbf528d82f508d69d8a81b), converted to LiteRT and quantised. The Robbeta Android app downloads it to compare outfit photos with the clothes in a wardrobe, on the phone.

File Bytes SHA-256
marqo-fashionclip-image-dynint8-v1.tflite 88 160 672 8aa00548e4f7768f97c9c11f1130d646bd6c1cf6b980451b5e9609a8295427b6

A later model gets a new file name; this file never changes.

Input and output

  • Input [1, 3, 224, 224] float32, NCHW, RGB.
  • Output [1, 512] float32, not normalised. Normalise it (L2) before taking cosines.
  • Preprocessing, as open_clip does for this model: resize the shorter side to 224 with Pillow bicubic, then centre-crop to 224 x 224 (Robbeta pads to a square first, so nothing is cropped), scale to 0..1, subtract mean (0.48145466, 0.4578275, 0.40821073) and divide by std (0.26862954, 0.26130258, 0.27577711).

How it was made

litert-torch 0.9.4 converted model.visual from open_clip 3.3.0 (torch 2.13.0) to a float32 .tflite; ai-edge-quantizer 0.9.0 then applied the dynamic_wi8_afp32 recipe (int8 weights, float activations, dynamic-range int8 matmuls). The script is tools/matcher/convert.py in the Robbeta repository.

Licence

Apache-2.0, the licence of the original model (see LICENSE). This file is a modified version of Marqo-FashionCLIP's weights: converted to LiteRT and quantised to int8. Marqo-FashionCLIP is by Marqo; see the original model card for its training data and evaluation.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jmigual/robbeta-models

Finetuned
(1)
this model