Repasar Lens
Repasar Lens is built on Google EmbeddingGemma 2, fine-tuned by Repasar for photos of second-hand items listed on Repasar, a Singapore marketplace. It runs on the phone with ONNX Runtime and suggests a listing category (116 category/subcategory pairs) and, for 28 brands, the brand โ each only when it is clearly sure.
It is not a Google product and is not endorsed by Google. EmbeddingGemma 2 is ยฉ Google DeepMind,
released under the Apache License 2.0; this repository keeps that licence (see LICENSE).
What Repasar changed
- LoRA on the vision tower (rank 16, alpha 32, on q/k/v/o and gate/up/down of all 16 vision layers, 4.5 M parameters), trained on ~7,000 labelled photos at 70 image tokens per photo.
- Two small heads on the 768-number embedding: categories (116 ร 768) and brands (28 + "No brand").
- Export to ONNX of the vision tower + embed_vision, 4-bit (MatMulNBits, block 32), with every LoRA matrix as an overridable input so the LoRA ships as a separate ONNX Runtime adapter file.
The text part (onnx/model_q4.onnx) is the unchanged 4-bit text model from
onnx-community/embeddinggemma-2-ONNX (Apache 2.0).
Files
| File | What |
|---|---|
onnx/vision_base_q4.onnx (+ .data) |
Google's EmbeddingGemma 2 vision tower, 4-bit, with empty LoRA slots |
onnx/lens_lora.onnx_adapter |
Repasar's LoRA (ONNX Runtime adapter, fp16) โ the part a retrain replaces |
onnx/model_q4.onnx (+ _data) |
EmbeddingGemma 2 text part, 4-bit (onnx-community) |
heads/category_head.f32 + .json |
116 ร 768 float32 category vectors, rows in pairs order |
heads/brand_head.f32 + .json |
brand vectors, last row "No brand" |
Score = L2-normalised embedding ยท row. A category or brand is used only when the best score beats the runner-up by 0.05.
Inputs
Photos are prepared exactly like transformers' Gemma4ImageProcessor with max_soft_tokens=70:
aspect-preserving bicubic resize to sides divisible by 48 within 630 patches of 16 ร 16, RGB scaled to
0..1, patches row by row with (column, row) positions, padded to 630 at position (-1, -1). The text
model takes <bos><|image> + image token ร n + <image|><eos> (ids 2, 255999, 258880 ร n, 258882, 1).
import onnxruntime as ort
vision = ort.InferenceSession("onnx/vision_base_q4.onnx")
adapter = ort.LoraAdapter(); adapter.Load("onnx/lens_lora.onnx_adapter")
run = ort.RunOptions(); run.add_active_adapter(adapter)
features = vision.run(None, {"pixel_values": pv, "pixel_position_ids": pos}, run)[0]
Evaluation (held-out photos never used in training)
| Exact category/subcategory | Category | |
|---|---|---|
| EmbeddingGemma 2 LiteRT (Google, int4) + Repasar head | 87.7 % | 93.0 % |
| Repasar Lens, ONNX 4-bit, LoRA merged (600 photos) | 90.0 % | 95.3 % |
The published layout keeps the LoRA as a separate fp16 adapter on the 4-bit base (cosine 0.97-0.98 to the merged model on spot checks); its own held-out score is being measured and will be added here.
Brands (296 held-out photos, sure-only): 82.1 % named correctly, 0.5 % wrong brand, 0 % brands invented for unbranded items. On a OnePlus 11 (CPU, 4 threads) the model takes ~1.9 s per photo, 412 MB peak.
Training data
Only the trained numbers are published, never the photos.
- Open Images V7 โ images listed as CC BY 2.0 (only CC BY images kept, per image)
- Amazon Berkeley Objects โ CC BY 4.0
- Open Food Facts โ product photos CC BY-SA 3.0 (brand layer)
- Wikimedia Commons โ CC0 / public domain / CC BY files (brand layer)
- Stanford Online Products โ research dataset of eBay listing photos
Per-photo credits for the CC BY and CC BY-SA images are kept with the training data.
Limitations
- 76 of 116 categories had training photos; the rest rely on the base model alone.
- Catalogue and stock photos dominate; real second-hand photos with busy backgrounds (an item in a bag, a sachet on a table) are harder โ the sure-only rule hides those answers rather than guessing.
- Brands are limited to the 28 trained; printed brand names are read by the app separately.
- It names categories and brands only; it does not write titles or read model numbers.
Model tree for Repasar-Labs/repasar-lens
Base model
google/embeddinggemma-2