Repasar Lens

Repasar Lens is built on Google EmbeddingGemma 2, fine-tuned by Repasar for photos of second-hand items listed on Repasar, a Singapore marketplace. It runs on the phone with ONNX Runtime and suggests a listing category (116 category/subcategory pairs) and, for 28 brands, the brand โ€” each only when it is clearly sure.

It is not a Google product and is not endorsed by Google. EmbeddingGemma 2 is ยฉ Google DeepMind, released under the Apache License 2.0; this repository keeps that licence (see LICENSE).

What Repasar changed

  • LoRA on the vision tower (rank 16, alpha 32, on q/k/v/o and gate/up/down of all 16 vision layers, 4.5 M parameters), trained on ~7,000 labelled photos at 70 image tokens per photo.
  • Two small heads on the 768-number embedding: categories (116 ร— 768) and brands (28 + "No brand").
  • Export to ONNX of the vision tower + embed_vision, 4-bit (MatMulNBits, block 32), with every LoRA matrix as an overridable input so the LoRA ships as a separate ONNX Runtime adapter file.

The text part (onnx/model_q4.onnx) is the unchanged 4-bit text model from onnx-community/embeddinggemma-2-ONNX (Apache 2.0).

Files

File What
onnx/vision_base_q4.onnx (+ .data) Google's EmbeddingGemma 2 vision tower, 4-bit, with empty LoRA slots
onnx/lens_lora.onnx_adapter Repasar's LoRA (ONNX Runtime adapter, fp16) โ€” the part a retrain replaces
onnx/model_q4.onnx (+ _data) EmbeddingGemma 2 text part, 4-bit (onnx-community)
heads/category_head.f32 + .json 116 ร— 768 float32 category vectors, rows in pairs order
heads/brand_head.f32 + .json brand vectors, last row "No brand"

Score = L2-normalised embedding ยท row. A category or brand is used only when the best score beats the runner-up by 0.05.

Inputs

Photos are prepared exactly like transformers' Gemma4ImageProcessor with max_soft_tokens=70: aspect-preserving bicubic resize to sides divisible by 48 within 630 patches of 16 ร— 16, RGB scaled to 0..1, patches row by row with (column, row) positions, padded to 630 at position (-1, -1). The text model takes <bos><|image> + image token ร— n + <image|><eos> (ids 2, 255999, 258880 ร— n, 258882, 1).

import onnxruntime as ort
vision = ort.InferenceSession("onnx/vision_base_q4.onnx")
adapter = ort.LoraAdapter(); adapter.Load("onnx/lens_lora.onnx_adapter")
run = ort.RunOptions(); run.add_active_adapter(adapter)
features = vision.run(None, {"pixel_values": pv, "pixel_position_ids": pos}, run)[0]

Evaluation (held-out photos never used in training)

Exact category/subcategory Category
EmbeddingGemma 2 LiteRT (Google, int4) + Repasar head 87.7 % 93.0 %
Repasar Lens, ONNX 4-bit, LoRA merged (600 photos) 90.0 % 95.3 %

The published layout keeps the LoRA as a separate fp16 adapter on the 4-bit base (cosine 0.97-0.98 to the merged model on spot checks); its own held-out score is being measured and will be added here.

Brands (296 held-out photos, sure-only): 82.1 % named correctly, 0.5 % wrong brand, 0 % brands invented for unbranded items. On a OnePlus 11 (CPU, 4 threads) the model takes ~1.9 s per photo, 412 MB peak.

Training data

Only the trained numbers are published, never the photos.

Per-photo credits for the CC BY and CC BY-SA images are kept with the training data.

Limitations

  • 76 of 116 categories had training photos; the rest rely on the base model alone.
  • Catalogue and stock photos dominate; real second-hand photos with busy backgrounds (an item in a bag, a sachet on a table) are harder โ€” the sure-only rule hides those answers rather than guessing.
  • Brands are limited to the 28 trained; printed brand names are read by the app separately.
  • It names categories and brands only; it does not write titles or read model numbers.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Repasar-Labs/repasar-lens

Adapter
(1)
this model