colSmol-256M-6bit

An MLX copy of vidore/colSmol-256M, a page retriever: it turns a page picture, or a text query, into one 128-number vector per token, and pages are ranked against a query by late interaction (MaxSim).

It does not chat. Use it to find the right page in a set of documents.

How this copy was made

  1. The LoRA adapter in vidore/colSmol-256M was merged into the weights of its base, vidore/ColSmolVLM-Instruct-256M-base.
  2. The config's model_type was set to colidefics3, the type mlx-vlm loads.
  3. The merged weights were quantized to 6 bits (affine, group size 64) with mlx-vlm 0.7.6.

Nothing was retrained. The merge was checked against the original PyTorch weights: MaxSim scores agree within 0.6%. Six-bit rounding moves the vectors by about 0.006 on average.

Use

With mlx-vlm in Python, load it as any colidefics3 checkpoint.

In Swift, Different-Productions/mlx-swift-lm loads it through MLXEmbedders (ColIdefics3Model, ColIdefics3Processor), and matches mlx-vlm on the same files: the same token ids and ranking, scores within 0.7% on PNG pages.

License and credit

MIT, as the original. The model is the work of the ViDoRe team at Illuin Technology; see the original card for training data, benchmarks and how to cite it. Its backbone, SmolVLM-256M-Instruct by Hugging Face, is Apache 2.0.

This repository changes the original in the three ways listed above and no others.

Downloads last month
18
Safetensors
Model size
0.2B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DifferentProductions/colSmol-256M-6bit