nomic-embed-vision-v1.5 โ€” GGUF for CrispEmbed

GGUF conversion of nomic-ai/nomic-embed-vision-v1.5 (Apache-2.0, ยฉ Nomic AI) for CrispEmbed. It embeds images into the same space as nomic-embed-text-v1.5, so text can search images.

file size precision
nomic-embed-vision-v1.5-f16.gguf 188 MB F16 weights (norms, biases, position embeddings in F32)
nomic-embed-vision-v1.5-f32.gguf 372 MB F32
crispembed -m nomic-embed-vision-v1.5 --image photo.jpg          # 768-d, L2-normalised
crispembed -m nomic-embed-text-v1.5 "search_query: a photo of a cat"

For cross-modal scoring, the upstream model card applies a parameter-free LayerNorm (subtract the mean, divide by the standard deviation) to the text embedding before L2 normalisation.

Parity

Checked against transformers (NomicVisionModel, trust_remote_code) with tools/ci-heavy/nomic_vision.py. The cosine between C++ and transformers embeddings was 1.000000 on all four test images, for both files. That holds both for the processor's exact pixels and through CrispEmbed's own image loader.

Converted with models/convert-nomic-vision-to-gguf.py. The original model's licence and attribution apply.

Downloads last month
123
GGUF
Model size
92.9M params
Architecture
vit
Hardware compatibility
Log In to add your hardware

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cstr/nomic-embed-vision-v1.5-GGUF

Quantized
(1)
this model