most-embed-de GGUF
GGUF format of malteos/most-embed-de for use with CrispEmbed.
MOST Embed DE โ German customer-support retrieval model, with query: / passage: prompts and normalized 2048-d embeddings.
Files
| File | Quantization | Size |
|---|---|---|
| most-embed-de-q4_k-attn-q8.gguf | Q4_K | 875 MB |
| most-embed-de-q8_0.gguf | Q8_0 | 1163 MB |
Parity vs HuggingFace reference
Cosine similarity vs the upstream sentence-transformers reference on a fixed test set (text):
| Quant | Text |
|---|---|
| f16 | 1.0000 |
| q8_0 | 0.9998 |
| q4_k | 0.9872 |
Note: below the 0.99 retrieval-quality bar โ text: q4_k (0.987). Embeddings are still functionally usable (>0.9 = directionally correct for similarity ranking) but expect small differences in nearest-neighbor results vs the upstream f32 reference.
License and provenance
The fine-tune is distributed under CC-BY-NC-4.0; commercial use is not permitted without separate authorization from the fine-tune author. It is derived from NVIDIA's Nemotron-3-Embed-1B-BF16, whose Model Materials are distributed under OpenMDW-1.1. Redistribution must retain the OpenMDW agreement plus all applicable copyright and origin notices. Both sets of terms and the upstream model cards must be reviewed and preserved.
Quick Start
# Download
huggingface-cli download cstr/most-embed-de-GGUF most-embed-de-q4_k-attn-q8.gguf --local-dir .
# Run with CrispEmbed
./crispembed -m most-embed-de-q4_k-attn-q8.gguf "Hello world"
# Or with auto-download
./crispembed -m most-embed-de "Hello world"
Model Details
| Property | Value |
|---|---|
| Architecture | Ministral3 bidirectional encoder |
| Parameters | 1.14B |
| Embedding Dimension | 2048 |
| Layers | 16 |
| Pooling | mean |
| Tokenizer | Tekken ByteLevel BPE |
| Base Model | malteos/most-embed-de |
Verification
Compared with the original Transformers implementation. F16 reaches cosine 1.000000 at every dumped transformer boundary and on the final embedding. Q8_0 reaches final cosine 0.999818. The compact Q4_K artifact keeps token embeddings and attention at Q8_0; over eight German query/document texts it has minimum cosine 0.987191, preserves every top-1 retrieval result, and reduces maximum similarity-score error to 0.02595.
Usage with CrispEmbed
CrispEmbed is a lightweight C/C++ text embedding inference engine using ggml. No Python runtime, no ONNX. Supports BERT, XLM-R, Qwen3, and Gemma3 architectures.
# Build CrispEmbed
git clone https://github.com/CrispStrobe/CrispEmbed
cd CrispEmbed
cmake -S . -B build && cmake --build build -j
# Encode
./build/crispembed -m most-embed-de-q4_k-attn-q8.gguf "query text"
# Server mode
./build/crispembed-server -m most-embed-de-q4_k-attn-q8.gguf --port 8080
curl -X POST http://localhost:8080/v1/embeddings \
-d '{"input": ["Hello world"], "model": "most-embed-de"}'
Credits
- Original model: malteos/most-embed-de
- Inference engine: CrispEmbed (ggml-based)
- Conversion:
convert-decoder-embed-to-gguf.py
- Downloads last month
- 127
8-bit
Model tree for cstr/most-embed-de-GGUF
Base model
mistralai/Ministral-3-3B-Base-2512