bge-m3-8bit-MLX — BAAI/bge-m3 in 8-bit (gs64), tuned for oMLX on Apple Silicon

English · 中文

BAAI/bge-m3 (multilingual dense embeddings: XLM-RoBERTa-large, 568M parameters, 1024-d, CLS pooling, L2-normalised) in the exact form we run every day on a Mac Studio M1 Ultra 64 GB with oMLX 0.7.0: 8-bit affine quantization, group size 64 — 612 MB instead of 2.27 GB, per-query latency on par with FP32, vector cosine ≥ 0.999 against an independent PyTorch FP32 reference.

What we changed compared with the upstream repository (all measured, all reproducible — see Rebuild recipe):

  1. Format. Upstream ships only pytorch_model.bin. We converted it without torch and without executing pickle code (restricted unpickler that only allows tensor types), then quantized it to the MLX layout. Rebuilding with the scripts in recipe/ reproduces model.safetensors byte for byte.
  2. 8-bit affine (group size 64) on the word-embedding table and all Linear weights of the 24 encoder layers; LayerNorm, biases, position/type embeddings and pooler stay FP16.
  3. Serving settings that were tuned on oMLX: context cap 2048 tokens, embedding_batch_size 16, idle TTL 1800 s — and the traps we hit on the way (FP16 weights are 3.3× slower on oMLX 0.7.0; the default context cap reaches past the position table).

Only the dense head is included. The sparse (lexical weights) and ColBERT (multi-vector) heads of BGE-M3 are not used by oMLX /v1/embeddings and are not shipped here.

8-bit vs other variants

At a glance

Item Value
Architecture XLM-RoBERTa-large: 24 layers, hidden 1024, 16 heads, 568M parameters, vocabulary 250,002, 8194 position slots (≤ 8192 tokens usable)
Output 1024-d dense vector, CLS pooling, L2-normalised (cosine = dot product). No instruction / prefix needed for queries
Quantization 8-bit affine, group size 64 (word embeddings + all encoder Linear weights); everything else FP16
Size 612 MB on disk (upstream FP32: 2.27 GB); ≈ 0.6 GB resident in oMLX
Languages 100+ as upstream; zh↔en cross-lingual retrieval tested
Speed (M1 Ultra) single query p50 11.8 ms in-process (FP32: 13.1), 20 ms over HTTP; index batch of 363 texts / 74.7K tokens in 5.7 s ≈ 13K tok/s over HTTP; cold start 1.9–4.4 s
Memory index batch +2.0 GB above the weights; worst case 24 texts × 2048 tokens +7.9 GB (batch 16); one 8192-token input ≈ 14–16 GB (naive O(L²) attention in oMLX 0.7.0)
License MIT (upstream)
Not included sparse and ColBERT heads

Lineage: BAAI/bge-m3 revision 5617a9f61b028005a4858fdac845db406aefb181 (2024-07-03; pytorch_model.bin SHA-256 b5e0ce34… matches the official LFS hash; tokenizer files match too) → recipe/bin2st.py → recipe/make_bgem3_q8.py → this repository (model.safetensors SHA-256 b30faac883ee1ced82f43a9564216c081985d991f21d3e985eadcb0d6af85fd5).

Measured results

All numbers: Mac Studio M1 Ultra (64-core GPU) 64 GB, macOS 15.8, oMLX 0.7.0 (MLX 0.32.2, mlx-embeddings 0.1.0), one machine. "Engine path" = the same code path the oMLX server uses (including the cache clear after every forward pass). Reference vectors: an independent PyTorch FP32 implementation (transformers AutoModel + CLS + L2).

Quantization variants

Variant Weights Single query p50 Batch-32 tok/s Peak (batch 32, incl. weights) Cosine vs FP32 mean / min Top-10 overlap MRR@10
FP32 (upstream) 2271 MB 13.1 ms 16.2K 4.16 GB 1 / 1 1.000 0.582
FP16 1136 MB 43.3 ms 15.9K 3.06 GB 1.00000 / 0.99997 1.000 0.582
8-bit gs64 (this repo) 612 MB 11.8 ms 14.1K 2.51 GB 0.99969 / 0.99924 0.988 0.589
embeddings 4-bit + encoder 8-bit 485 MB 12.2 ms 14.1K 2.38 GB 0.99908 / 0.99785 0.950 0.584
4-bit 334 MB — — — 0.95154 / 0.93266 0.819 0.585

(Fidelity corpus: 371 texts — 315 document chunks, 48 queries, 8 long texts of 1K–7K tokens. "Top-10 overlap" compares the query rankings with FP32.) 8-bit costs ~13 % of the large-batch throughput (dequantization) but is faster for single queries, 3.7× smaller and 1.65 GB lighter at the peak. 4-bit is rejected: its MRR looks fine on 48 queries, but vectors drift (cosine 0.95, neighbours overlap only 81 %) — a small ranking test cannot see that, a larger corpus would.

FP16 is a trap on oMLX 0.7.0 (and what fixes it)

oMLX's native XLM-R builds the attention mask in FP32, which up-casts FP16 activations; the engine also clears the MLX buffer pool after each request, so every request pays for those temporary buffers again. Single-query p50 with a cache clear after each forward: FP16 44.5 ms vs FP32 13.3 ms vs 8-bit 12.4 ms (×3.3). A run-time-only probe that made the mask FP16 (nothing installed or modified on disk) gave 11.5 ms for FP16 and 18.5–18.7K tok/s at batch 32 (16 % faster than FP32). That is upstream PR #4168 (in oMLX v0.7.1.dev1; the stable release is still 0.7.0). Until a stable release ships it, do not store this model in FP16 for oMLX — we will re-test then.

FP16 trap

Batch size, context cap and memory

  • embedding_batch_size 16 (oMLX default was 32). Index throughput −3 to −4 % (363 texts: 5.55 s vs 5.39 s in-process), but a single-query request that arrives while a batch is running waits 256 ms instead of 467 ms (head-of-line blocking is linear in batch size), and the worst long-text batch peaks at +7.9 GB instead of +15.6 GB. Batch 8 gives 119 ms / −14 % throughput.
  • Context cap 2048 tokens (max_context_window). Without it oMLX takes the model's max_position_embeddings = 8194, which is past the 8192 usable positions: no error, no NaN, but the vectors are no longer the same function of the text (cosine 0.9959 after one extra token). Attention is naive O(L²): peak memory above the weights for B texts × L tokens — L=512: 1.6–2.3 GB; L=1024: 1.4–4.7 GB; L=2048: 4.0 / 6.3 / 7.9 / 15.6 GB for B = 4 / 8 / 16 / 32.
  • With a 28.5 GB LLM resident, an index batch (363 texts) added +2.0 GB with no memory-pressure event; the worst case (24 × 2048 tokens) added +8.0 GB, crossed oMLX's soft threshold for about a second and recovered without evicting the LLM.

batch size, head-of-line blocking, memory

Retrieval quality (small test — read the caveats)

48 hand-written queries (zh→zh 20, en→en 14, zh→en 7, en→zh 7) over 315 chunks of technical documents. MRR@10: PyTorch FP32 0.582, 8-bit 0.589 (paired bootstrap ΔMRR +0.007, 95 % CI [−0.010, +0.032]; 45 of 48 queries have the same gold rank), 4-bit 0.585. For reference, Qwen3-Embedding-0.6B (8-bit) scored 0.463 without and 0.523 with a retrieval instruction on the same set; BGE-M3's lead over the instructed variant is not statistically significant here (+0.066, CI [−0.051, +0.181]). The set is small and built from one project's documents, so it has low power and says little about general retrieval — we rely on the fidelity numbers above.

retrieval quality

Recommended oMLX settings

omlx/model_settings.entry.json (key = directory name):

{ "max_context_window": 2048, "ttl_seconds": 1800, "model_alias": "BGE-M3" }

Global (oMLX settings.json → scheduler, or the admin API POST /admin/api/global-settings {"embedding_batch_size": 16}): embedding_batch_size = 16. Not pinned: it loads in 1.9–4.4 s on first use and sits at ≈ 0.6 GB. If you keep a large LLM resident and memory is tight, lower max_context_window to 1024 (worst case ≈ +3.6 GB). A client-side max_length overrides the cap — never send more than 8192.

Quick start

HF_ENDPOINT=https://hf-mirror.com hf download YCF-AI/bge-m3-8bit-MLX --local-dir ~/models/bge-m3-8bit
# oMLX: put the folder under a model_dirs entry, merge omlx/model_settings.entry.json (key = folder name),
# set scheduler.embedding_batch_size = 16, reload. Then:
curl http://127.0.0.1:8000/v1/embeddings -H "Authorization: Bearer $OMLX_API_KEY" -H "Content-Type: application/json" \
     -d '{"model": "BGE-M3", "input": ["什么是向量检索?", "What is vector retrieval?"]}'

Without oMLX, with mlx-embeddings (tested with 0.1.0; it reproduced the oMLX server's vectors with cosine 1.0000000):

import mlx.core as mx
from mlx_embeddings.utils import load

model, tokenizer = load("YCF-AI/bge-m3-8bit-MLX")          # or a local folder
inputs = tokenizer(["什么是向量检索?", "What is vector retrieval?"], return_tensors="mlx",
                   padding=True, truncation=True, max_length=512)
out = model(inputs["input_ids"], attention_mask=inputs["attention_mask"])
emb = out.last_hidden_state[:, 0, :]                        # CLS pooling (1_Pooling/config.json)
emb = emb / mx.linalg.norm(emb, axis=-1, keepdims=True)     # L2-normalise: cosine == dot product

This is an MLX-quantized checkpoint (.scales / .biases tensors): transformers / sentence-transformers cannot load it — use BAAI/bge-m3 for those.

Rebuild recipe

# 1. fetch the pinned upstream revision (HF mirror, no proxy) and check the hash
HF_ENDPOINT=https://hf-mirror.com hf download BAAI/bge-m3 --revision 5617a9f61b028005a4858fdac845db406aefb181 --local-dir up
shasum -a 256 up/pytorch_model.bin      # b5e0ce3470abf5ef3831aa1bd5553b486803e83251590ab7ff35a117cf6aad38
# 2. .bin -> FP32 safetensors (no torch, no arbitrary pickle code)
python recipe/bin2st.py up/pytorch_model.bin fp32/model.safetensors
cp up/{config.json,config_sentence_transformers.json,modules.json,sentence_bert_config.json,special_tokens_map.json,tokenizer_config.json,sentencepiece.bpe.model,tokenizer.json} fp32/ && cp -R up/1_Pooling fp32/
# 3. 8-bit affine, group size 64 (needs mlx; verified with mlx 0.32.2 on Apple Silicon)
python recipe/make_bgem3_q8.py fp32 out
shasum -a 256 out/model.safetensors     # b30faac883ee1ced82f43a9564216c081985d991f21d3e985eadcb0d6af85fd5

Verified 2026-10-10: this reproduces model.safetensors, config.json and all tokenizer/config files of this repository byte for byte. After changing weights under the same oMLX model ID, purge that model's SSD-cache entries.

Known limits

  • Dense retrieval only (no sparse / ColBERT). Query and passage are encoded symmetrically — no instruction prefix.
  • Long inputs are memory-hungry in oMLX 0.7.0 (naive attention): cap the context, keep batches small. Behaviour beyond 8192 tokens is undefined.
  • Vectors are very close to FP32 BGE-M3 vectors (cosine ≥ 0.999; query top-10 overlap 98.8 %), so this build can stand in for a hosted FP16/FP32 BGE-M3 (we use it as a local fallback) — but vectors of other embedding models are not compatible; changing the embedding model requires re-indexing. Re-validate on your own data.
  • One machine, one small English/Chinese technical test set; numbers are not a promise for yours.

License and attribution

MIT, inherited from BAAI/bge-m3 (see README-upstream.md for the original card). Please cite the original work: BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation (Chen et al., 2024, arXiv:2402.03216). Part of a set of notes on running a local model stack on an M1 Ultra: tuning notes and data · reranker companion: bge-reranker-v2-m3-fp16emb.


中文说明

English · 中文

BAAI/bge-m3(多语言稠密向量:XLM-RoBERTa-large,5.68 亿参数,1024 维,CLS 池化 + L2 归一化)的 8-bit(group size 64)MLX 版,就是我们在 Mac Studio M1 Ultra 64 GB + oMLX 0.7.0 上每天使用的那一份:权重 612 MB(上游 2.27 GB),单条延迟与 FP32 持平,向量与独立 PyTorch FP32 参考的余弦 ≥ 0.999。

相对上游做了三件事(均有实测、可复现):

  1. 格式:上游只有 pytorch_model.bin。我们用受限反序列化(只放行张量类型,不依赖 torch、不执行 pickle 代码)转成 safetensors,再量化成 MLX 格式;用 recipe/ 里的脚本重建,model.safetensors 逐字节一致。
  2. 8-bit affine(group size 64):词嵌入表 + 24 层编码器全部 Linear 权重;LayerNorm、偏置、位置/类型嵌入、pooler 保持 FP16。
  3. oMLX 服务参数:上下文上限 2048、embedding_batch_size 16、空闲 TTL 1800 s,以及途中踩过的坑(oMLX 0.7.0 上 FP16 权重反而慢 3.3 倍;默认上下文上限会越过位置表)。

只含稠密头;BGE-M3 的稀疏(词权重)与 ColBERT(多向量)头 oMLX 的 /v1/embeddings 不用,未随包提供。

要点

项 值
输出 1024 维稠密向量,已 L2 归一化(余弦 = 点积);查询无需指令前缀
体积 / 常驻 612 MB(上游 2.27 GB);oMLX 内常驻约 0.6 GB
速度(M1 Ultra) 单条 p50 进程内 11.8 ms(FP32 13.1),HTTP 20 ms;363 条 / 7.47 万 token 的索引批 5.7 s ≈ 13K tok/s;冷启动 1.9–4.4 s
内存 索引批在权重之上 +2.0 GB;最坏情形(24 条 × 2048 token,批 16)+7.9 GB;单条 8192 token 约 14–16 GB(oMLX 0.7.0 的注意力为朴素 O(L²) 实现)
保真 对 PyTorch FP32:余弦均值 0.99969 / 最小 0.99924,查询 top-10 重叠 98.8 %;48 查询 MRR@10 0.589 vs 0.582(配对 bootstrap 无可测差异)
许可 MIT(上游)

关键实测结论

  • 8-bit 胜出:比 FP32 小 3.7 倍,单条更快;大批吞吐低约 13 %(反量化开销)。4-bit 否决:48 查询的 MRR 看不出差别(0.585),但向量保真明显下降(余弦 0.95,近邻 top-10 重叠仅 81 %)——小评测集发现不了,大语料会累积。
  • FP16 在 oMLX 0.7.0 上是陷阱:原生 XLM-R 的注意力掩码是 FP32,FP16 权重每次前向都要升精度,而引擎每个请求后清缓存,单条延迟 44.5 ms vs FP32 13.3 ms(×3.3)。运行时仅把掩码改成 FP16 的对照实验(磁盘与应用均未改动):11.5 ms,批 32 吞吐 18.5–18.7K tok/s(比 FP32 快 16 %)。这就是上游 PR #4168(含于 oMLX v0.7.1.dev1,稳定版仍是 0.7.0);稳定版发布前不要把这个模型存成 FP16,发布后我们会复测。
  • 批大小 16(oMLX 默认 32):索引吞吐 −3~4 %,但批在跑时到来的单条查询等待 256 ms(原 467 ms),最坏长文批峰值 +7.9 GB(原 +15.6 GB)。批 8:119 ms,吞吐 −14 %。
  • 上下文上限 2048:不设时 oMLX 取 max_position_embeddings = 8194,越过 8192 个可用位置——不报错、不出 NaN,但向量已不是同一个函数(多 1 个 token 余弦即 0.9959)。
  • 与 28.5 GB 的大模型同驻:索引批 +2.0 GB 无压力事件;最坏长文批 +8.0 GB 短暂越过 oMLX 软阈值约 1 秒后自愈,大模型未被驱逐。
  • 检索质量(48 条中英/跨语言查询、315 个技术文档块,小测试,请看限制):8-bit 与 FP32 在 45/48 条查询上金标名次相同;Qwen3-Embedding-0.6B(8-bit)同一评测集 0.463(无指令)/ 0.523(带指令),BGE-M3 对带指令版的领先不显著。评测集小、取自同一项目的文档,统计功效低,我们主要依据保真度数据。

推荐 oMLX 设置

见上方 JSON 与 omlx/model_settings.entry.json(键名=目录名):max_context_window 2048、ttl_seconds 1800、别名 BGE-M3;全局 scheduler.embedding_batch_size 设为 16。不 pinned(首次使用加载 1.9–4.4 s,常驻约 0.6 GB)。若同时常驻大模型且内存紧张,把 max_context_window 降到 1024(最坏约 +3.6 GB)。请求里的 max_length 优先于上限,不要超过 8192。

快速开始与复现

下载、oMLX 调用与 mlx-embeddings 代码见上方英文部分(mlx-embeddings 代码已实测:与 oMLX 服务端向量余弦 1.0000000)。这是 MLX 量化检查点(含 .scales / .biases),transformers / sentence-transformers 无法加载,请用 BAAI/bge-m3。 复现:固定上游 revision 下载(国内用 HF_ENDPOINT=https://hf-mirror.com、不走代理)→ 校验 pytorch_model.bin SHA-256 → recipe/bin2st.py → recipe/make_bgem3_q8.py;2026-10-10 验证可逐字节复现本仓库的 model.safetensors(b30faac8…)、config.json 与全部分词器/配置文件。同一 oMLX 模型 ID 下权重变化后,请清理该模型的 SSD 缓存。

已知限制

  • 仅稠密检索;查询与文档对称编码,无指令前缀。
  • oMLX 0.7.0 对长输入很耗内存(朴素注意力):务必限上下文、用小批;超过 8192 token 行为未定义。
  • 向量与 FP32 的 BGE-M3 非常接近(余弦 ≥ 0.999),可作为托管 FP16/FP32 BGE-M3 的本地兜底(我们就这样用),但与其他嵌入模型的向量不兼容,换模型必须重建索引;请在自己的数据上复核。
  • 单机、小规模中英技术文本评测,数字不代表你的环境。

许可与致谢

MIT,继承自 BAAI/bge-m3(原卡见 README-upstream.md)。请引用原论文:BGE M3-Embedding(Chen et al., 2024,arXiv:2402.03216)。 本模型属于一套 M1 Ultra 本地模型栈调优笔记:笔记与数据 · 配套重排器:bge-reranker-v2-m3-fp16emb。

Downloads last month
21
Safetensors
Model size
0.6B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YCF-AI/bge-m3-8bit-MLX

Base model

BAAI/bge-m3
Quantized
(303)
this model

Paper for YCF-AI/bge-m3-8bit-MLX