Instructions to use YCF-AI/bge-m3-8bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use YCF-AI/bge-m3-8bit-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download YCF-AI/bge-m3-8bit-MLX --local-dir bge-m3-8bit-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
bge-m3-8bit-MLX — BAAI/bge-m3 in 8-bit (gs64), tuned for oMLX on Apple Silicon
English · 中文
BAAI/bge-m3 (multilingual dense embeddings: XLM-RoBERTa-large, 568M parameters, 1024-d, CLS pooling, L2-normalised) in the exact form we run every day on a Mac Studio M1 Ultra 64 GB with oMLX 0.7.0: 8-bit affine quantization, group size 64 — 612 MB instead of 2.27 GB, per-query latency on par with FP32, vector cosine ≥ 0.999 against an independent PyTorch FP32 reference.
What we changed compared with the upstream repository (all measured, all reproducible — see Rebuild recipe):
- Format. Upstream ships only
pytorch_model.bin. We converted it without torch and without executing pickle code (restricted unpickler that only allows tensor types), then quantized it to the MLX layout. Rebuilding with the scripts inrecipe/reproducesmodel.safetensorsbyte for byte. - 8-bit affine (group size 64) on the word-embedding table and all Linear weights of the 24 encoder layers; LayerNorm, biases, position/type embeddings and pooler stay FP16.
- Serving settings that were tuned on oMLX: context cap 2048 tokens,
embedding_batch_size16, idle TTL 1800 s — and the traps we hit on the way (FP16 weights are 3.3× slower on oMLX 0.7.0; the default context cap reaches past the position table).
Only the dense head is included. The sparse (lexical weights) and ColBERT (multi-vector) heads of BGE-M3 are not used by oMLX /v1/embeddings and are not shipped here.
At a glance
| Item | Value |
|---|---|
| Architecture | XLM-RoBERTa-large: 24 layers, hidden 1024, 16 heads, 568M parameters, vocabulary 250,002, 8194 position slots (≤ 8192 tokens usable) |
| Output | 1024-d dense vector, CLS pooling, L2-normalised (cosine = dot product). No instruction / prefix needed for queries |
| Quantization | 8-bit affine, group size 64 (word embeddings + all encoder Linear weights); everything else FP16 |
| Size | 612 MB on disk (upstream FP32: 2.27 GB); ≈ 0.6 GB resident in oMLX |
| Languages | 100+ as upstream; zh↔en cross-lingual retrieval tested |
| Speed (M1 Ultra) | single query p50 11.8 ms in-process (FP32: 13.1), 20 ms over HTTP; index batch of 363 texts / 74.7K tokens in 5.7 s ≈ 13K tok/s over HTTP; cold start 1.9–4.4 s |
| Memory | index batch +2.0 GB above the weights; worst case 24 texts × 2048 tokens +7.9 GB (batch 16); one 8192-token input ≈ 14–16 GB (naive O(L²) attention in oMLX 0.7.0) |
| License | MIT (upstream) |
| Not included | sparse and ColBERT heads |
Lineage: BAAI/bge-m3 revision 5617a9f61b028005a4858fdac845db406aefb181 (2024-07-03; pytorch_model.bin SHA-256 b5e0ce34… matches the official LFS hash; tokenizer files match too) → recipe/bin2st.py → recipe/make_bgem3_q8.py → this repository (model.safetensors SHA-256 b30faac883ee1ced82f43a9564216c081985d991f21d3e985eadcb0d6af85fd5).
Measured results
All numbers: Mac Studio M1 Ultra (64-core GPU) 64 GB, macOS 15.8, oMLX 0.7.0 (MLX 0.32.2, mlx-embeddings 0.1.0), one machine. "Engine path" = the same code path the oMLX server uses (including the cache clear after every forward pass). Reference vectors: an independent PyTorch FP32 implementation (transformers AutoModel + CLS + L2).
Quantization variants
| Variant | Weights | Single query p50 | Batch-32 tok/s | Peak (batch 32, incl. weights) | Cosine vs FP32 mean / min | Top-10 overlap | MRR@10 |
|---|---|---|---|---|---|---|---|
| FP32 (upstream) | 2271 MB | 13.1 ms | 16.2K | 4.16 GB | 1 / 1 | 1.000 | 0.582 |
| FP16 | 1136 MB | 43.3 ms | 15.9K | 3.06 GB | 1.00000 / 0.99997 | 1.000 | 0.582 |
| 8-bit gs64 (this repo) | 612 MB | 11.8 ms | 14.1K | 2.51 GB | 0.99969 / 0.99924 | 0.988 | 0.589 |
| embeddings 4-bit + encoder 8-bit | 485 MB | 12.2 ms | 14.1K | 2.38 GB | 0.99908 / 0.99785 | 0.950 | 0.584 |
| 4-bit | 334 MB | — | — | — | 0.95154 / 0.93266 | 0.819 | 0.585 |
(Fidelity corpus: 371 texts — 315 document chunks, 48 queries, 8 long texts of 1K–7K tokens. "Top-10 overlap" compares the query rankings with FP32.) 8-bit costs ~13 % of the large-batch throughput (dequantization) but is faster for single queries, 3.7× smaller and 1.65 GB lighter at the peak. 4-bit is rejected: its MRR looks fine on 48 queries, but vectors drift (cosine 0.95, neighbours overlap only 81 %) — a small ranking test cannot see that, a larger corpus would.
FP16 is a trap on oMLX 0.7.0 (and what fixes it)
oMLX's native XLM-R builds the attention mask in FP32, which up-casts FP16 activations; the engine also clears the MLX buffer pool after each request, so every request pays for those temporary buffers again. Single-query p50 with a cache clear after each forward: FP16 44.5 ms vs FP32 13.3 ms vs 8-bit 12.4 ms (×3.3). A run-time-only probe that made the mask FP16 (nothing installed or modified on disk) gave 11.5 ms for FP16 and 18.5–18.7K tok/s at batch 32 (16 % faster than FP32). That is upstream PR #4168 (in oMLX v0.7.1.dev1; the stable release is still 0.7.0). Until a stable release ships it, do not store this model in FP16 for oMLX — we will re-test then.
Batch size, context cap and memory
embedding_batch_size16 (oMLX default was 32). Index throughput −3 to −4 % (363 texts: 5.55 s vs 5.39 s in-process), but a single-query request that arrives while a batch is running waits 256 ms instead of 467 ms (head-of-line blocking is linear in batch size), and the worst long-text batch peaks at +7.9 GB instead of +15.6 GB. Batch 8 gives 119 ms / −14 % throughput.- Context cap 2048 tokens (
max_context_window). Without it oMLX takes the model'smax_position_embeddings= 8194, which is past the 8192 usable positions: no error, no NaN, but the vectors are no longer the same function of the text (cosine 0.9959 after one extra token). Attention is naive O(L²): peak memory above the weights for B texts × L tokens — L=512: 1.6–2.3 GB; L=1024: 1.4–4.7 GB; L=2048: 4.0 / 6.3 / 7.9 / 15.6 GB for B = 4 / 8 / 16 / 32. - With a 28.5 GB LLM resident, an index batch (363 texts) added +2.0 GB with no memory-pressure event; the worst case (24 × 2048 tokens) added +8.0 GB, crossed oMLX's soft threshold for about a second and recovered without evicting the LLM.
Retrieval quality (small test — read the caveats)
48 hand-written queries (zh→zh 20, en→en 14, zh→en 7, en→zh 7) over 315 chunks of technical documents. MRR@10: PyTorch FP32 0.582, 8-bit 0.589 (paired bootstrap ΔMRR +0.007, 95 % CI [−0.010, +0.032]; 45 of 48 queries have the same gold rank), 4-bit 0.585. For reference, Qwen3-Embedding-0.6B (8-bit) scored 0.463 without and 0.523 with a retrieval instruction on the same set; BGE-M3's lead over the instructed variant is not statistically significant here (+0.066, CI [−0.051, +0.181]). The set is small and built from one project's documents, so it has low power and says little about general retrieval — we rely on the fidelity numbers above.
Recommended oMLX settings
omlx/model_settings.entry.json (key = directory name):
{ "max_context_window": 2048, "ttl_seconds": 1800, "model_alias": "BGE-M3" }
Global (oMLX settings.json → scheduler, or the admin API POST /admin/api/global-settings {"embedding_batch_size": 16}): embedding_batch_size = 16. Not pinned: it loads in 1.9–4.4 s on first use and sits at ≈ 0.6 GB.
If you keep a large LLM resident and memory is tight, lower max_context_window to 1024 (worst case ≈ +3.6 GB). A client-side max_length overrides the cap — never send more than 8192.
Quick start
HF_ENDPOINT=https://hf-mirror.com hf download YCF-AI/bge-m3-8bit-MLX --local-dir ~/models/bge-m3-8bit
# oMLX: put the folder under a model_dirs entry, merge omlx/model_settings.entry.json (key = folder name),
# set scheduler.embedding_batch_size = 16, reload. Then:
curl http://127.0.0.1:8000/v1/embeddings -H "Authorization: Bearer $OMLX_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "BGE-M3", "input": ["什么是向量检索?", "What is vector retrieval?"]}'
Without oMLX, with mlx-embeddings (tested with 0.1.0; it reproduced the oMLX server's vectors with cosine 1.0000000):
import mlx.core as mx
from mlx_embeddings.utils import load
model, tokenizer = load("YCF-AI/bge-m3-8bit-MLX") # or a local folder
inputs = tokenizer(["什么是向量检索?", "What is vector retrieval?"], return_tensors="mlx",
padding=True, truncation=True, max_length=512)
out = model(inputs["input_ids"], attention_mask=inputs["attention_mask"])
emb = out.last_hidden_state[:, 0, :] # CLS pooling (1_Pooling/config.json)
emb = emb / mx.linalg.norm(emb, axis=-1, keepdims=True) # L2-normalise: cosine == dot product
This is an MLX-quantized checkpoint (.scales / .biases tensors): transformers / sentence-transformers cannot load it — use BAAI/bge-m3 for those.
Rebuild recipe
# 1. fetch the pinned upstream revision (HF mirror, no proxy) and check the hash
HF_ENDPOINT=https://hf-mirror.com hf download BAAI/bge-m3 --revision 5617a9f61b028005a4858fdac845db406aefb181 --local-dir up
shasum -a 256 up/pytorch_model.bin # b5e0ce3470abf5ef3831aa1bd5553b486803e83251590ab7ff35a117cf6aad38
# 2. .bin -> FP32 safetensors (no torch, no arbitrary pickle code)
python recipe/bin2st.py up/pytorch_model.bin fp32/model.safetensors
cp up/{config.json,config_sentence_transformers.json,modules.json,sentence_bert_config.json,special_tokens_map.json,tokenizer_config.json,sentencepiece.bpe.model,tokenizer.json} fp32/ && cp -R up/1_Pooling fp32/
# 3. 8-bit affine, group size 64 (needs mlx; verified with mlx 0.32.2 on Apple Silicon)
python recipe/make_bgem3_q8.py fp32 out
shasum -a 256 out/model.safetensors # b30faac883ee1ced82f43a9564216c081985d991f21d3e985eadcb0d6af85fd5
Verified 2026-10-10: this reproduces model.safetensors, config.json and all tokenizer/config files of this repository byte for byte. After changing weights under the same oMLX model ID, purge that model's SSD-cache entries.
Known limits
- Dense retrieval only (no sparse / ColBERT). Query and passage are encoded symmetrically — no instruction prefix.
- Long inputs are memory-hungry in oMLX 0.7.0 (naive attention): cap the context, keep batches small. Behaviour beyond 8192 tokens is undefined.
- Vectors are very close to FP32 BGE-M3 vectors (cosine ≥ 0.999; query top-10 overlap 98.8 %), so this build can stand in for a hosted FP16/FP32 BGE-M3 (we use it as a local fallback) — but vectors of other embedding models are not compatible; changing the embedding model requires re-indexing. Re-validate on your own data.
- One machine, one small English/Chinese technical test set; numbers are not a promise for yours.
License and attribution
MIT, inherited from BAAI/bge-m3 (see README-upstream.md for the original card). Please cite the original work:
BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation (Chen et al., 2024, arXiv:2402.03216).
Part of a set of notes on running a local model stack on an M1 Ultra: tuning notes and data · reranker companion: bge-reranker-v2-m3-fp16emb.
中文说明
English · 中文
BAAI/bge-m3(多语言稠密向量:XLM-RoBERTa-large,5.68 亿参数,1024 维,CLS 池化 + L2 归一化)的 8-bit(group size 64)MLX 版,就是我们在 Mac Studio M1 Ultra 64 GB + oMLX 0.7.0 上每天使用的那一份:权重 612 MB(上游 2.27 GB),单条延迟与 FP32 持平,向量与独立 PyTorch FP32 参考的余弦 ≥ 0.999。
相对上游做了三件事(均有实测、可复现):
- 格式:上游只有
pytorch_model.bin。我们用受限反序列化(只放行张量类型,不依赖 torch、不执行 pickle 代码)转成 safetensors,再量化成 MLX 格式;用recipe/里的脚本重建,model.safetensors逐字节一致。 - 8-bit affine(group size 64):词嵌入表 + 24 层编码器全部 Linear 权重;LayerNorm、偏置、位置/类型嵌入、pooler 保持 FP16。
- oMLX 服务参数:上下文上限 2048、
embedding_batch_size16、空闲 TTL 1800 s,以及途中踩过的坑(oMLX 0.7.0 上 FP16 权重反而慢 3.3 倍;默认上下文上限会越过位置表)。
只含稠密头;BGE-M3 的稀疏(词权重)与 ColBERT(多向量)头 oMLX 的 /v1/embeddings 不用,未随包提供。
要点
| 项 | 值 |
|---|---|
| 输出 | 1024 维稠密向量,已 L2 归一化(余弦 = 点积);查询无需指令前缀 |
| 体积 / 常驻 | 612 MB(上游 2.27 GB);oMLX 内常驻约 0.6 GB |
| 速度(M1 Ultra) | 单条 p50 进程内 11.8 ms(FP32 13.1),HTTP 20 ms;363 条 / 7.47 万 token 的索引批 5.7 s ≈ 13K tok/s;冷启动 1.9–4.4 s |
| 内存 | 索引批在权重之上 +2.0 GB;最坏情形(24 条 × 2048 token,批 16)+7.9 GB;单条 8192 token 约 14–16 GB(oMLX 0.7.0 的注意力为朴素 O(L²) 实现) |
| 保真 | 对 PyTorch FP32:余弦均值 0.99969 / 最小 0.99924,查询 top-10 重叠 98.8 %;48 查询 MRR@10 0.589 vs 0.582(配对 bootstrap 无可测差异) |
| 许可 | MIT(上游) |
关键实测结论
- 8-bit 胜出:比 FP32 小 3.7 倍,单条更快;大批吞吐低约 13 %(反量化开销)。4-bit 否决:48 查询的 MRR 看不出差别(0.585),但向量保真明显下降(余弦 0.95,近邻 top-10 重叠仅 81 %)——小评测集发现不了,大语料会累积。
- FP16 在 oMLX 0.7.0 上是陷阱:原生 XLM-R 的注意力掩码是 FP32,FP16 权重每次前向都要升精度,而引擎每个请求后清缓存,单条延迟 44.5 ms vs FP32 13.3 ms(×3.3)。运行时仅把掩码改成 FP16 的对照实验(磁盘与应用均未改动):11.5 ms,批 32 吞吐 18.5–18.7K tok/s(比 FP32 快 16 %)。这就是上游 PR #4168(含于 oMLX v0.7.1.dev1,稳定版仍是 0.7.0);稳定版发布前不要把这个模型存成 FP16,发布后我们会复测。
- 批大小 16(oMLX 默认 32):索引吞吐 −3~4 %,但批在跑时到来的单条查询等待 256 ms(原 467 ms),最坏长文批峰值 +7.9 GB(原 +15.6 GB)。批 8:119 ms,吞吐 −14 %。
- 上下文上限 2048:不设时 oMLX 取
max_position_embeddings= 8194,越过 8192 个可用位置——不报错、不出 NaN,但向量已不是同一个函数(多 1 个 token 余弦即 0.9959)。 - 与 28.5 GB 的大模型同驻:索引批 +2.0 GB 无压力事件;最坏长文批 +8.0 GB 短暂越过 oMLX 软阈值约 1 秒后自愈,大模型未被驱逐。
- 检索质量(48 条中英/跨语言查询、315 个技术文档块,小测试,请看限制):8-bit 与 FP32 在 45/48 条查询上金标名次相同;Qwen3-Embedding-0.6B(8-bit)同一评测集 0.463(无指令)/ 0.523(带指令),BGE-M3 对带指令版的领先不显著。评测集小、取自同一项目的文档,统计功效低,我们主要依据保真度数据。
推荐 oMLX 设置
见上方 JSON 与 omlx/model_settings.entry.json(键名=目录名):max_context_window 2048、ttl_seconds 1800、别名 BGE-M3;全局 scheduler.embedding_batch_size 设为 16。不 pinned(首次使用加载 1.9–4.4 s,常驻约 0.6 GB)。若同时常驻大模型且内存紧张,把 max_context_window 降到 1024(最坏约 +3.6 GB)。请求里的 max_length 优先于上限,不要超过 8192。
快速开始与复现
下载、oMLX 调用与 mlx-embeddings 代码见上方英文部分(mlx-embeddings 代码已实测:与 oMLX 服务端向量余弦 1.0000000)。这是 MLX 量化检查点(含 .scales / .biases),transformers / sentence-transformers 无法加载,请用 BAAI/bge-m3。
复现:固定上游 revision 下载(国内用 HF_ENDPOINT=https://hf-mirror.com、不走代理)→ 校验 pytorch_model.bin SHA-256 → recipe/bin2st.py → recipe/make_bgem3_q8.py;2026-10-10 验证可逐字节复现本仓库的 model.safetensors(b30faac8…)、config.json 与全部分词器/配置文件。同一 oMLX 模型 ID 下权重变化后,请清理该模型的 SSD 缓存。
已知限制
- 仅稠密检索;查询与文档对称编码,无指令前缀。
- oMLX 0.7.0 对长输入很耗内存(朴素注意力):务必限上下文、用小批;超过 8192 token 行为未定义。
- 向量与 FP32 的 BGE-M3 非常接近(余弦 ≥ 0.999),可作为托管 FP16/FP32 BGE-M3 的本地兜底(我们就这样用),但与其他嵌入模型的向量不兼容,换模型必须重建索引;请在自己的数据上复核。
- 单机、小规模中英技术文本评测,数字不代表你的环境。
许可与致谢
MIT,继承自 BAAI/bge-m3(原卡见 README-upstream.md)。请引用原论文:BGE M3-Embedding(Chen et al., 2024,arXiv:2402.03216)。
本模型属于一套 M1 Ultra 本地模型栈调优笔记:笔记与数据 · 配套重排器:bge-reranker-v2-m3-fp16emb。
- Downloads last month
- 21
8-bit
Model tree for YCF-AI/bge-m3-8bit-MLX
Base model
BAAI/bge-m3


