Sakura EmbeddingGemma 2 — AutoRound W4A16

Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's official embeddinggemma-2 architecture.


Overview

This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.

It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.

  • Upstream Model: google/embeddinggemma-2
  • Pinned Upstream Commit: 914f7f89142e33e77833254d9c9b90c3cef7303b
  • Base Architecture: EmbeddingGemma2ForSequenceClassification / EmbeddingGemma2TextModel
  • Model File Size: model.safetensors: 1,299,010,632 bytes / 1,299.01 MB / 1,238.83 MiB. The pinned BF16 original model file is 1,488,915,288 bytes / 1,488.92 MB / 1,419.94 MiB: 12.75% smaller on disk. This compares model files, not total directories or measured runtime memory.
  • Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
  • Quantization Framework: AutoRound 0.16.0, as recorded in the published quantization_config.json.
  • Quantization Scheme: W4A16 symmetric (bits=4, group_size=64, iters=200), calibrated on 234 mixed retrieval texts (build/calibration_dual256.json); recorded packing format auto_round:auto_gptq.
  • Version: v2 (2026-10-07). Re-quantized with a broader calibration set, more iterations and a smaller group size. The v1 release (group size 128, 100 iterations, 64 calibration texts) is kept in this repository's git history. v2 improves BF16 fidelity and external retrieval modestly; see the tables below.
  • Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)

Architectural Breakdown: What is W4 vs. What Remains BF16

We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:

Component Precision Details
Text Backbone Linear Layers W4A16 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 64, symmetric.
Token Embeddings BF16 language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse.
Embedding Projection Head BF16 language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry.
Normalization Layers BF16 All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping.
Vision & Audio Towers BF16 Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized.

Fidelity & Benchmark Results

All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.

Release Gate Status

The candidate was evaluated against strict quality criteria:

  • Empirical Gate Result: RELEASE GO / Early Community Release
  • Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement is not fully reached (v2 achieved: 0.9897 Mean Cosine, 89.0 % Top-5 Agreement; v1: 0.9878 / 87.6 %). With 100.0 % Top-1 Agreement, 100.0 % Recall@5 and 0.9814 Spearman Correlation, these benchmark results show retained retrieval success alongside measurable BF16-fidelity loss; no NaN or Inf anomaly was observed.

The companion GGUF HQ releases achieve higher reported BF16 fidelity on the same existing BF16 reference and text/code benchmark. They are separate quantization runs; their results do not change this Safetensors model's release-gate outcome.

GGUF companion tier MiB Combined Mean Cosine (768d) Combined Spearman (768d) Text Top-5 Text/Code Recall@5
Q4 HQ 169.79 0.99400270 0.98857239 90.00% 100% / 100%
Q5 HQ — recommended balance 199.88 0.99828082 0.99664205 95.40% 100% / 100%
Q6 HQ 231.75 0.99940187 0.99879852 96.80% 100% / 100%

Values are quoted from the GGUF README's Side-by-Side Comparison. Aggregation differs: this Safetensors section reports alignment over 200 text query/document vectors; the GGUF combined Mean Cosine/Spearman include all 300 text-and-code vectors. Text Top-5 uses the 100-pair text retrieval benchmark in both. These are not identically aggregated mean scores, and no all-metric comparison with Unsloth is implied.

MRL (Matryoshka Representation Learning) Dimensions

EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:

MRL Dimension Mean Cosine Sim Min Cosine Sim Spearman Rank Corr Mean L2 Drift Top-1 Agreement Top-5 Agreement Recall@5 nDCG@10
768d (Full) 0.9897 0.9662 0.9814 0.1404 100.0 % 89.0 % 100.0 % 0.9645
512d 0.9899 0.9670 0.9815 0.1390 99.0 % 88.6 % 100.0 % 0.9590
256d 0.9910 0.9701 0.9832 0.1309 100.0 % 85.6 % 100.0 % 0.9579
128d 0.9934 0.9807 0.9828 0.1120 99.0 % 86.6 % 100.0 % 0.9454

Language & Domain Subsets (768d)

Subset Mean Cosine Sim Top-1 Agreement Top-5 Agreement Recall@5
English Queries & Docs 0.9897 100.0 % 88.4 % 100.0 %
German Queries & Docs 0.9898 100.0 % 87.6 % 100.0 %
Code Retrieval 0.9893 100.0 % 83.2 % 100.0 %

External retrieval (BEIR / MTEB test sets, measured)

Same texts and prefixes for all rows, texts truncated to 128 tokens; 500-document corpora (all positives + sampled distractors). BF16 = unquantized google/embeddinggemma-2.

Benchmark Model MRR nDCG@10 Recall@10 MRR vs BF16 nDCG@10 vs BF16
SciFact (300 queries) BF16 reference 0.8971 0.9137 0.9700 100 % 100 %
W4A16 v2 (this release) 0.8921 0.9105 0.9733 99.4 % 99.7 %
W4A16 v1 (previous release) 0.8892 0.9051 0.9667 99.1 % 99.1 %
NFCorpus (249 queries) BF16 reference 0.4609 0.3099 0.6386 100 % 100 %
W4A16 v2 (this release) 0.4509 0.2971 0.6265 97.8 % 95.9 %
W4A16 v1 (previous release) 0.4466 0.2973 0.6145 96.9 % 95.9 %

The differences between v1 and v2 are small and not uniform. v2 is better in mean cosine and Spearman correlation in every MRL dimension and subset, in Top-5 agreement at 768d (+1.4 points), 128d, English and German, and in SciFact/NFCorpus MRR. It is slightly worse in Top-5 agreement at 256d (-0.6 points) and on the code subset (-4.0 points: 83.2 % vs. 87.2 %), and nDCG@10 on the text benchmark is 0.001-0.003 lower; NFCorpus nDCG@10 is unchanged. NFCorpus has only 249 queries, so differences of about one point are within noise. Overall v2 is a modest improvement, not a step change.


Runtime & Community Format Comparison

Distribution / Format Runtime Compatibility Direct Transformers / ST Support Notes
Sakura AutoRound W4A16 (This repo) Python, PyTorch, Transformers, SentenceTransformers Yes (Plug-and-play) Runs directly in existing Python AI pipelines without needing custom binary builds.
GGUF Community Releases (unsloth, ggml-org) llama.cpp No (requires llama-server or bindings) EmbeddingGemma 2 support is upstream in llama.cpp via PR #30054, merged 2026-10-06. The earlier unknown model architecture result came from an older local build; use a build that includes this support. Community GGUF controls have since been loaded and benchmarked with a supported build.
ONNX Community Releases (onnx-community) ONNX Runtime / Transformers.js No (ONNX graph format) Modular multi-graph export tailored for WebGPU/JavaScript execution.
Sakura AutoRound GGUF HQ (Companion) llama.cpp Via llama-server / bindings Q4 HQ 169.79 MiB, Q5 HQ 199.88 MiB (recommended balance), Q6 HQ 231.75 MiB. This Safetensors/ST package is for native Python pipelines; the standalone text GGUF variants are for llama.cpp.

Safetensors vs. GGUF HQ: separate quantization runs

These are separate optimization runs, not two exports of one quantized state. The published Safetensors quantization_config.json records a different scheme and iteration count from the preserved GGUF cal256 build configurations/logs. Both use AutoRound 0.16.0 and the same base model, but that does not imply identical optimized weights or calibration.

Setting This Safetensors W4A16 release GGUF HQ cal256 runs
Transformer weight types 4-bit symmetric W4A16 Q4_K/Q5_K/Q6_K base with 48 Q8_0 block-PLE linears
Weight grouping Group size 64 Q4_K/Q5_K: 32 weights per subgroup, 8 subgroups per 256-weight superblock; Q6_K: 16, 16 subgroups per 256-weight superblock. Q8_0 uses 32-weight blocks.
Iterations 200, published configuration 50, saved build configurations/logs
Calibration selection 234 real retrieval texts from 13 public datasets (queries with the EmbeddingGemma query prefix, passages with the document prefix; EN/DE/code/science/web), manifest build/calibration_dual256.json (SHA256 d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd), built with build/requant.py 256 synthetic retrieval samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap
Optimization/export path AutoRound W4A16; published packing format auto_round:auto_gptq; Safetensors for Transformers/ST Native AutoRound SignRoundV2 (enable_alg_ext=True), matching GGUF optimized-state packing, then selected embedding-path precision overrides
AutoRound version 0.16.0, published configuration 0.16.0, preserved builds
Scope Full multimodal checkpoint, vision/audio and sensitive components retained in BF16 Standalone text GGUF; CPU/native llama.cpp runtime

Provenance of this release (v2): the calibration texts, the quantization script (build/requant.py, run with AutoRound 0.16.0 on CPU) and the evaluation results are published or reproducible from this repository. The v1 release (git history) was built by the earlier pipeline script with up to 64 texts from a local calibration file whose exact historical content was not preserved. The calibration texts come from public datasets (some, e.g. MS MARCO, carry research-use terms); only the texts used for calibration are included.

The GGUF cal256 runs do retain explicit sample count, calibration SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8, resolved per-layer configuration, native optimized-state packing evidence and training logs. These differences are enough to rule out a shared optimized quantization state. Public GGUF settings and runtime provenance are linked from the companion README; the inspected W4A16 pipeline settings are summarized here without claiming historical inputs that were not preserved.


Quickstart & Usage

1. With SentenceTransformers (Recommended)

from sentence_transformers import SentenceTransformer
import torch

# Load the quantized model
model = SentenceTransformer(
    "webmp3/Sakura-EmbeddingGemma-2-AutoRound",
    model_kwargs={"torch_dtype": torch.bfloat16}
)

# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")

# Documents (unprompted)
docs = [
    "Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
    "The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)

# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)

2. Matryoshka Dimension Truncation (MRL)

To reduce memory and storage footprint, simply slice the vector and re-normalize:

import torch
import torch.nn.functional as F

# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)

3. With Hugging Face Transformers

from transformers import AutoModel, AutoTokenizer
import torch

model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)

inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

Limitations & Honest Disclosure

  1. Text Precision and File Size: The 216 text-backbone linear layers use 4-bit packed weights. model.safetensors is 1,299,010,632 bytes (1,299.01 MB / 1,238.83 MiB), compared with 1,488,915,288 bytes (1,488.92 MB / 1,419.94 MiB) for the pinned full BF16 original. The 12.75% model-file saving is modest: token embeddings and multimodal towers remain BF16. These on-disk measurements do not establish a runtime-memory reduction; the former ~5.5 GB BF16 comparison is not supported by the checked model files.
  2. Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
  3. Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.

SHA256 Checksums

Every file in this release has been cryptographically verified:

1be8f9083684a77b24317163310782efd0df30e5ce5cc2c097d447b3a5402ce1  model.safetensors
48d4e0ca37c329ae9fc6582f7de5224f6d5ce9a3b81ea00dfe61eaf7a622c45c  config.json
072b3e5dc502e1beabac6414e8a663516995b75abe1c03fedf1ddb696eb49483  quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5  config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228  sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd  modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78  preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c  processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4  tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874  tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e  chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9  1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524  2_Normalize/config.json
d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd  build/calibration_dual256.json

Citation & Acknowledgements


中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura EmbeddingGemma 2 — AutoRound W4A16

社区量化版本。并非 Google 官方发布。
由 Sakura(webmp3)在 Google 官方 embeddinggemma-2 架构之上,使用 Intel AutoRound 优化进行量化和评估。


概览

本仓库提供 google/embeddinggemma-2 的优化 AutoRound W4A16(4 比特权重,16 比特激活)量化版本。

它被打包为标准的 Hugging Face 仓库,包含打包好的 Safetensors 权重,可以在 sentence-transformers 和 transformers 中 原生地、开箱即用地 运行。

  • 上游模型: google/embeddinggemma-2
  • 固定的上游提交: 914f7f89142e33e77833254d9c9b90c3cef7303b
  • 基础架构: EmbeddingGemma2ForSequenceClassification / EmbeddingGemma2TextModel
  • 模型文件大小: model.safetensors:1,299,010,632 字节 / 1,299.01 MB / 1,238.83 MiB。固定的 BF16 原始模型文件为 1,488,915,288 字节 / 1,488.92 MB / 1,419.94 MiB:**磁盘上小 12.75%**。这比较的是模型文件,而不是整个目录或实测的运行时内存。
  • 验证状态: 已验证可公开访问 Hub;本地发布冒烟测试通过;已为 Transformers 和 SentenceTransformers 打包。
  • 量化框架: AutoRound 0.16.0,记录在已发布的 quantization_config.json 中。
  • 量化方案: W4A16 对称(bits=4、group_size=64、iters=200),在 234 条混合检索文本(build/calibration_dual256.json)上校准;记录的打包格式为 auto_round:auto_gptq。
  • 版本: **v2(2026-10-07)**。使用更广泛的校准集、更多的迭代次数和更小的分组大小重新量化。v1 版本(分组大小 128,迭代 100 次,64 条校准文本)保留在本仓库的 git 历史中。v2 对 BF16 保真度和外部检索有适度改善;见下表。
  • 发布决定: RELEASE GO(基于实测保真度保持情况的早期社区发布)

架构拆解:哪些是 W4,哪些保持 BF16

我们明确 不 声称是“完整的 INT4/Q4”。高保真的嵌入模型需要谨慎对待敏感组件:

组件 精度 详情
文本主干线性层 W4A16 language_model.layers.0 到 language_model.layers.23 中的 216 个线性层(q_proj、k_proj、v_proj、o_proj、gate_proj、up_proj、down_proj)。分组大小 64,对称。
Token 嵌入 BF16 language_model.embed_tokens(256,000 词表)保持 bfloat16,以避免语义词表坍缩。
嵌入投影头 BF16 language_model.embedding_projection(768 维输出头)保持 bfloat16,以保留精确的方向几何。
归一化层 BF16 所有 RMSNorm 和 LayerNorm 模块保持 bfloat16,以防止激活尺度被截断。
视觉与音频塔 BF16 多模态编码器(vision_tower、audio_tower)和多模态投影头保持未量化。

保真度与基准结果

所有评估都是针对未量化的 bfloat16 参照 进行的,涵盖 100 个查询/文档检索对(英语和德语)以及 50 个代码检索对。

发布关卡状态

候选版本依据严格的质量标准进行了评估:

  • 实测关卡结果: RELEASE GO / 早期社区发布
  • 关卡背景: 最初的理论目标是 $\ge 0.99$ 的 Mean Cosine Similarity 和 $\ge 95%$ 的 Top-5 检索一致率,但并未完全达到(v2 达到:0.9897 Mean Cosine,89.0 % Top-5 一致率;v1:0.9878 / 87.6 %)。100.0 % 的 Top-1 一致率、100.0 % 的 Recall@5 和 0.9814 的 Spearman 相关系数 表明,这些基准结果显示检索成功率得以保持,同时存在可测量的 BF16 保真度损失;没有观察到 NaN 或 Inf 异常。

配套的 GGUF HQ 发布版本 在相同的现有 BF16 参照和文本/代码基准上取得了更高的 BF16 保真度。它们是独立的量化运行;其结果 不会 改变这个 Safetensors 模型的发布关卡结论。

GGUF 配套档位 MiB 综合 Mean Cosine (768d) 综合 Spearman (768d) Text Top-5 Text/Code Recall@5
Q4 HQ 169.79 0.99400270 0.98857239 90.00% 100% / 100%
Q5 HQ — 推荐的平衡之选 199.88 0.99828082 0.99664205 95.40% 100% / 100%
Q6 HQ 231.75 0.99940187 0.99879852 96.80% 100% / 100%

数值引自 GGUF README 的 并排比较。汇总方式不同: 本 Safetensors 部分报告的是 200 个文本查询/文档向量上的对齐程度;GGUF 的综合 Mean Cosine/Spearman 包含全部 300 个文本和代码向量。两者的 Text Top-5 都使用 100 对的文本检索基准。这些不是以相同方式汇总的平均分,也不暗示与 Unsloth 在所有指标上的比较。

MRL(Matryoshka 表示学习)维度

EmbeddingGemma 2 支持先截断维度、再做 L2 重新归一化。W4A16 模型在所有标准 MRL 截断下都表现出稳健的稳定性:

MRL 维度 Mean Cosine Sim Min Cosine Sim Spearman Rank Corr Mean L2 Drift Top-1 一致率 Top-5 一致率 Recall@5 nDCG@10
768d (完整) 0.9897 0.9662 0.9814 0.1404 100.0 % 89.0 % 100.0 % 0.9645
512d 0.9899 0.9670 0.9815 0.1390 99.0 % 88.6 % 100.0 % 0.9590
256d 0.9910 0.9701 0.9832 0.1309 100.0 % 85.6 % 100.0 % 0.9579
128d 0.9934 0.9807 0.9828 0.1120 99.0 % 86.6 % 100.0 % 0.9454

语言与领域子集 (768d)

子集 Mean Cosine Sim Top-1 一致率 Top-5 一致率 Recall@5
英文查询与文档 0.9897 100.0 % 88.4 % 100.0 %
德文查询与文档 0.9898 100.0 % 87.6 % 100.0 %
代码检索 0.9893 100.0 % 83.2 % 100.0 %

外部检索(BEIR / MTEB 测试集,实测)

所有行使用相同的文本和前缀,文本截断为 128 个 token;500 篇文档的语料(所有正例 + 采样的干扰项)。BF16 = 未量化的 google/embeddinggemma-2。

基准 模型 MRR nDCG@10 Recall@10 相对 BF16 的 MRR 相对 BF16 的 nDCG@10
SciFact(300 个查询) BF16 参照 0.8971 0.9137 0.9700 100 % 100 %
W4A16 v2(本次发布) 0.8921 0.9105 0.9733 99.4 % 99.7 %
W4A16 v1(上一次发布) 0.8892 0.9051 0.9667 99.1 % 99.1 %
NFCorpus(249 个查询) BF16 参照 0.4609 0.3099 0.6386 100 % 100 %
W4A16 v2(本次发布) 0.4509 0.2971 0.6265 97.8 % 95.9 %
W4A16 v1(上一次发布) 0.4466 0.2973 0.6145 96.9 % 95.9 %

v1 与 v2 之间的差异很小,而且并不一致。v2 在每个 MRL 维度和子集的 mean cosine 和 Spearman 相关性上更好,在 768d(+1.4 个百分点)、128d、英文和德文的 Top-5 一致率上更好,在 SciFact/NFCorpus 的 MRR 上也更好。它在 256d 的 Top-5 一致率(-0.6 个百分点)和代码子集(-4.0 个百分点:83.2 % 对 87.2 %)上略差,文本基准的 nDCG@10 低 0.001-0.003;NFCorpus 的 nDCG@10 没有变化。NFCorpus 只有 249 个查询,因此约一个百分点的差异处于噪声范围内。总体而言,v2 是适度的改进,不是一次飞跃。


运行时与社区格式对比

发行版 / 格式 运行时兼容性 是否直接支持 Transformers / ST 说明
Sakura AutoRound W4A16 (本仓库) Python、PyTorch、Transformers、SentenceTransformers 是(即插即用) 可直接在现有的 Python AI 流水线中运行,无需自定义二进制构建。
GGUF 社区发布版本(unsloth、ggml-org) llama.cpp 否(需要 llama-server 或绑定) EmbeddingGemma 2 在 llama.cpp 中通过 PR #30054 获得上游支持,于 2026-10-06 合并。较早出现的 unknown model architecture 结果来自较旧的本地构建;请使用包含此支持的构建。此后,社区 GGUF 对照版本已用受支持的构建加载并做了基准测试。
ONNX 社区发布版本(onnx-community) ONNX Runtime / Transformers.js 否(ONNX 图格式) 为 WebGPU/JavaScript 执行量身定制的模块化多图导出。
Sakura AutoRound GGUF HQ (配套) llama.cpp 通过 llama-server / 绑定 Q4 HQ 169.79 MiB,Q5 HQ 199.88 MiB(推荐的平衡之选),Q6 HQ 231.75 MiB。本 Safetensors/ST 包适用于原生 Python 流水线;独立的文本 GGUF 变体适用于 llama.cpp。

Safetensors 与 GGUF HQ:独立的量化运行

这些是独立的优化运行,而不是同一个量化状态的两种导出。 已发布的 Safetensors quantization_config.json 所记录的方案和迭代次数,与保存下来的 GGUF cal256 构建配置/日志不同。两者都使用 AutoRound 0.16.0 和相同的基础模型,但这并不意味着优化后的权重或校准相同。

设置 本 Safetensors W4A16 发布版本 GGUF HQ cal256 运行
Transformer 权重类型 4 比特对称 W4A16 Q4_K/Q5_K/Q6_K 基础,加 48 个 Q8_0 block-PLE 线性层
权重分组 分组大小 64 Q4_K/Q5_K:每个子组 32 个权重,每个 256 权重的超级块含 8 个子组;Q6_K:16,每个 256 权重的超级块含 16 个子组。Q8_0 使用 32 权重的块。
迭代次数 200,已发布的配置 50,已保存的构建配置/日志
校准选择 来自 13 个公开数据集的 234 条真实检索文本(带 EmbeddingGemma 查询前缀的查询、带文档前缀的段落;EN/DE/代码/科学/网页),清单 build/calibration_dual256.json(SHA256 d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd),由 build/requant.py 构建 256 个合成检索样本:48 个英文查询、48 个英文文档、48 个德文查询、48 个德文文档、32 个代码查询和 32 个代码片段;与基准完全没有重合
优化/导出路径 AutoRound W4A16;已发布的打包格式 auto_round:auto_gptq;供 Transformers/ST 使用的 Safetensors 原生 AutoRound SignRoundV2(enable_alg_ext=True),匹配的 GGUF 优化状态打包,然后是选定的嵌入路径精度覆盖
AutoRound 版本 0.16.0,已发布的配置 0.16.0,保存下来的构建
范围 完整的多模态检查点,视觉/音频和敏感组件保持 BF16 独立的文本 GGUF;CPU/原生 llama.cpp 运行时

本次发布(v2)的来源: 校准文本、量化脚本(build/requant.py,在 CPU 上使用 AutoRound 0.16.0 运行)和评估结果已发布,或可从本仓库复现。v1 版本(git 历史)由较早的流水线脚本构建,使用来自一个本地校准文件的最多 64 条文本,该文件的确切历史内容没有被保留。校准文本来自公开数据集(其中一些,例如 MS MARCO,带有仅限研究使用的条款);只包含用于校准的文本。

GGUF cal256 运行确实保留了明确的样本数量、校准 SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8、已解析的逐层配置、原生优化状态打包的证据以及训练日志。这些差异足以排除共享同一个优化后的量化状态。公开的 GGUF 设置和运行时来源已在配套 README 中链接;经过检查的 W4A16 流水线设置在此处做了概述,而没有声称那些未被保留的历史输入。


快速开始与用法

1. 使用 SentenceTransformers(推荐)

from sentence_transformers import SentenceTransformer
import torch

# Load the quantized model
model = SentenceTransformer(
    "webmp3/Sakura-EmbeddingGemma-2-AutoRound",
    model_kwargs={"torch_dtype": torch.bfloat16}
)

# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")

# Documents (unprompted)
docs = [
    "Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
    "The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)

# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)

2. Matryoshka 维度截断(MRL)

要减小内存和存储占用,只需切片向量并重新归一化:

import torch
import torch.nn.functional as F

# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)

3. 使用 Hugging Face Transformers

from transformers import AutoModel, AutoTokenizer
import torch

model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)

inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

局限与坦率的披露

  1. 文本精度与文件大小: 216 个文本主干线性层使用 4 比特打包权重。model.safetensors 为 1,299,010,632 字节(1,299.01 MB / 1,238.83 MiB),而固定的完整 BF16 原始文件为 1,488,915,288 字节(1,488.92 MB / 1,419.94 MiB)。12.75% 的模型文件节省是有限的:token 嵌入和多模态塔仍为 BF16。这些磁盘上的测量并不能证明运行时内存有所减少;此前提到的约 5.5 GB 的 BF16 对比,在核对过的模型文件上并不成立。
  2. 多模态塔: 视觉和音频塔保持 16 比特精度。如果你不使用视觉或音频输入,只加载文本语言模型即可把内存消耗降到最低。
  3. 执行设备: 最佳执行需要现代 CPU(AVX-512 / VNNI)或支持加速 bfloat16 和打包 int4 内核的 GPU 环境。

SHA256 校验和

本次发布中的每个文件都已经过加密校验:

1be8f9083684a77b24317163310782efd0df30e5ce5cc2c097d447b3a5402ce1  model.safetensors
48d4e0ca37c329ae9fc6582f7de5224f6d5ce9a3b81ea00dfe61eaf7a622c45c  config.json
072b3e5dc502e1beabac6414e8a663516995b75abe1c03fedf1ddb696eb49483  quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5  config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228  sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd  modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78  preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c  processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4  tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874  tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e  chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9  1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524  2_Normalize/config.json
d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd  build/calibration_dual256.json

引用与致谢

Downloads last month
35
Safetensors
Model size
0.6B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound

Quantized
(43)
this model