EmbeddingGemma 2 β AX650 / AX8850 (LLM-8850) NPU
google/embeddinggemma-2 (text, images and audio β one 768-d embedding space) converted with Axera Pulsar2 7.0-patch1 to run on the AX650 / AX8850 NPU, tested on the M5Stack LLM-8850 PCIe card (AXCL). The repository is a ready-to-run bundle: compiled axmodels, host-side assets, and a small host runtime that needs only numpy, Pillow and tokenizers (no PyTorch). The same code runs on Windows and Linux hosts.
νκ΅μ΄ μλ΄λ μλμ μμ΅λλ€.
Accuracy and speed
Measured on an LLM-8850 card (PCIe Gen2 x2) against the float32 model run through its own processor (the sentence-transformers pipeline: mean pooling + L2 normalization).
| Input | Cosine to float32 (min) | Notes |
|---|---|---|
| Text (STS-B and KorSTS, 1,000 sentences; retrieval pairs) | 0.9996 | STS-B Spearman 0.9477 (float32 0.9478), KorSTS 0.8980 (0.8984) |
| Long text (up to 1024 tokens) | 0.9997 | |
| Images (12 photos) | 0.9994 | text β photo retrieval top-1 100% |
| Audio (10 Korean / English clips) | 0.9994 | transcript β clip retrieval top-1 100% |
| Text + image + audio in one input | 0.9994 |
| Model | NPU time | One input end to end |
|---|---|---|
| Text, 128 tokens | 15.4 ms | short sentence β 17 ms |
| Text, 512 tokens | 78.4 ms | |
| Text, 1024 tokens | 246.7 ms | β 255 ms |
| Vision (three graphs, 2520 patches) | β 1.34 s | β 1.6 s per image |
| Audio (up to 11.2 s) | 62 ms | 5.8 s clip β 0.15 s |
selftest.py repeats this check against the float32 references in refs/ and prints PASS.
Contents
| Path | Description |
|---|---|
assets/eg2_text_l128/l512/l1024.axmodel |
Text encoder for 128 / 512 / 1024 tokens; the shortest one that fits is used |
assets/eg2_vision_p2520_a/b1/b2.axmodel |
Vision encoder, 16 layers split 8 + 4 + 4 (up to 280 soft tokens per image, aspect ratio kept) |
assets/eg2_audio_f1120.axmodel |
Audio encoder, up to 1120 log-mel frames (11.2 s β 280 soft tokens) |
assets/*.npy, embed_tokens.bf16.bin, tokenizer.json, host_meta.json |
Host assets (token embeddings, vision position table, projection, mel filters, metadata) |
assets/config.json, embeddinggemma2_tokenizer.txt |
Config for text-only serving with axllm (see below) |
eg2_host.py |
Preprocessing, tokenization, multimodal sequence assembly and NPU calls |
axcl_session.py |
ctypes wrapper of the AXCL runtime (libaxcl_rt.dll on Windows, libaxcl_rt.so on Linux) |
eg2_cli.py, eg2_server.py |
Command line tool; OpenAI-compatible /v1/embeddings server (text, images, audio) |
selftest.py, refs/ |
Self test against float32 references |
install_ubuntu.sh, build_axcl_driver.sh, eg2-embed.service |
Ubuntu setup, AXCL driver build for kernel 7.x, systemd unit |
Quick start
Requirements: an AX650 / AX8850 AXCL card with the AXCL driver and runtime installed (axcl-smi shows the card), Python 3.10+.
git lfs install
git clone https://huggingface.co/<this repo> && cd embeddinggemma-2-AX650
pip install -r requirements.txt # numpy, tokenizers, pillow, soundfile, scipy
python3 selftest.py # -> PASS
python3 eg2_cli.py --prompt query "text:a photo of a cat" "image:refs/files/cat_0.jpeg" "image:refs/files/dog_0.jpeg"
python3 eg2_server.py --host 0.0.0.0 --port 8010
curl -s localhost:8010/v1/embeddings -H 'Content-Type: application/json' \
-d '{"input": ["weather in Seoul", "weather in Busan"], "input_type": "query", "dimensions": 256}'
from eg2_host import Eg2
eg = Eg2("assets")
q = eg.encode(text="a red city bus", prompt="query")
d = eg.encode(images=["refs/files/bus.jpg"])
mixed = eg.encode(text="A red city bus <|image|> and its announcement <|audio|>",
images=["refs/files/bus.jpg"], audios=["refs/files/en_bus.wav"], prompt="document")
print(float(q @ d))
eg.close()
prompt/input_typeselects the task prefix of the original model card:query,document,sts,classification,clustering,code. Usequeryfor queries anddocumentfor documents in retrieval; images and audio take no prefix.dim/dimensions: 768, 512, 256 or 128 (Matryoshka; re-normalized after truncation).- Server inputs: a string, a list of strings, or objects
{"text": "...<|image|>...", "image": "<base64>", "audio": "<base64 wav>"}.
Text-only serving with axllm
JonPark0/ax-llm (branch windows) adds model_type: embedding_gemma2 to axllm serve. Point it at assets/:
axllm serve assets --port 8000 # OpenAI-compatible /v1/embeddings, text only
Limitations
- At most 1024 tokens per input (text and soft tokens together); longer inputs are truncated keeping BOS and the final EOS. The original model accepts 8K.
- Audio: up to 11.2 s per clip (the original processor's cap of 280 tokens). Images: up to 280 soft tokens (the original default). Video is not wired up.
- One request at a time per card.
- Image preprocessing uses Pillow's bicubic resize instead of torchvision's; on float32 this alone gives cosine β₯ 0.99997 to the original.
- Tested on Windows 11 with the AXCL Windows driver. Linux hosts use the same code; the kernel 7.x driver build in
build_axcl_driver.shcompiles but has not been load-tested on hardware.
How it was converted
- Each encoder was re-implemented in PyTorch with export-friendly operations and checked against the Hugging Face model (cosine 1.0) before export: static shapes, padding masks applied as multiply-and-renormalize after softmax, RMSNorm in rsqrt form, no dynamic shapes or boolean tensors in the graph.
- Pulsar2 7.0-patch1,
U16activations for all layers, NPU3, calibration on 32β64 real inputs (English, Korean, Chinese and code text; photos; speech). - Host-side parts: token embedding lookup (Γ β512), vision patch position embeddings, 2-D RoPE tables and the 3Γ3 pooling matrix (so one graph serves every aspect ratio), the final vision RMSNorm and 768β512 projection (its squares overflow the NPU float path), audio log-mel features, L2 normalization.
License
Apache License 2.0, as the original model. See LICENSE and NOTICE (attribution and list of modifications).
νκ΅μ΄
Google embeddinggemma-2(ν
μ€νΈΒ·μ΄λ―Έμ§Β·μμ± β νλμ 768μ°¨μ λ²‘ν° κ³΅κ°)λ₯Ό Axera Pulsar2 7.0-patch1λ‘ λ³νν΄ AX650 / AX8850 NPUμμ λ리λ λ¬Άμμ
λλ€. M5Stack LLM-8850 PCIe μΉ΄λ(AXCL)μμ μννμ΅λλ€. torch μμ΄ numpyΒ·PillowΒ·tokenizersλ§μΌλ‘ λμκ°λ©°, Windowsμ Linuxμμ κ°μ μ½λκ° λμν©λλ€.
- μ νλ: μλ³Έ(float32) λλΉ μ½μ¬μΈ μ΅μ 0.9994(ν μ€νΈ 0.9996). STS-B Spearman 0.9477(μλ³Έ 0.9478), KorSTS 0.8980(0.8984). ν μ€νΈΒ·μ¬μ§Β·μμ± κ²μ μ λ΅λ₯ λͺ¨λ 100%.
- μλ: μ§§μ λ¬Έμ₯ μ½ 17 ms, 1024ν ν° μ½ 255 ms, μ΄λ―Έμ§ 1μ₯ μ½ 1.6μ΄, 5.8μ΄ μμ± μ½ 0.15μ΄.
- μ¬μ©:
pip install -r requirements.txtβpython3 selftest.py(PASSνμΈ) βeg2_cli.pyλλeg2_server.py(OpenAI νΈν/v1/embeddings). ν μ€νΈλ§ μΈ λλ JonPark0/ax-llmwindowsλΈλμΉμaxllm serve assetsλ μΈ μ μμ΅λλ€. - μμ
μ λμ΄:
query,document,sts,classification,clustering,code. κ²μμμλ μ§μμquery, λ¬Έμμdocumentλ₯Ό μλλ€. μ΄λ―Έμ§Β·μμ±μλ λΆμ΄μ§ μμ΅λλ€. - μ ν: μ λ ₯ νλλΉ 1024ν ν°(μλ³Έμ 8K), μμ±μ ν΄λ¦½λΉ 11.2μ΄, μ΄λ―Έμ§λ μ₯λΉ 280ν ν°, μμμ λ―Έμ§μ, μμ²μ ν λ²μ νλ.
- Ubuntu μ€μΉ, 컀λ 7.xμ© AXCL λλΌμ΄λ² λΉλ, systemd μλΉμ€ λ± μμΈν νκ΅μ΄ μλ΄λ
GUIDE_ko.mdμ μμ΅λλ€. - λΌμ΄μ μ€: μλ³Έκ³Ό κ°μ Apache 2.0(
LICENSE,NOTICE).
Model tree for jonpark0/embeddinggemma-2-AX650
Base model
google/embeddinggemma-2