EmbeddingGemma 2 β€” AX650 / AX8850 (LLM-8850) NPU

google/embeddinggemma-2 (text, images and audio β†’ one 768-d embedding space) converted with Axera Pulsar2 7.0-patch1 to run on the AX650 / AX8850 NPU, tested on the M5Stack LLM-8850 PCIe card (AXCL). The repository is a ready-to-run bundle: compiled axmodels, host-side assets, and a small host runtime that needs only numpy, Pillow and tokenizers (no PyTorch). The same code runs on Windows and Linux hosts.

ν•œκ΅­μ–΄ μ•ˆλ‚΄λŠ” μ•„λž˜μ— μžˆμŠ΅λ‹ˆλ‹€.

Accuracy and speed

Measured on an LLM-8850 card (PCIe Gen2 x2) against the float32 model run through its own processor (the sentence-transformers pipeline: mean pooling + L2 normalization).

Input Cosine to float32 (min) Notes
Text (STS-B and KorSTS, 1,000 sentences; retrieval pairs) 0.9996 STS-B Spearman 0.9477 (float32 0.9478), KorSTS 0.8980 (0.8984)
Long text (up to 1024 tokens) 0.9997
Images (12 photos) 0.9994 text β†’ photo retrieval top-1 100%
Audio (10 Korean / English clips) 0.9994 transcript β†’ clip retrieval top-1 100%
Text + image + audio in one input 0.9994
Model NPU time One input end to end
Text, 128 tokens 15.4 ms short sentence β‰ˆ 17 ms
Text, 512 tokens 78.4 ms
Text, 1024 tokens 246.7 ms β‰ˆ 255 ms
Vision (three graphs, 2520 patches) β‰ˆ 1.34 s β‰ˆ 1.6 s per image
Audio (up to 11.2 s) 62 ms 5.8 s clip β‰ˆ 0.15 s

selftest.py repeats this check against the float32 references in refs/ and prints PASS.

Contents

Path Description
assets/eg2_text_l128/l512/l1024.axmodel Text encoder for 128 / 512 / 1024 tokens; the shortest one that fits is used
assets/eg2_vision_p2520_a/b1/b2.axmodel Vision encoder, 16 layers split 8 + 4 + 4 (up to 280 soft tokens per image, aspect ratio kept)
assets/eg2_audio_f1120.axmodel Audio encoder, up to 1120 log-mel frames (11.2 s β†’ 280 soft tokens)
assets/*.npy, embed_tokens.bf16.bin, tokenizer.json, host_meta.json Host assets (token embeddings, vision position table, projection, mel filters, metadata)
assets/config.json, embeddinggemma2_tokenizer.txt Config for text-only serving with axllm (see below)
eg2_host.py Preprocessing, tokenization, multimodal sequence assembly and NPU calls
axcl_session.py ctypes wrapper of the AXCL runtime (libaxcl_rt.dll on Windows, libaxcl_rt.so on Linux)
eg2_cli.py, eg2_server.py Command line tool; OpenAI-compatible /v1/embeddings server (text, images, audio)
selftest.py, refs/ Self test against float32 references
install_ubuntu.sh, build_axcl_driver.sh, eg2-embed.service Ubuntu setup, AXCL driver build for kernel 7.x, systemd unit

Quick start

Requirements: an AX650 / AX8850 AXCL card with the AXCL driver and runtime installed (axcl-smi shows the card), Python 3.10+.

git lfs install
git clone https://huggingface.co/<this repo> && cd embeddinggemma-2-AX650
pip install -r requirements.txt          # numpy, tokenizers, pillow, soundfile, scipy
python3 selftest.py                      # -> PASS

python3 eg2_cli.py --prompt query "text:a photo of a cat" "image:refs/files/cat_0.jpeg" "image:refs/files/dog_0.jpeg"
python3 eg2_server.py --host 0.0.0.0 --port 8010
curl -s localhost:8010/v1/embeddings -H 'Content-Type: application/json' \
  -d '{"input": ["weather in Seoul", "weather in Busan"], "input_type": "query", "dimensions": 256}'
from eg2_host import Eg2
eg = Eg2("assets")
q = eg.encode(text="a red city bus", prompt="query")
d = eg.encode(images=["refs/files/bus.jpg"])
mixed = eg.encode(text="A red city bus <|image|> and its announcement <|audio|>",
                  images=["refs/files/bus.jpg"], audios=["refs/files/en_bus.wav"], prompt="document")
print(float(q @ d))
eg.close()
  • prompt / input_type selects the task prefix of the original model card: query, document, sts, classification, clustering, code. Use query for queries and document for documents in retrieval; images and audio take no prefix.
  • dim / dimensions: 768, 512, 256 or 128 (Matryoshka; re-normalized after truncation).
  • Server inputs: a string, a list of strings, or objects {"text": "...<|image|>...", "image": "<base64>", "audio": "<base64 wav>"}.

Text-only serving with axllm

JonPark0/ax-llm (branch windows) adds model_type: embedding_gemma2 to axllm serve. Point it at assets/:

axllm serve assets --port 8000     # OpenAI-compatible /v1/embeddings, text only

Limitations

  • At most 1024 tokens per input (text and soft tokens together); longer inputs are truncated keeping BOS and the final EOS. The original model accepts 8K.
  • Audio: up to 11.2 s per clip (the original processor's cap of 280 tokens). Images: up to 280 soft tokens (the original default). Video is not wired up.
  • One request at a time per card.
  • Image preprocessing uses Pillow's bicubic resize instead of torchvision's; on float32 this alone gives cosine β‰₯ 0.99997 to the original.
  • Tested on Windows 11 with the AXCL Windows driver. Linux hosts use the same code; the kernel 7.x driver build in build_axcl_driver.sh compiles but has not been load-tested on hardware.

How it was converted

  • Each encoder was re-implemented in PyTorch with export-friendly operations and checked against the Hugging Face model (cosine 1.0) before export: static shapes, padding masks applied as multiply-and-renormalize after softmax, RMSNorm in rsqrt form, no dynamic shapes or boolean tensors in the graph.
  • Pulsar2 7.0-patch1, U16 activations for all layers, NPU3, calibration on 32–64 real inputs (English, Korean, Chinese and code text; photos; speech).
  • Host-side parts: token embedding lookup (Γ— √512), vision patch position embeddings, 2-D RoPE tables and the 3Γ—3 pooling matrix (so one graph serves every aspect ratio), the final vision RMSNorm and 768β†’512 projection (its squares overflow the NPU float path), audio log-mel features, L2 normalization.

License

Apache License 2.0, as the original model. See LICENSE and NOTICE (attribution and list of modifications).


ν•œκ΅­μ–΄

Google embeddinggemma-2(ν…μŠ€νŠΈΒ·μ΄λ―Έμ§€Β·μŒμ„± β†’ ν•˜λ‚˜μ˜ 768차원 벑터 곡간)λ₯Ό Axera Pulsar2 7.0-patch1둜 λ³€ν™˜ν•΄ AX650 / AX8850 NPUμ—μ„œ λŒλ¦¬λŠ” λ¬ΆμŒμž…λ‹ˆλ‹€. M5Stack LLM-8850 PCIe μΉ΄λ“œ(AXCL)μ—μ„œ μ‹œν—˜ν–ˆμŠ΅λ‹ˆλ‹€. torch 없이 numpyΒ·PillowΒ·tokenizers만으둜 λŒμ•„κ°€λ©°, Windows와 Linuxμ—μ„œ 같은 μ½”λ“œκ°€ λ™μž‘ν•©λ‹ˆλ‹€.

  • 정확도: 원본(float32) λŒ€λΉ„ 코사인 μ΅œμ†Œ 0.9994(ν…μŠ€νŠΈ 0.9996). STS-B Spearman 0.9477(원본 0.9478), KorSTS 0.8980(0.8984). ν…μŠ€νŠΈΒ·μ‚¬μ§„Β·μŒμ„± 검색 μ •λ‹΅λ₯  λͺ¨λ‘ 100%.
  • 속도: 짧은 λ¬Έμž₯ μ•½ 17 ms, 1024토큰 μ•½ 255 ms, 이미지 1μž₯ μ•½ 1.6초, 5.8초 μŒμ„± μ•½ 0.15초.
  • μ‚¬μš©: pip install -r requirements.txt β†’ python3 selftest.py(PASS 확인) β†’ eg2_cli.py λ˜λŠ” eg2_server.py(OpenAI ν˜Έν™˜ /v1/embeddings). ν…μŠ€νŠΈλ§Œ μ“Έ λ•ŒλŠ” JonPark0/ax-llm windows 브랜치의 axllm serve assets도 μ“Έ 수 μžˆμŠ΅λ‹ˆλ‹€.
  • μž‘μ—… 접두어: query, document, sts, classification, clustering, code. κ²€μƒ‰μ—μ„œλŠ” μ§ˆμ˜μ— query, λ¬Έμ„œμ— documentλ₯Ό μ”λ‹ˆλ‹€. μ΄λ―Έμ§€Β·μŒμ„±μ—λŠ” 뢙이지 μ•ŠμŠ΅λ‹ˆλ‹€.
  • μ œν•œ: μž…λ ₯ ν•˜λ‚˜λ‹Ή 1024토큰(원본은 8K), μŒμ„±μ€ 클립당 11.2초, μ΄λ―Έμ§€λŠ” μž₯λ‹Ή 280토큰, μ˜μƒμ€ 미지원, μš”μ²­μ€ ν•œ λ²ˆμ— ν•˜λ‚˜.
  • Ubuntu μ„€μΉ˜, 컀널 7.x용 AXCL λ“œλΌμ΄λ²„ λΉŒλ“œ, systemd μ„œλΉ„μŠ€ λ“± μžμ„Έν•œ ν•œκ΅­μ–΄ μ•ˆλ‚΄λŠ” GUIDE_ko.md에 μžˆμŠ΅λ‹ˆλ‹€.
  • λΌμ΄μ„ μŠ€: 원본과 같은 Apache 2.0(LICENSE, NOTICE).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for jonpark0/embeddinggemma-2-AX650

Quantized
(34)
this model