ruri-v3-130m-coreml

CoreML conversion of cl-nagoya/ruri-v3-130m (a Japanese ModernBERT sentence-embedding model) for running the encoder on the Apple Neural Engine (ANE).

What this is

The ANE needs a fixed-shape, batch=1 model, and one model per sequence length. Those lengths live in a single bundle as four CoreML functions over one shared copy of the weights:

buckets-64-128-256-512.mlpackage    # seq_64, seq_128, seq_256, seq_512
tokenizer.json
config.json
1_Pooling/config.json

Only the encoder forward ((input_ids, attention_mask) -> last_hidden_state) is in CoreML; tokenization, prefixing, mean pooling and L2 normalization stay in the host application. Output matches the reference f32 path to fp16 rounding — measured 1 - cosine of 5.2e-6 (median over Japanese and English texts up to 512 tokens).

Requirements

  • macOS 15 or newer. The conversion targets macOS15 / iOS 18 opsets.
  • Apple Silicon, for the ANE to exist at all.
  • For the kohagi path: kohagi 0.5.0 or newer, built with --features coreml. Earlier versions look for seq-<N>.mlpackage and will not find anything here.

Usage

kohagi

kohagi does the tokenizing, bucketing, pooling and normalization:

kohagi --device coreml --coreml-model-id takahashim/ruri-v3-130m-coreml \
       --prefix "検索文書: " < texts.jsonl

The first run compiles the bundle for the Neural Engine, which takes roughly 20 seconds; the result is cached under ~/Library/Caches/kohagi/coreml, so later runs start in well under a second.

You do not need this repository at all if you have kohagi 0.5.0: kohagi --device coreml converts cl-nagoya/ruri-v3-130m itself and caches the result. This repo is the shortcut — 260 MB of converted model against 503 MB of safetensors plus a 20-second conversion.

coremltools

Each function takes input_ids and attention_mask ([1, N] int32) and returns hidden ([1, N, 512] fp16). Pick the function by name and pad to its length.

import coremltools as ct, numpy as np
m = ct.models.MLModel("buckets-64-128-256-512.mlpackage", function_name="seq_128")
ids  = np.zeros((1, 128), np.int32); ids[0, :n]  = token_ids
mask = np.zeros((1, 128), np.int32); mask[0, :n] = 1
hidden = m.predict({"input_ids": ids, "attention_mask": mask})["hidden"][0]  # (128, 512)
vec = (hidden * mask[0, :, None]).sum(0) / mask.sum()
vec /= np.linalg.norm(vec)

Ruri expects a task prefix — "検索文書: " for documents, "検索クエリ: " for queries — prepended to the text before tokenizing.

Changes from the base model

  • Converted to a CoreML ML Program with fp16 weights, targeting CPU_AND_NE.
  • Four fixed sequence lengths as CoreML functions in one bundle, batch size 1.
  • Weights and forward math are otherwise unchanged from cl-nagoya/ruri-v3-130m (revision e3114c6).

Converted with kohagi 0.5.0's own emitter (coreml-convert), which for this model produces byte-identical output to the coremltools route and the same Neural Engine placement.

License and attribution

Apache-2.0, inherited from the base model cl-nagoya/ruri-v3-130m (itself built on ModernBERT-Ja). This is a format conversion (a derivative work); all credit for the model belongs to the original authors. Please cite Ruri as requested on the base model card.

Downloads last month
108
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for takahashim/ruri-v3-130m-coreml

Quantized
(11)
this model