ruri-v3-30m-lite: Zero-Torch & Ultra-Fast Japanese Text Embeddings

An ultra-optimized, lightweight Japanese text embedding model derived from cl-nagoya/ruri-v3-30m (ModernBERT-based, 256-dim).
Engineered strictly with sentencepiece_lite + onnxruntime + numpy โ€” 100% PyTorch-free (Zero-Torch) and zero compilation required (pre-built wheels provided).


๐Ÿ’ก Highlights

  • Ultra-Compact Footprint (Zero-Torch):
    • Runtime only requires sentencepiece_lite, onnxruntime (or onnxruntime-gpu), numpy, and huggingface_hub.
    • Eliminates gigabytes of PyTorch / Transformers dependencies. Ideal for serverless (AWS Lambda, Cloud Run) and edge environments with instant cold starts.
  • Lightning-Fast Multi-Threaded Tokenization:
    • Powered by Google's SentencePiece Lite with Safe Boundary Pre-tokenization (SBP) and FlatBuffers zero-copy mmap.
    • Over 13.4x faster than standard Hugging Face Fast Tokenizers (~3.4 ยตs per sentence).
  • Hardware-Adaptive Auto-Switching:
    • Dynamically queries GPU Compute Capability via ctypes without PyTorch overhead.
    • Automatically loads FP16 ONNX on modern GPUs (RTX 3060+, Ampere/Ada/Hopper) to leverage Tensor Cores, and FP32 ONNX on Pascal (GTX 1080) or CPU.
  • Mathematical Equivalence:
    • Validated cosine similarity of 1.0000 against official PyTorch outputs (Mean Pooling + L2 normalization).

๐Ÿš€ Installation

End users do not need PyTorch, Transformers, or C++ compilers:

# Install pre-built wheel and runtime dependencies
pip install https://huggingface.co/Chottokun/ruri-v3-30m-lite/resolve/main/wheels/sentencepiece_lite-0.1.0-cp311-cp311-linux_x86_64.whl \
            onnxruntime numpy huggingface_hub

# For GPU acceleration (CUDA)
pip install onnxruntime-gpu

๐Ÿ’ป Quickstart Inference

from ruri_v3_lite import RuriV3Lite
import numpy as np

# Download and initialize model (automatically chooses FP16 or FP32 based on hardware)
model = RuriV3Lite(repo_id="Chottokun/ruri-v3-30m-lite")

# Sentences formatted with official ruri-v3 task prefixes
queries = ["ๆคœ็ดขใ‚ฏใ‚จใƒช: ๆ—ฅๆœฌใฎ้ฆ–้ƒฝใฏใฉใ“ใงใ™ใ‹๏ผŸ"]
documents = [
    "ๆ–‡็ซ : ๆ—ฅๆœฌใฎ้ฆ–้ƒฝใฏๆฑไบฌ้ƒฝใงใ™ใ€‚",
    "ๆ–‡็ซ : ๅคง้˜ชใฏ้–ข่ฅฟๅœฐๆ–นใฎไธป่ฆ้ƒฝๅธ‚ใงใ™ใ€‚"
]

# Generate normalized embeddings (shape: (N, 256))
q_emb = model.encode(queries)
d_emb = model.encode(documents)

# Compute cosine similarity via dot product
similarities = np.dot(q_emb, d_emb.T)[0]

print(f"Query vs Tokyo: {similarities[0]:.4f}")
print(f"Query vs Osaka: {similarities[1]:.4f}")

โš–๏ธ License & Redistribution Terms (Apache License, Version 2.0)

This model repository, converted binaries, and wrapper scripts are distributed under the Apache License, Version 2.0 (the "License"). You may not use these files except in compliance with the License. You may obtain a copy of the License at:

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

Upstream Components & Notices of Modification (Section 4)

In compliance with Section 4 of the Apache License 2.0:

  1. Base Embedding Model:

    • Model: cl-nagoya/ruri-v3-30m
    • Copyright: Copyright 2024 Hayato Tsukagoshi and Ryohei Sasano (Nagoya University NLP Laboratory)
    • License: Apache License, Version 2.0
    • Prominent Notice of Modification (Section 4b):
      • Converted original PyTorch Safetensors weights into ONNX (FP32) and FP16 formats.
      • Converted official SentencePiece model into FlatBuffers serialization format (.spm.fb).
      • Added independent inference implementation (ruri_v3_lite.py) decoupling from PyTorch and Transformers.
  2. Tokenizer Engine:


๐Ÿ“š Citations

When utilizing or evaluating this model or tokenizer, please cite the corresponding works:

Ruri (Embedding Model)

@misc{Ruri,
  title={Ruri: Japanese General Text Embeddings}, 
  author={Hayato Tsukagoshi and Ryohei Sasano},
  year={2024},
  eprint={2409.07737},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2409.07737}, 
}

SentencePiece (Tokenizer)

@inproceedings{kudo-richardson-2018-sentencepiece,
  title = "{S}entence{P}iece: A simple and language independent subword tokenizer and detokenizer for {N}eural {T}ext {P}rocessing",
  author = "Kudo, Taku and Richardson, John",
  booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations",
  month = nov,
  year = "2018",
  address = "Brussels, Belgium",
  publisher = "Association for Computational Linguistics",
  url = "https://aclanthology.org/D18-2012",
  doi = "10.18653/v1/D18-2012",
  pages = "66--71",
}

๐Ÿ“„ License File

A full copy of the license is included in LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Chottokun/ruri-v3-30m-lite

Quantized
(11)
this model

Paper for Chottokun/ruri-v3-30m-lite