EXLLM-JPTOEN

EXLLM-JPTOEN is an experimental 0.005377824B-parameter Little Language Model for generating short English expressions from kana-oriented Japanese input. EXQ12 integer inference has been verified on a CASIO EX-word XD-B4800 (DATAPLUS 6).

This is a narrow task-specific model centered on 266 lexical entries, not a general translation model. The included GGUF is a separately trained Llama-compatible companion—not a conversion or quantization of the embedded EXLLM checkpoint.

日本語の概要はこちらです。

Highlights

  • 0.005377824B parameters; 0.000128M-token context; 0.000868M-token vocabulary
  • 0.100003620B processed non-padding training tokens, including extensive repetition and template-derived examples
  • EX-word integer inference verified at approximately 0.57–0.58 token/s in the recorded XD-B4800 runs
  • Project-specific evaluation: 266/266 bare known words, but only 34/266 fully unseen query templates
  • Separately trained 0.005441184B-parameter GGUF companion for LM Studio and llama.cpp
  • Apache-2.0

Model Details

Field Value
Model type decoder-only Transformer
Parameters 5,377,824 (0.005377824B)
Layers 6
Hidden size 288
Attention heads 9
FFN size 896
Context 128 tokens (0.000128M)
Vocabulary 868 tokens (0.000868M)
Tokenizer frequent Japanese characters + UTF-8 byte fallback, NFC
Embeddings tied token embedding and output head
License Apache-2.0

The parameter count excludes duplication from tied weights.

Training and Provenance

EXLLM-JPTOEN continues from an unpublished in-project checkpoint of the same architecture. No external pretrained checkpoint was used. The final checkpoint records 100,003,620 (0.100003620B) processed non-padding input tokens, including 97,102,074 continuation tokens over 11,488 steps with seed 40417001.

This count is not unique corpus size. Training repeatedly sampled deterministic examples derived from a project-authored 266-row lexicon, finite question and noise templates, and unknown-calibration lists. Most processed tokens are repeated or template-derived. The project owner assembled and edited these resources with generative-AI assistance; no third-party dictionary corpus or web crawl is recorded in this checkpoint lineage.

Exact hashes, token-count definitions, data rows, training mix, and lineage are recorded in DATA_PROVENANCE.md, the GitHub provenance document, and the machine-readable case-study manifest.

Evaluation

These are project-specific greedy-decoding regression suites, not standard translation benchmarks.

Suite Result
Bare known words 266 / 266
Learned-family query forms 262 / 266
Fully unseen query templates 34 / 266
Real-word unknown holdout 22 / 40
Ambiguous unknown 12 / 12
Nonce unknown 62 / 64
General OOD 5 / 5
EX-word regression set 6 / 6
Visible ASCII output 925 / 925

FP32 and dequantized INT8 produced the same results on these suites. The 34/266 fully unseen-template result is the clearest limitation: 0.100003620B processed tokens did not establish broad grammatical or general translation ability.

Real-word unknown holdout measures rejection of words outside the registered vocabulary contract, not translation accuracy. Semantically correct outputs such as たまご → egg or にく → meat therefore count as failures in that suite. Full cases and input hashes are available in common-eval-int8.json.

Physical EX-word Results

On the XD-B4800, model identification, loading, inference, and log capture were verified from both internal storage and the microSD registration cache. Recorded decode throughput was approximately 0.57–0.58 token/s. Correct answers, appropriate unknown responses, and incorrect English outputs all occurred; for example, ぎんこう produced temperature.

The complete five-prompt timing record and screenshots are referenced by the fixed-5M case study. Host-side regression results must not be interpreted as general free-form translation quality.

Usage

LM Studio / llama.cpp

EXLLM-JPTOEN-F16.gguf is a separately trained Llama-compatible companion created from the same JPTOEN data contract. It has different weights, tokenizer, architecture details, and runtime context from the embedded checkpoint.

llama-cli \
  -m EXLLM-JPTOEN-F16.gguf \
  -c 512 --single-turn \
  -p "おはようをえいごでいうと?"

An RTX 3060 llama.cpp smoke test generated good morning at approximately 415 token/s. This is a PC companion smoke-test measurement, not a directly comparable benchmark against the EX-word model. Its training record is in lmstudio/training_manifest.json and lmstudio/dataset-manifest.json.

PyTorch reference runtime

git clone https://github.com/ToTo-40417/exllm
cd exllm
python -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt

Load weights/EXLLM-JPTOEN.pt with this repository's config.json and tokenizer.json through the EXLLM reference runtime.

CASIO EX-word

Use exllm-exword v1.3.1 or later with weights/EXLLM-JPTOEN.q12. The file can be checked before deployment with exllm-model-check.

Release Artifacts

Artifact Size SHA-256 Purpose
EXLLM-JPTOEN.pt 21,536,780 bytes (20.54 MiB; 0.021536780 GB) 942ec9f1610a5566585a7597a1bac67b4128814532669dffc334c6629bd21fe1 FP32 PyTorch checkpoint
EXLLM-JPTOEN-int8.bin 5,443,105 bytes (5.19 MiB; 0.005443105 GB) 6f0d77b498e5e9f3bda8480832b6afdc66306205ab4fbdae2911203f935a0675 EXLLM8 INT8 artifact
EXLLM-JPTOEN.q12 5,443,105 bytes (5.19 MiB; 0.005443105 GB) ea5b59785e00b2a28b3c4eaa1e24785238f64ba3d45e300da7510aabf5d284eb EX-word EXQ12 artifact
EXLLM-JPTOEN-F16.gguf 10,916,864 bytes (10.41 MiB; 0.010916864 GB) 362a237ff01fec440472ceb239fe65b41195e971b55e2fd8609fae3b467f053c Separately trained LM Studio / llama.cpp companion
tokenizer.json 6,864 bytes 4c2a583b554e8165936cb3b11727f79d2ede43b73d416a48681f8d38385f5964 Embedded-model tokenizer

EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.

Intended Use

  • registered kana vocabulary to short English expressions;
  • prompt forms close to learned templates;
  • narrow unknown-rejection experiments;
  • EX-word and extremely small task-specific LM research.

Limitations

  • The task is centered on 266 lexical entries.
  • Unseen syntax, long-form text, free translation, and open-ended conversation are unreliable.
  • Unknown calibration is imperfect; the model may emit a known but unrelated English word.
  • Robustness to kana spelling variation, long vowels, small kana, and typos is not guaranteed.
  • The embedded model context is 128 tokens.
  • Results come from a single continuation seed; no multi-seed comparison was performed.
  • Host evaluation and the SH-4A fixed/integer runtime use different numerical paths.
  • Do not use it for medical, legal, financial, safety-critical, or professional translation.
  • The custom embedded checkpoint is not directly supported by Hugging Face Inference Providers.

Links

日本語

EXLLM-JPTOENは、かな中心の日本語入力から短い英語表現を生成する、5,377,824(0.005377824B)パラメータの実験的なLittle Language Modelです。266語を中心とする狭いモデルであり、一般翻訳モデルではありません。CASIO XD-B4800(DATAPLUS 6)上でEXQ12整数推論を確認しています。

学習量は100,003,620(0.100003620B)processed non-padding tokensですが、その大部分は反復または有限templateからの派生例です。未学習の質問形式は34/266に留まっており、広い文法理解や一般翻訳能力を示す結果ではありません。

同梱GGUFはEX-word用重みの変換ではなく、同じJPTOENデータ契約から別途学習したPC用companionです。詳しい評価、来歴、実行方法、制限事項は上の英語本文と関連リンクを参照してください。

Citation

@software{toto_exllm_jptoen_2026,
  author  = {ToTo},
  title   = {EXLLM-JPTOEN},
  year    = {2026},
  url     = {https://huggingface.co/ToTo-40417/EXLLM-JPTOEN}
}
Downloads last month
-
GGUF
Model size
5.44M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support