Okinawa Dialect Assistant (Shuri dictionary prototype)

A small LoRA fine-tune (merged into the base model for simple loading and hosted inference) intended to map short Japanese prompts to the corresponding Shuri Okinawan dictionary form. This is an experimental lexical translation aid, not a fluent Okinawan conversation model. It does not generate audio.

Training data

Training examples come from LoNebula/okinawa-ja-shuri-dictionary-translation, a filtered and attributed derivative of doraking/ryukyuan-okinawan-corpus. The source entries are marked as Shuri dialect and attributed to the National Institute for Japanese Language and Linguistics (NINJAL) dictionary under CC BY 4.0. The dataset contains dictionary entries, not parallel conversational sentences. Preserve the attribution and follow the dataset card and CC BY 4.0.

The source corpus spells Okinawan entries with a romanization and specialized symbols. Output should be treated as that dictionary notation, not standard Japanese kana orthography. Individual Okinawan varieties differ, and this model covers Shuri forms only.

Base model and method

  • Base: Qwen/Qwen2.5-0.5B-Instruct
  • Method: LoRA fine-tuning, rank 8, three epochs
  • Dataset: 9,031 filtered dictionary entries, split 90/10 (seed 42) for training and held-out dictionary loss
  • Base model license: Apache-2.0
  • The model repository follows the base model's Apache-2.0 license. The separately published training dataset is CC BY 4.0; retain its source attribution when reusing or redistributing that dataset.

Load and use with Python

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "LoNebula/okinawa-dialect-assistant"
tokenizer = AutoTokenizer.from_pretrained(model_id)
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16 if device == "cuda" else torch.float32,
).to(device)
messages = [
    {"role": "system", "content": "日本語に対応する首里方言の辞書形を答えてください。"},
    {"role": "user", "content": "ありがとう"},
]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(device)
output = model.generate(inputs, max_new_tokens=48, do_sample=False)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))

Install the runtime with pip install torch transformers. The repository contains merged weights, so peft is not needed to load it. The model can be loaded directly from its Hugging Face repository using transformers; GPU use is optional for this 0.5B model.

Limitations

  • Dictionary-level coverage only; sentence-level naturalness has not been established.
  • Held-out loss measures dictionary entry modeling only; it is not a sentence translation quality score.
  • The model may return romanized dictionary notation rather than kana or conversational speech.
  • It may hallucinate or choose an unsuitable lexical item. Have a fluent speaker review important uses.
  • No audio or speech synthesis model is included.
  • The published repository contains merged model weights and tokenizer, so it can load through ordinary Transformers or compatible hosted inference providers.

Attribution

National Institute for Japanese Language and Linguistics, Okinawa-go Jiten (1963; index), dataset packaging by doraking/ryukyuan-okinawan-corpus, CC BY 4.0. Source: https://mmsrv.ninjal.ac.jp/okinawago/.

Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LoNebula/okinawa-dialect-assistant

Adapter
(878)
this model