UnimeType Transliteration 4B

Type with familiar Latin characters. Get the writing system your message needs.

UnimeType Transliteration 4B is a focused language conversion model for everyday, mixed-language writing. It converts Pinyin, Romaji, Romanized Hindi, Hinglish, and Arabizi while keeping clear English words, product names, links, code, numbers, punctuation, and emoji in place.

It was created for the Convert experience from UnimeType.

Four writing systems, one familiar input

Input style Example input Model output
Chinese Pinyin jintian yao finish README 今天要 finish README
Japanese Romaji ashita Tokyo de meeting ga arimasu 明日 Tokyo で meeting があります
Romanized Hindi / Hinglish aaj meeting hai आज meeting है
Arabic Arabizi fee meeting bokra في meeting بكرة

These examples are direct outputs from the published 6-bit model in LM Studio.

Made for real messages

UnimeType Transliteration 4B is designed for text that does not fit into a traditional transliteration box:

  • Mixed-language sentences that combine romanized text with English terms
  • Short chat messages and conversational phrases
  • Product names, technical vocabulary, URLs, email addresses, and code spans
  • Punctuation, line breaks, Markdown, numbers, and emoji that should remain intact
  • Text that is already in the target writing system and should not be changed

The model returns the converted text directly, ready to review and use.

Supported conversion modes

Target Latin input Output
Simplified Chinese Pinyin and natural unmarked Pinyin Chinese characters
Japanese Romaji Kanji, Hiragana, and Katakana
Hindi Romanized Hindi and Hinglish Devanagari with preserved English
Arabic Arabizi and conversational romanized Arabic Arabic script with preserved English

Run locally

This repository contains an MLX model for Apple silicon Macs. The model can be downloaded from Hugging Face and loaded with MLX-LM or imported into an MLX-compatible local runtime such as LM Studio.

Recommended generation settings:

Setting Value
Thinking / reasoning Off
Temperature 0
Maximum output 128 tokens for short messages
Context used for validation 2048 tokens

Select one target language from chinese, japanese, hindi, or arabic, then provide the original text. The model is built for conversion, not translation, rewriting, polishing, or explanation.

Model details

Property Value
Base model Qwen3.5-4B
Parameters 4B class
Format MLX
Quantization 6-bit affine, group size 64
Model file size 3.42 GB
Primary task Context-aware transliteration and script conversion
License Apache 2.0

Evaluation

The published model was tested with exact final-text matching. A result passes only when the complete output matches an accepted answer, including preserved English, punctuation, code, and spacing.

Runtime Everyday conversion suite Reserved multilingual suite
Direct MLX 98 / 127 57 / 104
LM Studio 94 / 127 57 / 104

The reserved multilingual score measures exact agreement with its reference responses. Exact-match scoring is strict: a valid alternative spelling or wording can still count as a mismatch. Review ambiguous names, dialect terms, and very short fragments before sending them.

About UnimeType

UnimeType helps multilingual writers type naturally across languages and writing systems. Transliteration 4B is the dedicated local model behind that direction: focused on conversion, compact enough for local use, and designed around the way people actually mix languages in daily writing.

Downloads last month
-
Safetensors
Model size
0.9B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UnimeType/Transliteration-4B

Finetuned
Qwen/Qwen3.5-4B
Quantized
(401)
this model