KhmerLLM_instruct_e3

Model Description

KhmerLLM_instruct_e3 is a Khmer-language instruction-tuned model built on top of Qwen/Qwen2.5-0.5B (base, non-instruct). It was produced in two stages:

  1. Continual pretraining: The Qwen2.5-0.5B base model was further pretrained on ~5GB of Khmer text to adapt it to the Khmer language. The base Qwen2.5-0.5B model struggles to produce coherent Khmer word segmentation and grammar; after continual pretraining, the model generates significantly more coherent and fluent Khmer text.
  2. Instruction fine-tuning: The continually-pretrained model was then fine-tuned on ~50,000 Khmer instruction/response pairs, giving it the ability to follow instructions and behave as a conversational/instruct-style model.

Known limitations: Despite the fluency improvements, the model hallucinates frequently — likely due to the limited scale of both the pretraining corpus (5GB) and instruction dataset (50k pairs) relative to what's needed for a low-resource language like Khmer. Factual claims from this model should not be trusted without verification.

  • Developed by: AnotherPotatoCoder
  • Model type: Causal decoder-only language model (Qwen2 architecture)
  • Language(s): Khmer (km)
  • License: MIT
  • Finetuned from model: Qwen/Qwen2.5-0.5B (base, not instruct)

Uses

Direct Use

  • Khmer text generation and completion
  • Simple Khmer instruction-following / conversational assistant tasks
  • Research and experimentation on low-resource language adaptation

Out-of-Scope Use

  • Factual / knowledge-intensive tasks — the model hallucinates frequently and should not be used where factual accuracy matters (e.g., medical, legal, financial advice).
  • Production or safety-critical deployments without further evaluation and fine-tuning.
  • Tasks requiring strong reasoning or long-context understanding; the model is only 0.5B parameters and has limited capacity.

Bias, Risks, and Limitations

  • Hallucination: The model frequently generates plausible-sounding but factually incorrect or fabricated content, likely due to limited training data scale (~5GB pretraining, ~50k instruction pairs).
  • Data provenance: Training data was a mix of public and self-collected/scraped Khmer text and instruction pairs; it has not been rigorously audited for bias, toxicity, or duplication.
  • Small model size (0.5B params): Limits reasoning ability and knowledge capacity compared to larger models.
  • Language coverage: Optimized for Khmer; performance on other languages is not guaranteed and may be degraded relative to the original Qwen2.5-0.5B base.

Recommendations

Users should independently verify any factual claims generated by this model, especially for anything used outside casual/experimental contexts. This model is best suited for research, prototyping, and further fine-tuning rather than direct deployment.

How to Get Started with the Model

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "AnotherPotatoCoder/KhmerLLM_instruct_e3"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    dtype=torch.float16,
    device_map="auto",
)
model.eval()

prompt = "សួស្ដី! តើអ្នកឈ្មោះអ្វី?"

messages = [
    {"role": "system", "content": "អ្នកគឺជាជំនួយការដ៏ល្អម្នាក់។"},
    {"role": "user", "content": prompt}
]

input_text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

inputs = tokenizer(input_text, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=200,
        do_sample=True,
        temperature=1.0,
        top_k=100,
        top_p=0.8,
        no_repeat_ngram_size=3,
        pad_token_id=tokenizer.eos_token_id,
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training Details

Training Data

A mix of publicly available Khmer text/datasets and self-collected/scraped Khmer text:

  • Continual pretraining corpus: ~5GB of Khmer text
  • Instruction fine-tuning dataset: ~50,000 Khmer instruction/response pairs

Exact dataset sources are not fully itemized here; update this section with specific dataset names/links if you'd like full reproducibility and attribution.

Training Procedure

Two-stage training:

  1. Continual pretraining on ~5GB Khmer text corpus, starting from Qwen/Qwen2.5-0.5B (base).
  2. Instruction fine-tuning on ~50,000 Khmer instruction pairs, starting from the continually-pretrained checkpoint (stage 1 output).

Training Hyperparameters

  • Training regime: fp16/mixed precision (weights stored in fp16 on the Hub)
  • Epochs (instruct stage): 3
  • Base architecture: Qwen2 (0.5B parameters)

Speeds, Sizes, Times

  • Continual pretraining: ~4 hours, 3 epochs, on a single NVIDIA A100 (rented via vast.ai)
  • Instruction fine-tuning: performed on the continually-pretrained checkpoint (see above)

Evaluation

No formal benchmark evaluation has been conducted yet. Comparative examples against the base Qwen/Qwen2.5-0.5B-Instruct model are provided below to illustrate qualitative differences in Khmer fluency and coherence.

Qualitative Comparison: KhmerLLM_instruct_e3 vs. Qwen2.5-0.5B-Instruct

Prompt (Khmer) KhmerLLM_instruct_e3 Output Qwen2.5-0.5B-Instruct Output
សួស្ដី! តើអ្នកឈ្មោះអ្វី? សួស្តី! ខ្ញុំឈ្មោះ វាល្មិកិ។ 😊 បង្ហាញដែរ! ຫ្លោយទឹកចិត្តពីរនៅពេលដំណើរកាស់នឹងរបស់ខ្លួន: 1. 英語 (English) 2. 泰语 (Thai) 3. 马尼拉话 (Magana Language) 4. 环保英语 (Eco English) 5. 漢語 (Chinese) សហគមនុមខាង
សូមប្រាប់ខ្ញុំអំពីប្រទេសកម្ពុជា ប្រទេសកម្ពុជាគឺជាប្រទេសដែលចំណាស់ជាងគេនៅអាស៊ីអាគ្នេយ៍ដែលមានអាយុកាលជិត ៣០ ឆ្នាំ។ ប្រទេសកម្ពុជាស្ថិតនៅក្នុងតំបន់អាស៊ីអាគ្នេយ៍ (អ៊ីនassistant) និងអាស៊ីភាគឦសាន (ភូមា) ហើយវាជាទឹកដីនៃតំបន់មាស (អាស៊ី) ដែលរួមមានប្រទេសចំនួន ១០ គឺ៖ កម្ពុជា ឡាវ ភូមា ថៃ វៀតណាម ចិន និងថៃ។ ... ជាទើរជាមួញតិចតូចៗឲ្យស្លេងទៀត: - "មឹកស្រឡាញ់" (ក្នុងគីហ៊ូវ) - "ធ្វើដោយ" (គូងគ្រូគិត) - ប្សាផ្នែកមុខមែលមូប (រហូតទ
តើរាជធានីនៃប្រទេសបារាំងឈ្មោះអ្វី? រដ្ឋធានីនៃប្រទេសប៉ារីសគឺប៉ារីស។ សរសេរពណ៌ស្សូវភាពថាដែលខ្លា់បញ្ចប់ទ័ព្ទ់: 1.សម្រាប់ពិច្ចក្រោម សំខាន់: - 高度: Heights - 安全: Safety 2.ស៊ុនខាងក្នុងឆ្នា៉ែតឹមុខ (
តើអ្នកអាចជួយខ្ញុំសរសេរអ៊ីមែលបានទេ? ក្នុងនាមជាជំនួយការ AI ខ្ញុំមិនមានសមត្ថភាពសរសេរអ៊ីមែលទេ។ ទោះយ៉ាងណាក៏ដោយ ខ្ញុំអាចណែនាំគំនិតមួយចំនួនសម្រាប់ការសរសេរអ៊ីមែលប្រកបដោយប្រសិទ្ធភាព៖ 1. ជៀសវាងអ៊ីមែលវែង៖ ... អៃតូណាសង់សរោចក្រដើម្បីកាលពាណិភ្ជាពជាមួរដែនពេលដំឡើងមើលរបស់ព្រឹត្ថយហ៍ដៃទៅនេះ. ហើយប្រសោបអោយក្នុងសម័យតាមលំ
សូមសរសេរអ៊ីមែលសុំច្បាប់សម្រាកការងារមួយថ្ងៃ ប្រធានបទ៖ ការសម្រាកការលំហែកាយ និងលំហែកាយនៅថ្ងៃឈប់សម្រាក ជាទីគោរព [អ្នកទទួល] ខ្ញុំសង្ឃឹមថាអ៊ីមែលនេះរកឃើញអ្នកបានល្អ។ ខ្ញុំកំពុងសរសេរដើម្បីសុំឈប់សម្រាកមួយថ្ងៃពីការងារដែលខ្ញុំបានធ្វើកាលពីសប្តាហ៍មុន។ ... សមហត៍ទាច់នឹងប្រព័ន្ធផ្ទៃខាងតាំង, អង្លាបកាន់រឿង (Radar) បានផ្សេងៗ៖ "SAR Radar" សំរាប部落ិត្យរង្វង់ចាស់ហ៊ម (Orbital Command and Service Radar) នៅប្តូរកាង
សូមសរសេររឿងខ្លីមួយអំពីកុមារម្នាក់ដែលបានរកឃើញឆ្កែវង្វេង មានពេលមួយនៅក្នុងភូមិតូចមួយស្ថិតនៅចន្លោះភ្នំពីរ មានក្មេងប្រុសម្នាក់ឈ្មោះ Jack ។ Jack មានចិត្តសប្បុរស និងចង់ដឹងចង់ឃើញ ហើយគាត់តែងតែចំណាយពេលដើរលេងក្នុងទីរហោស្ថាន និងព្រៃឈើក្បែរនោះ។ ថ្ងៃមួយ គាត់បានជំពប់ដួលលើផែនទីចាស់ដែលឆ្លាក់តាមព្រៃ។ ... "រិទ្យាស៊ូ"
សូមតែងកំណាព្យខ្លីមួយអំពីរដូវវស្សា កំដៅព្រះអាទិត្យពេលអ្នកចេញទៅក្រៅ ធ្វើឱ្យស្បែករបស់អ្នកសើមដោយទឹកហូរ។ ទឹកហូរដូចជាអ្នកអាចបន្ធូរអារម្មណ៍បាន មិនមានការរំខាន និងស្ងប់ស្ងាត់នៅក្នុងចិត្តរបស់អ្នក។ ការឆ្លុះបញ្ចាំងនៅក្នុងព្រះអាទិត្យ ... បានធ្វើប្រទេសរបស់នាងឯកាណឡើងឡាប្រដុណ្ហៃប្តូរខ្មៅដែលមានភាពរបាញប្លែក. មិនត្រូndonេះជាងគ្រប់គ្នា៖ 1. សម្រាប់ផ្ដល់មាត
បើសិនជាខ្ញុំមានផ្លែប៉ោម ៥ ហើយឲ្យមិត្តភក្តិ ២ នាក់ៗម្នាក់មួយ ខ្ញុំនៅសល់ប៉ុន្មាន? ប្រសិនបើអ្នកមានផ្លែ pears ចំនួន 5 ហើយមិត្តរបស់អ្នកមានចំនួនពីរ អ្នកនឹងមានផ្លែស្ត្របឺរីសរុបចំនួន 10 ។ ចំនួននៃផ្លែ passistant នឹងមានចំនួនសរុបនៃផ្លែប៉ោម។ ស៊េរីកាលពឹងដែលចាំឆ្នាះថ្មូនបំផុតនៃអ្វីបាន់ទឹក: 1. 萨卡 (Sakura) - ในญี่ปุ่น, ត្រូវកាវមាជធម្មញ្ញ ដំណើរកាណាច់ មាស់គេមុ
តើខ្ញុំគួរធ្វើដូចម្តេចដើម្បីរៀនភាសាអង់គ្លេសឱ្យបានលឿន? នេះគឺជាគន្លឹះខ្លះៗដែលអាចជួយអ្នករៀនភាសាអាមេរិកបានយ៉ាងលឿន៖ 1. ចាប់ផ្តើមដំបូង៖ វាជារឿងសំខាន់ក្នុងការរៀនភាសាថ្មីឱ្យបានឆាប់តាមដែលអាចធ្វើទៅបាន។ ... សិទ្ធិពីបាត់ហ៊ូនមុខរង្វាល់សងៗ 2019, គឯណៅជាពែលកាន់ដោយប្រទេсяថ្មឹត្ថន្ទមានាទាផ្លូវគុណ. បាឌឈាងពិនិត្យសាច់ផ្លាស
សូមបកប្រែឃ្លានេះទៅជាភាសាអង់គ្លេស៖ 'ខ្ញុំស្រលាញ់ប្រទេសកម្ពុជា' "ខ្ញុំស្រលាញ់កម្ពុជា" សំរាប់សិក្សាមើម, "ខ្ដឹ交ស្ថាវតងី" (Kongfu) ហើយមានឈ្មោះកាយស៊េរិចក្នុងរវាងអំពីក្រុមហ៊ូស្មាណពិភពលោក។ 1. អង្គកាត់ឯក

Takeaway: KhmerLLM_instruct_e3 produces dramatically more fluent, grammatical, and topically-relevant Khmer text than the unmodified Qwen2.5-0.5B-Instruct, which struggles to produce coherent Khmer at all. However, the examples above also illustrate this model's core weakness — hallucination: e.g. the math prompt ("5 apples, give 2 to friends") is answered incorrectly ("10" instead of the correct "3"), the translation prompt echoes the Khmer phrase back instead of translating it to English, and longer factual answers (e.g. about Cambodia) drift into inaccurate claims partway through. Treat outputs as fluent but unreliable.

Note: KhmerLLM_instruct_e3 was fine-tuned from a continually-pretrained Khmer version of Qwen2.5-0.5B (base), while the comparison model here, Qwen2.5-0.5B-Instruct, is Qwen's own official instruct model with no Khmer-specific adaptation. This comparison is meant to demonstrate the effect of Khmer-specific continual pretraining + fine-tuning, not a general capability benchmark.

Hardware Used

  • Hardware Type: NVIDIA A100 (rented via vast.ai) for continual pretraining; NVIDIA P100 (Kaggle) for instruction fine-tuning
  • Hours used: ~4 hours (continual pretraining, A100) + ~3 hours (instruction fine-tuning on ~50,000 pairs, P100)
  • Cloud Provider: vast.ai (continual pretraining), Kaggle (instruction fine-tuning)

Technical Specifications

Model Architecture and Objective

Causal (autoregressive) decoder-only transformer, Qwen2 architecture, 0.5B parameters. Trained with a standard next-token prediction objective during continual pretraining, and instruction-tuned (supervised fine-tuning on instruction/response pairs) in the second stage.

Compute Infrastructure

  • Hardware: NVIDIA A100 GPU (vast.ai rental)
  • Software: 🤗 Transformers, PyTorch

Model Card Authors

AnotherPotatoCoder

Model Card Contact

Open an issue or discussion on this model's Hugging Face repository.

Downloads last month
48
Safetensors
Model size
0.5B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AnotherPotatoCoder/KhmerLLM_instruct_e3

Finetuned
(685)
this model