Jamjuri Mix7M

Clean Thai fine-tune mix for benchmark maxxing: 30,069 docs / 8,302,257 tokens (Qwen2.5 tokenizer, no special tokens).

Derived from SPAISS6F1/spai-ss6-llm-1b-thai-corpus โ€” check upstream terms (medical/exam subsets carry their own licenses) before commercial use. Research/eval use.

Composition

Source Docs Tokens Share
lst20 Thai QA 7,643 2,796,494 33.7%
exam QA 8,109 2,267,511 27.3%
synthetic_qa (10% sample) 1,324 1,200,296 14.5%
medical_qa, non-star (34% sample, CoT dropped, Q+A only) 5,579 800,210 9.6%
tourist instruction (9% sample) 2,454 600,278 7.2%
local_v2 instruction (9% sample) 3,641 500,035 6.0%
idioms (full) 1,152 113,001 1.4%
synonym (full) 167 24,432 0.3%

Cleaning

  • Dropped <think>-containing rows (found zero โ€” corpus already clean)
  • Min 20 chars, exact-md5 dedupe 137,708 โ†’ 111,563 before sampling
  • Stratified-random sample, seed 42

Format

mix7m.jsonl: one JSON per line โ€” {"text": ..., "source": ..., "ntok": ...}. MANIFEST.json: exact counts.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support