Jamjuri Mix7M
Clean Thai fine-tune mix for benchmark maxxing: 30,069 docs / 8,302,257 tokens (Qwen2.5 tokenizer, no special tokens).
Derived from SPAISS6F1/spai-ss6-llm-1b-thai-corpus โ check upstream terms (medical/exam subsets carry their own licenses) before commercial use. Research/eval use.
Composition
| Source | Docs | Tokens | Share |
|---|---|---|---|
| lst20 Thai QA | 7,643 | 2,796,494 | 33.7% |
| exam QA | 8,109 | 2,267,511 | 27.3% |
| synthetic_qa (10% sample) | 1,324 | 1,200,296 | 14.5% |
| medical_qa, non-star (34% sample, CoT dropped, Q+A only) | 5,579 | 800,210 | 9.6% |
| tourist instruction (9% sample) | 2,454 | 600,278 | 7.2% |
| local_v2 instruction (9% sample) | 3,641 | 500,035 | 6.0% |
| idioms (full) | 1,152 | 113,001 | 1.4% |
| synonym (full) | 167 | 24,432 | 0.3% |
Cleaning
- Dropped
<think>-containing rows (found zero โ corpus already clean) - Min 20 chars, exact-md5 dedupe 137,708 โ 111,563 before sampling
- Stratified-random sample, seed 42
Format
mix7m.jsonl: one JSON per line โ {"text": ..., "source": ..., "ntok": ...}.
MANIFEST.json: exact counts.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support