Laya Thai typed decisions
Fine-tune of convaiinnovations/laya (subfolder="multilingual") for Thai typed decisions. The model does not write text. One forward pass stamps choice, score, and noul (yes/no) answers and returns a probability for each.
The weights in this repo are epoch_02. The config is the calibrated one: same weights, softer confidence. A single temperature does not change the chosen label.
| Repo | team-od/laya-thai |
| Base | mmBERT-base, 322M parameters, context 1,024 |
| Library | laya 0.3.6 |
| Training device | Apple M4 Max, PyTorch, fp32 |
| License | CC BY-NC 4.0 |
| Date | 2026-09-26 |
Attribution
This release is published by the Hugging Face organization that owns the repo. No individual author is named.
The model is a fine-tune, not a new base:
| Piece | Credit | License of that piece |
|---|---|---|
| Base checkpoint | Convai Innovations, convaiinnovations/laya, multilingual subfolder |
Apache-2.0 |
| Encoder | jhu-clsp/mmBERT-base |
as shipped inside the Laya checkpoint |
| Thai and English intent rows | Amazon MASSIVE, mteb/amazon_massive_intent |
CC BY 4.0 |
| Thai entailment rows | facebook/xnli, config th |
CC BY-NC 4.0 |
| Thai news-topic rows | PyThaiNLP, prachathai-67k |
CC BY-NC on the Hugging Face card |
| Thai review rows | Wongnai, Wongnai/wongnai_reviews |
LGPL-3.0 |
Comparison numbers for iapp/OpenThai-SystemOne in the speed table were measured locally from those weights. iApp is not an author of this fine-tune.
What it is for
Route a Thai sentence to a label the caller already named.
- Intent among a menu of at most 20 names.
- News topic among the 12 Prachathai topics, and a yes/no "is this that topic?"
- A rough yes/no on whether a hypothesis follows a premise, and a 3-way entailment label.
- A rough 1–5 star guess on a Thai review. This is the weak stamp.
The English refund check still lands on billing. Short English intent menus are in scope. Long English work is not.
What it is not for
- Writing or extracting an answer span.
- A published star rating. Exact-rung accuracy on the held-out reviews is 0.599.
- Exams, math, code, handwriting, or images.
- Commercial use. See the license.
- A gate that treats an uncalibrated confidence of 0.99 as certainty. This repo already loads the softened temperatures. The 3-way choice and the yes/no stamp are still somewhat overconfident.
How to run it
import os
os.environ["USE_TF"] = "0"
import laya
agent = laya.Agent("team-od/laya-thai", device="mps")
out = agent.predict(
{"utterance": "คิดเงินซ้ำสองครั้ง กรุณาคืนเงินวันนี้"},
{
"department": {
"type": "choice",
"instructions": "Who should handle this?",
"criteria": {
"billing": "invoices, payments, refunds",
"technical": "bugs and outages",
"sales": "pricing and new plans",
},
}
},
)
print(out["answers"]["department"]["choice"])
choice returns the key it picked and a probability for every name on that slip. noul returns the probability the answer is yes. score returns an expected rung in 0 .. K-1 plus a probability per rung. A 1–5 star label in the training files is 1–5. The probability keys the model reads are "0" through "4".
Training
Three stages, each continuing from the previous weights. The loss is the upstream RLCD objective: Gaussian noise on the logits (group size 4), proper_reward with spherical weight 0.75, plus soft cross-entropy of weight 1.0. Optimizer is AdamW, encoder learning rate 2.5e-5, head learning rate 1e-4, weight decay 0.01, gradient clip 1.0. Effective batch size is 64 (microbatch 8, gradient accumulation 8). Training sequences were cut at 384 tokens. The saved config still advertises max_len 1024.
Noise starts at 0.4 and is scheduled down to 0.1 across a 4-epoch horizon. The last stage ran only two epochs, so it stopped at 0.30.
| Stage | Start | Train rows | Epochs kept |
|---|---|---|---|
| MASSIVE Thai only | multilingual base | 11,514 | epoch 6 of 8 |
| Mixed stamps | that epoch 6 | 83,028 rows, 139,752 questions | epoch 2 of 3 finished |
| Extra Wongnai reviews | that epoch 2 | 108,028 rows, 164,752 questions | epoch 2 of 2 |
The mixed file was MASSIVE Thai 11,514, XNLI Thai 20,000, Wongnai 15,000, Prachathai 25,000, and MASSIVE English 11,514. The last stage replaced the 15,000 Wongnai rows with all 40,000 labeled training reviews.
Rows the optimizer never saw:
- MASSIVE Thai test, 2,974. Menus are 20 names: the true intent plus 19 decoys, seed 13, Thai instruction. This is not the published 60-way number and not the 40-slip English-instruction probe.
- XNLI Thai validation, 2,490. The XNLI test split was not used.
- Prachathai validation, 6,721.
- Wongnai measure file, 6,203. This is the dataset's test split. There is no separate Wongnai validation split.
Evaluation
Held-out accuracy of epoch_02. A question is correct when the argmax matches the gold label. A row with two questions counts twice in n.
| Set | Questions | Accuracy |
|---|---|---|
| MASSIVE Thai test, 20-way intent | 2,974 | 0.888 |
| XNLI Thai validation, 3-way choice | 2,490 | 0.714 |
| XNLI Thai validation, yes/no | 2,490 | 0.820 |
| XNLI Thai validation, both | 4,980 | 0.767 |
| Prachathai validation, 12-way topic | 3,501 | 0.888 |
| Prachathai validation, yes/no | 13,119 | 0.881 |
| Prachathai validation, both | 16,620 | 0.882 |
| Wongnai measure, exact star | 6,203 | 0.599 |
Same files, stock multilingual Laya and iapp/OpenThai-SystemOne, one option order as written in the file:
| Job | Stock | Ours | iApp |
|---|---|---|---|
| Thai intent, 20-way | 0.395 | 0.888 | 0.905 |
| News topic, 12-way | 0.371 | 0.888 | 0.943 |
| "Is this that topic?" | 0.600 | 0.881 | 0.904 |
| "Does the hypothesis follow?" | 0.805 | 0.820 | 0.843 |
| 3-way entailment | 0.677 | 0.714 | 0.772 |
| Exact star | 0.255 | 0.599 | 0.630 |
Reference points for the MASSIVE test file only: the earlier MASSIVE-only fine-tune scored 0.878. Chance on that file is 0.05. Chance on the 12-way topic menu is 0.083, on the 3-way entailment menu 0.333, and on five stars 0.20.
Smoke, uncalibrated weights (temperatures at 1.0):
| Slip | Department | P(refund) |
|---|---|---|
| Thai double charge | billing | 1.000 |
| Thai outage | technical | 0.000 |
| English "Charged twice… refund today." | billing | 1.000 |
On this calibrated repo the English refund is still billing, and P(refund) is 0.949.
Speed
Milliseconds per row on an Apple M4 Max, Metal, after the weights are loaded. Laya was timed one row per call. iApp was timed in batches of 4, and batches of 2 on Prachathai.
| File | Rows | Stock | Ours | iApp |
|---|---|---|---|---|
| MASSIVE Thai test | 2,974 | 18 ms | 16 ms | 195 ms |
| XNLI Thai validation | 2,490 | 18 ms | 18 ms | 175 ms |
| Prachathai validation | 6,721 | 59 ms | 62 ms | 616 ms |
| Wongnai measure | 6,203 | 21 ms | 20 ms | 547 ms |
Calibration
Temperatures in temperature_by_options were fit on MASSIVE validation, XNLI validation, and Prachathai validation. The Wongnai measure file was not used, so score:3-5 stays at 1.0.
A temperature divides the logits of one question. It does not change the argmax. Laya clamps temperatures to [0.5, 5.0].
| Bucket | Questions | Temperature | ECE before → after |
|---|---|---|---|
choice:11+ (12- and 20-way) |
5,534 | 3.30 | 0.096 → 0.030 |
choice:3-5 |
2,490 | 5.00 | 0.251 → 0.092 |
noul:2 |
15,609 | 5.00 | 0.109 → 0.032 |
The 3-way and yes/no fits both sat on the ceiling of 5, so those confidences are still a bit sharp.
Limits
- Menus longer than 20 names were not trained. MASSIVE has 60 intents. The 0.888 number is the 20-name protocol.
- Star errors are exact-match. The card does not say whether the misses are off by one.
- Training saw at most 384 tokens. Longer inputs are cut by the library at 1,024.
- No Thai exam, toxicity, wiki answerability, Wisesight sentiment, or SIB-200 result is claimed.
- Confidence on
choice:3-5andnoul:2is improved and still not tight.
Picking the work back up
The weights to continue from are this repo, team-od/laya-thai, which is the calibrated epoch_02. Argmax matches the uncalibrated checkpoint. Locally that folder is checkpoints/wongnai_heavy_rlcd/epoch_02. The calibrated config is epoch_02_calibrated.
Do not train on these files. They are the scores in the table above:
data/massive_th.test.jsonl(2,974). Bar is 0.888.data/xnli_th.val.jsonl(2,490). Bar is 0.767.data/prachathai.val.jsonl(6,721). Bar is 0.882.data/wongnai_measure.jsonl(6,203). Bar is 0.599. This is the Wongnai test split. There is no separate validation split.
Train only from files whose names end in .train.jsonl. The heavy mix that produced this checkpoint is data/train_wongnai_heavy.jsonl (108,028 rows): MASSIVE Thai 11,514, XNLI Thai 20,000, all 40,000 Wongnai train reviews, Prachathai 25,000, MASSIVE English 11,514. The smaller data/train.jsonl used 15,000 Wongnai rows and is the previous mix, not this one.
Another epoch of that same heavy file is a poor bet. Wongnai moved from 0.589 to 0.599, and MASSIVE, XNLI, and Prachathai stayed flat. The next change, if the star number has to move, is the score loss: a 1-star miss and a 4-star miss are the same error today. Fit temperatures again only after a new scoreboard. Do not fit them on wongnai_measure.jsonl.
Score with:
uv run python eval_massive_th.py --ckpt checkpoints/wongnai_heavy_rlcd/epoch_02 --file data/massive_th.test.jsonl
uv run python eval_decisions.py --ckpt checkpoints/wongnai_heavy_rlcd/epoch_02 --file data/xnli_th.val.jsonl
uv run python eval_decisions.py --ckpt checkpoints/wongnai_heavy_rlcd/epoch_02 --file data/prachathai.val.jsonl
uv run python eval_decisions.py --ckpt checkpoints/wongnai_heavy_rlcd/epoch_02 --file data/wongnai_measure.jsonl
License
CC BY-NC 4.0. Non-commercial use only.
The base Laya checkpoint is Apache-2.0. This fine-tune also trained on XNLI (CC BY-NC 4.0) and on Prachathai, whose Hugging Face card says CC BY-NC. The Wongnai review corpus is LGPL-3.0. Those terms are why this repo is not Apache-2.0.
Model tree for team-od/laya-thai
Base model
convaiinnovations/laya