Instructions to use faall7479/laya-idjvsuen-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use faall7479/laya-idjvsuen-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="faall7479/laya-idjvsuen-v1")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("faall7479/laya-idjvsuen-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
laya-idjvsuen-v1
Bahasa / Language: Indonesia | English
A multilingual fine-tune of convaiinnovations/laya-multilingual
for calibrated typed decisions (choice / score / noul) on Indonesian (id), Javanese (jv),
Sundanese (su), English (en), and code-switched input.
Author: muhfalihr (github.com/muhfalihr)
| Model | Focus | Link |
|---|---|---|
| laya-idjvsuen-v1 | general multilingual (this repo) | this repo |
| laya-idjvsuen-v3 | v1 + 12-category ticket domain (recommended for ticket routing) | faall7479/laya-idjvsuen-v3 |
This is a non-autoregressive encoder-based decision model (mmBERT-base, 322M parameters): a single forward pass returns a typed answer plus a calibrated probability. The model never generates text and is not a chatbot. It builds on the work of ConvAI Innovations with the open-source NandhaKishorM/laya SDK (Apache-2.0).
Usage
pip install laya
import laya
agent = laya.load("faall7479/laya-idjvsuen-v1")
# Sentiment (choice, 3 options)
r = agent.predict("Pelayanane elek tenan, aku ora arep balik maneh", {
"sentiment": {"type": "choice",
"instructions": "What is the sentiment of the text?",
"criteria": {"negative": "the text expresses a negative opinion",
"neutral": "the text is neutral or factual",
"positive": "the text expresses a positive opinion"}}
})
# r["answers"]["sentiment"]["choice"] -> "negative" | "probabilities" | "answer_confidence"
# Intent (choice, 60 MASSIVE options)
opts = {"alarm_set": None, "calendar_set": None, "qa_factoid": None, ...} # 60 MASSIVE intents
r = agent.predict("bangunkan saya pukul lima pagi minggu ini", {
"intent": {"type": "choice",
"instructions": "Classify the user's utterance into the most likely intent.",
"criteria": opts}
})
For the full list of 60 intents in the exact option order used during training, see the
question_defs.json file in this model repo (or MASSIVE).
Training data
| Task | Languages | Train | Test | Source | Data license |
|---|---|---|---|---|---|
| Intent (60 classes) | id, en | 13,547 / lang | 2,000 / lang | MASSIVE 1.1 | CC-BY-4.0 |
| Intent (60 classes) | jv, su | 7,000 / lang | 2,000 / lang | machine translation idβjv/su (NLLB-200-distilled-600M) | see note |
| Sentiment (3 classes) | id, jv, su, en | 500 / lang | 400 / lang | NusaX-senti (native-speaker manual annotation) | CC-BY-SA-4.0 |
| Code-switching (eval only) | id+en, id+jv, id+su | β | 500 / combo | synthetic parallel half-splice | β |
43,094 training items in total; 400 items held out for temperature calibration fitting.
Training procedure
The RLCD recipe from the official Laya notebook, adapted to a single consumer GPU:
- Objective: policy gradient (REINFORCE, group-mean baseline, G=4, Ο 0.4β0.1) with a strictly proper scoring rule reward (log + spherical 0.75 + RPS 1.0) plus full soft cross-entropy supervision.
- Optimizer: 8-bit AdamW, wd 0.01 β encoder LR 2.5e-5, head LR 1e-4, cosine decay + 60-step warmup.
- Effective batch 64 (4 Γ 16 accumulation), bf16 autocast, gradient checkpointing.
- 3 epochs β 103 minutes on a single RTX 5050 Laptop 8 GB.
- Post-training: per-question-type temperature fitted on the hold-out β
choice3.52 (shipped inrl_agent_config.json, applied automatically by the SDK). - Config changes vs the base checkpoint:
max_len1024β768,head_max_len256β512 (so 60 intent options are no longer identically truncated).
Evaluation results
Accuracy on the test sets (evaluated through the SDK inference path with shipped temperatures):
| Test set | n | Baseline | This model | ECE base β FT |
|---|---|---|---|---|
| Intent id (MASSIVE) | 2,000 | 0.410 | 0.862 | 0.213 β 0.064 |
| Intent en (MASSIVE) | 2,000 | 0.479 | 0.857 | 0.203 β 0.056 |
| Intent jv (NLLB-translated) | 2,000 | 0.281 | 0.805 | 0.251 β 0.081 |
| Intent su (NLLB-translated) | 2,000 | 0.252 | 0.772 | 0.211 β 0.080 |
| Sentiment id (NusaX) | 400 | 0.710 | 0.860 | 0.159 β 0.121 |
| Sentiment en (NusaX) | 400 | 0.743 | 0.873 | 0.121 β 0.105 |
| Sentiment jv (NusaX) | 400 | 0.590 | 0.818 | 0.236 β 0.151 |
| Sentiment su (NusaX) | 400 | 0.415 | 0.757 | 0.385 β 0.213 |
| CS id+en (synthetic) | 500 | 0.376 | 0.820 | 0.267 β 0.083 |
| CS id+jv (synthetic) | 500 | 0.350 | 0.816 | 0.277 β 0.071 |
| CS id+su (synthetic) | 500 | 0.342 | 0.852 | 0.234 β 0.065 |
Macro-F1 and full details: HASIL-RETRAINING.en.md in the pipeline repo.
Benchmark vs Jev (OpenRouter) β ticket-domain 12 categories, identical 300 real samples
The latest domain fine-tune of this model family
(laya-idjvsuen-v3) reaches 93.3% accuracy /
ECE 0.015 on the identical samples, vs Jev 1.13 at 84.7% β running locally with zero API cost.
Gold = weak labels; methodology and honest caveats: BENCHMARK.en.md in the
pipeline repo.
Limitations
- Not a generative model β it only answers caller-defined typed questions (choice/score/noul) with caller-defined options.
- jv/su intent was trained and evaluated on machine-translated text (NLLB). Accuracy on authentically human-written jv/su is expected to be lower; the honest human-text numbers are the NusaX rows.
- Code-switching evaluation is synthetic (parallel half-splices), not natural social-media code-switching. No mature public EN-ID CS benchmark exists yet.
- Task coverage is narrow: MASSIVE intent + sentiment. Typed-decision capability on other tasks (e.g. Laya's original benchmarks: Banking77, SST-5, XNLI) was not re-measured after fine-tuning.
- Calibration was fitted on these two tasks' distribution; use
answer_confidencewith care out of distribution.
License & attribution
- These fine-tuned weights are released under CC-BY-SA-4.0 (NusaX data is share-alike).
- Base model convaiinnovations/laya-multilingual: Apache-2.0 Β© ConvAI Innovations; NandhaKishorM/laya SDK: Apache-2.0.
- MASSIVE 1.1 Β© Amazon: CC-BY-4.0. NusaX-senti Β© IndoNLP: CC-BY-SA-4.0.
- β οΈ The jv/su intent subset was produced with facebook/nllb-200-distilled-600M, which is licensed CC-BY-NC-4.0 (non-commercial). For commercial use, replace that subset with translations from a permissively-licensed MT model or human annotation and retrain (the full pipeline is available in the repo).
Citation
@misc{laya-idjvsuen-v1,
title = {laya-idjvsuen-v1: multilingual Laya decision model fine-tuned for Indonesian, Javanese, Sundanese, English, and code-switching},
author = {muhfalihr},
year = {2026},
note = {Fine-tune of convaiinnovations/laya-multilingual (Apache-2.0) on MASSIVE 1.1, NusaX-senti, and NLLB-generated translations},
url = {https://huggingface.co/faall7479/laya-idjvsuen-v1}
}
Model tree for faall7479/laya-idjvsuen-v1
Base model
convaiinnovations/laya-multilingual
