A small, fast decision model for Vietnamese · 307M parameters · runs on CPU · v0.1-beta
VIYA reads Vietnamese text and returns typed decisions with calibrated confidence in a single forward pass. It does not generate text: you describe the question and the options in plain language, and VIYA scores every option.
| Type | Returns | Example |
|---|---|---|
choice |
one of N options | Does the evidence support, refute, or not decide this claim? |
score |
a level on a scale | How many stars is this review? |
noul |
yes / no | Is this sentence a factual claim worth checking? |
Options are free text, so you can add or rename them without retraining.
Highlights
- Vietnamese fact-checking. 87.6 macro-F1 on ViFactCheck with gold evidence (human: 84.9). When it has to find the evidence itself in the full article, it scores 78.7, ahead of Gemini 1.5 Flash, XLM-R large and Mistral 7B.
- Full fact-checking pipeline. On ViWikiFC, VIYA finds the right evidence sentence and reaches the right verdict 75.7% of the time; the best published pipeline reaches 67.0%.
- Reads Vietnamese as people type it. Intent accuracy is 89.4 with diacritics and 88.7 without (
ko,dc, no tone marks). - One model, many tasks. The same weights handle fact-checking, NLI, intent, sentiment, moderation and English business decisions (75.9 on typed-decisions).
- Knows how sure it is. Confidence is calibrated per question type (ECE 0.03 to 0.05 on the main Vietnamese and fact-checking suites).
Quick start
pip install laya
import laya
viya = laya.load("hiepho/viya-base")
out = viya.predict(
{
"claim": "SAWACO ngưng cấp nước từ 12 giờ ngày 25-3.",
"evidence": "Thời gian thực hiện dự kiến từ 22 giờ ngày 25-3 đến 4 giờ ngày 26-3.",
},
{
"verdict": {
"type": "choice",
"instructions": "Đối chiếu `claim` với `evidence`: bằng chứng nói gì?",
"criteria": {
"supports": "bằng chứng xác nhận khẳng định là đúng",
"refutes": "bằng chứng cho thấy khẳng định là sai",
"nei": "bằng chứng không đủ để kết luận đúng hay sai",
},
},
},
)
print(out["answers"]["verdict"])
v0.1 runs on the open-source laya inference runtime. A native viya package will ship with the v0.2 architecture.
How it works
[CLS] question [SEP] [MASK] option 1 [MASK] option 2 … [SEP] your text [SEP]
│ │
└── each [MASK] is scored ──► softmax ──► calibrated probabilities
- Encoder:
jhu-clsp/mmBERT-base(22 layers, 8,192-token context, 1,800+ languages). - Decision head (v0.1): two transformer layers and a scorer over option markers, adopted from an open typed-decision design (see Credits). Input length is 1,024 tokens.
- Calibration: per-type temperature fitted on held-out data.
Training
- Vietnamese reading. 11 Vietnamese datasets turned into 18 kinds of decision questions. Each sentence is paired with its no-diacritics or teencode version, with a consistency loss so both get the same answer. Option names are randomly renamed so the model has to read the option descriptions.
- Continued pre-training. 7.8 hours of masked-language modelling on Vietnamese web, forum and Wikipedia text (FineWeb-2, tinhte, Wikipedia).
- Fact-checking. ViFactCheck, ViWikiFC, VitaminC, ClaimBuster and ReINTEL, covering three skills: which sentence needs checking, which sentence is evidence, and the verdict.
- Decision and anti-forgetting round. typed-decisions plus English spam, phishing, passage relevance and ticket data, with replay of the earlier stages.
Roadmap: VIYA's own architecture (v0.2)
- Read once. Claim and whole article in one sequence, with a marker before every sentence. An evidence head and a verdict head are trained jointly, so finding evidence and judging it happen in one pass instead of one pass per sentence.
- Self-confidence ("am I sure?"). A head trained on the model's own past mistakes. VIYA skims by default and switches to a careful sentence-by-sentence read only when it is unsure.
Evaluation
All VIYA numbers are measured on the official test sets (October 2026). Other models' numbers come from their papers, where each model is fine-tuned separately per task; VIYA is one model for all tasks.
Vietnamese fact-checking
| Model | Params | ViFactCheck, gold evidence (macro-F1) | ViFactCheck, full article (macro-F1) |
|---|---|---|---|
| Gemma 7B | 7B | 89.90 | 85.94 |
| Llama3 8B | 8B | 88.67 | 79.65 |
| Mistral 7B | 7B | 88.63 | 70.13 |
| XLM-R large | 550M | 88.02 | 75.42 |
| Gemini 1.5 Flash (prompting) | — | 74.88 | 76.26 |
| PhoBERT large | 370M | 79.76 | 62.93 |
| Human | 84.93 | ||
| VIYA v0.1-beta | 307M | 87.64 | 78.65 |
| ViWikiFC | InfoXLM large | VIYA v0.1-beta |
|---|---|---|
| Verdict with gold evidence (F1) | 86.51 | 85.20 |
| Full pipeline: right evidence and right verdict (strict accuracy) | 67.00 (with BM25) | 75.66 |
Vietnamese understanding
| Task | mmBERT starting checkpoint | VIYA v0.1-beta |
|---|---|---|
| MASSIVE-vi intent (20 options) | 36.9 | 89.4 |
| … typed without diacritics | 13.1 | 88.7 |
| ViNLI (accuracy) | 48.9 | 81.8 |
| XNLI-vi | 74.1 | 77.5 |
| ViOCD complaint, never trained (AUROC) | 65.8 | 90.6 |
| UIT-VSMEC emotion, never trained (accuracy) | 40.6 | 50.8 |
ViNLI reference points: CafeBERT 86.1, XLM-R large 86.0, PhoBERT large 80.7.
English decisions (typed-decisions, 2,000 decisions)
| Model | Accuracy | ECE |
|---|---|---|
| Laya fine-tune | 76.6 | 0.213 |
| VIYA v0.1-beta | 75.9 | 0.137 |
| TypeSafe Jev 1.13 | 72.7 | 0.144 |
Limitations
- Text only. Images, video and deepfakes are out of scope.
- Facts come from evidence, not memory. VIYA judges a claim against the evidence you give it. For breaking news with no published sources, the right answer is "not enough information".
- Many options. Choice questions work best with fewer than about 20 options; 77-way intent (Banking77) reaches only 53.8.
- Beta. A held-out "recent news" test and the v0.2 single-pass architecture are still in progress.
License and data
The model weights are released under Apache-2.0. Some training datasets (many UIT-NLP corpora, for example) are licensed for research use only; check each dataset's license before commercial use.
Credits
- mmBERT (JHU CLSP), the encoder.
- Laya (Convai Innovations, Apache-2.0): the v0.1 decision head and starting checkpoint (
laya-multilingual), and the typed-decisions benchmark harness. - The authors of ViFactCheck, ViWikiFC, ViNLI, UIT-ViHSD, UIT-VSFC, UIT-ViSFD, UIT-ViQuAD 2.0, MASSIVE, XNLI, VitaminC, ClaimBuster, ReINTEL and FineWeb-2.
Tiếng Việt
VIYA là model AI nhỏ (307 triệu tham số, chạy trên CPU) ra quyết định trên văn bản tiếng Việt: chọn một trong các lựa chọn, chấm mức độ, hoặc trả lời có/không, kèm độ tin cậy đã hiệu chỉnh. Model không viết văn; bạn mô tả câu hỏi và các lựa chọn bằng lời, VIYA chấm điểm từng lựa chọn.
Điểm nổi bật
- Kiểm chứng tin: 87,6 điểm F1 trên ViFactCheck khi có sẵn bằng chứng (con người: 84,9).
- Trọn quy trình ViWikiFC: 75,7%, so với 67,0% của hệ thống tốt nhất đã công bố.
- Đọc được tiếng Việt gõ không dấu: 88,7 so với 89,4 khi có dấu.
Phiên bản v0.1 dựng trên encoder mmBERT với đầu ra quyết định từ mã nguồn mở. Phiên bản v0.2 sẽ dùng kiến trúc riêng của VIYA:
- đọc cả bài trong một lượt, vừa tìm bằng chứng vừa kết luận;
- tự đánh giá "tôi có chắc không?" để quyết định khi nào cần đọc kỹ.
Model tree for hiepho/viya-base
Base model
jhu-clsp/mmBERT-base