XERON-0.1 ๐ฏ
XERON-0.1 is a fine-tuned typed-decision (System 1) model built on convaiinnovations/laya's
multilingual checkpoint (backbone: jhu-clsp/mmBERT-base, 322M params).
It answers typed questions over a state โ choice / score / noul (boolean) โ in a single forward pass,
returning calibrated probabilities. It never generates text, so it cannot hallucinate and cannot emit malformed schemas.
Trained by PIXELZX as the first release of the XERON family (training plan & data pipeline: https://github.com/PIXELZX0/XERON).
โจ What makes XERON-0.1 different
| Base | jhu-clsp/mmBERT-base (Laya multilingual subfolder) |
| Fine-tuned on | 148,281 sequences โ English typed-decisions + Korean (KLUE ynat/nli/sts) + browser/web decisions (BBC + AG News, SMS spam, phishing URLs) + web-agent actions (Mind2Web) + long-context (SCOTUS, 20 Newsgroups) |
| Training | RLCD โ strictly proper scoring rule + GRPO-style baseline; post-hoc temperature calibration. 4 epochs ยท 15,987 updates ยท ~5h on 1รGPU (bf16) |
| Context | 4,096 tokens trained (CTX_CAP up to 32,768 via RoPE extension) |
| Languages | 100+ via mmBERT (English + Korean trained explicitly) |
| License | Apache-2.0 |
๐ Head-to-head: Laya vs Jev vs XERON-0.1
A. ์ค์ธก ๋น๊ต (๋์ผ ํ๊ฐ์ ยท ๋์ผ ์คํฌ๋ฆฝํธ)
ํ๊ฐ์
LocalLLaMA/typed-decisions test split โ 400 cases = 600 choice + 400 score + 400 noul (1,400 decisions).
๋์ผํ scripts/evaluate.py / scripts/recompute_metrics.py, CPU ๋จ์ผ ํ๋ก์ธ์ค, 2026-09-22. ์๋ณธ: eval_comparison.md, eval_baseline_*.json.
| ๋ชจ๋ธ | ๋ฐฑ๋ณธ / ํ๋ผ๋ฏธํฐ | choice acc | soft acc | Brier | ECE โ | score MAE โ | noul acc |
|---|---|---|---|---|---|---|---|
| XERON-0.1 (ours) | mmBERT-base ยท 322M ยท fine-tuned | 0.7000 | 0.5171 | 0.4493 | 0.2143 | 0.4213 | 0.7817 |
laya-typed-decisions (Convai) |
ModernBERT-large ยท 421M ยท ๋ฒค๋ ํ๋ | 0.7333 | 0.4460 | 0.4669 | 0.2380 | 0.2963 | 0.8583 |
laya-multilingual (Convai) |
mmBERT-base ยท 322M ยท base(๋ฏธํ๋) | 0.2900 | 0.2758 | 0.9952 | 0.3282 | 1.0625 | 0.4983 |
| Jev (TypeSafe AI) | ๋น๊ณต๊ฐ ยท ํธ์คํ API | ์ธก์ ๋ถ๊ฐ (์จ์ดํธ ๋น๊ณต๊ฐ) | โ | โ | โ | โ | โ |
- ๋ฏธํ๋ Laya base(0.290) โ XERON-0.1(0.700) = +41.0%p. ํ์ธํ๋์ด ๊ฒฐ์ ์ (Laya base๋ ํ์คํฌ๋ณ FT ์ ์ ๋ผ๋ ๊ณต๊ฐ ์ค๋ช ๊ณผ ์ผ์น).
- ๋ฒค๋๊ฐ ์ด ๋ฒค์น๋งํฌ์ ์ง์ ํ๋ํ 421M ์์ด ๋ชจ๋ธ ๋๋น ํ๋ ๋ผ๋ฒจ -3.3%p, ๊ทธ๋ฌ๋ soft accuracy +7.1%p / ECE โ0.024 ๋ XERON ์ฐ์ โ ํ๋ฅ (์ ๋ขฐ๋) ํ์ง์ด ๋ ์ข์.
B. ์คํ ยท ๊ณต๊ฐ ์งํ ๋น๊ต
Laya/Jev ์์น๋ ๋ฒค๋ยท์ 3์ ๊ณต๊ฐ ์๋ฃ์์ ์ธ์ฉ (2026-09-22 ์กฐํ). XERON ์์น๋ ์ค์ธก.
| ํญ๋ชฉ | XERON-0.1 | Laya (Convai) | Jev (TypeSafe AI) |
|---|---|---|---|
| ๋ฐฐํฌ ํํ | ์คํ ์จ์ดํธ (self-host) | ์คํ ์จ์ดํธ (self-host) | ํธ์คํ API (closed) |
| ๋ผ์ด์ ์ค | Apache-2.0 | Apache-2.0 | ์์ฉ API |
| ๋ฐฑ๋ณธ / ํ๋ผ๋ฏธํฐ | mmBERT-base ยท 322M | EN 421M / multi 322M | ๋น๊ณต๊ฐ |
| ์ปจํ ์คํธ | 4,096 (cap 32,768) | 512 (EN) / 1,024 (multi) | 64k |
| ์ธ์ด | 100+ (mmBERT), ENยทKR ํ์ต | 100+ (multi ์ฒดํฌํฌ์ธํธ) | ๋ค๊ตญ์ด (๋น๊ณต๊ฐ) |
| ๊ฒฐ์ ํ๋ฆฌ๋ฏธํฐ๋ธ | choice / score / noul | choice / score / noul | choice / score / noul |
| ํ์ธํ๋ ํ์ | โ (์ด๋ฏธ ํ๋๋จ) | โ (base๋ FT ์ ์ ) | โ (์ ๋ก์ท) |
| ์ ํ๋ (๊ณต๊ฐ) | โ | in-task 0.753 ยท zero-shot 0.651 ยท typed-decisions 0.766 | JevBench v1.3.0 composite 74.4 (#1) |
| ๋์ด๋๋ณ (Easy/Std/Judge/Hard) | โ | 94.4 / 72.9 / 69.2 / 34.1 % | 100 / 99 / 94.5 / 74.1 % |
| ์บ๋ฆฌ๋ธ๋ ์ด์ (ECE) | 0.2143 (๋์ผ ํ๊ฐ์ ) | in-task 0.030 ยท zero-shot 0.204 ยท multi 0.081 (๋ฒค๋) | ๋ฏธ๊ณต๊ฐ |
| Brier | 0.4493 (๋์ผ ํ๊ฐ์ ) | 0.308 in-task / 0.532 zero-shot (๋ฒค๋) | ๋ฏธ๊ณต๊ฐ |
| ์ง์ฐ | 2.2 s/case (CPU, ๋ก์ปฌ ์ธก์ ) | p50 38.4 ms (1๋ฌธํญ, ๋ฒค๋) ยท 32.8 ms (T4) | ~150 ms (์ 3์ 236โ276 ms) |
| ๋น์ฉ | self-host (GPU ๋น์ฉ๋ง) | self-host (GPU ๋น์ฉ๋ง) | $0.042 / 1M input tokens |
| JevBench v1.3.0 ์์ | ๋ฏธ๋ฑ์ฌ | #33 (54.4) | #1 (74.4) |
โ ๏ธ A์ B๋ ๋ค๋ฅธ ํ๊ฐ์ ยท๋ค๋ฅธ ํ๋์จ์ด์ ๋๋ค. A๋ ๋์ผ ์กฐ๊ฑด ์ค์ธก(๋ชจ๋ธ ๊ฐ ์ง์ ๋น๊ต ๊ฐ๋ฅ), B๋ ๊ณต๊ฐ ์๋ฃ ์ธ์ฉ(์ฐธ๊ณ ์ฉ). Jev๋ ์จ์ดํธ๊ฐ ๊ณต๊ฐ๋์ง ์์ A์ ๋ฃ์ ์ ์์ต๋๋ค.
C. JevBench v1.3 ๋ฐฉ์ ํ๊ฐ (๊ณต๊ฐ ์์ดํ 231/534)
JevBench v1.3.0(Benchmark Heaven) ํ๋ค์คยท์ด๋ํฐยท์ฑ์ ์ฝ๋๋ฅผ ๊ทธ๋๋ก ์ฌ์ฉํด ์ธก์ ํ์ต๋๋ค. ๊ณต๊ฐ ์์ดํ ์ 534๊ฐ ์ค 231๊ฐ๋ฟ(judge tier๋ ์ ๋ถ ๋น๊ณต๊ฐ)์ด๋ฏ๋ก ๊ณต์ ๋ณด๋ ์์์ ์ง์ ๋น๊ตํ ์ ์์ต๋๋ค. ์ ์ฒด ๋ฐฉ๋ฒยทfamily๋ณ ๋ถ์ยท์ฌํ ์ฝ๋: https://github.com/PIXELZX0/XERON/tree/main/results/jevbench-public
๋์ผํ 231๊ฐ ๊ณต๊ฐ ์์ดํ ์์ 4๊ฐ ์์คํ ์ ๋น๊ต (Jev/Laya๋ ๊ณต์ per-task ์ํฐํฉํธ์ ๊ณต๊ฐ ์์ดํ ๊ฒฐ๊ณผ):
| ์์คํ | easy (48) | standard (72) | hard (111) | ์ ์ฒด (231) | Intelligence |
|---|---|---|---|---|---|
| Jev 1.13.0 (TypeSafe, API) | 1.000 | 0.986 | 0.730 | 0.866 | 82.2 |
| laya-typed-decisions (Convai, 421M, ๋ฒค๋ ํ๋) | 0.979 | 0.653 | 0.270 | 0.537 | 38.0 |
| XERON-0.1 (ours) | 0.875 | 0.444 | 0.306 | 0.468 | 23.3 |
| laya-multilingual (๋ฏธํ๋ base) | 0.896 | 0.403 | 0.324 | 0.468 | 21.5 |
์์งํ ๊ฒฐ๋ก
- XERON-0.1์ JevBench๋ฅ ํ์คํฌ(์ ์ฑ ๋ฌธ์ + ๋ฃจ๋ธ๋ฆญ ์ ํ)์์ ๋ฒค๋ Laya ํ๋ํ๋ณด๋ค ์ฝํฉ๋๋ค. ํ์ต ๋ฐ์ดํฐ(KLUEยท๋ธ๋ผ์ฐ์ ยทMind2WebยทSCOTUS)์ ๋๋ฉ์ธ์ด ๊ฒน์น์ง ์๊ธฐ ๋๋ฌธ์ ๋๋ค.
- ๋จ, ๋์ผ ๋ฐฑ๋ณธ ๋ฏธํ๋ base ๋๋น ํ์ธํ๋ ํจ๊ณผ๋ ๊ฒ์ฆ๋์ต๋๋ค (ECE hard 0.289 โ 0.211, ํ๋ฅ TVD 0.573 โ 0.389).
- hard tier์์๋ XERON์ด ๋ฒค๋ ํ๋ํ์ ์ด๊น๋๋ค (0.306 vs 0.270) โ adversarial / trap / judge_hard ๊ณ์ด.
- ๋ถ๋ถ JevBench Score(judge tier ์์, ์ฌ์ ๊ทํ): laya-typed-decisions 37.5 ยท XERON-0.1 11.5 ยท laya-multilingual 8.8. Intelligence<50 ํจ๋ํฐ
(I/50)ยฒ๊ฐ XERON ์ ์๋ฅผ ํฌ๊ฒ ๊น์ต๋๋ค. - ๋ค์ ๋จ๊ณ: JevBench ์คํ์ผ ๋ฐ์ดํฐ(๊ณต๊ฐ 231๊ฑด ๋๋ HF
Praveenrajus/jev-bench166k rows)๋ก ์ถ๊ฐ ํ์ธํ๋ ํ ์ฌํ๊ฐ.
์ฝ๋ ๋ฒ
- Jev = ์ ๋ก์ท ๊ฐ์ (ํ๋ ๋์ด๋ 74.1%).
- Laya / XERON = ์คํ ์จ์ดํธ, ํ์ธํ๋ ํ ์๊ธฐ ๋๋ฉ์ธ์์ ๊ฐํด์ง๋ ์ง์.
- XERON-0.1์ Laya multilingual๊ณผ ๊ฐ์ ๋ฐฑ๋ณธ + ํ๊ตญ์ดยท๋ธ๋ผ์ฐ์ ยท์น ์์ด์ ํธยท์ฅ๋ฌธ ๋ฐ์ดํฐ ์ถ๊ฐ ํ์ตํ.
๐ ์ฌ์ฉ๋ฒ
pip install laya
import laya
agent = laya.load("PIXELZX/XERON-0.1")
state = "Customer email: 'I was charged twice, please refund immediately.' tier=premium, sla=4h"
questions = {
"intent": {"type": "choice", "options": ["billing", "technical", "cancellation", "other"]},
"urgency": {"type": "score", "levels": ["0 โ no time pressure", "1 โ routine", "2 โ elevated", "3 โ critical"]},
"needs_refund": {"type": "noul"},
}
res = agent.predict(state, questions)
print(res["answers"])
Returned per question: selected key / expected level + full probability distribution + calibrated confidence.
๐ ํ๊ฐ ์ฌํ
python scripts/evaluate.py \
--model PIXELZX/XERON-0.1 \
--dataset LocalLLaMA/typed-decisions --split test \
--device cpu --output eval_results.json
python scripts/recompute_metrics.py eval_results.json # score MAE / ํ์
๋ณ ์ ํ๋
eval_results.json (per-decision predictions ํฌํจ)์ด ์ด ์ ์ฅ์์ ํฌํจ๋์ด ์์ต๋๋ค.
โ ๏ธ ํ๊ณ
- 4,096 ํ ํฐ ํ์ต โ ์ด์ฅ๋ฌธ์
CTX_CAPํ์ฅ ํ ์ฌํ์ต ํ์. - ํ๊ฐ๋ ์์ด typed-decisions split ์ค์ฌ. ํ๊ตญ์ด/๋ธ๋ผ์ฐ์ /์น ์์ด์ ํธ ํธ๋์ ๋ณ๋ ๋ฒค์น๋งํฌ ๋ฏธ๊ณต๊ฐ.
scoreํ์ ์ ์์ํ(ordinal) rubric ์ ์ฉ.
Citation
@misc{xeron01,
title = {XERON-0.1: a fine-tuned multilingual typed-decision model},
author = {PIXELZX},
year = {2026},
url = {https://huggingface.co/PIXELZX/XERON-0.1}
}
Built on Laya by Convai Innovations (Apache-2.0) and jhu-clsp/mmBERT-base.
- Downloads last month
- -