Instructions to use ThakiCloud/kd-4b-ko-v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ThakiCloud/kd-4b-ko-v0 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
kd-4b-ko-v0 β Korean typed-decision LoRA adapters on Qwen/Qwen3.5-4B-Base (weights only)
No single default adapter yet β choose by task (see the notes under the evaluation table). kotdaihubv2num-s1 is the broadest-coverage arm, not a validated default.
LoRA r16 + pointer-head adapters in the Kev typed-decision format (choice / noul / score questions over a text state, non-autoregressive
probability readout). Trained from the Kev-4B recipe reproduction on Qwen/Qwen3.5-4B-Base (Apache-2.0) with Korean typed-decision data built from
public Korean benchmarks (KoBEST, KLUE β CC-BY-SA-4.0; used for training only, not redistributed here), a small ThakiCloud synthetic
policy set, and β for the aihub-* / kotdaihub-* adapters β Korean public-sector datasets from AI Hub (νκ΅μ§λ₯μ 보μ¬νμ§ν₯μ μ¬μ
κ²°κ³Ό):
νμ λ¬Έμ λμ κΈ°κ³λ
ν΄ λ°μ΄ν° (569), κΈμ΅Β·λ²λ₯ λ¬Έμ κΈ°κ³λ
ν΄ λ°μ΄ν° (71610), λ²λ₯ /κ·μ ν
μ€νΈ λΆμ λ°μ΄ν° κ³ λν (71723), λ²λ₯ /κ·μ (νκ²°μ, μ½κ΄ λ±)
ν
μ€νΈ λΆμ λ°μ΄ν° (580). AI Hub data was used for model training only under its usage policy; no AI Hub data or derived text is redistributed
here. Weights only: no training or evaluation data is published in this repository.
Evaluation (internal, machine-reviewed; one question β 1.75pp on realdoc_v1, β 0.31pp on the sealed realdoc_v2)
| arm | aihub_dev_acc | aihub_dev_ece | ko_reviewed_acc | ko_reviewed_ece | kotd_dev_acc | kotd_dev_ece | kotd_dev_note | numeric_dev_acc | numeric_dev_ece | numeric_dev_in_acc | numeric_dev_in_ece | p50_ms | realdoc_v1_ci95_pp | realdoc_v1_consensus_acc | realdoc_v1_coverage_adjusted | realdoc_v1_delta_vs_init_pp | realdoc_v2_ci95_pp | realdoc_v2_consensus_acc | realdoc_v2_delta_vs_init_pp | transfer_v4_acc | transfer_v4_ece |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| init | 0.807 | 0.134 | 0.847 | 0.082 | 0.733 | 0.141 | 0.911 | 0.895 | 0.801 | ||||||||||||
| kotd-s1 | 0.839 | 0.103 | 0.880 | 0.071 | [0.0, 9.09] | 0.946 | 0.930 | 3.570 | [0.62, 6.29] | 0.835 | 3.420 | 0.814 | 0.119 | ||||||||
| kotdsyn-s1 | 0.935 | 0.049 | 0.878 | 0.081 | [0.0, 9.09] | 0.946 | 0.930 | 3.570 | 0.805 | 0.141 | |||||||||||
| aihubv2-s1 | 0.955 | 0.033 | 0.815 | 0.116 | 0.741 | 0.188 | [-9.09, 7.55] | 0.911 | 0.895 | 0.000 | 0.812 | 0.126 | |||||||||
| kotdaihubv2-s1 | 0.960 | 0.029 | 0.831 | 0.108 | 0.873 | 0.091 | [-3.45, 11.11] | 0.946 | 0.930 | 3.570 | 0.812 | 0.140 | |||||||||
| kotdaihubv2num-s1 | 0.959 | 0.028 | 0.839 | 0.135 | 0.879 | 0.087 | 0.470 | 0.371 | 0.999 | 0.001 | [1.69, 14.81] | 0.982 | 0.965 | 7.140 | [7.67, 14.87] | 0.913 | 11.180 | 0.808 | 0.139 | ||
| jev | 0.933 | 0.818 | first 1999 items only | 771.900 | 0.974 | 0.667 | |||||||||||||||
| laya | 0.476 | 0.460 | 691.000 | 0.375 | 0.368 | ||||||||||||||||
| qwen27b | 0.871 | 0.895 | first 600 items only | 522.700 | 0.804 | 0.789 |
jev= TypeSafe AI Jev (typesafe-ai/jev) called through Vercel AI Gateway/v1/evaluateon 2026-09-22 (same questions, zero-shot, one run). On the 39 statute questions both APIs accepted, kd-4b (kotd-s1) answered 39/39 and Jev 38/39 β a sample that cannot establish superiority or equivalence either way. The gateway rejectsscorequestions with more than 10 rungs, so Jev did not accept 18 of the 57 questions; that is an API-contract difference, not a model-capability difference, and it drives the coverage-adjusted column. Jev's p50 here is under free-tier rate limiting and must not be read as a latency comparison.laya= convaiinnovations/laya-multilingual (mmBERT-base 322M encoder, CPU, no Korean training) andqwen27b= in-house Qwen3.8-27B (NVFP4) prompted zero-shot in the jev-local style (one JSON answer per record, temperature 0) β non-trained comparison arms run on the same questions.init= Kev-4B recipe reproduced on Qwen3.5-4B-Base (English suites only); other arms are delta fine-tunes from it (lr 2e-5, batch 4 x accum 2, 2 epochs, seed 1).- Gains on kotd_dev / aihub_dev are in-distribution. On the statute test (57 machine-consensus questions, 56 Kev-supported) no adapter measurably changes accuracy: every delta is Β±1β2 questions with a document-clustered CI that includes 0. The test cannot detect effects below ~5pp.
- Sealed statute test and candidate default (2026-09-23).
realdoc_v2is a sealed test set: 115 verbatim Korean statute excerpts from 76 statutes / 25 domains, 345 questions, 322 scored where three model families (Claude, GPT, in-house Qwen) labeled blind and agreed with an adversarial-challenger veto; every document was checked verbatim against an independently fetched page and against the training set (sentence overlap β€5%, oracle absent); the benchmark file, the consensus gold, the 322-item mask and the scorer code are pinned by SHA-256 before any model was scored. Three seeds ofkotdaihubv2numscore 0.913 / 0.919 / 0.935 (mean 0.922) vs base 0.801: +12.1pp, 95% CI [8.8, 15.7]; vs a same-size control trained on 6,000 extra MRC records instead of the contrast pairs: +4.2pp, CI [2.3, 6.3], p=0.0001 (document-clustered paired bootstrap on the seed-averaged correctness; the published seed 1 alone: +11.2 [7.7, 14.9] vs base, +4.7 [2.1, 7.5] vs control). By question type the gap sits inscore(base 0.574 β control 0.691 β numeric 0.806, +11.4 [6.2, 17.3] vs control) whilechoice/noulare β₯0.98 for every trained arm and non-inferior (2pp margin) to the control. Promotion to default is on hold: the pre-registered gate also requires OOD non-inferiority within 1pp onko_reviewed, and with 124 OOD questions the CI ([β4.7, 3.5] vs base) cannot establish that either way; a larger OOD set is the prerequisite. Until thenkotdaihubv2num-s1is the candidate default and the broadest-coverage arm;kotd-s1keeps the best English retention,kotdsyn-s1the best ko_reviewed/ECE on short policy snippets,kotdaihubv2-s1AI Hub-style document QA. The sealed set is used only for this gate; data design and hyper-parameters are chosen on separate dev sets. A stricter post-seal oracle check flagged 2 of the 115 documents (6 questions) whose key sentence is quoted verbatim inside AI Hub court-case training records; the set was not modified, and excluding them leaves every conclusion unchanged (+12.3 [8.8, 15.9] vs base, +4.3 [2.4, 6.5] vs control). - Standard-terms (μ½κ΄) finding (2026-09-23): a semantic audit showed the AI Hub 580 judgment labels are not derivable from the excerpt alone (blind agreement 0.78, abstain 0.63), so a 'purity' retrain without that source was tried; across 3 seeds it regressed OOD by β4.6pp (CI [β8.6, β0.8]) and restoring the source recovered +3.2pp (CI [0.9, 6.1]). The source therefore stays in training as judgment-type data and is excluded from dev/test only. The
v4arm is not shipped. - No human validated any Korean label; statute-test labels are machine-consensus.
- Correction (2026-09-22, data audit): the
aihub_devcolumn is inflated by a label leak in one of its six sources. In the AI Hub 580 (μ½κ΄ μ λΆλ¦¬) source every 'μ 리' record carries a standard-clause section and no 'λΆλ¦¬' record does, and the v1 builder copied that section into the state, so its presence alone gives the label (aihub-s1 scored 1.000 on that source vs 0.365 for init). Excluding that source (n=1,297): init 0.877, aihub-s1 0.961, kotdaihub-s1 0.958. The two AI Hub multiple-choice sources also had the gold answer always in option position 0 in both train and dev; Kev permutes option order during training, and a same-run re-measurement with shuffled options (2026-09-22, H100) moved every arm by at most 0.6pp (e.g. init 0.781β0.779, aihub-s1 0.969β0.971), so no position bias is present and the multiple-choice numbers stand. The v2 builder (shuffled options, no standard-clause section, train/dev overlap removed) passes a deterministic audit (scripts/data_audit.py);aihubv2-s1andkotdaihubv2-s1are the v2 retrains (2026-09-23, H100): theirkotd_dev/aihub_devcells are measured on the deduplicated v2 dev sets (1,989 / 1,726 items), so they are not directly comparable to the v1 cells above them;realdoc_v1andko_reviewedare the same test sets for every arm. - Numeric-derivation adapter (2026-09-23):
kotdaihubv2num-s1= the v2 mix plus 6,000 code-generated contrast pairs whose gold is computed by a program (unit conversion μβλ§μ, day arithmetic, cap exceedance, reduction denominators; no human or model labels). It is the first arm whose statute-test gain overinithas a document-clustered 95% CI excluding zero (+7.1pp, [1.7, 14.8]); an equal-size control arm trained on extra machine-reading records instead gained +1.8pp ([β3.6, 7.8]). Caveat: the template families were designed after inspecting the statute test's residual errors, so that gain is a diagnostic, not an independent generalization claim; on a held-out numeric family the arm improves fractionβpercent questions by +18pp and does not transfer to interest-rate arithmetic. Single seed; no OOD regression (ko_reviewed 0.839, English retention 0.808). - Seed replication (2026-09-22):
kotdandkotdaihub(v1 data) were each retrained with seeds 2 and 3. kotd: statute-test accuracy 0.9464 in all three seeds, kotd_dev 0.878Β±0.002, ko_reviewed 0.836Β±0.005, transfer_v4 0.808Β±0.005. kotdaihub: ko_reviewed 0.807/0.823/0.839 (mean 0.823, sd 0.016) vs init 0.847 β the β4pp size was seed-1 specific, but all three seeds sit below init (3-seed paired mean β2.4pp, document-clustered 95% CI [β6.2, +1.1]); the direction remains, the size is not established. Only the seed-1 adapters are shipped.
Test descriptions: {"kotd_dev": "2,000 held-out validation items from KoBEST (boolq/copa/hellaswag/wic/sentineg) + KLUE (nli/sts/ynat), converted to typed decisions; in-distribution with training", "realdoc_v1": "57 machine-consensus questions over 20 verbatim Korean statute/ordinance excerpts (3 model families unanimous + adversarial challenger); out-of-training-window inference (state 2048); development diagnostic (1 q β 1.8pp)", "realdoc_v2": "SEALED statute test (2026-09-23): 322 machine-consensus questions over 115 verbatim excerpts from 76 Korean statutes / 25 domains, none sharing a statute family with realdoc_v1; consensus3 + challenger; guarded verbatim against independently fetched pages and against training-set sentence overlap; 1 q β 0.31pp; scored on the same 2048-token state window", "ko_reviewed": "124 reviewed questions on LLM-generated Korean policy snippets (same generator as the small synthetic training set)", "transfer_v4": "Kev English transfer-v4 suite (retention check)"}
How to use
Load with the Kev codebase (python -m kev.serve --run adapters/<arm>; base = Qwen/Qwen3.5-4B-Base). Each adapter directory carries
adapter_config.json, adapter_model.safetensors, head.pt, training_config.json. sha256_manifest.json pins every file.
Limits
Training window 384 state tokens; longer documents are served with the inference window (state 2048 / branch 4096) β out-of-training-window. Labels on the statute test are machine-consensus (three model families + adversarial challenger), not human gold.
Not affiliated with or endorsed by TypeSafe AI. "Jev" and "System One" are referenced only for comparative research; this repo uses a typed-decision research format and has NOT been contract-tested against any TypeSafe SDK.
- Downloads last month
- -
Model tree for ThakiCloud/kd-4b-ko-v0
Base model
Qwen/Qwen3.5-4B-Base