kd-4b-ko-v0 β€” Korean typed-decision LoRA adapters on Qwen/Qwen3.5-4B-Base (weights only)

No single default adapter yet β€” choose by task (see the notes under the evaluation table). kotdaihubv2num-s1 is the broadest-coverage arm, not a validated default.

LoRA r16 + pointer-head adapters in the Kev typed-decision format (choice / noul / score questions over a text state, non-autoregressive probability readout). Trained from the Kev-4B recipe reproduction on Qwen/Qwen3.5-4B-Base (Apache-2.0) with Korean typed-decision data built from public Korean benchmarks (KoBEST, KLUE β€” CC-BY-SA-4.0; used for training only, not redistributed here), a small ThakiCloud synthetic policy set, and β€” for the aihub-* / kotdaihub-* adapters β€” Korean public-sector datasets from AI Hub (ν•œκ΅­μ§€λŠ₯μ •λ³΄μ‚¬νšŒμ§„ν₯원 사업결과): ν–‰μ • λ¬Έμ„œ λŒ€μƒ 기계독해 데이터 (569), κΈˆμœ΅Β·λ²•λ₯  λ¬Έμ„œ 기계독해 데이터 (71610), 법λ₯ /κ·œμ • ν…μŠ€νŠΈ 뢄석 데이터 고도화 (71723), 법λ₯ /κ·œμ •(νŒκ²°μ„œ, μ•½κ΄€ λ“±) ν…μŠ€νŠΈ 뢄석 데이터 (580). AI Hub data was used for model training only under its usage policy; no AI Hub data or derived text is redistributed here. Weights only: no training or evaluation data is published in this repository.

Evaluation (internal, machine-reviewed; one question β‰ˆ 1.75pp on realdoc_v1, β‰ˆ 0.31pp on the sealed realdoc_v2)

arm aihub_dev_acc aihub_dev_ece ko_reviewed_acc ko_reviewed_ece kotd_dev_acc kotd_dev_ece kotd_dev_note numeric_dev_acc numeric_dev_ece numeric_dev_in_acc numeric_dev_in_ece p50_ms realdoc_v1_ci95_pp realdoc_v1_consensus_acc realdoc_v1_coverage_adjusted realdoc_v1_delta_vs_init_pp realdoc_v2_ci95_pp realdoc_v2_consensus_acc realdoc_v2_delta_vs_init_pp transfer_v4_acc transfer_v4_ece
init 0.807 0.134 0.847 0.082 0.733 0.141 0.911 0.895 0.801
kotd-s1 0.839 0.103 0.880 0.071 [0.0, 9.09] 0.946 0.930 3.570 [0.62, 6.29] 0.835 3.420 0.814 0.119
kotdsyn-s1 0.935 0.049 0.878 0.081 [0.0, 9.09] 0.946 0.930 3.570 0.805 0.141
aihubv2-s1 0.955 0.033 0.815 0.116 0.741 0.188 [-9.09, 7.55] 0.911 0.895 0.000 0.812 0.126
kotdaihubv2-s1 0.960 0.029 0.831 0.108 0.873 0.091 [-3.45, 11.11] 0.946 0.930 3.570 0.812 0.140
kotdaihubv2num-s1 0.959 0.028 0.839 0.135 0.879 0.087 0.470 0.371 0.999 0.001 [1.69, 14.81] 0.982 0.965 7.140 [7.67, 14.87] 0.913 11.180 0.808 0.139
jev 0.933 0.818 first 1999 items only 771.900 0.974 0.667
laya 0.476 0.460 691.000 0.375 0.368
qwen27b 0.871 0.895 first 600 items only 522.700 0.804 0.789
  • jev = TypeSafe AI Jev (typesafe-ai/jev) called through Vercel AI Gateway /v1/evaluate on 2026-09-22 (same questions, zero-shot, one run). On the 39 statute questions both APIs accepted, kd-4b (kotd-s1) answered 39/39 and Jev 38/39 β€” a sample that cannot establish superiority or equivalence either way. The gateway rejects score questions with more than 10 rungs, so Jev did not accept 18 of the 57 questions; that is an API-contract difference, not a model-capability difference, and it drives the coverage-adjusted column. Jev's p50 here is under free-tier rate limiting and must not be read as a latency comparison.
  • laya = convaiinnovations/laya-multilingual (mmBERT-base 322M encoder, CPU, no Korean training) and qwen27b = in-house Qwen3.8-27B (NVFP4) prompted zero-shot in the jev-local style (one JSON answer per record, temperature 0) β€” non-trained comparison arms run on the same questions.
  • init = Kev-4B recipe reproduced on Qwen3.5-4B-Base (English suites only); other arms are delta fine-tunes from it (lr 2e-5, batch 4 x accum 2, 2 epochs, seed 1).
  • Gains on kotd_dev / aihub_dev are in-distribution. On the statute test (57 machine-consensus questions, 56 Kev-supported) no adapter measurably changes accuracy: every delta is Β±1–2 questions with a document-clustered CI that includes 0. The test cannot detect effects below ~5pp.
  • Sealed statute test and candidate default (2026-09-23). realdoc_v2 is a sealed test set: 115 verbatim Korean statute excerpts from 76 statutes / 25 domains, 345 questions, 322 scored where three model families (Claude, GPT, in-house Qwen) labeled blind and agreed with an adversarial-challenger veto; every document was checked verbatim against an independently fetched page and against the training set (sentence overlap ≀5%, oracle absent); the benchmark file, the consensus gold, the 322-item mask and the scorer code are pinned by SHA-256 before any model was scored. Three seeds of kotdaihubv2num score 0.913 / 0.919 / 0.935 (mean 0.922) vs base 0.801: +12.1pp, 95% CI [8.8, 15.7]; vs a same-size control trained on 6,000 extra MRC records instead of the contrast pairs: +4.2pp, CI [2.3, 6.3], p=0.0001 (document-clustered paired bootstrap on the seed-averaged correctness; the published seed 1 alone: +11.2 [7.7, 14.9] vs base, +4.7 [2.1, 7.5] vs control). By question type the gap sits in score (base 0.574 β†’ control 0.691 β†’ numeric 0.806, +11.4 [6.2, 17.3] vs control) while choice/noul are β‰₯0.98 for every trained arm and non-inferior (2pp margin) to the control. Promotion to default is on hold: the pre-registered gate also requires OOD non-inferiority within 1pp on ko_reviewed, and with 124 OOD questions the CI ([βˆ’4.7, 3.5] vs base) cannot establish that either way; a larger OOD set is the prerequisite. Until then kotdaihubv2num-s1 is the candidate default and the broadest-coverage arm; kotd-s1 keeps the best English retention, kotdsyn-s1 the best ko_reviewed/ECE on short policy snippets, kotdaihubv2-s1 AI Hub-style document QA. The sealed set is used only for this gate; data design and hyper-parameters are chosen on separate dev sets. A stricter post-seal oracle check flagged 2 of the 115 documents (6 questions) whose key sentence is quoted verbatim inside AI Hub court-case training records; the set was not modified, and excluding them leaves every conclusion unchanged (+12.3 [8.8, 15.9] vs base, +4.3 [2.4, 6.5] vs control).
  • Standard-terms (μ•½κ΄€) finding (2026-09-23): a semantic audit showed the AI Hub 580 judgment labels are not derivable from the excerpt alone (blind agreement 0.78, abstain 0.63), so a 'purity' retrain without that source was tried; across 3 seeds it regressed OOD by βˆ’4.6pp (CI [βˆ’8.6, βˆ’0.8]) and restoring the source recovered +3.2pp (CI [0.9, 6.1]). The source therefore stays in training as judgment-type data and is excluded from dev/test only. The v4 arm is not shipped.
  • No human validated any Korean label; statute-test labels are machine-consensus.
  • Correction (2026-09-22, data audit): the aihub_dev column is inflated by a label leak in one of its six sources. In the AI Hub 580 (μ•½κ΄€ 유뢈리) source every '유리' record carries a standard-clause section and no '뢈리' record does, and the v1 builder copied that section into the state, so its presence alone gives the label (aihub-s1 scored 1.000 on that source vs 0.365 for init). Excluding that source (n=1,297): init 0.877, aihub-s1 0.961, kotdaihub-s1 0.958. The two AI Hub multiple-choice sources also had the gold answer always in option position 0 in both train and dev; Kev permutes option order during training, and a same-run re-measurement with shuffled options (2026-09-22, H100) moved every arm by at most 0.6pp (e.g. init 0.781β†’0.779, aihub-s1 0.969β†’0.971), so no position bias is present and the multiple-choice numbers stand. The v2 builder (shuffled options, no standard-clause section, train/dev overlap removed) passes a deterministic audit (scripts/data_audit.py); aihubv2-s1 and kotdaihubv2-s1 are the v2 retrains (2026-09-23, H100): their kotd_dev/aihub_dev cells are measured on the deduplicated v2 dev sets (1,989 / 1,726 items), so they are not directly comparable to the v1 cells above them; realdoc_v1 and ko_reviewed are the same test sets for every arm.
  • Numeric-derivation adapter (2026-09-23): kotdaihubv2num-s1 = the v2 mix plus 6,000 code-generated contrast pairs whose gold is computed by a program (unit conversion μ›β†”λ§Œμ›, day arithmetic, cap exceedance, reduction denominators; no human or model labels). It is the first arm whose statute-test gain over init has a document-clustered 95% CI excluding zero (+7.1pp, [1.7, 14.8]); an equal-size control arm trained on extra machine-reading records instead gained +1.8pp ([βˆ’3.6, 7.8]). Caveat: the template families were designed after inspecting the statute test's residual errors, so that gain is a diagnostic, not an independent generalization claim; on a held-out numeric family the arm improves fractionβ†’percent questions by +18pp and does not transfer to interest-rate arithmetic. Single seed; no OOD regression (ko_reviewed 0.839, English retention 0.808).
  • Seed replication (2026-09-22): kotd and kotdaihub (v1 data) were each retrained with seeds 2 and 3. kotd: statute-test accuracy 0.9464 in all three seeds, kotd_dev 0.878Β±0.002, ko_reviewed 0.836Β±0.005, transfer_v4 0.808Β±0.005. kotdaihub: ko_reviewed 0.807/0.823/0.839 (mean 0.823, sd 0.016) vs init 0.847 β€” the βˆ’4pp size was seed-1 specific, but all three seeds sit below init (3-seed paired mean βˆ’2.4pp, document-clustered 95% CI [βˆ’6.2, +1.1]); the direction remains, the size is not established. Only the seed-1 adapters are shipped.

Test descriptions: {"kotd_dev": "2,000 held-out validation items from KoBEST (boolq/copa/hellaswag/wic/sentineg) + KLUE (nli/sts/ynat), converted to typed decisions; in-distribution with training", "realdoc_v1": "57 machine-consensus questions over 20 verbatim Korean statute/ordinance excerpts (3 model families unanimous + adversarial challenger); out-of-training-window inference (state 2048); development diagnostic (1 q β‰ˆ 1.8pp)", "realdoc_v2": "SEALED statute test (2026-09-23): 322 machine-consensus questions over 115 verbatim excerpts from 76 Korean statutes / 25 domains, none sharing a statute family with realdoc_v1; consensus3 + challenger; guarded verbatim against independently fetched pages and against training-set sentence overlap; 1 q β‰ˆ 0.31pp; scored on the same 2048-token state window", "ko_reviewed": "124 reviewed questions on LLM-generated Korean policy snippets (same generator as the small synthetic training set)", "transfer_v4": "Kev English transfer-v4 suite (retention check)"}

How to use

Load with the Kev codebase (python -m kev.serve --run adapters/<arm>; base = Qwen/Qwen3.5-4B-Base). Each adapter directory carries adapter_config.json, adapter_model.safetensors, head.pt, training_config.json. sha256_manifest.json pins every file.

Limits

Training window 384 state tokens; longer documents are served with the inference window (state 2048 / branch 4096) β€” out-of-training-window. Labels on the statute test are machine-consensus (three model families + adversarial challenger), not human gold.

Not affiliated with or endorsed by TypeSafe AI. "Jev" and "System One" are referenced only for comparative research; this repo uses a typed-decision research format and has NOT been contract-tested against any TypeSafe SDK.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ThakiCloud/kd-4b-ko-v0

Adapter
(61)
this model

Collection including ThakiCloud/kd-4b-ko-v0