laya-pdpa-my-v4
Ten typed questions about a Malaysian call transcript, answered in one forward pass:
eight yes/no PDPA classes, one ordered severity, one choice for who disclosed the data.
Replaces a single fuzzy pdpa yes/no question that bundled NRIC, card numbers, OTPs,
addresses and health data into one boolean.
Results
96 held-out rows (32 conversation skeletons Γ 3 fabricated identities), split
fingerprint b047809f6f417e10. Every figure is quoted beside the majority-class
baseline, because a class whose baseline is 84% does not need a model to reach 84%.
| question | majority | v3 | v4 | gain | 95% CI |
|---|---|---|---|---|---|
pdpa_severity |
40.6 | 61.5 | 75.0 | +34.4 | 65.5 β 82.6 |
pdpa_actor |
34.4 | 59.4 | 63.5 | +29.1 | 53.6 β 72.5 |
pdpa_name |
68.8 | 90.6 | 96.9 | +28.1 | 91.2 β 98.9 |
pdpa_address |
81.2 | 99.0 | 100.0 | +18.8 | 96.2 β 100.0 |
pdpa_phone |
84.4 | 96.9 | 95.8 | +11.4 | 89.8 β 98.4 |
pdpa_health |
81.2 | 85.4 | 89.6 | +8.4 | 81.9 β 94.2 |
pdpa_email |
84.4 | 87.5 | 88.5 | +4.1 | 80.6 β 93.5 |
pdpa_mykad |
84.4 | 87.5 | 88.5 | +4.1 | 80.6 β 93.5 |
pdpa_credential_otp |
84.4 | 85.4 | 87.5 | +3.1 | 79.4 β 92.7 |
pdpa_bank_card |
84.4 | 84.4 | 80.2 | β4.2 | 71.1 β 86.9 |
| overall | 72.8 | 83.8 | 86.5 | +13.7 |
By language: bm 86.7% Β· zh 86.1% Β· en 79.5% Β· ta 79.0% Β· manglish 76.1%
Read the intervals before the gains. Six questions clear their baseline by more
than the interval width. Three β pdpa_email, pdpa_mykad, pdpa_credential_otp β
have intervals containing their own baseline, so at this sample size they are not
distinguishable from answering "no" every time. pdpa_bank_card is below its
baseline.
What this model is actually for
A page of regexes beats it on three classes: email 97.9% vs 88.5%, credential_otp 90.6% vs 87.5%, bank_card 82.3% vs 80.2%. Those are exactly the three whose intervals contain their baseline. If you only need to find fixed-format identifiers, use patterns β they are free, deterministic and better here.
What a pattern cannot express is the other two: pdpa_severity (+34.4) for ranking
a queue, and pdpa_actor (+29.1) for telling an agent soliciting an NRIC from a
caller volunteering one unprompted. Those are different compliance events, and that
distinction is the reason this model exists.
The sensible deployment is a cascade: regex for fixed formats, this model for the contextual judgements, an LLM for whatever is left and only if your data policy permits sending transcripts off-site.
Why not prompt an LLM instead
On raw accuracy you probably should. One argument survives, and for this job it decides: the transcript you want scored is the one containing the MyKad number and the spoken OTP. Sending it to a third-party API is itself a transfer you would have to justify under the law you are trying to comply with. This runs locally, offline, and in a browser tab.
Training data is entirely synthetic
You cannot collect a corpus of real PDPA breaches without creating the exposure you are
trying to detect. An LLM wrote conversation skeletons containing only {{SLOT}}
placeholders; local deterministic code substituted fabricated Malaysian identities. The
LLM never saw a MyKad number because there was none to see, and labels are a by-product
of substitution rather than a model's judgement.
No real call has ever been labelled. Every number above describes how well the model handles text our own generator wrote. That is the load-bearing limitation of this work and no amount of re-splitting fixes it.
Known corrections
Published figures for this checkpoint were wrong twice, both times in the evaluation rather than the model.
pdpa_severity was reported at 37.3%, then 22.9%, and called "below chance". It
scores 75.0%. laya returns an ordered score as a continuous expectation, so a
confident "high" arrives as score: 2.84 with 87% of the mass on class 3, and the
evaluation did int(ans["score"]) β truncating to 2, "medium". Wrong on 56 of 96 rows,
all in the same direction. Severity is this model's second-best question.
v4 was reported as a regression against v3. It is not: 86.5% against 83.8% on the same rows, better on eight of ten questions. The earlier comparison read v4's score on one test split beside v3's on another.
Calibration
Temperatures [1.032, 1.025, 1.100] (choice, score, noul), fitted on validation. v3
shipped [10.0, 10.0, 7.916], which laya clamps to 5 on load with a warning β its
confidences are unusable, though its accuracy is unaffected because temperature does not
move argmax. If you need calibrated probabilities rather than labels, use v4.
Browser build
Suarify/laya-pdpa-onnx-v4-32k β vocabulary pruned 256k β 32k and block-quantised to
4-bit, 189 MB, runs in a tab under onnxruntime-web with WebGPU. Nothing leaves the page.