laya-pdpa-my-v4

Ten typed questions about a Malaysian call transcript, answered in one forward pass: eight yes/no PDPA classes, one ordered severity, one choice for who disclosed the data. Replaces a single fuzzy pdpa yes/no question that bundled NRIC, card numbers, OTPs, addresses and health data into one boolean.

Results

96 held-out rows (32 conversation skeletons Γ— 3 fabricated identities), split fingerprint b047809f6f417e10. Every figure is quoted beside the majority-class baseline, because a class whose baseline is 84% does not need a model to reach 84%.

question majority v3 v4 gain 95% CI
pdpa_severity 40.6 61.5 75.0 +34.4 65.5 – 82.6
pdpa_actor 34.4 59.4 63.5 +29.1 53.6 – 72.5
pdpa_name 68.8 90.6 96.9 +28.1 91.2 – 98.9
pdpa_address 81.2 99.0 100.0 +18.8 96.2 – 100.0
pdpa_phone 84.4 96.9 95.8 +11.4 89.8 – 98.4
pdpa_health 81.2 85.4 89.6 +8.4 81.9 – 94.2
pdpa_email 84.4 87.5 88.5 +4.1 80.6 – 93.5
pdpa_mykad 84.4 87.5 88.5 +4.1 80.6 – 93.5
pdpa_credential_otp 84.4 85.4 87.5 +3.1 79.4 – 92.7
pdpa_bank_card 84.4 84.4 80.2 βˆ’4.2 71.1 – 86.9
overall 72.8 83.8 86.5 +13.7

By language: bm 86.7% Β· zh 86.1% Β· en 79.5% Β· ta 79.0% Β· manglish 76.1%

Read the intervals before the gains. Six questions clear their baseline by more than the interval width. Three β€” pdpa_email, pdpa_mykad, pdpa_credential_otp β€” have intervals containing their own baseline, so at this sample size they are not distinguishable from answering "no" every time. pdpa_bank_card is below its baseline.

What this model is actually for

A page of regexes beats it on three classes: email 97.9% vs 88.5%, credential_otp 90.6% vs 87.5%, bank_card 82.3% vs 80.2%. Those are exactly the three whose intervals contain their baseline. If you only need to find fixed-format identifiers, use patterns β€” they are free, deterministic and better here.

What a pattern cannot express is the other two: pdpa_severity (+34.4) for ranking a queue, and pdpa_actor (+29.1) for telling an agent soliciting an NRIC from a caller volunteering one unprompted. Those are different compliance events, and that distinction is the reason this model exists.

The sensible deployment is a cascade: regex for fixed formats, this model for the contextual judgements, an LLM for whatever is left and only if your data policy permits sending transcripts off-site.

Why not prompt an LLM instead

On raw accuracy you probably should. One argument survives, and for this job it decides: the transcript you want scored is the one containing the MyKad number and the spoken OTP. Sending it to a third-party API is itself a transfer you would have to justify under the law you are trying to comply with. This runs locally, offline, and in a browser tab.

Training data is entirely synthetic

You cannot collect a corpus of real PDPA breaches without creating the exposure you are trying to detect. An LLM wrote conversation skeletons containing only {{SLOT}} placeholders; local deterministic code substituted fabricated Malaysian identities. The LLM never saw a MyKad number because there was none to see, and labels are a by-product of substitution rather than a model's judgement.

No real call has ever been labelled. Every number above describes how well the model handles text our own generator wrote. That is the load-bearing limitation of this work and no amount of re-splitting fixes it.

Known corrections

Published figures for this checkpoint were wrong twice, both times in the evaluation rather than the model.

pdpa_severity was reported at 37.3%, then 22.9%, and called "below chance". It scores 75.0%. laya returns an ordered score as a continuous expectation, so a confident "high" arrives as score: 2.84 with 87% of the mass on class 3, and the evaluation did int(ans["score"]) β€” truncating to 2, "medium". Wrong on 56 of 96 rows, all in the same direction. Severity is this model's second-best question.

v4 was reported as a regression against v3. It is not: 86.5% against 83.8% on the same rows, better on eight of ten questions. The earlier comparison read v4's score on one test split beside v3's on another.

Calibration

Temperatures [1.032, 1.025, 1.100] (choice, score, noul), fitted on validation. v3 shipped [10.0, 10.0, 7.916], which laya clamps to 5 on load with a warning β€” its confidences are unusable, though its accuracy is unaffected because temperature does not move argmax. If you need calibrated probabilities rather than labels, use v4.

Browser build

Suarify/laya-pdpa-onnx-v4-32k β€” vocabulary pruned 256k β†’ 32k and block-quantised to 4-bit, 189 MB, runs in a tab under onnxruntime-web with WebGPU. Nothing leaves the page.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Suarify/laya-pdpa-my-v4

Finetuned
(68)
this model
Quantizations
1 model