SecondHand Laya (round 2, int8 ONNX)
A fine-tune of convaiinnovations/laya that fills food-assistance form questions from an applicant's saved facts. The SecondHand desktop app runs it on the applicant's own computer with onnxruntime-node. Nothing about the applicant leaves the machine.
It never writes text. For each candidate answer it returns a calibrated probability that the candidate is correct. SecondHand fills an answer only when the best candidate clears a confidence bar and beats "None of these, or the facts don't say"; otherwise the question is left to the applicant.
Files
| File | Bytes | SHA-256 |
|---|---|---|
model.onnx |
3,820,399 | 4fba842d827c73f596b8f17be4a98700d258cfc43365e32898f577b365314fb8 |
model.onnx.data |
421,294,080 | af1f87d7d95ff5c72f205b81c91f414cd53bc4632d41972397b99a43a13cd741 |
tokenizer/tokenizer.json |
3,583,228 | 6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30 |
tokenizer/tokenizer_config.json |
308 | 50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12 |
rl_agent_config.json |
1,019 | 7f4cb9dd484a70cd6352eb1efb804e5ebcc9b346e352215318aff486ced58c36 |
The files are int8 ONNX, about 429 MB in total. They run on the CPU.
Input
Every decision is one Laya noul question with a fixed instruction, the same one the model was trained on:
Given the facts about the household, is the candidate the correct answer to the form question?
- Answering a choice or yes/no question. The state is
{ "facts": …, "question": …, "candidate": … }:factsis a short list of plain sentences built by code from the saved profile, for example "The applicant is 41 years old. The household has 3 people. The applicant lives in Iowa." All arithmetic happens in that code, never in the model.- Each option is scored in its own pass, plus the candidate "None of these, or the facts don’t say".
- Matching a text box to a saved field. The state is
{ "question": …, "candidate": "Saved answer: <field description>" }. There are no facts, and the saved value itself is never given to the model.
The probability is answers.correct.noul, after the per-bucket temperature calibration in rl_agent_config.json.
Training
- Method: LoRA (rank 16, alpha 32, dropout 0.05, top 4 layers trained fully), with the proper-scoring objective and balanced class weighting.
- Schedule: 1 epoch of 4,217 updates (batch 8, gradient accumulation 2), learning rate 2e-4, trained in bfloat16 on an Apple M4 Max with LayaStudio.
- Validation: loss 0.027, accuracy 0.993.
- Calibration: expected calibration error 0.0045 before and 0.0013 after temperature fitting.
- Data: 780 questions copied from 39 public food-assistance forms (Google Forms, Jotform, PDF and web intake forms), plus 757 synthetic rewordings used only for training. Every label is computed by code from each question's answer rule and fictional households. No real person's data was used.
- Splits: by form. 7 test forms and 7 held-out forms collected after the first model are never trained on. A test prevents synthetic questions from repeating a test or held-out label.
Results (confidence bar 0.9)
| Task | Set | Precision | Coverage | Wrong fills |
|---|---|---|---|---|
| Answering | 2,016 test questions | 0.749 raw | 0.780 | 43 raw, 0 genuinely wrong |
| Answering | 368 held-out questions | 0.565 raw | 0.765 | 10 raw, 0 genuinely wrong |
| Matching (at 0.95) | 169 test boxes | 0.909 | 0.854 | 7 |
| Matching (at 0.95) | 78 held-out boxes | 0.902 | 0.787 | 4 |
Every "wrong" answering fill is a correct answer, by the household's facts, to one of three questions the frozen answer key can't express:
- "Are there other members in your household in addition to yourself?"
- "Does your household have more than 8 members?"
- An "Apply for?" row, when the applicant is applying for SNAP.
Limitations
- Boxes that belong to someone else. A family member's "Name" or "Date of Birth" in a repeated section can still match the applicant's fields, because the model sees the label but not its section.
- Combined boxes. Boxes like "City and Zip Code" match one of their parts.
- Precision. Matching precision is below 0.95.
- Many options. Temperatures for questions with 11 or more options are clamped, so their confidence is uncalibrated.
- Speed. Each candidate is a separate pass, about 26 ms for a text-box candidate and 80–120 ms for a choice candidate on an M4 Max CPU, so long pages take seconds.
- Language. English only.
License
Apache-2.0, as is the base model convaiinnovations/laya. This model is a fine-tune of it.
Model tree for JacobTDang/secondhand-laya
Base model
convaiinnovations/laya