laya-multilingual · split ONNX for the browser (WebGPU → WASM)
An unmodified ONNX export of the multilingual checkpoint of
convaiinnovations/laya (mmBERT-base encoder + decision head,
322M params, 100+ languages), in the split format that laya-ts
loads with ONNX Runtime Web:
| file | size | runs on |
|---|---|---|
encoder.onnx + encoder.onnx.data |
1.23 GB (fp32) | WebGPU, falls back to WASM |
head.onnx + head.onnx.data |
60 MB | WASM |
tokenizer.json, rl_agent_config.json |
34 MB | — |
Exported at laya commit 6d942c92081fbc139e736bbd9ac0023223c29b7f with
python laya-ts/scripts/export_onnx.py --repo convaiinnovations/laya --subfolder multilingual --out-dir ./model-ml
which checks torch vs ONNX outputs agree within 1e-4 (measured max diff 2.9e-6).
Use it
import { Agent } from "laya-ts";
const agent = await Agent.load("https://huggingface.co/Steven10429/laya-multilingual-webgpu/resolve/main/");
const r = await agent.predict("我打开设置页面应用就闪退。", {
department: { type: "choice", instructions: "Which department should handle this?",
criteria: { billing: "invoices, payments, refunds", technical: "bugs, crashes", other: "everything else" } },
});
Measured in Chrome on an Apple M4 Pro (3 questions, 163 input tokens): WebGPU p50 ≈ 148 ms, forced WASM p50 ≈ 808 ms, identical answers. Works inside a cross-site iframe (e.g. an itch.io embed) — Hugging Face serves these files with CORS.
Caveats (from upstream)
The base checkpoints are close to chance zero-shot on the typed-decisions benchmark; fine-tune on your own decisions for real accuracy. Probabilities are not interchangeable with other System One models — recalibrate thresholds.
License & attribution
Apache-2.0, same as the original weights by ConvAI Innovations / Nandakishor M (NandhaKishorM/laya). This repository only changes the file format.
Model tree for Steven10429/laya-multilingual-webgpu
Base model
convaiinnovations/laya