YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Laya, packaged for the browser

Every build laya.voronkov.club has served, in one place, so a demo does not depend on several separate repositories staying where they are. Each folder is self-contained: the graphs, the tokenizer that produced their token ids, and the config carrying the sequence limits and temperatures.

All Apache-2.0, as upstream. Credits are per variant below; the base models are not ours.

folder what it is
multilingual-tuned-q8/v1/ Fine-tuned, and the one the demo serves by default. laya-multilingual fine-tuned on typed-decisions, then exported weight-only int8. Runs on wasm, so it needs no GPU. Accepts a batch and answers the same (EQUIVALENT). 100% argmax agreement with the PyTorch reference; worst single probability shift 0.024. 574 MB.
multilingual-tuned-fp32-batched/v1/ The same fine-tune at full precision, re-traced so it accepts a batch. Nothing is lost to quantization: 100% argmax and a worst shift of 0.0000011. Needs WebGPU. 1.3 GB. Use it when the exact numbers matter; on every case measured it makes the same decisions as the int8 build above.
multilingual-tuned-fp32/v1/ The first full-precision export of the fine-tune, before the batched re-trace. Identical fidelity, but refuses a batch greater than 1. Kept so the difference the re-trace made is still inspectable; prefer multilingual-tuned-fp32-batched.
multilingual-fp16/v1/ The general multilingual checkpoint, untuned, at half precision. No quantization loss, and it batches. Scores 0.342 on typed decisions, below the 0.461 of always answering the most common option -- it is here to show what the base model does, not to be relied on. No fitted temperatures. Needs WebGPU. 647 MB.
multilingual-int8/v1/ The same untuned checkpoint in dynamic int8. The fastest build here and the least faithful: activation scales are derived per tensor at run time, which measured 93.8% argmax agreement with full precision and a worst shift of 0.169. It accepts a batch and answers differently when it does, by up to 21 points. 326 MB.
english-q8/v1/ Laya English base (ModernBERT-large, 421M), weight-only int8. 100% argmax agreement with fp32, worst shift 0.016. 512-token context. Refuses a batch. No longer served by the demo, which is multilingual only.
typed-decisions-q8/v1/ Laya fine-tuned on typed decisions (ModernBERT-large, 421M, English only). Scores 0.766, the highest here, but reads English alone. Weight-only int8 exported by us: 100% argmax, worst shift 0.009. Refuses a batch. Note that its temperature_by_options is inherited from the base checkpoint and overrides the temperatures fitted for it -- upstream says so too.

Where the numbers come from

argmax agreement and worst shift are measured against the PyTorch model the export came from, over a fixed set of 26 questions, by export/gate.py in laya-web-poc. Whether a build accepts a batch is measured too, by export/batch_probe.py, because it is a property of the export rather than of the checkpoint: three of these refuse one outright and a fourth accepts one and silently answers differently.

The benchmark scores are accuracy on the test split of LocalLLaMA/typed-decisions, 400 cases and 2,000 decisions. For scale: always answering the most common option scores 0.461, the teachers who wrote the answers agree with themselves 0.735 of the time, and the fine-tune here scores 0.748.

Credits

Base models: convaiinnovations/laya and its family, by Nandakishor M / Convai Innovations. Benchmark by LocalLLaMA.

The v1 segment

Browsers cache by URL. Re-quantizing publishes to v2 rather than overwriting v1, or a returning visitor keeps weights the app no longer believes it is running.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support