Laya decision models, compiled for the SiMa.ai MLSoC Modalix

Laya by Convai Innovations is a "System-1" decision model: an encoder (ModernBERT-large, or mmBERT-base for the multilingual one) with a decision head that answers a typed question about a piece of text in one forward pass. It picks one of several options, scores on a scale, or says yes or no.

These are those checkpoints compiled for the Modalix MLA. Every graph is a single MLA stage, with no layers on the CPU, and a decision takes about 19 ms at 128 tokens (8 ms for the multilingual model).

folder model precision latency per decision, by tokens same decision as PyTorch size
general Laya BF16 128: 19.5 ms, 256: 35.2 ms, 512: 87.5 ms 100 / 100 3.27 GB
general-int8 Laya INT8 A_BF16_W_INT8 128: 14 ms 98 / 100 0.67 GB
typed-decisions Laya typed-decisions BF16 128: 19.4 ms, 512: 87.5 ms, 1024: 282 ms 99 / 100 4.35 GB
multilingual Laya multilingual BF16 128: 8.1 ms, 256: 15.3 ms, 512: 33.9 ms, 1024: 116.5 ms 100 / 100 2.61 GB
dino Laya-dino BF16 64: 18 ms 30 / 30 game states 0.95 GB

Latency is one decision on a Modalix DevKit, measured inside the runtime. "Same decision" is against the fp32 PyTorch model on 100 decisions with the 128-token graph. Precision BF16 is weights and activations in bfloat16; A_BF16_W_INT8 keeps bfloat16 activations and stores the weights as int8.

What is in a folder

file what it is
laya_s<N>_stage1_mla.elf the compiled graph for sequences of up to N tokens
token_embeddings.bf16 the embedding table, looked up on the CPU
act_tail.f32 the last layer of the act head, run on the CPU
tokenizer.json the checkpoint's tokenizer
laya_config.json token ids, calibration temperatures, and which graphs there are

models.json lists every folder with file sizes and SHA-256 sums.

Using them

The files run on a Modalix board with the runtime and web app from neat-laya-studio: its Models page downloads a model from this repository onto the board and loads it on the MLA. By hand:

hf download TDoSiMa/sima-laya --include "general/*" --local-dir .
./laya run general --state "We were billed twice for March." \
    --question '{"type": "noul", "instructions": "Is this a billing problem?"}'

They are not usable with PyTorch or ONNX Runtime; for that, use the original checkpoints at convaiinnovations/laya.

License and attribution

Apache-2.0. The models are Convai Innovations' Laya checkpoints (https://github.com/NandhaKishorM/laya, Apache-2.0), compiled without retraining. The exception is dino, whose decision head was fine-tuned on top of the English checkpoint for a browser game. Zero-shot quality is the checkpoints': the port reproduces PyTorch's answers, including its wrong ones.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TDoSiMa/sima-laya

Finetuned
(144)
this model