Laya Multilingual (Verdict-ready)

The Laya Multilingual System One decision model (mmBERT-base, 100+ languages), exported to ONNX fp32 and packaged for Verdict -- a high-performance serving engine for System One decision models. Serve it with one command:

# Hugging Face (default hub; set HF_ENDPOINT to use a mirror):
verdict serve verdict-models/laya-multilingual

# ModelScope:
VERDICT_USE_MODELSCOPE=1 verdict serve sgyaqing/laya-multilingual

For English-only workloads, use verdict-models/laya.

Endpoints: the standard decision protocols POST /v1/systemone and POST /v1/decisions, native POST /v1/choice | /v1/score | /v1/noul, plus an OpenAI-compatible layer (/v1/chat/completions, /v1/models).

Performance

Measured on one RTX 3080:

concurrency 32 Throughput P50
FastAPI + torch 46.0 rps 691 ms
Verdict 1059 rps 30 ms

Files

File Purpose
model.json manifest: graph file name, budgets, calibration, special-token names
laya-multilingual.onnx the ONNX graph (fp32)
tokenizer/tokenizer.json (+ tokenizer_config.json) the tokenizer
SHA256SUMS checksums for every file above

Provenance & fidelity

Weights are unchanged from convaiinnovations/laya-multilingual; this repo adds only the ONNX export and the serving manifest. Conversion and serving are verified against the official model: identical answers on CPU, within 1e-4 on CUDA.

License

Apache-2.0, inherited from the base model. Credit for the model itself belongs to the original authors; this repository only repackages it for serving.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for verdict-models/laya-multilingual

Quantized
(37)
this model