laya-p150
Derived from convaiinnovations/laya. Weights revision 7b928d82. This represents the model implementation on Tenstorrent hardware. See the original model card for license, training, and evaluation details.
Laya is a non-LLM calibrated decision model: a ModernBERT-large encoder with an option-marker decision head (421M parameters). One forward pass answers typed noul, choice and score questions about a text or JSON state with calibrated probabilities; it never generates text. This package serves the Jev-compatible /v1/systemone API of the authors' laya-serve on one Blackhole p150 chip, with a live demo at /demo/.
Runs on p150 or p150x4 β see the serve profiles below.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | ModernBERT-large encoder (28 layers, 1024 hidden, GeGLU, dual-theta RoPE, one global attention layer per three) + a 2-layer decision head with an option-marker scorer and an act head |
| Hardware | p150, p150x4 |
| License | apache-2.0 |
| Status | Experimental community bring-up |
Intended use
Direct use: Calibrated typed decisions over text states (routing, triage, guardrails, relevance, scoring).
Out-of-scope use: Text generation; non-English input (use the multilingual checkpoint); more than about 20 options at the default head budget.
Quickstart
uv tool install tenstorrent # once β the Tenstorrent CLI, `tt`
tt model pull tt-hous/laya-p150
tt serve tt-hous/laya-p150
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the convaiinnovations/laya weights at 7b928d828b7b0e022f929d9bd2e44165aa270148 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Without tt-cli β tt-model alone does the whole job:
tt-model pull tt-hous/laya-p150 --with-weights
tt-model serve tt-hous/laya-p150
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh |
|---|---|---|
p150 (default) |
p150 | P150 |
p150x4 |
p150x4 | P150x4 |
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API β its request and response shapes are the model's own. See the author's notes above for the payload it expects.
POST /v1/systemone with {"state": ..., "questions": {id: {"type": "choice"|"score"|"noul", "instructions": ..., "criteria": ...}}, "max_len", "head_max_len", "min_confidence"} returns {"model", "answers": {id: {...}}, "usage": {input_tokens, output_tokens, state_tokens, state_tokens_dropped, truncated, truncated_questions}}. POST /v1/systemone/batch takes {"states": [...], "questions": {...}} (1 to 64 states sharing one question set) and returns {"results": [...], "total_usage": {...}}. GET /v1/health reports backend, precision, mesh, buckets and the startup sanity check. Every /v1 response carries X-Inference-Time-Ms, X-Laya-Device-Ms and X-Laya-Batch. Existing laya and Jev clients work by changing the base URL. The demo page is served at /demo/. shim/laya_tt_backend.py installs a TtBackend on a pip laya Agent that routes agent.model.forward to this server's /v1/forward (served because LAYA_RAW_FORWARD=1 in this package).
Expected performance
Measured from the packaged image served on one p150 (build 5, image tt-model/laya-p150:dd4da8467626, code sha256 7b36a18faed4998a61c5c6e355eb2bd62cdd9ca9f8ec408bf489b6c30def831b; builds 6 and 7 carried the same code and build 7 reproduced every cell within 0.7 ms; this build changes only the demo page files (demo/, server/demo.py and their tests), not the model or the API path; result files under /home/hous/dev/laya/evals/results/package_p150_b5_20261006T024919Z, release notes in doc/release/RUN_NOTES.md). Latency per call (client median, warm, authors' bench_latency protocol), p150 versus the authors' Tesla T4: 1 question 10.9 ms (T4 39.5), 5 questions 24.6 ms (84.5), 10 questions 42.5 ms (158.6), 50 questions 197.5 ms (771); device forward 9.2 / 22.7 / 40.3 / 191.6 ms. Batched /v1/systemone/batch: 203 to 250 questions per second on one p150 (T4 103 to 332); the p150x4 profile 349 to 940. Typed-decisions test set (400 cases, 2,000 decisions, zero-shot): accuracy 0.359, soft accuracy 0.332, Brier 0.311, ECE 0.172, score MAE 0.689 (published 0.362 / 0.332 / 0.316 / 0.175 / 0.694; CPU fp32 on this host 0.3615 / 0.3315 / 0.3155 / 0.1747 / 0.6937). AG News 0.955 and DAIR Emotion 0.593 accuracy (authors' CPU run 0.950 / 0.595; Jev 0.910 / 0.480). Parity against the CPU fp32 reference on 488 decisions (wire and tensor paths): 403 of 403 confident decisions (reference margin at least 0.10) agree, 476 of 488 argmax agree, median max abs delta p 0.008, p95 0.031, max 0.12, scorer-logit PCC 0.994; on the 800 single-row suite decisions 777 of 779 confident decisions agree, median 0.001, max 0.27.
Limitations
English root checkpoint only (max_len 512, head_max_len 192). Raw calibration as shipped: the authors' ECE of 0.081 is after domain temperature fitting with an unpublished protocol (their raw value is 0.466). The base checkpoint is near chance zero-shot on typed-decisions; the published 0.766 belongs to the fine-tuned sibling. action.act_probability carries no usable signal (authors' issue #185); gate on answer_confidence. Temperatures are clamped to [0.5, 5.0] as pip laya 0.3.27 does, so choice questions with more than 10 options use 0.5 instead of the shipped 0.1006. Requests are padded to 128, 256 or 512 tokens x 1 to 64 rows (exact buckets at 5, 10 and 50 rows); longer sequences are refused with 422. Agreement with the CPU fp32 reference depends on the evidence set: on the 488-decision parity corpus (5-row calls, typed-decisions and benchmark-ticket questions) 12 decisions flip, all with a reference margin under 0.10, and none of the 403 confident decisions flips; on the 400 single-row six-option DAIR Emotion decisions 6 flip and 2 of 386 confident decisions flip (reference margins 0.125 and 0.369), on the 400 single-row four-option AG News decisions 2 flip and none of the 393 confident decisions. Short single-row inputs with several near-tied options are the most sensitive shape. The investigation (doc/full_model/single_row_investigation.md) shows the deviation is a property of the bfp8 numerics on those inputs, not of the bucket or batch placement; on the seven Emotion texts with the largest moves and on the 200-decision gate subset, LAYA_PRECISION=bf16_hifi4 keeps every decision within 0.11 of the reference with no flip, at about 11 percent more latency (not measured over all 800 single-row decisions). Placement invariance (the same question answered alone or inside a batch): max probability spread 0.009 on the gate corpus; 0.019 on 16 short single-row questions measured at the 512-token bucket, with the same answer in every placement. One device process serializes requests. p150x4 status: data parallel over rows, bit-identical to one chip at equal per-chip batch, served from the packaged image (build 5: healthy in 20 s, 68.5 ms for 50 questions; parity smoke 97 of 100 decisions and 82 of 82 confident). Not reproduced: MASSIVE, XNLI, 45 of 51 languages, post-temperature ECE.
Risks and safety considerations
bf8w_hifi3_erf numerics move probabilities against the fp32 reference by up to 0.12 (median 0.008) on the parity corpus, up to 0.056 on AG News and up to 0.27 on DAIR Emotion (7 of 400 decisions above 0.12), see limitations; gate on confidence, not act_probability. Long option texts are truncated to 48 tokens silently, as in the reference. The shipped temperatures are over-confident per the authors; refit on your own data.
Licensing
Laya weights and the vendored rl_common.py, rl_agent_api.py and email_utils.py: Apache-2.0 (Convai Innovations); ModernBERT-large: Apache-2.0 (Answer.AI); TTNN port and serving code: Apache-2.0 (Tenstorrent AI ULC); demo feed cases: LocalLLaMA/typed-decisions, Apache-2.0.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/laya-p150/discussions β that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from β code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout β commit not published |
code/ digest |
42ab6916a82d7364 (sha256, first 16 hex digits) |
| built | 2026-10-06T13:48:28+00:00 by tt-model 0.1.0 |
Model tree for tt-hous/laya-p150
Base model
answerdotai/ModernBERT-large