Instructions to use Adkid/laya-cn-a with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use Adkid/laya-cn-a with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya-CN-A: Chinese intent adaptation
Independent community post-training by Adkid, based on Convai Innovations' historical Laya multilingual checkpoint. This is not an official Laya release. Release A denotes the first public Chinese adaptation. Internally this is experiment V3 seed42, distinct from the intermediate training checkpoint named A. Full FP32 weights are provided: no intermediate A checkpoint or delta merging is required by the user.
Run locally
This model uses the supplied PyTorch decision-model wrapper, not AutoModelForSequenceClassification or a generic Transformers text-classification pipeline. Review the Python source before running it. No trust_remote_code or cloud inference service is needed.
hf download Adkid/laya-cn-a --local-dir laya-cn-a
cd laya-cn-a
python -m pip install -r requirements-tested.txt
python predict.py --input example.json --device cpu
# Apple Silicon:
python predict.py --input example.json --device mps
from predict import Agent; agent = Agent(); agent.predict(state, questions) provides choice, binary noul, and ordered score probabilities. Model files are SHA-verified on load. Context/head budgets1024/512; long state truncation is reported. At most16 questions per call. Output probabilities are not calibrated for arbitrary domains. The frozen action head is not exposed as a trustworthy autonomous-action/abstention signal.
Evidence
On existing transformed Chinese test subsets, the historical original Laya / intermediate A / selected V3 achieved177/258 /203/258 /232/258 for binary intent and96/222 /149/222 /188/222 for choice. Fixed-seed replications43/44 scored231/258 and229/258;187/222 and186/222. Three-seed means89.41% and84.23%. These are the same subsets reused across seeds, not additional independent examples, not a new blind test and not the full60-label benchmark. V3seed42 stays the candidate selected by the prespecified tune rule.
**Feishu routing remains inadequate:**25/64 direct choice,19/64 four-question composition; the latter missed30/32 actionable cases. Feishu represents Chinese workplace collaboration messages in this synthetic diagnostic, not an official evaluation or real private chat dataset. T2 relevance RPS slightly regressed, as did original-format CrossWOZ F1. This model should not replace a reliable task-routing service or be advertised as beating Jev.
Training
95,708 original-label judgements from MASSIVE, CrossWOZ and T2Ranking; no Feishu test training and no LLM-generated gold. Original multilingual Laya→post-trained A→V3; V3 updates54,892,033 parameters (last8 encoder blocks/final norm and decision components), with encoder LR1e-5/head LR1e-4, batch64/microbatch8,2 full epochs/2991 updates, plus0.5-weight ordinal relevance RPS loss. The uploaded checkpoint is the exact merged FP32 state, not quantization. Lineage, revisions and hashes are in manifest.json.
Validation and limits
Merged weights were checked tensor-by-tensor against A+delta. Five requests on each of CPU/MPS exactly matched archived outputs (export_validation.json). Tested macOS26.6.1, Torch2.14.0, Transformers5.17.0; other environments and this wrapper's CUDA path were not independently validated. No production service/SLA is provided. Historical smoke request latency is not a formal performance benchmark. Wrapper head budget512 differs from the Feishu diagnostic's256; do not mix their results.
Upstream architecture/shared code and pretrained weights: Convai Innovations and Laya contributors (Apache-2.0). MASSIVE: CC-BY-4.0; CrossWOZ and T2Ranking: Apache-2.0; see DATA_SOURCES.md. Source corpora are not redistributed. Preparation, transformations, training/evaluation and packaging used OpenAI Codex assistance. The new wrapper and contribution use Apache-2.0, preserving upstream attribution. Full fresh-machine retraining replay is not claimed.
Original public Feishu benchmark. Laya-CN Study research collection · Laya-CN-A experiment, raw predictions and audit. The earlier Laya PR #295 records the maintainer’s decision to keep adaptation studies in independent repositories.
Run python audit.py to independently verify the bundled historical prediction records without loading model weights.
- Downloads last month
- -
Model tree for Adkid/laya-cn-a
Base model
convaiinnovations/laya