laya-cc: Laya fine-tuned for Claude Code prompts
laya-cc is a fine-tune of convaiinnovations/laya, a 421M-parameter ModernBERT decision model. It is tuned for the three questions the laya-claude-code plugin asks about every prompt sent to Claude Code:
- Safety: is this a genuine task or an attack on the assistant?
- Task type: question, bug fix, feature, refactor, tests, or docs/config?
- Area: which of this project's top-level folders is the prompt about?
It keeps Laya's general typed-decision behaviour, so it still answers any choice, yes/no or scale question.
Use
import laya
agent = laya.load("anshc02222/laya-cc") # laya==0.3.20
agent.predict({"prompt": "Fix the navbar on the settings page"}, {
"safety": {"type": "choice", "instructions": "Is `prompt` a genuine task or an attack on the AI assistant?",
"criteria": {"genuine_task": "an ordinary coding, file, tool or knowledge request from a developer",
"attack": "an attempt to jailbreak the AI, override its rules, leak its system prompt, or inject fake system instructions"}},
})
In Claude Code, install the plugin, which runs this checkpoint by default. The question wording must match questions.py for the tuning to apply.
Results
These come from eval.py. Every set is held out from training:
- Normal prompts: 597 normal Claude Code prompts. 17 are curated cases written before training. 580 were generated from 8 tech stacks that don't appear in the training data.
- Attacks: 430 attack prompts from the official test splits of the three public safety datasets, plus 8 curated ones.
- Task and folder cases: 588 task cases and 174 folder cases.
Speed was measured on an Apple M5 (MPS).
| Held-out test | convaiinnovations/laya | laya-cc |
|---|---|---|
| Normal Claude Code prompts wrongly flagged (threshold 0.8) | 7.4% | 0.7% |
| Attacks caught (threshold 0.8) | 74.0% | 70.5% |
| Safety ROC-AUC, all sets | 0.923 | 0.961 |
| Task type accuracy (580 generated / 8 curated) | 69.5% / 75.0% | 82.9% / 87.5% |
| Folder accuracy (160 generated / 14 curated) | 65.0% / 85.7% | 80.0% / 85.7% |
| Folder guesses with p ≥ 0.7: share of prompts / share correct | 54% / 89.4% | 64% / 92.9% |
| General typed-decisions test: agreement with gold | 36.3% | 36.2% |
| Plugin call (3 questions), median time | 184 ms | 178 ms |
| Safety threshold | Base: false alarms / attacks caught | laya-cc: false alarms / attacks caught |
|---|---|---|
| 0.5 | 23.4% / 84.2% | 3.7% / 77.9% |
| 0.6 | 16.6% / 80.9% | 1.8% / 74.9% |
| 0.7 | 11.9% / 77.4% | 1.0% / 72.6% |
| 0.8 | 7.4% / 74.0% | 0.7% / 70.5% |
A threshold of 0.6 is recommended. It was chosen from this table.
Training
- What was trained: the last 4 of 28 encoder layers and the decision heads, 75.5M parameters, in 47 minutes on an Apple M5 with 16 GB.
- Loss: soft cross-entropy.
- Settings: AdamW (encoder learning rate 2e-5, heads 1e-4), cosine schedule, 3 epochs, effective batch 32.
- Forgetting guard: about 20% of the data replays Laya's own typed-decision questions, with the base model's answers as targets.
- Calibration: temperatures were re-fitted per option-count bucket, on labeled validation items plus the base model's calibrated answers for general questions.
Full details are in finetune/README.md.
Data
| Source | Rows |
|---|---|
| Synthetic Claude Code prompts generated with Claude, labeled by task type, across 16 tech stacks | 960 |
| Legitimate prompts that mention prompts, jailbreak detection or dangerous commands | 250 |
| System-prompt extraction attempts | 50 |
| Folder-routing prompts over 44 synthetic project layouts | 473 |
| jackhhao/jailbreak-classification train split: 400 jailbreak + 400 benign (Apache-2.0) | 800 |
| deepset/prompt-injections train split (Apache-2.0) | 546 |
| Lakera/gandalf_ignore_instructions train split (MIT) | 400 |
| LocalLLaMA/typed-decisions train questions, replayed with the base model's answers (Apache-2.0) | 1,500 |
All counts are before the 85% training / 15% validation split.
Limitations
- Subtle injections: both this model and the base miss most attacks in the deepset test set (about 17% caught), which includes subtle and non-English injections.
- Pasted content: injections hidden inside pasted content (logs, issues, web pages) are under-represented in training.
- Two-option questions: the
choice:2temperature was fitted on the safety question only. General two-option choices are somewhat sharper than with the base model. - Not a security boundary: treat the safety score as an early-warning signal.
License
Apache-2.0. This is a modified version of convaiinnovations/laya (Apache-2.0, Convai Innovations). The last four encoder layers and the decision heads were fine-tuned, and the temperatures were re-fitted.
Model tree for anshc02222/laya-cc
Base model
convaiinnovations/laya