MiniCPM5 1B + Automaticity V9 LoRA

Rank-16 response-only LoRA trained for one epoch on the private Automaticity V9 friendly direct-tool corpus. This is a strong validation candidate, not a production-promoted autonomous router.

The model routes one current thought to at most one available tool, or makes no tool call. Training used MiniCPM5's native XML tool calls with thinking disabled and loss only on the assistant turn.

Training

  • Base: openbmb/MiniCPM5-1B
  • Base/tokenizer revision: 4e9de7a0778dc1c362e983e6858f0e77542cbdca
  • Rows: 4,900; dataset SHA-256: f3421604542d8f333576db814b471751900bc6aa2cfe109084c68d0f9ddf9c20
  • Context: 2,048 tokens; no truncation; maximum rendered row 2,036 tokens
  • Precision: BF16 LoRA, not QLoRA
  • LoRA: rank 16, alpha 32, dropout 0.05; attention and MLP projections
  • Epochs: 1; cosine learning-rate schedule; 3% warmup
  • Peak learning rate: 2e-4
  • Effective batch: 16 (1 x 16 gradient accumulation)
  • Seed: 3407
  • Loss: native assistant response only
  • Trainer runtime: 1,020.55 seconds
  • Adapter SHA-256: 8c0f24b5fce0063237b6f89ea22558b013b9e454bb3009895b0f81c1f8a65209

Frozen validation result

Evaluation used 1,050 private validation rows with normal five-tool retrieval, no gold injection, 100% action-gold retrieval recall, thinking disabled, and no decoding constraint. The validation dataset SHA-256 is 85094c96ca7fa2f96cbb0f7f85bd08510d56b9d4639646156d8806680bca9715.

Metric Result
End-to-end exact 90.86%
Routing 98.38%
Action exact 69.55%
No-tool precision 100%
No-tool recall 99.86%
Listed-tool rate 100%
Valid-call rate 100%
Latency average 0.272 s
Latency p50 0.141 s
Latency p95 0.910 s
No-tool latency average / p95 0.130 s / 0.177 s
Action latency average / p95 0.607 s / 1.671 s

The untuned base on the identical CUDA validation condition scored 16.29% end-to-end exact, 28.38% routing, 52.24% action exact, and 1.08% no-tool recall.

Limitations

This adapter is not yet promoted for autonomous execution. The frozen validation set still contains 17 wrong-tool rows and 79 wrong-argument rows; action exact is 69.55%. Nested XML arguments are a recurring failure mode. Use strict listed-name and schema validation or constrained decoding and reject invalid calls at runtime. Constraints cannot repair semantically wrong listed tools or schema-valid wrong arguments.

The private dataset and row-level evaluation repository is turnercore/automaticity-v9.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for turnercore/minicpm5-1b-automaticity-v9-lora

Adapter
(44)
this model