Instructions to use turnercore/minicpm5-1b-automaticity-v9-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use turnercore/minicpm5-1b-automaticity-v9-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("openbmb/MiniCPM5-1B") model = PeftModel.from_pretrained(base_model, "turnercore/minicpm5-1b-automaticity-v9-lora") - Transformers
How to use turnercore/minicpm5-1b-automaticity-v9-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="turnercore/minicpm5-1b-automaticity-v9-lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("turnercore/minicpm5-1b-automaticity-v9-lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use turnercore/minicpm5-1b-automaticity-v9-lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "turnercore/minicpm5-1b-automaticity-v9-lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/minicpm5-1b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/turnercore/minicpm5-1b-automaticity-v9-lora
- SGLang
How to use turnercore/minicpm5-1b-automaticity-v9-lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "turnercore/minicpm5-1b-automaticity-v9-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/minicpm5-1b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "turnercore/minicpm5-1b-automaticity-v9-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/minicpm5-1b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use turnercore/minicpm5-1b-automaticity-v9-lora with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for turnercore/minicpm5-1b-automaticity-v9-lora to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for turnercore/minicpm5-1b-automaticity-v9-lora to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for turnercore/minicpm5-1b-automaticity-v9-lora to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="turnercore/minicpm5-1b-automaticity-v9-lora", max_seq_length=2048, ) - Docker Model Runner
How to use turnercore/minicpm5-1b-automaticity-v9-lora with Docker Model Runner:
docker model run hf.co/turnercore/minicpm5-1b-automaticity-v9-lora
MiniCPM5 1B + Automaticity V9 LoRA
Rank-16 response-only LoRA trained for one epoch on the private Automaticity V9 friendly direct-tool corpus. This is a strong validation candidate, not a production-promoted autonomous router.
The model routes one current thought to at most one available tool, or makes no tool call. Training used MiniCPM5's native XML tool calls with thinking disabled and loss only on the assistant turn.
Training
- Base:
openbmb/MiniCPM5-1B - Base/tokenizer revision:
4e9de7a0778dc1c362e983e6858f0e77542cbdca - Rows: 4,900; dataset SHA-256:
f3421604542d8f333576db814b471751900bc6aa2cfe109084c68d0f9ddf9c20 - Context: 2,048 tokens; no truncation; maximum rendered row 2,036 tokens
- Precision: BF16 LoRA, not QLoRA
- LoRA: rank 16, alpha 32, dropout 0.05; attention and MLP projections
- Epochs: 1; cosine learning-rate schedule; 3% warmup
- Peak learning rate: 2e-4
- Effective batch: 16 (1 x 16 gradient accumulation)
- Seed: 3407
- Loss: native assistant response only
- Trainer runtime: 1,020.55 seconds
- Adapter SHA-256:
8c0f24b5fce0063237b6f89ea22558b013b9e454bb3009895b0f81c1f8a65209
Frozen validation result
Evaluation used 1,050 private validation rows with normal five-tool retrieval,
no gold injection, 100% action-gold retrieval recall, thinking disabled, and no
decoding constraint. The validation dataset SHA-256 is
85094c96ca7fa2f96cbb0f7f85bd08510d56b9d4639646156d8806680bca9715.
| Metric | Result |
|---|---|
| End-to-end exact | 90.86% |
| Routing | 98.38% |
| Action exact | 69.55% |
| No-tool precision | 100% |
| No-tool recall | 99.86% |
| Listed-tool rate | 100% |
| Valid-call rate | 100% |
| Latency average | 0.272 s |
| Latency p50 | 0.141 s |
| Latency p95 | 0.910 s |
| No-tool latency average / p95 | 0.130 s / 0.177 s |
| Action latency average / p95 | 0.607 s / 1.671 s |
The untuned base on the identical CUDA validation condition scored 16.29% end-to-end exact, 28.38% routing, 52.24% action exact, and 1.08% no-tool recall.
Limitations
This adapter is not yet promoted for autonomous execution. The frozen validation set still contains 17 wrong-tool rows and 79 wrong-argument rows; action exact is 69.55%. Nested XML arguments are a recurring failure mode. Use strict listed-name and schema validation or constrained decoding and reject invalid calls at runtime. Constraints cannot repair semantically wrong listed tools or schema-valid wrong arguments.
The private dataset and row-level evaluation repository is
turnercore/automaticity-v9.
- Downloads last month
- -
Model tree for turnercore/minicpm5-1b-automaticity-v9-lora
Base model
openbmb/MiniCPM5-1B