Instructions to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-1.2B-Instruct") model = PeftModel.from_pretrained(base_model, "turnercore/lfm2.5-1.2b-automaticity-v9-lora") - Transformers
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="turnercore/lfm2.5-1.2b-automaticity-v9-lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("turnercore/lfm2.5-1.2b-automaticity-v9-lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "turnercore/lfm2.5-1.2b-automaticity-v9-lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/lfm2.5-1.2b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/turnercore/lfm2.5-1.2b-automaticity-v9-lora
- SGLang
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "turnercore/lfm2.5-1.2b-automaticity-v9-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/lfm2.5-1.2b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "turnercore/lfm2.5-1.2b-automaticity-v9-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "turnercore/lfm2.5-1.2b-automaticity-v9-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for turnercore/lfm2.5-1.2b-automaticity-v9-lora to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for turnercore/lfm2.5-1.2b-automaticity-v9-lora to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for turnercore/lfm2.5-1.2b-automaticity-v9-lora to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="turnercore/lfm2.5-1.2b-automaticity-v9-lora", max_seq_length=2048, ) - Docker Model Runner
How to use turnercore/lfm2.5-1.2b-automaticity-v9-lora with Docker Model Runner:
docker model run hf.co/turnercore/lfm2.5-1.2b-automaticity-v9-lora
LFM2.5 1.2B + Automaticity V9 LoRA
Rank-16 response-only LoRA trained for one epoch on the private Automaticity V9 friendly direct-tool corpus. This is the strongest current V9 validation candidate, not a production-promoted autonomous router.
The model routes one current thought to at most one available tool, or makes no tool call. Training used LFM2.5's native marked Python-call-list format and loss only on the assistant turn.
Training
- Base:
LiquidAI/LFM2.5-1.2B-Instruct - Base/tokenizer revision:
868df74dd56ff8a0c2ac5dbf281690c2dbebe4c9 - Rows: 4,900; dataset SHA-256:
3fb79e5fe3cf762b3258c5674a806903e310aebb35d8ed153a525b0377b3bd8f - Context: 2,048 tokens; no truncation; maximum rendered row 1,978 tokens
- Precision: ROCm BF16 LoRA, not QLoRA
- LoRA: rank 16, alpha 16, dropout 0;
q/k/v/out/in_projandw1/w2/w3 - Epochs: 1; linear learning-rate schedule; 3% warmup
- Peak learning rate: 2e-4; weight decay: 0.001
- Effective batch: 16 (4 x 4 gradient accumulation)
- Seed: 3407
- Loss: native assistant response only
- Trainer runtime: 4,037 seconds
- Adapter SHA-256:
e81bdda7e1a684ae3a3f8d952303446d974c9ededf0cf383b8c76791112340ea
Frozen validation result
Evaluation used 1,050 private validation rows with normal five-tool retrieval,
no gold injection, 100% action-gold retrieval recall, and no decoding constraint.
The validation dataset SHA-256 is
85094c96ca7fa2f96cbb0f7f85bd08510d56b9d4639646156d8806680bca9715.
| Metric | Result |
|---|---|
| End-to-end exact | 95.33% |
| Routing | 98.00% |
| Action exact | 84.94% |
| No-tool precision | 99.86% |
| No-tool recall | 99.73% |
| Argument schema validity | 99.90% |
| Listed-tool rate | 99.90% |
| Valid-call rate | 100% |
| Latency average | 1.242 s |
| Latency p50 | 0.518 s |
| Latency p95 | 4.520 s |
| No-tool latency average / p95 | 0.471 s / 0.661 s |
| Action latency average / p95 | 3.065 s / 8.413 s |
The untuned base on the identical ROCm validation condition scored 32.29% end-to-end exact, 45.52% routing, 29.17% action exact, and 33.60% no-tool recall.
Limitations
This adapter is not yet promoted for autonomous execution. The frozen validation set still contains 20 wrong-tool rows, 28 wrong-argument rows, and one unlisted call. Action p95 latency is 8.413 seconds, 6.3% slower than the untuned action p95 even though aggregate latency improved substantially. Use strict listed-name and schema validation or constrained decoding and reject invalid calls at runtime. Constraints cannot repair semantically wrong listed tools or schema-valid wrong arguments.
The private dataset and row-level evaluation repository is
turnercore/automaticity-v9.
- Downloads last month
- 15
Model tree for turnercore/lfm2.5-1.2b-automaticity-v9-lora
Base model
LiquidAI/LFM2.5-1.2B-Base