Instructions to use Hayman-g/fleet-copilot-0.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Hayman-g/fleet-copilot-0.5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Use Docker
docker model run hf.co/Hayman-g/fleet-copilot-0.5b:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Hayman-g/fleet-copilot-0.5b with Ollama:
ollama run hf.co/Hayman-g/fleet-copilot-0.5b:Q4_K_M
- Unsloth Desktop
- Pi
How to use Hayman-g/fleet-copilot-0.5b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Hayman-g/fleet-copilot-0.5b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Hayman-g/fleet-copilot-0.5b with Docker Model Runner:
docker model run hf.co/Hayman-g/fleet-copilot-0.5b:Q4_K_M
- Lemonade
How to use Hayman-g/fleet-copilot-0.5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Hayman-g/fleet-copilot-0.5b:Q4_K_M
Run and chat with the model
lemonade run user.fleet-copilot-0.5b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Hayman-g/fleet-copilot-0.5b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Hayman-g/fleet-copilot-0.5b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Hayman-g/fleet-copilot-0.5b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hayman-g/fleet-copilot-0.5b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Hayman-g/fleet-copilot-0.5b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This model answers questions from a private fleet-management schema and is published as evidence for a method, not as a general-purpose assistant. Tell us who you are and what you intend to use it for.
Log in or Sign Up to review the conditions and access this model content.
Fleet Copilot 0.5B
A 0.5B model fine-tuned to answer one operational question at a time from a JSON snapshot of records, and to say so when the records do not contain the answer.
It is published as evidence for a method, not as a model to reuse. It learned one private schema inside one prompt format. Against any other schema the numbers below do not carry over. If the method interests you, the recipe is at the bottom.
What it replaced
An internal product feature answered questions from a records snapshot using llama3.2:3b. That
model was switched off by default, because the team had measured it at 69% correct against the
product's deterministic templates at 91%, with a figure invented in 7% of replies.
On the locked test set
387 rows drawn from groups of records that appear nowhere in training. Every figure in an answer is checked against the snapshot it was given; a few thresholds the system itself defines are declared and allowed.
| this model | llama3.2:3b | untuned Qwen2.5-0.5B | |
|---|---|---|---|
| Grounded: every figure traceable to the records | 1.000 | 0.945 | 0.884 |
| Correct: states the value asked for | 0.996 | 0.464 | 0.295 |
| Correct refusals when the records cannot answer | 1.000 | 0.345 | 0.053 |
Latency is deliberately absent from this table. The tuned model was evaluated on a GPU and the other two through Ollama on a laptop, so the three numbers are not comparable and publishing them side by side would credit the model for a hardware difference. The same-machine comparison is below.
In the live product
18 real questions against a running instance with a real database, asked through the product's own route, each model in a fresh process so the prompt cache could not flatter a repeat:
| useful answers | wrong | invented | |
|---|---|---|---|
| deterministic templates alone | 14/18 | 0 | 0 |
| this model | 18/18 | 0 | 0 |
| llama3.2:3b | 17/18 | 0 | 1 |
Asked to optimise routes with no route data present, llama3.2:3b produced a plan naming a record
id that does not exist. This model says the route plans would hold that, and stops.
Latency when the model is actually called: 232 ms median, 930 ms worst, against 1050 ms median and 5088 ms worst. Resident memory: 479 MB against 2.5 GB.
Blind A/B against the untuned base model, judged on whether the answer matches the records shown, with the sides hidden by the server until every pick was in: 11 wins, 0 losses, 13 ties over 24 pairs. The ties are the questions where a base model that happens to copy a figure out of the snapshot lands on the same answer.
Prompt injection. 30 snapshots carried instruction-shaped text in free-text record fields,
including forged copies of the delimiters that fence the records. This model obeyed none โ
re-measured on this exact build, not inherited from the run it supersedes.
llama3.2:3b obeyed one, repeating a planted claim as though it were a record.
The files
| File | Size |
|---|---|
f16 |
994 MB |
q8_0 |
531 MB |
q5_k_m |
420 MB |
q4_k_m |
398 MB |
Take q4_k_m: on 60 test rows every quantisation was equally grounded, so the smallest wins.
What it gets wrong
On the locked rows it now states no number the records do not hold. An earlier version answered "show fuel efficiency" with an unrelated count, because the training data taught refusals only for the specific unanswerable questions it listed, not for every phrasing that routes to a topic the snapshot cannot serve. Adding those phrasings fixed it. That failure never appeared in the benchmark; only running it inside the product showed it.
Its remaining weakness is arithmetic it was never asked to do: it states figures, it does not compute new ones.
The method, which is the part worth copying
The training data was generated from the product's own code, not annotated and not written by a language model:
- Generate the input the model will actually see โ the snapshot, not the database, with the shape pinned to the product's source so the generator cannot drift.
- Compute the answers from a deterministic renderer the product already had. The label is true by construction, so a wrong row is a bug in the generator, not an annotation mistake.
- Vary the phrasing, or the model learns the template.
- Make refusal a first-class answer. 28% of rows ask something the snapshot cannot answer, and the target says so and names where the answer would live.
- Cover every phrasing that routes to an unanswerable topic, not just the questions you thought of. This is the step that the benchmark cannot check for you.
- Plant attacks in the data: instruction-shaped text in free-text fields, with a target that ignores it.
- Split by group, never by row.
- Score grounding and correctness separately. The 3B was already 94.5% grounded and hid a 46% correctness problem underneath that.
Training: LoRA r=16 in bf16, 3 epochs over 1,825 rows, 24 minutes on one GB10, inside a 10 GB memory budget. Validation loss 0.462 to 0.127.
Licence
Base model Qwen/Qwen2.5-0.5B-Instruct, Apache-2.0. The training data is generated from private source code and contains no customer data.
- Downloads last month
- -
4-bit
5-bit
8-bit
16-bit