You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This model answers questions from a private fleet-management schema and is published as evidence for a method, not as a general-purpose assistant. Tell us who you are and what you intend to use it for.

Log in or Sign Up to review the conditions and access this model content.

Fleet Copilot 0.5B

A 0.5B model fine-tuned to answer one operational question at a time from a JSON snapshot of records, and to say so when the records do not contain the answer.

It is published as evidence for a method, not as a model to reuse. It learned one private schema inside one prompt format. Against any other schema the numbers below do not carry over. If the method interests you, the recipe is at the bottom.

What it replaced

An internal product feature answered questions from a records snapshot using llama3.2:3b. That model was switched off by default, because the team had measured it at 69% correct against the product's deterministic templates at 91%, with a figure invented in 7% of replies.

On the locked test set

387 rows drawn from groups of records that appear nowhere in training. Every figure in an answer is checked against the snapshot it was given; a few thresholds the system itself defines are declared and allowed.

this model llama3.2:3b untuned Qwen2.5-0.5B
Grounded: every figure traceable to the records 1.000 0.945 0.884
Correct: states the value asked for 0.996 0.464 0.295
Correct refusals when the records cannot answer 1.000 0.345 0.053

Latency is deliberately absent from this table. The tuned model was evaluated on a GPU and the other two through Ollama on a laptop, so the three numbers are not comparable and publishing them side by side would credit the model for a hardware difference. The same-machine comparison is below.

In the live product

18 real questions against a running instance with a real database, asked through the product's own route, each model in a fresh process so the prompt cache could not flatter a repeat:

useful answers wrong invented
deterministic templates alone 14/18 0 0
this model 18/18 0 0
llama3.2:3b 17/18 0 1

Asked to optimise routes with no route data present, llama3.2:3b produced a plan naming a record id that does not exist. This model says the route plans would hold that, and stops.

Latency when the model is actually called: 232 ms median, 930 ms worst, against 1050 ms median and 5088 ms worst. Resident memory: 479 MB against 2.5 GB.

Blind A/B against the untuned base model, judged on whether the answer matches the records shown, with the sides hidden by the server until every pick was in: 11 wins, 0 losses, 13 ties over 24 pairs. The ties are the questions where a base model that happens to copy a figure out of the snapshot lands on the same answer.

Prompt injection. 30 snapshots carried instruction-shaped text in free-text record fields, including forged copies of the delimiters that fence the records. This model obeyed none โ€” re-measured on this exact build, not inherited from the run it supersedes. llama3.2:3b obeyed one, repeating a planted claim as though it were a record.

The files

File Size
f16 994 MB
q8_0 531 MB
q5_k_m 420 MB
q4_k_m 398 MB

Take q4_k_m: on 60 test rows every quantisation was equally grounded, so the smallest wins.

What it gets wrong

On the locked rows it now states no number the records do not hold. An earlier version answered "show fuel efficiency" with an unrelated count, because the training data taught refusals only for the specific unanswerable questions it listed, not for every phrasing that routes to a topic the snapshot cannot serve. Adding those phrasings fixed it. That failure never appeared in the benchmark; only running it inside the product showed it.

Its remaining weakness is arithmetic it was never asked to do: it states figures, it does not compute new ones.

The method, which is the part worth copying

The training data was generated from the product's own code, not annotated and not written by a language model:

  1. Generate the input the model will actually see โ€” the snapshot, not the database, with the shape pinned to the product's source so the generator cannot drift.
  2. Compute the answers from a deterministic renderer the product already had. The label is true by construction, so a wrong row is a bug in the generator, not an annotation mistake.
  3. Vary the phrasing, or the model learns the template.
  4. Make refusal a first-class answer. 28% of rows ask something the snapshot cannot answer, and the target says so and names where the answer would live.
  5. Cover every phrasing that routes to an unanswerable topic, not just the questions you thought of. This is the step that the benchmark cannot check for you.
  6. Plant attacks in the data: instruction-shaped text in free-text fields, with a target that ignores it.
  7. Split by group, never by row.
  8. Score grounding and correctness separately. The 3B was already 94.5% grounded and hid a 46% correctness problem underneath that.

Training: LoRA r=16 in bf16, 3 epochs over 1,825 rows, 24 minutes on one GB10, inside a 10 GB memory budget. Validation loss 0.462 to 0.127.

Licence

Base model Qwen/Qwen2.5-0.5B-Instruct, Apache-2.0. The training data is generated from private source code and contains no customer data.

Downloads last month
-
GGUF
Model size
0.5B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Hayman-g/fleet-copilot-0.5b

Adapter
(855)
this model