Catbird

KOS-V5-Instruct · "Catbird"

Developed by

University of Kentucky

University of Louisville

A 3.72B-parameter medical language model trained from scratch. It is not distilled, not pruned and not continued-pretrained from a general base. KOS-V5 (codename Catbird) is the fifth-generation Kentucky Open Science model line. This repository holds the instruction-tuned head of that line: the KOS-V5-Base pretraining checkpoint, taken through SFT and two GRPO reinforcement-learning legs.

Unlike the base, this model follows instructions and calls tools. It is the downstream SFT/RL artifact that KOS-V5-Base was built to initialise.

Code name: Catbird. The KOS-V5 series is nicknamed Catbird; native to Kentucky, the Gray Catbird (Dumetella carolinensis) is a medium-sized songbird famous for its distinct, cat-like "meow" call. This LLM was trained completely from scratch by teams from the University of Kentucky (Cat) and University of Louisville (Bird), so the code name is fitting.

⚠️ Research use only. This model is provided for research purposes only and must not be used for any commercial, clinical, legal, or production-grade application. The user assumes all risks associated with its use.


Its instruction ability comes from GRPO reinforcement learning against the official IFEval verifier, and its tool-calling ability from a second GRPO leg against the official BFCL AST checker, on a base that ranks first of 17 at modelling held-out clinical text.

IFEval reported as strict-avg = (prompt-level strict + instruction-level strict) / 2 — the exact metric the Hugging Face Open LLM Leaderboard publishes as "IFEval."

IFEval strict-avg model who built it, and how
72.19 KOS-V5-Instruct (ours) University research team, 235B tokens, from scratch
64.7 Qwen2.5-3B-Instruct Alibaba, ~18 trillion tokens
61.6 KOS-V4-Instruct (previous generation) University research team, 180B tokens, 24 GPUs
55.9 GPT-3.5-turbo-1106 (the original ChatGPT) OpenAI, ~10,000-GPU supercomputer

KOS-V5-Instruct improves on KOS-V4-Instruct across every benchmark measured: IFEval strict-avg 61.6 → 72.19 (+10.6), MMLU 0.2782 → 0.4512 (+17.3), medical QA (PubMedQA 0.7060, MedQA 0.3802, MedMCQA 0.3648 — all up on V4), and official BFCL function-calling 72.75/73.00/60.50 → 85.00/84.00/80.50 (+12.3 / +11.0 / +20.0). It clears the original GPT-3.5-turbo generation and the commercially trained Qwen2.5-3B on instruction following, and its tool calling now runs above the Qwen3-4B-Instruct-2507 peer.

Core specifications

Attribute Detail
Architecture Decoder-only Transformer (Qwen3ForCausalLM), Grouped-Query Attention
Parameters 3.715 B
Hidden / Layers 2560 / 36
Attention 32 query / 8 KV heads (GQA 4:1), head_dim 128, per-head QK-RMSNorm
Feed-forward SwiGLU, intermediate 9728
Vocabulary 32,000, custom medical byte-level BPE
Context length 32,768
Position encoding RoPE, θ = 25,000
Embeddings tied
Precision bfloat16 (7.43 GB, single shard)

Pre-training (the KOS-V5 base)

Fine-tuned from KOS-V5-Base — the from-scratch pretrained foundation, a complete single-epoch run over 235.2B tokens. See that card for corpus composition and disclosed pretraining issues.

Post-training (this model)

Three stages on top of the base — no LoRA, no distillation, no reward model, no LLM judge.

SFT — one shuffled full-parameter pass over a 736,990-record / 1.32B-token audited instruction mix (ChatML). The mix was deduplicated, instruction-collision canonicalized, structurally validated, per-record decontaminated and BFCL-closure scanned; clinical FHIR records were dropped and tool-record system prompts de-welded so tool use binds to the request rather than to a fixed frame.

RL leg 1 — instruction following (GRPO via verl) — deterministic verifiable reward. The reward is the official lm_eval IFEval instruction registry — the same checkers the benchmark scores with, not a re-implementation. One 8-GPU wall, 164 steps, KL 0.001 (low_var_kl), rollout n=16.

RL leg 2 — tool calling (GRPO via verl) — a second leg seeded from leg 1. The reward is the official BFCL ast_checker (bfcl_eval). Each prompt renders its tool schemas through the model's own chat template (byte-exact to the official tools= rendering), and the prompt set is filtered to only rows the official checker can grade. 8-GPU wall, KL 0.001, rollout n=16; this repository ships the step-40 checkpoint, selected for the best BFCL / abstention balance and least policy drift. BFCL rose 76.8/71.0/69.5 → 85.0/84.0/80.5 with instruction following, grounded abstention and knowledge all held.

Forgetting control — out-of-distribution broad-holdout perplexity at 0.99× the pre-RL base (8.88 vs 8.97), measured on a web crawl postdating the training corpus. No measurable forgetting.

The medical foundation

This is a medical model. KOS-V5-Instruct inherits a base trained on a 54-source medical/biomedical corpus — not a general-purpose model with medical fine-tuning bolted on.

The strongest evidence is bits-per-byte on held-out medical text, which is tokenizer-agnostic and therefore the only strictly fair cross-model comparison. In a 17-model pool — including dedicated biomedical specialists BioMedLM (300B PubMed tokens), Meditron-7B, PMC-LLaMA-7B and MedGemma-4B — the KOS-V5 base ranks first:

medical text (BPB, lower is better) KOS-V5-Base rank
5-corpus mean, held-out medical text 0.4635 1 / 17
clinical narratives 0.4179 1 / 17
radiology 0.5132 1 / 17
chest X-ray reports 0.6688 1 / 17
BIOSSES biomedical sentence similarity (Pearson / Spearman) 0.7097 / 0.7014 1 / 17
BLURB biomedical probe mean 0.7268 2 / 17

Every comparator in that pool was trained on 1.3–153× more data (0.3–36T tokens vs our 0.235T). See KOS-V5-Base for the full 96-metric evaluation.

Medical MMLU (from the 57-subject run above)

The 9 medical subjects of MMLU, extracted from the same official 5-shot run:

medical subject KOS-V5-Instruct KOS-V4-Instruct
high-school biology 0.5774 0.2387
clinical knowledge 0.5623 0.3170
nutrition 0.5359 0.2843
college biology 0.5347 0.2917
medical genetics 0.5100 0.2700
anatomy 0.4815 0.3185
professional medicine 0.4375 0.2132
college medicine 0.4046 0.2486
virology 0.3795 0.2952
medical-9 mean 0.4915 0.2752

+21.6 points over KOS-V4-Instruct, and above the model's own full-MMLU average (0.4512) — the medical subjects are where it is strongest.

Medical QA benchmarks (official suites)

Official medqa_4options, medmcqa and pubmedqa from the same pristine lm-evaluation-harness, 5-shot, loglikelihood, metric acc.

medical benchmark KOS-V5-Instruct KOS-V4-Instruct YuLan-Mini-Instruct marin-8b-instruct Qwen3-4B-Instruct-2507
params 3.7B 3.0B 2.4B 8.0B 4.0B
PubMedQA 0.7060 0.6860 0.6960 0.7500 0.7720
MedQA (USMLE, 4-option) 0.3802 0.2820 0.3511 0.4878 0.6159
MedMCQA 0.3648 0.2778 0.3856 0.4961 0.5804

KOS-V5-Instruct improves on KOS-V4-Instruct on all three (+9.8 MedQA, +8.7 MedMCQA, +2.0 PubMedQA).

PubMedQA is the standout: 0.7060, ahead of YuLan-Mini and within reach of Stanford's Marin-8B at less than half the parameters. PubMedQA tests comprehension of biomedical literature — the closest of these three to what the base was actually trained on. The USMLE-style exam MCQs (MedQA, MedMCQA) are where the answer-letter bottleneck below bites hardest.

⚠️ Why the MCQ numbers understate this model. Our own measurements show KOS models place very little probability mass on MCQ answer letters: the format, not the knowledge, is the bottleneck. A model that ranks 1 of 17 at modelling clinical text while scoring modestly on multiple-choice is exhibiting exactly that gap. Read the BPB results as the medical signal and the MCQ results as a floor, not a ceiling.

Evaluation & benchmarks

Official suites only, EleutherAI lm-evaluation-harness 0.4.12.dev0 at upstream commit c1c4bea, run from a pristine clone with stock, unmodified task definitions.

  • IFEval — stock ifeval task, 0-shot, greedy (do_sample=false, temperature=0.0), task-default max_gen_toks=1280, --apply_chat_template, seed 0. Constraint checking by the harness's vendored Google verifier (instructions_registry, 25 instruction types).
  • MMLU — stock mmlu group, official 57 subjects / 14,042 test items, 5-shot from dev (first_n), loglikelihood over A–D, metric acc (not acc_norm), no chat template.
  • Medical QA — stock medqa_4options, medmcqa, pubmedqa tasks, 5-shot, loglikelihood, metric acc, no chat template.
benchmark KOS-V5-Instruct KOS-V4-Instruct Δ
IFEval strict-avg 72.19 61.6 +10.6
IFEval prompt-strict 0.6728 0.5471 +0.126
IFEval inst-strict 0.7710 0.6655 +0.106
IFEval prompt-loose 0.6932 0.5693 +0.124
IFEval inst-loose 0.7878 0.6882 +0.100
MMLU (57-subj, 5-shot, acc) 0.4512 0.2782 +0.173

Tool / function calling — official BFCL

Measured with the official bfcl_eval suite in FC (function-calling) mode, non-live categories, the model prompted in its own native tool format and served via vLLM. Tool calling is a trained objective of this model — the second GRPO leg optimised the official BFCL AST checker directly.

BFCL (official, FC mode, non-live AST) KOS-V5-Instruct Qwen3-4B-Instruct-2507 (peer) KOS-V4-Instruct
simple (334/400) 85.00 83.20 72.75
multiple (157/200) 84.00 79.00 73.00
parallel (147/200) 80.50 73.50 60.50

KOS-V5-Instruct is above the Qwen3-4B-Instruct-2507 peer on all three BFCL categories, and far above the previous KOS-V4-Instruct. This is the axis the tool-calling GRPO leg was built to move, and it moved.

Scope. These are the non-live AST categories only (simple_python, multiple, parallel). The live, multi-turn, web-search and memory categories were not run, so no BFCL overall score is reported here — the suite's aggregate column is not meaningful when most categories are unrun.

Engine note. These BFCL numbers come from the official bfcl_eval harness on a vLLM backend, whereas the IFEval and MMLU figures on this card come from the HuggingFace backend of a pristine lm-evaluation-harness. Both are official suites; they are not the same inference stack, and that is stated rather than blurred.

Cross-harness reproduction. IFEval strict-avg measured 72.19 (pristine HF harness) and 72.0 (our RL-evaluation harness) in two independent runs — a 0.19-point agreement across two harness builds, far below the benchmark's own ±2.14-point standard error on 541 prompts, so they are the same measurement.

Harness validation. The identical pipeline scored the peer mark Qwen3-4B-Instruct-2507 at 84.71 IFEval strict-avg on the same pristine harness, and independently reproduced KOS-V4-Instruct's MMLU to four decimal places (0.2782). A score of 0.0 on this pipeline would therefore be a model property, not a harness failure.

Grounded abstention & robustness — official RGB

Measured on the official RGB harness (retrieval-augmented generation benchmark).

RGB (official) KOS-V5-Instruct Qwen3-4B-Instruct-2507 (peer)
negative rejection (declines the unanswerable) 57.33 39.0
noise robustness 64.0 93.67

Grounded abstention is a genuine strength: neg-reject 57.33 vs the peer's 39.0 — this model declines to answer unanswerable questions far more often than it invents an answer. Noise-robustness (64.0) improved over an earlier revision (58.67) but remains below the 70 threshold we treat as a pass.

IFEval in context (strict-avg)

model weights company params IFEval strict-avg
GPT-4o-mini Proprietary OpenAI 8B + 79 *
Llama-3.2-3B-Instruct Open Meta 3.2B 73.9
KOS-V5-Instruct (ours) Open Univ. of Kentucky / Louisville 3.7B 72.19
Qwen2.5-3B-Instruct Open Alibaba 3.0B 64.7
Phi-3-medium-4k-instruct Open Microsoft 14.0B 64.2
Mistral-Large Proprietary Mistral AI 46.7B + 63 *
KOS-V4-Instruct (previous gen) Open Univ. of Kentucky 3.0B 61.6
Yi-1.5-9B-Chat Open 01.AI 8.8B 60.5
Phi-3.5-mini-instruct Open Microsoft 3.8B 57.7
GPT-3.5-turbo-0613 Proprietary OpenAI 20B + 57 *
Phi-3-mini-4k-instruct Open Microsoft 3.8B 56.1
GPT-3.5-turbo-1106 Proprietary OpenAI 20B + 55.9
Mistral-7B-Instruct-v0.2 Open Mistral AI 7.2B 55.0
Llama-3.1-8B-Instruct Open Meta 8.0B 44.3
Llama-2-13b-chat Open Meta 13.0B 39.8

* strict estimate — no official IFEval strict sub-metrics published; estimated from published AVG4 or prompt-strict (loose metrics run ~2–4 pts above strict). + unofficial params.

Against university-built instruction models

Measured by us on the identical pristine harness, same protocol:

model institution params IFEval strict-avg MMLU
KOS-V5-Instruct (ours) UK / UofL 3.7B 72.19 0.4512
marin-8b-instruct Stanford 8.0B 70.83 0.6112
YuLan-Mini-Instruct Renmin 2.4B 61.51 0.5278
KOS-V4-Instruct (ours) UK 3.0B 60.63 0.2782
LLäMmlein-7B-chat Würzburg 7.0B 54.07 0.5252
Poro-34B-chat U Turku 34.2B 34.63
Minerva-7B-instruct Sapienza 7.4B 21.51 0.4071
CroissantLLMChat CentraleSupélec 1.3B 19.94 0.2401
Tucano-2b4-Instruct U Bonn 2.4B 14.95 0.2589

On instruction following KOS-V5-Instruct now places first among nine university-built instruct models, ahead of Stanford's Marin-8B (70.83) at less than half its parameters, and of Poro-34B at 9× its parameter count. Note that several of these models are non-English-first (Finnish, Italian, French, Portuguese, German) and are being measured on English benchmarks, which understates their designed capability. On parametric knowledge (MMLU) the larger, more heavily trained models still lead.

Retrieval & embeddings

Beyond generation, KOS-V5-Instruct also serves as a dense text retriever. A companion LoRA adapterKOS-V5-Retriever — converts this model into an embedding model (llm2vec-style: bidirectional attention + mean-pooling + a contrastively-trained rank-32 LoRA), with these base weights frozen and unchanged.

On the official BEIR SciFact benchmark (the beir library + pytrec_eval — the public-leaderboard scorer), zero-shot (training excluded SciFact, verified clean), it scores NDCG@10 = 0.7007 (Recall@10 0.864) — a strong dense retriever, above BM25 (0.665) and within the GTR/E5/BGE band (0.70–0.76). The adapter is hot-swappable: attach it for retrieval, detach it for generation. See the adapter card for the encode recipe and the full retrieval details.

Known regressions and limitations

These are disclosed deliberately. A high benchmark score does not make this checkpoint production-ready.

  • Tool / function calling is a trained strength — and now measured against the peer. The second GRPO leg optimised the official BFCL AST checker directly; BFCL is 85.0/84.0/80.5, above the peer and far above the previous KOS generation. Earlier internal KOS tool-calling figures are deliberately omitted: several were measured against benchmark data the model had been trained on and are recorded in our own audit as invalid.
  • Grounded abstention / fabrication IS measured — and strong. On the official RGB harness, negative-rejection is 57.33 (peer 39.0). Earlier revisions of this card stated fabrication was NOT independently measured; it now is. Noise-robustness (64.0) is still below our 70 pass bar, so retrieval-noise handling remains a known gap.
  • The peer leads on knowledge and raw instruction following. IFEval 72.19 vs the peer's 84.71; MMLU 0.4512 vs 0.7266; medical QA below the peer. This is a from-scratch 3.7B model on 235B tokens against one trained on orders of magnitude more data — strong for its scale, not state-of-the-art in absolute terms.
  • MMLU 0.4512 is above chance (0.25) but modest. This is not a knowledge model; it should not be used as a medical question-answering authority.
  • Not a medical-MCQ model. As with KOS-V4, do not benchmark or deploy it as one; read the BPB results (base card) as the medical signal and the MCQ results as a floor.

Data contamination

  • IFEval: CLEAN (verbatim). The SFT mix and the IFEval RL prompt set were exact-containment scanned against IFEval's official 541 test prompts — 0 exact containments. Exact matching cannot detect paraphrase or reformatting.
  • Tool calling (BFCL): CLEAN (verbatim). The tool-calling RL prompt pool was exact-containment scanned against 5,437 full-length official BFCL prompts — 0 exact containments.
  • MMLU / PubMedQA / MedQA / MedMCQA: UNCHECKED. Contamination against these four has not been scanned for this checkpoint. Those numbers should be read with that caveat.

Prompt / chat format (ChatML)

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Write a haiku about Kentucky. Do not use any commas.<|im_end|>
<|im_start|>assistant

Tool / function calling uses the model's native <tools></tools> schema block and <tool_call></tool_call> response format; pass your function schemas via the tokenizer's apply_chat_template(..., tools=[...]).

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Kentucky-Open-Science/KOS-V5-Instruct"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "Write a haiku about Kentucky. Do not use any commas."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

Deployment

Precision Approx. VRAM Notes
bfloat16 ~9 GB native weights (7.43 GB) + activations; a single 16 GB GPU is comfortable

Related models

  • KOS-V5-Retriever — a LoRA retrieval/embedding adapter for this model (llm2vec; official BEIR SciFact NDCG@10 0.70, zero-shot).
  • KOS-V5-Base — the from-scratch pretrained foundation this model is tuned from (3.72B, 235.2B tokens).
  • KOS-V4-Instruct — previous generation (3.0B, IFEval 61.6), the public release.
  • KOS-V4-Base — previous-generation foundation (3.015B, 180.3B tokens).

Intended use & limitations

Research use only. English only. Not for clinical, commercial, legal, or production-grade use. Outputs may be factually wrong or fabricated. This model must not be used to make or inform medical decisions.

Naming

The program is KOS (KOS-V1..V6). Earlier internal names are not used.

Downloads last month
7
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kentucky-Open-Science/KOS-V5-Instruct

Finetuned
(1)
this model
Adapters
1 model

Collection including Kentucky-Open-Science/KOS-V5-Instruct