Valid but Vacuous
Error guarantees for small medical language models keep their promise by staying silent.
Koushik Swarna Β· Invariance Labs Β· koushik@invariancelabs.org
π Paper (PDF)
A safety certificate lets a model answer only when it is confident, and promises that the error rate on its answers stays under a budget Ξ± in all but Ξ΄ = 5% of deployments. We prove three short facts, test them on 10 small open models, 4 medical benchmarks, and 40,000 simulated deployments, and find that all three hold.
| Fact | Statement | In the data |
|---|---|---|
| 1. Silence hides failure | Once the model speaks, a valid certificate promises only P(break | speaks) β€ Ξ΄ / s, and no better bound exists. | Among deployments that speak, 7.4% to 39% break the promise. |
| 2. Strict promises force silence | A check must see at least nβ = ln(1/Ξ΄) / ln(1/(1βΞ±)) β 3/Ξ± answers before it can pass. | At Ξ± = 5%, the model never speaks in 40,000 deployments. |
| 3. Passing is luck | Luck is at most β(ln(1/p) / 2n), and the rarer the pass, the bigger the luck. | New questions went worse than the check said in 68% to 98% of deployments that spoke. |
What a report should include: how often the model speaks, the failure rate when it speaks, how many questions it answers compared with a rule without a guarantee, and a matched control for any shift test.
Repository layout
ValidButVacuous/
βββ README.md β this card
βββ paper/
β βββ Valid_but_Vacuous.pdf β compiled paper
β βββ main.tex, refs.bib β LaTeX source
β βββ figs/ β all figures (PDF + PNG)
βββ vbv/ β the library
β βββ theory.py β Facts 1β3 as functions (Ξ΄/s, floor nβ, luck ceiling, g(p))
β βββ certificate.py β HoeffdingβBentkus p-value, split + fixed-sequence certificate,
β β no-guarantee / oracle reference rules
β βββ simulate.py β within-benchmark and transport deployments, Wilson CIs, summaries
βββ scripts/
β βββ score_models.py β score MCQ benchmarks with a small model (fp16 or 4-bit NF4)
β βββ run_within.py β 1,000 deployments per (model, benchmark) β pooled report
β βββ theory_tables.py β prints every closed-form number in the paper
βββ figures/
β βββ make_figures.py β rebuilds Figures 1β4 into paper/figs/
β βββ paper_data.py β counts behind every figure
β βββ data/ β per-cell points for Fig. 3c and 4d
βββ examples/synthetic_demo.py β all three facts on a CPU in about 1 second, no GPU needed
βββ tests/ β checks the code against the paper's stated numbers
Quick start
pip install -r requirements.txt
python scripts/theory_tables.py # floors 59/29/14/9/6 and 117/57/27/17/12, g(p), ceilings
python examples/synthetic_demo.py # the silence pattern on a synthetic model
python -m pytest -q tests # 11 checks against the paper
cd figures && python make_figures.py # rebuild Figures 1β4
Use the certificate on your own scores:
from vbv import Certificate, run_within, summarize
lam, info = Certificate(alpha=0.2).fit(conf_half1, correct_half1, conf_half2, correct_half2)
# lam is None β the model stays silent
report = summarize(run_within(conf, correct, n_splits=1000))
# each row: speaking_rate, broke_all, broke_when_spoke, limit_delta_over_s, luck gap, coverage
Full pipeline
- Score:
python scripts/score_models.py --model Qwen/Qwen2.5-1.5B-Instruct --questions medqa.jsonl --out scores/qwen1.5b_medqa.npz. Add--nf4for 4-bit models. Confidence is the softmax over the option letters' next-token logits, and each question is scored again with shuffled options. - Deploy:
python scripts/run_within.py --scores "scores/*.npz" --splits 1000. Each split uses 560 practice questions (two halves of 280) and 140 new ones, at Ξ± β {5, 10, 20, 30, 40}%, with Ξ΄ = 5% and G = 20 cut-offs. - Plot: update
figures/paper_data.pyand runmake_figures.py.
Models in the paper: Qwen2.5 (0.5B, 1.5B, 3B, 7B), SmolLM2-1.7B, Granite-3.1-2B, Mistral-7B-v0.3, and BioMistral-7B. The 7B models run in NF4, and Qwen 1.5B and 3B run in both fp16 and NF4. Benchmarks: MedMCQA, MedQA, MMLU (medical and biology subsets), and PubMedQA, with 700 questions each. Everything fits on one 16 GB GPU (a T4).
Citation
@article{swarna2026validvacuous,
title = {Valid but Vacuous: Error Guarantees for Small Medical Language Models
Keep Their Promise by Staying Silent},
author = {Swarna, Koushik},
year = {2026},
note = {Invariance Labs}
}
License
Code: MIT. Paper and figures: CC BY 4.0.