gliner_agents_v3

Fine-tune of urchade/gliner_large-v2.1 for data-use mention extraction from document text.

Annotation workflow

Labels were produced with a locally run LLM through the passage-annotation harness, with human-in-the-loop passage review and adjudication under the project doctrine.

  • dataset: rafmacalaba/datause-agents-v3 (config gliner)
  • holdout: 2,108 passages, 1,918 spans
  • origins scored: fcv_pads_east_africa, prwp, reliefweb, umar_pads

Labels

  • NAMED_DATA — a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA — a source described in words but not named
  • VAGUE_DATA — generic data wording with no identifiable source

Entity matching is on span text, so these classes never enter the metrics above: they describe what the model was asked to separate, not what it is scored on.

Training

  • base model: urchade/gliner_large-v2.1
  • dataset: rafmacalaba/datause-agents-v3 (config gliner)
  • label mode: keep (prompt: NAMED_DATA, DESCRIPTIVE_DATA, VAGUE_DATA)
  • corpus: all
  • epochs: 5
  • learning rate: 5e-06
  • batch size: 16
  • precision: bf16

Evaluation (holdout)

thr tp fp fn precision recall f0.5 f1
0.10 1886 1933 27 0.4938 0.9859 0.5486 0.6581
0.20 1865 1426 48 0.5667 0.9749 0.6185 0.7168
0.30 1845 1090 68 0.6286 0.9645 0.6757 0.7611
0.40 1806 765 107 0.7025 0.9441 0.7403 0.8055
0.50 1734 381 179 0.8199 0.9064 0.8358 0.8610
0.60 1441 126 472 0.9196 0.7533 0.8807 0.8282
0.70 924 48 989 0.9506 0.4830 0.7964 0.6406

Best F0.5: 0.8807 (thr=0.6) Best F1: 0.8610 (thr=0.5)

Evaluation breakdown (holdout)

group examples spans thr precision recall f0.5 f1
overall 2108 1918 0.60 0.9196 0.7533 0.8807 0.8282
None 0 0 0.60 0.9196 0.7533 0.8807 0.8282
fcv_pads_east_africa 49 7 0.70 1.0000 0.2857 0.6667 0.4444
prwp 670 510 0.60 0.9313 0.7219 0.8802 0.8133
reliefweb 112 56 0.60 0.8636 0.6786 0.8190 0.7600
umar_pads 1277 1345 0.60 0.9190 0.7692 0.8846 0.8375

Per-label (overall)

label examples spans thr precision recall f0.5 f1
NAMED_DATA 2108 1360 0.60 0.8123 0.7919 0.8081 0.8019
DESCRIPTIVE_DATA 2108 548 0.60 0.7143 0.3193 0.5726 0.4414
VAGUE_DATA 2108 10 0.10 0.0000 0.0000 0.0000 0.0000
Origin breakdown (per-origin metrics)
origin examples spans thr precision recall f0.5 f1
fcv_pads_east_africa 49 7 0.70 1.0000 0.2857 0.6667 0.4444
prwp 670 510 0.60 0.9313 0.7219 0.8802 0.8133
reliefweb 112 56 0.60 0.8636 0.6786 0.8190 0.7600
umar_pads 1277 1345 0.60 0.9190 0.7692 0.8846 0.8375
Per-label details
label examples spans thr precision recall f0.5 f1
NAMED_DATA 2108 1360 0.60 0.8123 0.7919 0.8081 0.8019
DESCRIPTIVE_DATA 2108 548 0.60 0.7143 0.3193 0.5726 0.4414
VAGUE_DATA 2108 10 0.10 0.0000 0.0000 0.0000 0.0000

Where it fails (holdout, score threshold 0.60)

Entity-level counts at the best-F0.5 threshold, matched with the same Jaccard >= 0.5 rule as the tables above — so these are the spans those numbers were computed from.

tp fp fn redundant (same entity found twice)
1441 126 472 7

False positives, by how close each comes to a gold span: spurious 113 · partial 10 · near_miss 3

Misses, by span length: 1-2 tokens 158 · 3-5 tokens 242 · 6+ tokens 72

Label errors on matched spans: 194 of 1441 (13.5%) — boundary right, class wrong.

Suppression: 940 passages carry no gold span at all; the model still emitted 66 spans on them.

label gold clusters recalled recall
DESCRIPTIVE_DATA 548 313 0.5712
NAMED_DATA 1355 1125 0.8303
VAGUE_DATA 10 3 0.3
Most frequent false positives
span passages
Administrative data 2
World Bank Open Data 2
FERTIMAP 2
Crop Growth Monitoring system 2
2021 data as per Atlas methodology 2
UETCL BNG monitoring database 1
nutrition data 1
Bangladesh demographic and health survey 2007 1
WB Data Transparency Data 1
Malawi De velopmental Assessment Test 1
Most frequent misses
span passages
Human Development Index 4
Gender Inequality Index 3
GHS 2
household survey data 2
Ministry of Finance 2
UN data 2
DHS 2
Statistical Capacity Index 2
Gender Development Index 2
NEGU Statistics 2

Per-passage detail is in holdout_predictions.jsonl — one record per holdout passage with the passage text, every prediction and its score (kept marks the ones above the threshold above), and which of them matched. Re-derive any threshold from it. Gold spans are token indices; prediction spans are character offsets into text (GLiNER's units).

Downloads last month
35
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support