gliner25_agents_v3

Fine-tune of fastino/gliner2.5-base-v1 (GLiNER2, boundary architecture) for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).

Annotation workflow

Labels were produced with a locally run LLM through the passage-annotation harness, with human-in-the-loop passage review and adjudication under the project doctrine. Holdout F1 is therefore agreement with the annotating agent, not owner accuracy.

  • dataset: rafmacalaba/datause-agents-v3 (config gliner2)
  • holdout: 2,108 passages, 1,916 spans
  • origins scored: fcv_pads_east_africa, prwp, reliefweb, umar_pads

Labels

  • NAMED_DATA — a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA — a source described in words but not named
  • VAGUE_DATA — generic data wording with no identifiable source

Training

  • base model: fastino/gliner2.5-base-v1
  • dataset: rafmacalaba/datause-agents-v3 (gliner2 config)
  • epochs: 5
  • encoder LR: 5e-06
  • task LR: 0.0001
  • batch size: 32
  • precision: bf16

Evaluation (holdout, label-agnostic)

2 gold mention(s) in the holdout sit inside a run the word splitter glues into one token (a URL, EM-DAT_), so no GLiNER2-family model can address them as a span; they are removed from every split before training and scoring. Listed in holdout_metrics.json.

thr tp fp fn precision recall f0.5 f1
0.10 1449 977 464 0.5973 0.7574 0.6237 0.6679
0.20 1388 606 525 0.6961 0.7256 0.7018 0.7105
0.30 1322 469 591 0.7381 0.6911 0.7282 0.7138
0.40 1254 367 659 0.7736 0.6555 0.7467 0.7097
0.50 1186 289 727 0.8041 0.6200 0.7590 0.7001
0.60 1090 212 823 0.8372 0.5698 0.7653 0.6781
0.70 961 165 952 0.8535 0.5024 0.7488 0.6324
0.80 778 82 1135 0.9047 0.4067 0.7267 0.5611
0.90 452 32 1461 0.9339 0.2363 0.5872 0.3771

Best F0.5: 0.7653 (thr=0.6) Best F1: 0.7138 (thr=0.3)

Where it fails (holdout, score threshold 0.60)

Entity-level counts at the reported threshold, matched with the same Jaccard >= 0.5 rule as the table above — these are the spans those numbers were computed from.

tp fp fn redundant (same entity found twice)
1090 212 823 165

False positives, by closeness to a gold span: near_miss 4 · partial 29 · spurious 199

Misses, by span length: 1-2 tokens 272 · 3-5 tokens 425 · 6+ tokens 126

Misses, by cause: 464 never emitted at all (proposal / abstention — no threshold moves these) · 359 emitted but scored below the cutoff (calibration).

Label errors on matched spans: 147 of 1090 — boundary right, class wrong.

Suppression: 941 passages carry no gold span at all; the model still emitted 103 spans on 105 of them.

label gold clusters recalled recall
DESCRIPTIVE_DATA 548 161 0.2938
NAMED_DATA 1355 926 0.6834
VAGUE_DATA 10 3 0.3000

Per-passage detail is in holdout_predictions.jsonl — one record per holdout passage with the passage text, every prediction with its score and label (kept marks the ones above the threshold above), and which of them matched. Re-derive any threshold from it; gold and predictions are both span text plus character offsets.

Downloads last month
49
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support