gliner_datause

Fine-tune of urchade/gliner_large-v2.1 for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).

Labels

  • NAMED_DATA โ€” a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA โ€” a source described in words but not named
  • VAGUE_DATA โ€” generic data wording with no identifiable source

Training

  • base model: urchade/gliner_large-v2.1
  • dataset: rafmacalaba/data-use-mentions (gliner config)
  • epochs: 5
  • learning rate: 5e-06
  • batch size: 16
  • precision: bf16

Evaluation (holdout)

thr tp fp fn precision recall f0.5 f1
0.10 7147 5212 199 0.5783 0.9729 0.6293 0.7254
0.20 7082 3928 264 0.6432 0.9641 0.6891 0.7716
0.30 7017 3234 329 0.6845 0.9552 0.7256 0.7975
0.40 6901 2549 445 0.7303 0.9394 0.7643 0.8217
0.50 6704 1874 642 0.7815 0.9126 0.8046 0.8420
0.60 6212 1196 1134 0.8386 0.8456 0.8400 0.8421
0.70 4923 645 2423 0.8842 0.6702 0.8311 0.7624

Best F0.5: 0.8400 (thr=0.6) Best F1: 0.8421 (thr=0.6)

Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support