arxivist

This model is a fine-tuned version of google-bert/bert-base-uncased on data collected from arxiv. See the blog post on this model for more information. This was used for my own educational purposes.

It achieves the following results on the evaluation set:

  • Loss: 0.4750
  • Roc Auc: 0.9709

Model description

The model is a BertForSequenceClassification model fine-tuned using the Hugging Face transformers library. The base model used was google-bert/bert-base-uncased.

Intended uses & limitations

This model is fine tuned to predict paper abstracts into either "Artificial Intelligence", "Information Retrieval" and "Robotics".

Training and evaluation data

See the blog post on this model for more information.

Training procedure

The model was trained for 5 epochs with a learning rate of 1e-4 and a batch size of 16 for training and 8 for evaluation. Dynamic padding was used during training.

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0001
  • train_batch_size: 16
  • eval_batch_size: 8
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 5

Training results

Training Loss Epoch Step Validation Loss Roc Auc
0.461 1.0 90 0.6019 0.9568
0.2395 2.0 180 0.2899 0.9775
0.1343 3.0 270 0.3588 0.9808
0.0481 4.0 360 0.4495 0.9771
0.0264 5.0 450 0.4750 0.9709

Framework versions

  • Transformers 4.55.0
  • Pytorch 2.6.0+cu124
  • Datasets 4.0.0
  • Tokenizers 0.21.4
Downloads last month
4
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mdh266/arxivist

Finetuned
(6996)
this model