Text Generation
PEFT
Safetensors
Transformers
Spanish
English
lora

dparcon

This model is a fine-tuned version of gplsi/Aitana-2B-S-base on the dataset showed in my account (descriptions about fruits and vegetables). It achieves the following results on the evaluation set:

  • Loss: 1.8984

Model description

The gplsi/Aitana-2B-S-base consists on a generative language model on multilingual data (Spanish, Valencian and English). In this case, we have fine-tuning this model to answer information about certain fruits and vegetables.

Intended uses & limitations

More information needed

Training and evaluation data

We have used the following dataset

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0002
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Use OptimizerNames.PAGED_ADAMW_8BIT with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 2
  • num_epochs: 10
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss
2.1405 1.0 3 2.1311
2.0658 2.0 6 2.0559
1.9534 3.0 9 2.0014
1.8923 4.0 12 1.9646
1.8391 5.0 15 1.9432
1.7374 6.0 18 1.9290
1.7135 7.0 21 1.9172
1.6700 8.0 24 1.9076
1.6643 9.0 27 1.9014
1.6647 10.0 30 1.8984

Framework versions

  • PEFT 0.19.1
  • Transformers 5.14.1
  • Pytorch 2.11.0+cu128
  • Datasets 5.0.0
  • Tokenizers 0.22.2
Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dparra19/dparcon

Adapter
(1)
this model

Dataset used to train dparra19/dparcon