YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Model Card for t5_small Summarization Model

Model Details

This model is based on the t5-small architecture, which is a small version of the T5 (Text-To-Text Transfer Transformer) model. In this case, the model has been fine-tuned for summarization tasks, specifically to generate summaries for long articles such as those from the CNN/DailyMail dataset.

  • Model type: T5-small
  • Number of parameters: 60 million
  • Pretrained on: Various NLP tasks, pre-fine-tuning

Training Data

The model was fine-tuned using the CNN/DailyMail dataset (version 3.0.0). This dataset contains news articles with associated highlights that serve as summaries.

The dataset consists of:

  • Training set: 287,113 samples
  • Validation set: 13,368 samples
  • Test set: 11,490 samples

The dataset contains long articles as inputs and corresponding highlights as target summaries.

Training Procedure

  • Optimizer: AdamW optimizer with weight decay
  • Learning rate: 2e-5
  • Batch size: 4 per device for both training and evaluation
  • Number of epochs: 1
  • Evaluation strategy: The model was evaluated after every epoch using ROUGE and BLEU metrics to ensure the quality of the generated summaries.
  • Hardware used: The model was trained using GPUs with mixed-precision (fp16) to accelerate training and reduce memory usage.

How to Use

To use this model for summarization, load the model and tokenizer using the transformers library from Hugging Face. Here’s an example in Python:

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained("path/to/saved/model")
model = AutoModelForSeq2SeqLM.from_pretrained("path/to/saved/model")

# Summarization example

summarizer = pipeline("summarization", model=model, tokenizer=tokenizer)
summary = summarizer("Input text here", max_length=150, min_length=30, num_beams=4, do_sample=False)
print(summary)

Evaluation

The model was evaluated using both ROUGE and BLEU metrics. Here are the scores on the validation dataset:

ROUGE-1: 43.25 ROUGE-2: 20.85 ROUGE-L: 39.18 BLEU-1: 31.5 BLEU-2: 21.7 BLEU-4: 10.3 These scores indicate that the model performs well at summarizing long texts while maintaining key information.

Limitations

The model is trained on news articles and may not generalize well to other domains such as legal, scientific, or conversational text. Since the model uses beam search, the generated summaries might occasionally be repetitive or overly verbose.

Ethical Considerations

Bias in data: The model was trained on news articles, which may contain inherent biases based on the sources of the data. Hallucination: As with most generative models, the T5 model may sometimes generate inaccurate or misleading information in the summary, especially if the input text contains ambiguous or conflicting information. Data privacy: This model was trained on publicly available datasets, and no private data was used in the fine-tuning process.

Downloads last month
1
Safetensors
Model size
60.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support