YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Model Card for t5_small Summarization Model
Model Details
This model is based on the t5-small architecture, which is a small version of the T5 (Text-To-Text Transfer Transformer) model.
In this case, the model has been fine-tuned for summarization tasks,
specifically to generate summaries for long articles such as those from the CNN/DailyMail dataset.
- Model type: T5-small
- Number of parameters: 60 million
- Pretrained on: Various NLP tasks, pre-fine-tuning
Training Data
The model was fine-tuned using the CNN/DailyMail dataset (version 3.0.0). This dataset contains news articles with associated highlights that serve as summaries.
The dataset consists of:
- Training set: 287,113 samples
- Validation set: 13,368 samples
- Test set: 11,490 samples
The dataset contains long articles as inputs and corresponding highlights as target summaries.
Training Procedure
- Optimizer: AdamW optimizer with weight decay
- Learning rate: 2e-5
- Batch size: 4 per device for both training and evaluation
- Number of epochs: 1
- Evaluation strategy: The model was evaluated after every epoch using ROUGE and BLEU metrics to ensure the quality of the generated summaries.
- Hardware used: The model was trained using GPUs with mixed-precision (fp16) to accelerate training and reduce memory usage.
How to Use
To use this model for summarization, load the model and tokenizer using the transformers library from Hugging Face. Here’s an example in Python:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("path/to/saved/model")
model = AutoModelForSeq2SeqLM.from_pretrained("path/to/saved/model")
# Summarization example
summarizer = pipeline("summarization", model=model, tokenizer=tokenizer)
summary = summarizer("Input text here", max_length=150, min_length=30, num_beams=4, do_sample=False)
print(summary)
Evaluation
The model was evaluated using both ROUGE and BLEU metrics. Here are the scores on the validation dataset:
ROUGE-1: 43.25 ROUGE-2: 20.85 ROUGE-L: 39.18 BLEU-1: 31.5 BLEU-2: 21.7 BLEU-4: 10.3 These scores indicate that the model performs well at summarizing long texts while maintaining key information.
Limitations
The model is trained on news articles and may not generalize well to other domains such as legal, scientific, or conversational text. Since the model uses beam search, the generated summaries might occasionally be repetitive or overly verbose.
Ethical Considerations
Bias in data: The model was trained on news articles, which may contain inherent biases based on the sources of the data. Hallucination: As with most generative models, the T5 model may sometimes generate inaccurate or misleading information in the summary, especially if the input text contains ambiguous or conflicting information. Data privacy: This model was trained on publicly available datasets, and no private data was used in the fine-tuning process.
- Downloads last month
- 1