YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Dataset Card for Custom Text Dataset
Dataset Name
Custom Text Dataset based on CNN/DailyMail
Overview
This dataset contains articles and corresponding summaries for training and evaluating summarization models. The articles are taken from the CNN/DailyMail dataset, and additional custom examples are included.
Composition
- Train Set: 1 example from custom data, plus numerous articles from CNN/DailyMail.
- Test Set: 100 examples taken from the CNN/DailyMail test set.
Collection Process
The dataset was compiled using the CNN/DailyMail dataset version 3.0.0. The articles were chosen based on their relevance to summarization tasks. Custom examples were added to provide additional training data.
Preprocessing
- Articles were tokenized and normalized.
- Labels were formatted to match the summarization format.
- No special characters were retained in the summaries.
How to Use
from datasets import load_from_disk
# Load the custom dataset
train_dataset = load_from_disk('./results/custom_dataset/train')
test_dataset = load_from_disk('./results/custom_dataset/test')
Evaluation
Limitations
Ethical Considerations
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support