YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Dataset Card for Custom Text Dataset
Dataset Name
Custom text dataset
Overview
This dataset contains text data for training abstractive summarization models. The data is collected from CNN/Daily mail dataset
Composition
- Number of records
- Train set: 1
- Test set: 100
- Fields:
sentence,labels - Size
- Train set: 24KB
- Test set: 360KB
Collection Process
The data was collected from the CNN/DailyMail dataset.
Preprocessing
- Train test dataset split
- Limit the test data size
How to Use
from datasets import load_dataset
dataset = load_dataset("path")
for ex in dataset['train']:
print(ex['sentence'], ex['labels'], sep='
')
Evaluation
This dataset is designed for evaluating summarization models. The common evaluation metrics are ROUGE, BLEU
Limitations
The dataset may contain some potentially sensitive information. Users should be aware of this when using the dataset. And the dataset is very small. So, It is hard to use this dataset for training robust models.
Ethical Considerations
- Privacy: Ensure that the data does not contain personal information.
- Bias: Be aware of potetial biases in the data.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support