YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Dataset Card for Custom Text Dataset
Dataset Name
Custom CNN/Daily mail dataset
Overview
Modified an English dataset containing over 300,000 unique news articles written by journalists from CNN and the Daily Mail.
Composition
Train Data:
- Articles: A manually provided article about the Palestinian Authority joining the International Criminal Court.
- Labels: A manually written summary for the corresponding article.
Test Data:
- Articles: 100 articles extracted from the original CNN/Daily Mail test set.
- Labels: 100 corresponding summaries from the original dataset.
Collection Process
The train data in this custom dataset was created manually by selecting a specific article and writing a corresponding summary. For the test data, 100 articles and their summaries were sampled from the CNN/Daily Mail dataset.
Preprocessing
For the train set, no additional preprocessing was applied as the data was manually created. The test set was selected directly from the CNN/Daily Mail dataset without modification.
How to Use
from datasets import load_from_disk
# Load the custom train and test datasets
train_dataset = load_from_disk('./results/custom_dataset/train')
test_dataset = load_from_disk('./results/custom_dataset/test')
# Example of accessing the data
print(train_dataset['train'][0]['sentence'])
print(test_dataset['test'][0]['sentence'])
Evaluation
Evaluation of this dataset could be done using text summarization models
Limitations
The training dataset contains only one example, so it is not sufficient for training a model.
Ethical Considerations
- Privacy: Ensure that the data does not contain personal information.
- Bias: Be aware of potential biases in the data.