YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Dataset Card for Custom Text Dataset

Dataset Name

Custom text dataset

Overview

This dataset contains text data for training abstractive summarization models. The data is collected from CNN/Daily mail dataset

Composition

  • Number of records
    • Train set: 1
    • Test set: 100
  • Fields: sentence, labels
  • Size
    • Train set: 24KB
    • Test set: 360KB

Collection Process

The data was collected from the CNN/DailyMail dataset.

Preprocessing

  • Train test dataset split
  • Limit the test data size

How to Use

from datasets import load_dataset
dataset = load_dataset("path")

for ex in dataset['train']:
    print(ex['sentence'], ex['labels'], sep='
')

Evaluation

This dataset is designed for evaluating summarization models. The common evaluation metrics are ROUGE, BLEU

Limitations

The dataset may contain some potentially sensitive information. Users should be aware of this when using the dataset. And the dataset is very small. So, It is hard to use this dataset for training robust models.

Ethical Considerations

  • Privacy: Ensure that the data does not contain personal information.
  • Bias: Be aware of potetial biases in the data.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support