YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Dataset Card for Custom Text Dataset

Dataset Name

Custom Text Dataset based on CNN/DailyMail

Overview

This dataset contains articles and corresponding summaries for training and evaluating summarization models. The articles are taken from the CNN/DailyMail dataset, and additional custom examples are included.

Composition

  • Train Set: 1 example from custom data, plus numerous articles from CNN/DailyMail.
  • Test Set: 100 examples taken from the CNN/DailyMail test set.

Collection Process

The dataset was compiled using the CNN/DailyMail dataset version 3.0.0. The articles were chosen based on their relevance to summarization tasks. Custom examples were added to provide additional training data.

Preprocessing

  • Articles were tokenized and normalized.
  • Labels were formatted to match the summarization format.
  • No special characters were retained in the summaries.

How to Use

from datasets import load_from_disk

# Load the custom dataset
train_dataset = load_from_disk('./results/custom_dataset/train')
test_dataset = load_from_disk('./results/custom_dataset/test')

Evaluation

Limitations

Ethical Considerations

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support