YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

πŸ“Œ Model Description

This model is a fine-tuned version of google/t5-small trained on a custom text-to-text dataset (derived from cleaned news articles). The objective was to build a baseline encoder-decoder model that can transform raw article text into simplified outputs (summaries / cleaned text).

  • Architecture: T5-small (60M parameters)
  • Task: Text-to-Text Generation (text2text-generation)
  • Framework: Hugging Face Transformers (PyTorch backend)
  • Training Logs: Tracked with Weights & Biases (W&B)

πŸ›  Intended Uses & Limitations

βœ… Intended uses

  • Text simplification
  • News article summarization
  • As a baseline for experimenting with encoder-decoder fine-tuning

⚠️ Limitations

  • Small model size β†’ may hallucinate or underperform on complex tasks
  • Not suitable for production-level summarization without further fine-tuning
  • Dataset size was limited (baseline experiment), so generalization may be weak

πŸ“Š Training and Evaluation

Dataset

The model was fine-tuned on a custom dataset of news articles scraped and cleaned in [Week 1’s pipeline](link to repo).

  • Input: Raw article text
  • Output: Processed / summarized text

πŸ‘‰ Future versions will be trained on larger, more diverse datasets.

Hyperparameters

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • gradient_accumulation_steps: 2
  • optimizer: adamw_torch_fused
  • num_epochs: 20
  • mixed_precision_training: Native AMP

Results

Epoch Validation Loss
1 9.7087
5 4.3784
10 2.3668
15 1.9624
20 1.8684
  • Final Validation Loss: 1.8684
  • Qualitative evaluation: outputs were shorter, cleaner summaries compared to baseline t5-small.

πŸš€ How to Use

from transformers import pipeline

model = "yakshithk/t5-small-baseline"
summarizer = pipeline("text2text-generation", model=model)

input_text = "The quick brown fox jumped over the lazy dog."
output = summarizer(input_text, max_length=50, do_sample=False)

print(output[0]["generated_text"])

πŸ” Future Work

  • Fine-tune larger models (flan-t5-base, flan-t5-large)
  • Train on domain-specific datasets (finance, healthcare)
  • Deploy via Hugging Face Spaces for live demo

πŸ“š Citation

If you use this model in your research or project:

@misc{yakshithk_t5small_baseline,
  author = {Yakshith K},
  title = {t5-small-baseline: Fine-tuned T5-small on custom dataset},
  year = {2025},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/yakshithk/t5-small-baseline}},
}

Downloads last month
9
Safetensors
Model size
60.5M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support