YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
π Model Description
This model is a fine-tuned version of google/t5-small trained on a custom text-to-text dataset (derived from cleaned news articles).
The objective was to build a baseline encoder-decoder model that can transform raw article text into simplified outputs (summaries / cleaned text).
- Architecture: T5-small (60M parameters)
- Task: Text-to-Text Generation (
text2text-generation) - Framework: Hugging Face Transformers (PyTorch backend)
- Training Logs: Tracked with Weights & Biases (W&B)
π Intended Uses & Limitations
β Intended uses
- Text simplification
- News article summarization
- As a baseline for experimenting with encoder-decoder fine-tuning
β οΈ Limitations
- Small model size β may hallucinate or underperform on complex tasks
- Not suitable for production-level summarization without further fine-tuning
- Dataset size was limited (baseline experiment), so generalization may be weak
π Training and Evaluation
Dataset
The model was fine-tuned on a custom dataset of news articles scraped and cleaned in [Week 1βs pipeline](link to repo).
- Input: Raw article text
- Output: Processed / summarized text
π Future versions will be trained on larger, more diverse datasets.
Hyperparameters
- learning_rate:
5e-05 - train_batch_size:
8 - eval_batch_size:
8 - gradient_accumulation_steps:
2 - optimizer:
adamw_torch_fused - num_epochs:
20 - mixed_precision_training: Native AMP
Results
| Epoch | Validation Loss |
|---|---|
| 1 | 9.7087 |
| 5 | 4.3784 |
| 10 | 2.3668 |
| 15 | 1.9624 |
| 20 | 1.8684 |
- Final Validation Loss: 1.8684
- Qualitative evaluation: outputs were shorter, cleaner summaries compared to baseline
t5-small.
π How to Use
from transformers import pipeline
model = "yakshithk/t5-small-baseline"
summarizer = pipeline("text2text-generation", model=model)
input_text = "The quick brown fox jumped over the lazy dog."
output = summarizer(input_text, max_length=50, do_sample=False)
print(output[0]["generated_text"])
π Future Work
- Fine-tune larger models (
flan-t5-base,flan-t5-large) - Train on domain-specific datasets (finance, healthcare)
- Deploy via Hugging Face Spaces for live demo
π Citation
If you use this model in your research or project:
@misc{yakshithk_t5small_baseline,
author = {Yakshith K},
title = {t5-small-baseline: Fine-tuned T5-small on custom dataset},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/yakshithk/t5-small-baseline}},
}
- Downloads last month
- 9
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support