Instructions to use karencardiel/topic-classification-bert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use karencardiel/topic-classification-bert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="karencardiel/topic-classification-bert")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("karencardiel/topic-classification-bert") model = AutoModelForSequenceClassification.from_pretrained("karencardiel/topic-classification-bert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
BERT Topic Classification
Training Data
This model was fine-tuned on the AG News dataset (fancyzhx/ag_news) for four-class news topic classification:
- World
- Sports
- Business
- Sci/Tech
The dataset was divided into 108,000 training examples, 12,000 validation examples, and 7,600 test examples. A random seed of 42 was used.
Metrics
Test-set results:
| Metric | Score |
|---|---|
| Accuracy | 94.49% |
| Macro F1 | 94.49% |
| Loss | 0.3470 |
Intended Use
This model is intended for English news topic classification into the four AG News categories: World, Sports, Business, and Sci/Tech.
It was developed for educational purposes and experimentation with BERT adaptation methods.
Example Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "karencardiel/topic-classification-bert"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
text = "Researchers developed a new artificial intelligence system for analyzing genomic data."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
padding=True
)
with torch.no_grad():
outputs = model(**inputs)
prediction = torch.argmax(outputs.logits, dim=1).item()
label = model.config.id2label[prediction]
print("Predicted topic:", label)
Example output:
Predicted topic: Sci/Tech
Limitations
- The model is trained on English news data and may not generalize well to other domains or languages.
- It only supports the four categories present in AG News.
- Performance on real-world data may differ from the reported test-set results.
- The model is not intended for high-stakes decision-making.
References
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https://arxiv.org/abs/1810.04805
Zhang, X., Zhao, J., & LeCun, Y. (2015). Character-level Convolutional Networks for Text Classification. https://arxiv.org/abs/1509.01626
Tunstall, L., von Werra, L., & Wolf, T. Natural Language Processing with Transformers. O'Reilly Media, Chapter 2.
- Downloads last month
- 31
Model tree for karencardiel/topic-classification-bert
Base model
google-bert/bert-base-uncased