library_name: transformers tags: - text-classification - distilbert - goodreads - mlops datasets: - ucsd-goodreads metrics: - accuracy - f1 pipeline_tag: text-classification

DistilBERT Goodreads Genre Classifier

This model is a fine-tuned version of distilbert-base-cased on the UCSD Goodreads book reviews dataset to classify books into 7 distinct genres. It was developed as part of an MLOps assignment for IIT Jodhpur.

Model Details

Model Description

  • Developed by: Er. Abhishek Kumar (IIT Jodhpur)
  • Model type: Transformer-based Text Classification (DistilBERT)
  • Language(s) (NLP): English
  • Finetuned from model: distilbert-base-cased
  • Associated Workspace: IIT Jodhpur MLOps Assignment Dashboard

Model Sources

Uses

Direct Use

This model is intended to analyze short or long-form book reviews and predict the corresponding book genre. It can be integrated into digital libraries, book recommendation engines, or cataloging systems.

Target Genres

The model classifies text into one of the following 7 labels:

  1. comics_graphic
  2. fantasy_paranormal
  3. history_biography
  4. mystery_thriller_crime
  5. poetry
  6. romance
  7. young_adult

How to Get Started with the Model

Use the code below to load the model and tokenizer for inference:

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_name = "g25ait2004/DistilBERT_Goodreads"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

review = "The world-building was absolutely breathtaking, full of dark magic, hidden ancient kingdoms, and dragons."

inputs = tokenizer(review, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    outputs = model(**inputs)
    
probs = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_class_id = outputs.logits.argmax().item()

print(f"Predicted class ID: {predicted_class_id}")

### Downstream Use [optional]

This model is well-suited for integration into digital libraries, book discovery applications, and automated content cataloging pipelines. It can be used as a backend service to automatically tag user-generated reviews with relevant genre nodes or to power downstream recommendation engines based on text sentiment and genre alignment.

### Out-of-Scope Use

* **Non-Review Text Processing:** The model is not intended to classify full-length manuscripts, legal copy, news articles, or technical code repositories. 
* **Multilingual Input:** It was fine-tuned purely on English text reviews; feeding it non-English text will result in highly unreliable performance.
* **Automated Moderation:** This model should not be used to flag or filter out explicit or harmful text, as its objective is strictly genre classification.

## Bias, Risks, and Limitations

The dataset relies heavily on self-reported, user-generated content from the UCSD Goodreads Graph, introducing self-selection bias and uneven structural syntax in text inputs. 

A significant technical limitation observed during evaluation is the model's performance variability across genres. While it exhibits strong predictive power for unique stylistic categories like **poetry (0.79 F1)** and **comics_graphic (0.81 F1)**, it experiences high confusion rates on structurally overlapping genres such as **young_adult (0.28 F1)** and **fantasy_paranormal (0.41 F1)**.

### Recommendations

Direct and downstream users should expect lower classification fidelity when analyzing books targeting young adult or genre-bending fantasy audiences. We recommend using a confidence threshold (e.g., softmax probability greater than 70%) or falling back to a human-in-the-loop validation model for these ambiguous categories.

## How to Get Started with the Model

Use the code below to quickly load the model and its tokenizer for basic inference:

```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

# Initialize model and tokenizer
model_name = "g25ait2004/DistilBERT_Goodreads"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

# Sample review text
review_text = "The world-building was absolutely breathtaking, full of dark magic and deep mystery."

# Tokenize and predict
inputs = tokenizer(review_text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    outputs = model(**inputs)

# Extract predicted class
probabilities = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_class = outputs.logits.argmax().item()
print(f"Predicted Class ID: {predicted_class}")


**BibTeX:**
```bibtex
@inproceedings{sanh2019distilbert,
  title={DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
  author={Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
  booktitle={NeurIPS EMC^2 Workshop},
  year={2019}
}

@article{lacoste2019quantifying,
  title={Quantifying the carbon emissions of machine learning},
  author={Lacoste, Alexandre and Alexandra, Luccioni and Schmidt, Victor and Dandres, Thomas},
  journal={arXiv preprint arXiv:1910.09700},
  year={2019}
}
APA:

Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. NeurIPS EMC^2 Workshop.

Lacoste, A., Alexandra, L., Schmidt, V., & Dandres, T. (2019). Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700.

Glossary [optional]
Distillation: A structural compression mechanism where a smaller student model (DistilBERT) attempts to recreate the output probability distribution vectors of a much bulkier teacher network (BERT-Base).

Macro F1-Score: The unweighted mean of individual F1-scores across all 7 classes. This treats all genre categories equally, regardless of variations in local test support records.

Mixed Precision (fp16): An optimization method where models calculate gradients inside a 16-bit float structure to maximize processing speed and lower memory usage, while storing base weights in 32-bit floats.

More Information [optional]
This repository belongs to the educational coursework sequences submitted under student assignment benchmarks for the Indian Institute of Technology Jodhpur (IIT Jodhpur) curriculum.

Model Card Authors [optional]
Er. Abhishek Kumar (M.Tech Data Science & Engineering Track Student)

Model Card Contact
For development inquiries, pipeline tracking issues, or alternative checkpoint requests, please submit an issue ticket directly inside your active GitHub project dashboard: GitHub Abhishek Repo Management.
Downloads last month
1
Safetensors
Model size
65.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for g25ait2004/DistilBERT_Goodreads