Indonesian E-commerce Review Sentiment Analysis

This model is a fine-tuned version of xlm-roberta-base for the task of sentiment analysis on Indonesian e-commerce product reviews.

Model Description

The model was trained on the dipawidia/ecommerce-product-reviews-sentiment dataset, which consists of product reviews. The model classifies reviews into two categories: POSITIVE and NEGATIVE.

Intended uses & limitations

This model is intended for sentiment analysis of product reviews in the Indonesian language. It is a good starting point for a Business Analyst to understand customer feedback at scale. The primary limitation is that it was trained for only one epoch, so while its performance is high, it may not be as robust as a model trained for multiple epochs.

Training and evaluation data

The model was fine-tuned using the dipawidia/ecommerce-product-reviews-sentiment dataset. The dataset's review column was used as the input, and the sentimen column was used as the label. The sentimen column was mapped to 0 for negative reviews and 1 for positive reviews.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 500
  • num_epochs: 1

Training results

Training Loss Epoch Step Validation Loss Accuracy
0.3998 1.0 1306 0.2992 0.9173

How to Use

You can use this model directly with the Hugging Face pipeline function.

from transformers import AutoModelForSequenceClassification, AutoTokenizer, pipeline

# Define the label mapping
label_map = {0: "NEGATIVE", 1: "POSITIVE"}

# Load the model directly from your profile
model = AutoModelForSequenceClassification.from_pretrained(
    "MAwaisM/results",
    num_labels=2,
    id2label=label_map
)

# Load the tokenizer
tokenizer = AutoTokenizer.from_pretrained("xlm-roberta-base")

# Create the pipeline
classifier = pipeline("sentiment-analysis", model=model, tokenizer=tokenizer)

# Test with a positive Indonesian review
text = "Pengiriman sangat cepat, saya sangat senang dengan produknya."
print(classifier(text))

# Test with a negative Indonesian review
text = "Pelayanan pelanggan sangat buruk, saya tidak akan membeli lagi."
print(classifier(text))
Downloads last month
7
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MAwaisM/results

Finetuned
(4227)
this model