YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Regression Model

A fine-tuned RoBERTa model for sequence classification with regression problem type.

Model Details

  • Architecture: RobertaForSequenceClassification
  • Model Type: RoBERTa
  • Problem Type: Regression
  • Base Model: RoBERTa
  • Transformers Version: 4.51.3

Model Specifications

  • Hidden Size: 768
  • Number of Layers: 12
  • Number of Attention Heads: 12
  • Intermediate Size: 3072
  • Max Position Embeddings: 514
  • Vocabulary Size: 50,265
  • Hidden Activation: GELU
  • Dropout: 0.1 (attention and hidden layers)

Tokenizer

  • Tokenizer Type: RobertaTokenizer
  • Model Max Length: 512 tokens
  • Special Tokens: <s>, </s>, <pad>, <unk>, <mask>

Files

  • model.safetensors - Model weights in SafeTensors format
  • config.json - Model configuration
  • tokenizer.json - Tokenizer model file
  • tokenizer_config.json - Tokenizer configuration
  • vocab.json - Vocabulary file
  • merges.txt - BPE merge rules
  • special_tokens_map.json - Special tokens mapping
  • training_args.bin - Training arguments (binary)

Usage

Installation

pip install transformers torch

Loading the Model

from transformers import RobertaForSequenceClassification, RobertaTokenizer
import torch

# Load tokenizer and model
tokenizer = RobertaTokenizer.from_pretrained("./")
model = RobertaForSequenceClassification.from_pretrained("./")

# Set model to evaluation mode
model.eval()

Inference Example

# Example text
text = "Your input text here"

# Tokenize input
inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=512,
    padding=True
)

# Get predictions
with torch.no_grad():
    outputs = model(**inputs)
    predictions = outputs.logits

# For regression, the output is a continuous value
predicted_value = predictions.item()
print(f"Predicted value: {predicted_value}")

Batch Inference

# Multiple texts
texts = ["Text 1", "Text 2", "Text 3"]

# Tokenize
inputs = tokenizer(
    texts,
    return_tensors="pt",
    truncation=True,
    max_length=512,
    padding=True
)

# Predict
with torch.no_grad():
    outputs = model(**inputs)
    predictions = outputs.logits

# Get predictions for each text
for i, text in enumerate(texts):
    print(f"Text: {text}")
    print(f"Predicted value: {predictions[i].item()}")

Notes

  • This model is configured for regression tasks, meaning it outputs continuous values rather than discrete class labels
  • The model accepts sequences up to 512 tokens in length
  • Inputs longer than 512 tokens will be truncated
  • The model uses SafeTensors format for efficient and safe model loading

Requirements

  • Python 3.7+
  • PyTorch
  • Transformers >= 4.51.3
  • safetensors (for loading model weights)

License

Please refer to the original RoBERTa model license and any additional terms specified for this fine-tuned model.

Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support