distilBERT-based Big-5 personality scorer

This is a sentence-transformers model finetuned from distilbert/distilbert-base-multilingual-cased on the ola-owo/big-five-personality-traits dataset. It maps sentences & paragraphs to a 5-dimensional dense vector space and can be used for feature extraction.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'DistilBertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
  (2): Dense({'in_features': 768, 'out_features': 5, 'bias': True, 'activation_function': 'torch.nn.modules.linear.Identity', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("ola-owo/distilbert-bigfive-sentence-transformer")
# Run inference
sentences = [
    'Shows little interest in exploring unconventional viewpoints or imaginative scenarios.',
    'Can be outgoing in familiar settings but also enjoys solitude.',
    'They require constant reassurance, feel persecuted by minor feedback, and oscillate between rage and hopelessness.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 5]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.9112, 0.9027],
#         [0.9112, 1.0000, 0.9922],
#         [0.9027, 0.9922, 1.0000]])

Training Details

Training Dataset

ola-owo/big-five-personality-traits

  • Dataset: ola-owo/big-five-personality-traits
  • Size: 2,250 training samples
  • Columns: sentence and label
  • Approximate statistics based on the first 100 samples:
    sentence label
    type string list
    modality text
    details
    • min: 10 tokens
    • mean: 21.44 tokens
    • max: 40 tokens
    • size: 5 elements
  • Samples:
    sentence label
    They feel exhausted after group travel, wait for others to approach them, and enjoy peaceful nature walks alone. [0.5, 0.5, 0.0, 0.5, 0.5]
    Mantiene un enfoque constante y una atención meticulosa al detalle. [0.5, 0.85, 0.5, 0.5, 0.5]
    Participates in social activities as opportunities arise but does not actively seek them out. [0.5, 0.5, 0.5, 0.5, 0.5]
  • Loss: main.MultiLabelBCEWithLogitsLoss

Evaluation Dataset

ola-owo/big-five-personality-traits

  • Dataset: ola-owo/big-five-personality-traits
  • Size: 250 evaluation samples
  • Columns: sentence and label
  • Approximate statistics based on the first 100 samples:
    sentence label
    type string list
    modality text
    details
    • min: 11 tokens
    • mean: 22.88 tokens
    • max: 37 tokens
    • size: 5 elements
  • Samples:
    sentence label
    Keeps tasks organized and well-planned. [0.5, 0.85, 0.5, 0.5, 0.5]
    Tiende a ser espontáneo y desorganizado, con poca atención a los planes a largo plazo o a los detalles. [0.5, 0.0, 0.5, 0.5, 0.5]
    Is comfortable with alone time, but also enjoys occasional socializing. [0.5, 0.5, 0.15000000000000002, 0.5, 0.5]
  • Loss: main.MultiLabelBCEWithLogitsLoss

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 64
  • per_device_eval_batch_size: 64
  • learning_rate: 0.001
  • num_train_epochs: 10
  • warmup_steps: 0.1
  • fp16: True
  • load_best_model_at_end: True
  • push_to_hub: True
  • hub_revision: main

All Hyperparameters

Click to expand
  • do_predict: False
  • prediction_loss_only: True
  • per_device_train_batch_size: 64
  • per_device_eval_batch_size: 64
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 0.001
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 10
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_ratio: None
  • warmup_steps: 0.1
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • enable_jit_checkpoint: False
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • use_cpu: False
  • seed: 42
  • data_seed: None
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: -1
  • ddp_backend: None
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch_fused
  • optim_args: None
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • push_to_hub: True
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: main
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • auto_find_batch_size: False
  • full_determinism: False
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • use_cache: False
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss Validation Loss
1.0 36 0.6944 0.7037

Training Time

  • Training: 8.6 seconds

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.5.1
  • Transformers: 5.0.0
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.13.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
Downloads last month
18
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ola-owo/distilbert-bigfive-sentence-transformer

Finetuned
(440)
this model

Dataset used to train ola-owo/distilbert-bigfive-sentence-transformer

Paper for ola-owo/distilbert-bigfive-sentence-transformer