SentenceTransformer based on BAAI/bge-m3

This is a sentence-transformers model finetuned from BAAI/bge-m3. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: BAAI/bge-m3
  • Maximum Sequence Length: 8192 tokens
  • Output Dimensionality: 1024 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'XLMRobertaModel'})
  (1): Pooling({'embedding_dimension': 1024, 'pooling_mode': 'cls', 'include_prompt': True})
  (2): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
queries = [
    'mắm chưng trứng vịt muối',
]
documents = [
    'Mắm Chưng Trứng Vịt Muối Vissan Hộp 150G | Mắm Chưng Trứng Vịt Muối Vissan hộp 150G là sự kết hợp hài hòa giữa nạc heo, mỡ heo, mắm cá sặc đậm đà và trứng vịt muối béo ngậy. Sản phẩm được bổ sung thêm hành, tỏi, ớt, tiêu tạo nên hương vị thơm ngon, hấp dẫn. Với thành phần chất lượng, an toàn cho sức khỏe, Mắm Chưng Vissan mang đến bữa ăn tiệ | Danh mục: Thực Phẩm Đóng Hộp | Thương hiệu: Việt Nam | Giá: 29,900 VNĐ',
    "Pizza 4P's Half Gà Teriyaki Hộp 150G | Thưởng thức hương vị Pizza 4P's Gà Teriyaki thơm ngon, đậm đà chuẩn Nhật Bản ngay tại nhà với hộp 150G tiện lợi. Vỏ bánh được làm từ bột mì mềm mịn, phủ lớp sốt Teriyaki đặc trưng, thịt gà mềm ngọt, phô mai béo ngậy quyện cùng rong biển và lá tía tô thanh mát. Sản phẩm được chế biến sẵn, chỉ cần vài | Danh mục: Thực Phẩm Chế Biến Sẵn | Thương hiệu: Việt Nam | Giá: 77,300 VNĐ",
    'Nước Súc Miệng No Brand Strong Coolmint Hương Bạc Hà Chai 800Ml | Nước Súc Miệng No Brand Strong Coolmint Hương Bạc Hà Chai 800Ml mang đến giải pháp chăm sóc răng miệng toàn diện, cho hơi thở thơm mát suốt cả ngày. Sản phẩm chứa các thành phần như Ethanol, Glycerin, Xylitol, Sodium Fluoride, cùng chiết xuất từ các loại thảo mộc tự nhiên như đinh hương, bạch chỉ, h | Danh mục: Chăm Sóc Cá Nhân | Thương hiệu: Hàn Quốc | Giá: 95,000 VNĐ',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 1024] [3, 1024]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.8572, 0.1578, 0.2286]])

Evaluation

Metrics

Information Retrieval

Metric Value
cosine_accuracy@1 0.4029
cosine_accuracy@3 0.5904
cosine_accuracy@5 0.6775
cosine_accuracy@10 0.7672
cosine_precision@1 0.4029
cosine_precision@3 0.1968
cosine_precision@5 0.1355
cosine_precision@10 0.0767
cosine_recall@1 0.4029
cosine_recall@3 0.5904
cosine_recall@5 0.6775
cosine_recall@10 0.7672
cosine_ndcg@10 0.5768
cosine_mrr@10 0.5166
cosine_map@100 0.5247

Training Details

Training Dataset

Unnamed Dataset

  • Size: 20,441 training samples
  • Columns: query and positive
  • Approximate statistics based on the first 100 samples:
    query positive
    type string string
    modality text text
    details
    • min: 7 tokens
    • mean: 10.26 tokens
    • max: 16 tokens
    • min: 115 tokens
    • mean: 138.31 tokens
    • max: 157 tokens
  • Samples:
    query positive
    dinh dưỡng cho trẻ từ 1 đến 10 tuổi Combo 2 Thực phẩm bổ sung dinh dưỡng dành cho trẻ từ 1 đến 10 tuổi Kid Essentials Nutritionally Complete vị vani | Thực phẩm bổ sung dinh dưỡng Kid Essentials vị vani là giải pháp dinh dưỡng y học toàn diện, được thiết kế đặc biệt cho trẻ em từ 1 đến 10 tuổi biếng ăn. Sản phẩm cung cấp đầy đủ các dưỡng chất cần thiết, hỗ trợ trẻ phát triển khỏe mạnh và cải thiện tình trạng biếng ăn. Với hương vani thơm ngon, dễ | Danh mục: Sữa bột cao cấp | Thương hiệu: Kid Essentials | Giá: 765,000 VNĐ
    thực phẩm dinh dưỡng y học ensure gold Thực phẩm dinh dưỡng y học Ensure Gold 800g (bao bì cũ 850g, giao bao bì ngẫu nhiên) | Ensure Gold 800g là thực phẩm dinh dưỡng y học tiên tiến, bổ sung HMB giúp duy trì và phát triển khối cơ, cùng hệ dưỡng chất YBG tăng cường miễn dịch hiệu quả. Sản phẩm giàu Omega-3 và 12 loại vitamin thiết yếu, hỗ trợ cải thiện sức khỏe toàn diện chỉ sau 8 tuần sử dụng. Ensure Gold là lựa chọn lý t | Danh mục: Sữa bột cao cấp | Thương hiệu: Ensure | Giá: 935,000 VNĐ
    sữa abbott grow 2+ cho bé Sữa Abbott Grow 2+ 1,6kg (trên 2 tuổi) (tên cũ: Abbott Grow 4 1,7kg, giao bao bì ngẫu nhiên) | Sữa Abbott Grow 2+ 1,6kg là lựa chọn tuyệt vời cho bé trên 2 tuổi, hỗ trợ phát triển chiều cao vượt trội. Sản phẩm được thiết kế đặc biệt với công thức tiên tiến, cung cấp đầy đủ dưỡng chất cần thiết cho sự phát triển thể chất và trí tuệ của trẻ. Với hương vị thơm ngon, dễ uống, Abbott Grow 2+ sẽ là | Danh mục: Sữa bột cao cấp | Thương hiệu: Abbott Grow | Giá: 625,000 VNĐ
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 2
  • num_train_epochs: 2
  • learning_rate: 2e-05
  • gradient_accumulation_steps: 4
  • fp16: True
  • load_best_model_at_end: True

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 2
  • num_train_epochs: 2
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 4
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss valid_cosine_ndcg@10
0.0391 100 0.0033 -
0.0783 200 0.0048 -
0.1174 300 0.0046 -
0.1565 400 0.0053 -
0.1957 500 0.0014 -
0.2348 600 0.0047 -
0.2739 700 0.0052 -
0.3131 800 0.0108 -
0.3522 900 0.0021 -
0.3914 1000 0.0019 -
0.4305 1100 0.0024 -
0.4696 1200 0.0065 -
0.5088 1300 0.0032 -
0.5479 1400 0.0013 -
0.5870 1500 0.0015 -
0.6262 1600 0.0081 -
0.6653 1700 0.0035 -
0.7044 1800 0.0031 -
0.7436 1900 0.0020 -
0.7827 2000 0.0023 -
0.8218 2100 0.0017 -
0.8610 2200 0.0012 -
0.9001 2300 0.0088 -
0.9392 2400 0.0009 -
0.9784 2500 0.0037 -
1.0 2556 - 0.5521
1.0172 2600 0.0025 -
1.0564 2700 0.0047 -
1.0955 2800 0.0012 -
1.1346 2900 0.0044 -
1.1738 3000 0.0021 -
1.2129 3100 0.0032 -
1.2520 3200 0.0054 -
1.2912 3300 0.0097 -
1.3303 3400 0.0016 -
1.3694 3500 0.0037 -
1.4086 3600 0.0008 -
1.4477 3700 0.0024 -
1.4868 3800 0.0005 -
1.5260 3900 0.0034 -
1.5651 4000 0.0004 -
1.6042 4100 0.0027 -
1.6434 4200 0.0078 -
1.6825 4300 0.0003 -
1.7217 4400 0.0008 -
1.7608 4500 0.0029 -
1.7999 4600 0.0070 -
1.8391 4700 0.0002 -
1.8782 4800 0.0007 -
1.9173 4900 0.0071 -
1.9565 5000 0.0022 -
1.9956 5100 0.0055 -
2.0 5112 - 0.5768

Training Time

  • Training: 1.5 hours

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.6.0
  • Transformers: 5.13.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.14.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
30
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for htdht/bge_m3_finetuned

Base model

BAAI/bge-m3
Finetuned
(523)
this model

Papers for htdht/bge_m3_finetuned

Evaluation results