SentenceTransformer

This model was finetuned with Unsloth.

based on unsloth/Qwen3-Embedding-4B

This is a sentence-transformers model finetuned from unsloth/Qwen3-Embedding-4B on the json dataset. It maps sentences & paragraphs to a 2560-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: unsloth/Qwen3-Embedding-4B
  • Maximum Sequence Length: 512 tokens
  • Output Dimensionality: 2560 dimensions
  • Similarity Function: Cosine Similarity
  • Training Dataset:
    • json

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'PeftModelForFeatureExtraction'})
  (1): Pooling({'word_embedding_dimension': 2560, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': True, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    '错误率',
    '错误率e',
    'Hyper Text Transfer Protocol (HTTP)',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 2560]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.9531, -0.0085],
#         [ 0.9531,  1.0000, -0.0060],
#         [-0.0085, -0.0060,  1.0000]])

Training Details

Training Dataset

json

  • Dataset: json
  • Size: 73,671 training samples
  • Columns: anchor and positive
  • Approximate statistics based on the first 1000 samples:
    anchor positive
    type string string
    details
    • min: 2 tokens
    • mean: 18.47 tokens
    • max: 166 tokens
    • min: 2 tokens
    • mean: 18.35 tokens
    • max: 133 tokens
  • Samples:
    anchor positive
    SURVEILLANCE APPROACH Surveillance Approach
    运输:客机顶层总功能,涵盖执行乘客和货物运行及提供地面运动。 Aviation Transportation:航空运输,作为安全关键系统,是数据稀缺性挑战的典型场景。
    序列到序列LSTM-AE模型:Sequence-to-Sequence LSTM-AE Model,一种用于轨迹预测的深度学习功能模块,包含编码和解码两个核心处理流程。 LSTM-RNN自编码器:一种基于长短时记忆网络(LSTM)和循环神经网络(RNN)的自编码器模型,用于处理无标签数据。
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 32
  • learning_rate: 3e-05
  • max_steps: 4600
  • lr_scheduler_type: constant_with_warmup
  • warmup_ratio: 0.03
  • bf16: True
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: no
  • prediction_loss_only: True
  • per_device_train_batch_size: 32
  • per_device_eval_batch_size: 8
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 3e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 3.0
  • max_steps: 4600
  • lr_scheduler_type: constant_with_warmup
  • lr_scheduler_kwargs: None
  • warmup_ratio: 0.03
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • bf16: True
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch_fused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss
0.0434 100 0.1728
0.0868 200 0.071
0.1303 300 0.0454
0.1737 400 0.0494
0.2171 500 0.0342
0.0434 100 0.0165
0.0868 200 0.0127
0.1303 300 0.0078
0.1737 400 0.01
0.2171 500 0.0077
0.2605 600 0.0173
0.3040 700 0.028
0.3474 800 0.0208
0.3908 900 0.0253
0.4342 1000 0.0169
0.4776 1100 0.0141
0.5211 1200 0.0163
0.5645 1300 0.0168
0.6079 1400 0.0192
0.6513 1500 0.0156
0.6947 1600 0.0142
0.7382 1700 0.014
0.7816 1800 0.0117
0.8250 1900 0.0116
0.8684 2000 0.0076
0.9119 2100 0.009
0.9553 2200 0.0094
0.9987 2300 0.0114
1.0421 2400 0.0082
1.0855 2500 0.0054
1.1290 2600 0.0059
1.1724 2700 0.0071
1.2158 2800 0.0048
1.2592 2900 0.0083
1.3026 3000 0.007
1.3461 3100 0.0071
1.3895 3200 0.0095
1.4329 3300 0.0057
1.4763 3400 0.0044
1.5198 3500 0.0037
1.5632 3600 0.009
1.6066 3700 0.0055
1.6500 3800 0.0053
1.6934 3900 0.0071
1.7369 4000 0.005
1.7803 4100 0.0058
1.8237 4200 0.0065
1.8671 4300 0.0059
1.9106 4400 0.009
1.9540 4500 0.0071
1.9974 4600 0.0041

Framework Versions

  • Python: 3.12.3
  • Sentence Transformers: 5.2.0
  • Transformers: 4.57.6
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.12.0
  • Datasets: 4.3.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
Downloads last month
67
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ExceedZhang/Qwen3-Embedding-4B-0815-merged

Finetuned
(7)
this model

Papers for ExceedZhang/Qwen3-Embedding-4B-0815-merged