SentenceTransformer based on sentence-transformers/all-distilroberta-v1

This is a sentence-transformers model finetuned from sentence-transformers/all-distilroberta-v1 on the ai-job-embedding-finetuning dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: RobertaModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("banik/distilroberta-ai-job-embeddings")
# Run inference
sentences = [
    'Data cleansing techniques, Power Query (M Language, DAX), ERP systems (JDE)',
    'requirements and objectives. Collect, cleanse, and validate data from various sources to ensure accuracy and consistency. Develop and implement data cleaning processes to identify and resolve errors, duplicates, and inconsistencies in datasets. Create and maintain data dictionaries, documentation, and metadata to facilitate data understanding and usage. Design and execute data transformation and normalization processes to prepare raw data for analysis. Design, standardize, and maintain data hierarchy for business functions within the team. Perform exploratory data analysis to identify trends, patterns, and outliers in the data. Develop and maintain automated data cleansing pipelines to streamline the data preparation process. Provide insights and recommendations to improve data quality, integrity, and usability. Stay updated on emerging trends, best practices, and technologies in data cleansing and data management. QualificationsQualifications: Bachelor’s degree required in computer science, Statistics, Mathematics, or related field. Proven experience (2 years) as a Data Analyst, Data Engineer, or similar role, with a focus on data cleansing and preparation. Competencies: Strong analytical and problem-solving skills with the ability to translate business requirements into technical solutions. Proficiency in Power Query (M Language, DAX) for data transformation and cleansing within Microsoft Excel and Power BI environments. Proficiency in SQL and data manipulation tools (e.g., Python and R). Experience with data visualization tools (e.g., Tableau, Power BI) is a plus. Experience with ERP systems, particularly JDE (JD Edwards), and familiarity with its data structures and modules for sales orders related tables. Experience working with large-scale datasets and data warehousing technologies (e.g., iSeries IBM). Attention to detail and a commitment to data accuracy and quality. Excellent communication and collaboration skills with the ability to work effectively in a team environment. Additional InformationWhy work for Cornerstone Building Brands?The US base salary range for this full-time position is $85,000 to $95,000 + medical, dental, vision benefits starting day 1 + 401k and PTO. Our salary ranges are determined by role, level, and location. Individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. (Full-time is defined as regularly working 30+ hours per week.)Our teams are at the heart of our purpose to positively contribute to the communities where we live, work and play. Full-time* team members receive** medical, dental and vision benefits starting day 1. Other benefits include PTO, paid holidays, FSA, life insurance, LTD, STD, 401k, EAP, discount programs, tuition reimbursement, training, and professional development. You can also join one of our Employee Resource Groups which help support our commitment to providing a diverse and inclusive work environment.*Full-time is defined as regularly working 30+ hours per week. **Union programs may vary depending on the collective bargaining agreement.All your information will be kept confidential according to',
    "experiences? Join us as a Remote Data Scientist and play a key role in optimizing our delivery operations. We're seeking a talented individual with expertise in SQL, MongoDB, and cloud computing services to help us analyze data, uncover insights, and improve our delivery processes.\n\nRequirements:\n- Advanced degree in Computer Science, Statistics, Mathematics, or a related field.\n- Proven experience in applying machine learning techniques to real-world problems.\n- Proficiency in programming languages such as Python, R, or Julia.\n- Strong understanding of SQL and experience with relational databases.\n- Familiarity with MongoDB and NoSQL database concepts.\n- Basic knowledge of cloud computing services, with experience in AWS, Azure, or Google Cloud Platform preferred.\n- Excellent analytical and problem-solving skills, with a keen eye for detail.\n- Outstanding communication skills and the ability to convey complex ideas effectively.\n\nPerks:\n- Exciting opportunities to work on cutting-edge projects with global impact.\n- Remote-friendly environment with flexible work hours.\n- Competitive salary and comprehensive benefits package.\n- Access to top-of-the-line tools and resources to fuel your creativity and innovation.\n- Supportive team culture that values collaboration, diversity, and personal growth.\n\nJoin Us:\nIf you're ready to make a difference in the delivery industry and be part of a dynamic team that's shaping the future of delivery services, we want to hear from you! OPT and H1B candidates are welcome to apply.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

Evaluation

Metrics

Triplet

Metric ai-job-validation ai-job-test
cosine_accuracy 1.0 0.9804

Training Details

Training Dataset

ai-job-embedding-finetuning

  • Dataset: ai-job-embedding-finetuning at f83afcb
  • Size: 810 training samples
  • Columns: query, job_description_pos, and job_description_neg
  • Approximate statistics based on the first 810 samples:
    query job_description_pos job_description_neg
    type string string string
    details
    • min: 7 tokens
    • mean: 15.0 tokens
    • max: 42 tokens
    • min: 7 tokens
    • mean: 345.84 tokens
    • max: 512 tokens
    • min: 7 tokens
    • mean: 348.67 tokens
    • max: 512 tokens
  • Samples:
    query job_description_pos job_description_neg
    Deep learning frameworks expertise, custom algorithm design, image segmentation solutions Experience required.
    Key requirements and translate them into innovative machine learning solutions.- Conduct ongoing research to stay abreast of the latest developments in machine learning, deep learning, and data science, and apply this knowledge to enhance project outcomes. Required Qualifications:- Bachelor’s or Master’s degree in Computer Science, Applied Mathematics, Engineering, or a related field.- Minimum of 12 years of experience in machine learning or data science, with a proven track record of developing custom, complex solutions.- Extensive experience with machine learning frameworks like PyTorch and TensorFlow.- Demonstrated ability in designing algorithms from the ground up, as indicated by experience with types of algorithms like Transformers, FCNN, RNN, GRU, Sentence Embedders, and Auto-Encoders, rather than plug-and-play approaches.- Strong coding skills in Python and familiarity with software engineering best practices.Preferred Skills:- Previous experience as a soft...
    experience using ETL and platforms like Snowflake. If you are a Senior data engineer who thrives in a transforming organization where an impact can be made apply today! This role is remote, but preference will be given to local candidates. This role does not support C2C or sponsorship at this time.
    Job Description:Managing the data availability, data integrity, and data migration needsManages and continually improves the technology used between campuses and software systems with regard to data files and integration needs.Provides support for any data storage and/or retrieval issues, as well as develops and maintains relevant reports for the department.This role will be responsible for how the organization plans, specifies, enables, creates, acquires, maintains, uses, archives, retrieves, controls and purges data.This position is also expected to be able to create databases, stored procedures, user-defined functions, and create data transformation processes via ETL tools such as Informa...
    Snowflake data warehousing, Python design patterns, AWS tools expertise Requirements:
    - Good communication; and problem-solving abilities- Ability to work as an individual contributor; collaborating with Global team- Strong experience with Data Warehousing- OLTP, OLAP, Dimension, Facts, Data Modeling- Expertise implementing Python design patterns (Creational, Structural and Behavioral Patterns)- Expertise in Python building data application including reading, transforming; writing data sets- Strong experience in using boto3, pandas, numpy, pyarrow, Requests, Fast API, Asyncio, Aiohttp, PyTest, OAuth 2.0, multithreading, multiprocessing, snowflake python connector; Snowpark- Experience in Python building data APIs (Web/REST APIs)- Experience with Snowflake including SQL, Pipes, Stream, Tasks, Time Travel, Data Sharing, Query Optimization- Experience with Scripting language in Snowflake including SQL Stored Procs, Java Script Stored Procedures; Python UDFs- Understanding of Snowflake Internals; experience in integration with Reporting; UI applications- Stron...
    experience:

    GS-14:

    Supervisory/Managerial Organization Leadership

    Supervises an assigned branch and its employees. The work directed involves high profile data science projects, programs, and/or initiatives within other federal agencies.Provides expert advice in the highly technical and specialized area of data science and is a key advisor to management on assigned/delegated matters related to the application of mathematics, statistical analysis, modeling/simulation, machine learning, natural language processing, and computer science from a data science perspective.Manages workforce operations, including recruitment, supervision, scheduling, development, and performance evaluations.Keeps up to date with data science developments in the private sector; seeks out best practices; and identifies and seizes opportunities for improvements in assigned data science program and project operations.


    Senior Expert in Data Science

    Recognized authority for scientific data analysis using advanc...
    Clinical Operations data analysis, eTMF system expertise, Data Transfer Agreements documentation requirements, and objectives for Clinical initiatives Technical SME for system activities for the clinical system(s), enhancements, and integration projects. Coordinates support activities across vendor(s) Systems include but are not limited to eTMF, EDC, CTMS and Analytics Interfaces with external vendors at all levels to manage the relationship and ensure the proper delivery of services Document Data Transfer Agreements for Data Exchange between BioNTech and Data Providers (CRO, Partner Organizations) Document Data Transformation logic and interact with development team to convert business logic into technical details

    What you have to offer:

    Bachelor’s or higher degree in a scientific discipline (e.g., computer science/information systems, engineering, mathematics, natural sciences, medical, or biomedical science) Extensive experience/knowledge of technologies and trends including Visualizations /Advanced Analytics Outstanding analytical skills and result orientation Ab...
    Requirements

    Typically requires 13+ years of professional experience and 6+ years of diversified leadership, planning, communication, organization, and people motivation skills (or equivalent experience).

    Critical Skills

    12+ years of experience in a technology role; proven experience in a leadership role, preferably in a large, complex organization.8+ years Data Engineering, Emerging Technology, and Platform Design experience4+ years Leading large data / technical teams – Data Engineering, Solution Architects, and Business Intelligence Engineers, encouraging a culture of innovation, collaboration, and continuous improvement.Hands-on experience building and delivering Enterprise Data SolutionsExtensive market knowledge and experience with cutting edge Data, Analytics, Data Science, ML and AI technologiesExtensive professional experience with ETL, BI & Data AnalyticsExtensive professional experience with Big Data systems, data pipelines and data processingDeep expertise in Data Archit...
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim"
    }
    

Evaluation Dataset

ai-job-embedding-finetuning

  • Dataset: ai-job-embedding-finetuning at f83afcb
  • Size: 101 evaluation samples
  • Columns: query, job_description_pos, and job_description_neg
  • Approximate statistics based on the first 101 samples:
    query job_description_pos job_description_neg
    type string string string
    details
    • min: 10 tokens
    • mean: 15.7 tokens
    • max: 29 tokens
    • min: 27 tokens
    • mean: 386.39 tokens
    • max: 512 tokens
    • min: 14 tokens
    • mean: 341.18 tokens
    • max: 512 tokens
  • Samples:
    query job_description_pos job_description_neg
    machine learning model optimization, big data technologies, cloud computing platforms requirements and deliver innovative solutionsPerform data cleaning, preprocessing, and feature engineering to improve model performanceOptimize and fine-tune machine learning models for scalability and efficiencyEvaluate and improve existing ML algorithms, frameworks, and toolkitsStay up-to-date with the latest trends and advancements in the field of machine learning
    RequirementsBachelor's degree in Computer Science, Engineering, or a related fieldStrong knowledge of machine learning algorithms and data modeling techniquesProficiency in Python and its associated libraries such as TensorFlow, PyTorch, or scikit-learnExperience with big data technologies such as Hadoop, Spark, or Apache KafkaFamiliarity with cloud computing platforms such as AWS or Google CloudExcellent problem-solving and analytical skillsStrong communication and collaboration abilitiesAbility to work effectively in a fast-paced and dynamic environment
    skills.50% of the time candidate will need to manage and guide a team of developers and the other 50% of the time will be completing the technical work (hands on). Must have previous experience with this (i.e., technical lead)Code review person. Each spring. Coders will do developing then candidate will be reviewing code and auditing the code to ensure its meeting the standard (final eye)Migrating to a data warehouse.
    Required Skills:Informatica, IICS data pipeline development experienceCloud Datawarehouse (Snowflake preferred), on-prem to cloud migration experience.Ability to perform peer SIT testing with other Cloud Data EngineersDatabase - MS SQL Server, Snowflake
    Nice to have:Medium priority: Informatica PowerCenter (high priority)Analytical reporting - Tableau / Qlik Sense / SAS / R (migrating existing reports - mostly Tableau / moving from Qlik View to Qlik Sense)Kafka, KubernetesFinance, Lease / Loan or Automotive experience is a plus.
    Candidate can expect a panel interview with...
    Senior Data Scientist GenAI NLP MLOps Kubernetes experience building enterprise level GenAI applications, designed and developed MLOps pipelines . The ideal candidate should have deep understanding of the NLP field, hands on experience in design and development of NLP models and experience in building LLM-based applications. Excellent written and verbal communication skills with the ability to collaborate effectively with domain experts and IT leadership team is key to be successful in this role. We are looking for candidates with expertise in Python, Pyspark, Pytorch, Langchain, GCP, Web development, Docker, Kubeflow etc. Key requirements and transition plan for the next generation of AI/ML enablement technology, tools, and processes to enable Walmart to efficiently improve performance with scale. Tools/Skills (hands-on experience is must):• Ability to transform designs ground up and lead innovation in system design• Deep understanding of GenAI applications and NLP field• Hands on experience in the design and development of NLP mode... skills, education, experience, and other qualifications.
    Featured Benefits:
    Medical Insurance in compliance with the ACA401(k)Sick leave in compliance with applicable state, federal, and local laws
    Description: Responsible for performing routine and ad-hoc analysis to identify actionable business insights, performance gaps and perform root cause analysis. The Data Analyst will perform in-depth research across a variety of data sources to determine current performance and identify trends and improvement opportunities. Collaborate with leadership and functional business owners as well as other personnel to understand friction points in data that cause unnecessary effort, and recommend gap closure initiatives to policy, process, and system. Qualification: Minimum of three (3) years of experience in data analytics, or working in a data analyst environment.Bachelor’s degree in Data Science, Statistics, Applied Math, Computer Science, Business, or related field of study from an accredited ...
    Energy policy analysis, regulatory impact modeling, distributed energy resource management. skills, modeling, energy data analysis, and critical thinking are required for a successful candidate. Knowledge of energy systems and distributed solar is required.

    Reporting to the Senior Manager of Government Affairs, you will work across different teams to model data to inform policy advocacy. The ability to obtain data from multiple sources, including regulatory or legislative hearings, academic articles, and reports, are fundamental to the role.

    A willingness to perform under deadlines and collaborate within an organization is required. Honesty, accountability, and integrity are a must.

    Energy Policy & Data Analyst Responsibilities

    Support Government Affairs team members with energy policy recommendations based on data modelingEvaluate relevant regulatory or legislative filings and model the impacts to Sunnova’s customers and businessAnalyze program proposals (grid services, incentives, net energy metering, fixed charges) and develop recommendations that align with Sunnova’s ...
    Experience working in AWS environment (S3, Snowflake, EC2, APIs)Skilled in coding languages (Python, SQL, Spark)Ability to thrive in a fast-paced, evolving work environment Experience with BI tools like Tableau, QuicksightPrevious experience building and executing tools to monitor and report on data quality
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim"
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • eval_strategy: steps
  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • learning_rate: 2e-05
  • num_train_epochs: 1
  • warmup_ratio: 0.1
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: steps
  • prediction_loss_only: True
  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 2e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 1
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: {}
  • warmup_ratio: 0.1
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • use_ipex: False
  • bf16: False
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • dispatch_batches: None
  • split_batches: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: False
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • eval_use_gather_object: False
  • average_tokens_across_devices: False
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional

Training Logs

Epoch Step ai-job-validation_cosine_accuracy ai-job-test_cosine_accuracy
0 0 0.9406 -
1.0 51 1.0 0.9804

Framework Versions

  • Python: 3.11.14
  • Sentence Transformers: 3.3.1
  • Transformers: 4.48.0
  • PyTorch: 2.9.0
  • Accelerate: 1.11.0
  • Datasets: 4.4.1
  • Tokenizers: 0.21.4

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
Downloads last month
21
Safetensors
Model size
82.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for banik/distilroberta-ai-job-embeddings

Finetuned
(55)
this model

Dataset used to train banik/distilroberta-ai-job-embeddings

Papers for banik/distilroberta-ai-job-embeddings

Evaluation results