banik/ai-job-embedding-finetuning
Viewer • Updated • 1.01k • 12
How to use banik/distilroberta-ai-job-embeddings with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("banik/distilroberta-ai-job-embeddings")
sentences = [
"Senior Data Analyst Pricing, pricing strategies, data governance, B2B & B2C analytics.",
"skills.50% of the time candidate will need to manage and guide a team of developers and the other 50% of the time will be completing the technical work (hands on). Must have previous experience with this (i.e., technical lead)Code review person. Each spring. Coders will do developing then candidate will be reviewing code and auditing the code to ensure its meeting the standard (final eye)Migrating to a data warehouse.\nRequired Skills:Informatica, IICS data pipeline development experienceCloud Datawarehouse (Snowflake preferred), on-prem to cloud migration experience.Ability to perform peer SIT testing with other Cloud Data EngineersDatabase - MS SQL Server, Snowflake\nNice to have:Medium priority: Informatica PowerCenter (high priority)Analytical reporting - Tableau / Qlik Sense / SAS / R (migrating existing reports - mostly Tableau / moving from Qlik View to Qlik Sense)Kafka, KubernetesFinance, Lease / Loan or Automotive experience is a plus.\nCandidate can expect a panel interview with the hiring manager and members of the team.Potential for 2nd interview to be scheduled\nWFH:This person will be onsite 100 percent of the time during training. If the candidate shows they are can work independently and productively, some flexibility could be offered to work from home. This is up to the hiring manager.\nEducation:Bachelor’s Degree in Information technology or like degree plus 5 years of IT work experience.\nexperienced data pipeline builder and data wrangler who enjoys optimizing data systems and building them from the ground up. During various aspects of this process, you should collaborate with co workers to ensure that your approach meets the needs of each project.To ensure success as a data engineer, you should demonstrate flexibility, creativity, and the capacity to receive and utilize constructive criticism. A formidable data engineer will demonstrate unsatiated curiosity and outstanding interpersonal skills.\nKey accountabilities of the function Leading Operations for Assigned Systems:Designing, implementing, and operating assigned cloud technology platforms as the technical expert.Leading internal and external resources in the appropriate utilization of cloud technology platforms.Executing ITSM/ITIL processes to ensure ongoing stable operations and alignment with SLAs.Steering providers in the execution of tier 2 and 3 support tasks and SLAs.Resolving escalated support issues.Performing routine maintenance, administering access and security levels.Driving System Management & Application Monitoring.Ensuring monitoring and correct operation of the assigned system.Ensuring changes to the system are made for ongoing run and support.Ensuring consolidation of emergency activities into regular maintenance.Analyzing system data (system logs, performance metrics, performance counters) to drive performance improvement.Supporting Agility & Customer Centricity.Supporting the end user with highly available systems.Participating in the support rotation.Performing other duties as assigned by management Additional skills: special skills / technical ability etc.Demonstrated experience in vendor and partner management.Technically competent with various business applications, especially Financial Management systems.Experience at working both independently and in a team-oriented, collaborative environment is essential.Must be able to build and maintain strong relationships in the business and Global IT organization.Ability to elicit cooperation from a wide variety of sources, including central IT, clients, and other departments.Strong written and oral communication skills.Strong interpersonal skills.\nQualifications:This position requires a Bachelor's Degree in Computer Science or a related technical field, and 5+ years of relevant employment experience.2+ years of work experience with ETL and Data Modeling on AWS Cloud Databases.Expert-level skills in writing and optimizing SQL.Experience operating very large data warehouses or data lakes.3+ years SQL Server.3+ years of Informatica or similar technology.Knowledge of Financial Services industry.\nPREFERRED QUALIFICATIONS:5+ years of work experience with ETL and Data Modeling on AWS Cloud Databases.Experience migrating on-premise data processing to AWS Cloud.Relevant AWS certification (AWS Certified Data Analytics, AWS Certified Database, etc.).Expertise in ETL optimization, designing, coding, and tuning big data processes using Informatica Data Management Cloud or similar technologies.Experience with building data pipelines and applications to stream and process datasets at low latencies.Show efficiency in handling data - tracking data lineage, ensuring data quality, and improving discoverability of data.Sound knowledge of data management and knows how to optimize the distribution, partitioning, and MPP of high-level data structures.Knowledge of Engineering and Operational Excellence using standard methodologies.\n HKA Enterprises is a global workforce solutions firm. If you're seeking a new career opportunity or project experience, our recruiters will work to understand your qualifications, experience, and personal goals. At HKA, we recognize the importance of matching employee goals with those of the employer. We strive to seek credibility, satisfaction, and endorsement from all of our applicants. We invite you to take time and search for your next career experience with us! HKA is an",
"Skills You BringBachelor’s or Master’s Degree in a technology related field (e.g. Engineering, Computer Science, etc.) required with 6+ years of experienceInformatica Power CenterGood experience with ETL technologiesSnaplogicStrong SQLProven data analysis skillsStrong data modeling skills doing either Dimensional or Data Vault modelsBasic AWS Experience Proven ability to deal with ambiguity and work in fast paced environmentExcellent interpersonal and communication skillsExcellent collaboration skills to work with multiple teams in the organization",
"experience, an annualized transactional volume of $140 billion in 2023, and approximately 3,200 employees located in 12+ countries, Paysafe connects businesses and consumers across 260 payment types in over 40 currencies around the world. Delivered through an integrated platform, Paysafe solutions are geared toward mobile-initiated transactions, real-time analytics and the convergence between brick-and-mortar and online payments. Further information is available at www.paysafe.com.\n\nAre you ready to make an impact? Join our team that is inspired by a unified vision and propelled by passion.\n\nPosition Summary\n\nWe are looking for a dynamic and flexible, Senior Data Analyst, Pricing to support our global Sales and Product organizations with strategic planning, analysis, and commercial pricing efforts . As a Senior Data Analyst , you will be at the frontier of building our Pricing function to drive growth through data and AI-enabled capabilities. This opportunity is high visibility for someone hungry to drive the upward trajectory of our business and be able to contribute to their efforts in the role in our success.\n\nYou will partner with Product Managers to understand their commercial needs, then prioritize and work with a cross-functional team to deliver pricing strategies and analytics-based solutions to solve and execute them. Business outcomes will include sustainable growth in both revenues and gross profit.\n\nThis role is based in Jacksonville, Florida and offers a flexible hybrid work environment with 3 days in the office and 2 days working remote during the work week.\n\nResponsibilities\n\n Build data products that power the automation and effectiveness of our pricing function, driving better quality revenues from merchants and consumers. Partner closely with pricing stakeholders (e.g., Product, Sales, Marketing) to turn raw data into actionable insights. Help ask the right questions and find the answers. Dive into complex pricing and behavioral data sets, spot trends and make interpretations. Utilize modelling and data-mining skills to find new insights and opportunities. Turn findings into plans for new data products or visions for new merchant features. Partner across merchant Product, Sales, Marketing, Development and Finance to build alignment, engagement and excitement for new products, features and initiatives. Ensure data quality and integrity by following and enforcing data governance policies, including alignment on data language. \n\n Qualifications \n\n Bachelor’s degree in a related field of study (Computer Science, Statistics, Mathematics, Engineering, etc.) required. 5+ years of experience of in-depth data analysis role, required; preferably in pricing context with B2B & B2C in a digital environment. Proven ability to visualize data intuitively, cleanly and clearly in order to make important insights simplified. Experience across large and complex datasets, including customer behavior, and transactional data. Advanced in SQL and in Python, preferred. Experience structuring and analyzing A/B tests, elasticities and interdependencies, preferred. Excellent communication and presentation skills, with the ability to explain complex data insights to non-technical audiences. \n\n Life at Paysafe: \n\nOne network. One partnership. At Paysafe, this is not only our business model; this is our mindset when it comes to our team. Being a part of Paysafe means you’ll be one of over 3,200 members of a world-class team that drives our business to new heights every day and where we are committed to your personal and professional growth.\n\nOur culture values humility, high trust & autonomy, a desire for excellence and meeting commitments, strong team cohesion, a sense of urgency, a desire to learn, pragmatically pushing boundaries, and accomplishing goals that have a direct business impact.\n\n \n\nPaysafe provides equal employment opportunities to all employees, and applicants for employment, and prohibits discrimination of any type concerning ethnicity, religion, age, sex, national origin, disability status, sexual orientation, gender identity or expression, or any other protected characteristics. This policy applies to all terms and conditions of recruitment and employment. If you need any reasonable adjustments, please let us know. We will be happy to help and look forward to hearing from you."
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model finetuned from sentence-transformers/all-distilroberta-v1 on the ai-job-embedding-finetuning dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: RobertaModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("banik/distilroberta-ai-job-embeddings")
# Run inference
sentences = [
'Data cleansing techniques, Power Query (M Language, DAX), ERP systems (JDE)',
'requirements and objectives. Collect, cleanse, and validate data from various sources to ensure accuracy and consistency. Develop and implement data cleaning processes to identify and resolve errors, duplicates, and inconsistencies in datasets. Create and maintain data dictionaries, documentation, and metadata to facilitate data understanding and usage. Design and execute data transformation and normalization processes to prepare raw data for analysis. Design, standardize, and maintain data hierarchy for business functions within the team. Perform exploratory data analysis to identify trends, patterns, and outliers in the data. Develop and maintain automated data cleansing pipelines to streamline the data preparation process. Provide insights and recommendations to improve data quality, integrity, and usability. Stay updated on emerging trends, best practices, and technologies in data cleansing and data management. QualificationsQualifications: Bachelor’s degree required in computer science, Statistics, Mathematics, or related field. Proven experience (2 years) as a Data Analyst, Data Engineer, or similar role, with a focus on data cleansing and preparation. Competencies: Strong analytical and problem-solving skills with the ability to translate business requirements into technical solutions. Proficiency in Power Query (M Language, DAX) for data transformation and cleansing within Microsoft Excel and Power BI environments. Proficiency in SQL and data manipulation tools (e.g., Python and R). Experience with data visualization tools (e.g., Tableau, Power BI) is a plus. Experience with ERP systems, particularly JDE (JD Edwards), and familiarity with its data structures and modules for sales orders related tables. Experience working with large-scale datasets and data warehousing technologies (e.g., iSeries IBM). Attention to detail and a commitment to data accuracy and quality. Excellent communication and collaboration skills with the ability to work effectively in a team environment. Additional InformationWhy work for Cornerstone Building Brands?The US base salary range for this full-time position is $85,000 to $95,000 + medical, dental, vision benefits starting day 1 + 401k and PTO. Our salary ranges are determined by role, level, and location. Individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. (Full-time is defined as regularly working 30+ hours per week.)Our teams are at the heart of our purpose to positively contribute to the communities where we live, work and play. Full-time* team members receive** medical, dental and vision benefits starting day 1. Other benefits include PTO, paid holidays, FSA, life insurance, LTD, STD, 401k, EAP, discount programs, tuition reimbursement, training, and professional development. You can also join one of our Employee Resource Groups which help support our commitment to providing a diverse and inclusive work environment.*Full-time is defined as regularly working 30+ hours per week. **Union programs may vary depending on the collective bargaining agreement.All your information will be kept confidential according to',
"experiences? Join us as a Remote Data Scientist and play a key role in optimizing our delivery operations. We're seeking a talented individual with expertise in SQL, MongoDB, and cloud computing services to help us analyze data, uncover insights, and improve our delivery processes.\n\nRequirements:\n- Advanced degree in Computer Science, Statistics, Mathematics, or a related field.\n- Proven experience in applying machine learning techniques to real-world problems.\n- Proficiency in programming languages such as Python, R, or Julia.\n- Strong understanding of SQL and experience with relational databases.\n- Familiarity with MongoDB and NoSQL database concepts.\n- Basic knowledge of cloud computing services, with experience in AWS, Azure, or Google Cloud Platform preferred.\n- Excellent analytical and problem-solving skills, with a keen eye for detail.\n- Outstanding communication skills and the ability to convey complex ideas effectively.\n\nPerks:\n- Exciting opportunities to work on cutting-edge projects with global impact.\n- Remote-friendly environment with flexible work hours.\n- Competitive salary and comprehensive benefits package.\n- Access to top-of-the-line tools and resources to fuel your creativity and innovation.\n- Supportive team culture that values collaboration, diversity, and personal growth.\n\nJoin Us:\nIf you're ready to make a difference in the delivery industry and be part of a dynamic team that's shaping the future of delivery services, we want to hear from you! OPT and H1B candidates are welcome to apply.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
ai-job-validation and ai-job-testTripletEvaluator| Metric | ai-job-validation | ai-job-test |
|---|---|---|
| cosine_accuracy | 1.0 | 0.9804 |
query, job_description_pos, and job_description_neg| query | job_description_pos | job_description_neg | |
|---|---|---|---|
| type | string | string | string |
| details |
|
|
|
| query | job_description_pos | job_description_neg |
|---|---|---|
Deep learning frameworks expertise, custom algorithm design, image segmentation solutions |
Experience required. |
experience using ETL and platforms like Snowflake. If you are a Senior data engineer who thrives in a transforming organization where an impact can be made apply today! This role is remote, but preference will be given to local candidates. This role does not support C2C or sponsorship at this time. |
Snowflake data warehousing, Python design patterns, AWS tools expertise |
Requirements: |
experience: |
Clinical Operations data analysis, eTMF system expertise, Data Transfer Agreements documentation |
requirements, and objectives for Clinical initiatives Technical SME for system activities for the clinical system(s), enhancements, and integration projects. Coordinates support activities across vendor(s) Systems include but are not limited to eTMF, EDC, CTMS and Analytics Interfaces with external vendors at all levels to manage the relationship and ensure the proper delivery of services Document Data Transfer Agreements for Data Exchange between BioNTech and Data Providers (CRO, Partner Organizations) Document Data Transformation logic and interact with development team to convert business logic into technical details |
Requirements |
MultipleNegativesRankingLoss with these parameters:{
"scale": 20.0,
"similarity_fct": "cos_sim"
}
query, job_description_pos, and job_description_neg| query | job_description_pos | job_description_neg | |
|---|---|---|---|
| type | string | string | string |
| details |
|
|
|
| query | job_description_pos | job_description_neg |
|---|---|---|
machine learning model optimization, big data technologies, cloud computing platforms |
requirements and deliver innovative solutionsPerform data cleaning, preprocessing, and feature engineering to improve model performanceOptimize and fine-tune machine learning models for scalability and efficiencyEvaluate and improve existing ML algorithms, frameworks, and toolkitsStay up-to-date with the latest trends and advancements in the field of machine learning |
skills.50% of the time candidate will need to manage and guide a team of developers and the other 50% of the time will be completing the technical work (hands on). Must have previous experience with this (i.e., technical lead)Code review person. Each spring. Coders will do developing then candidate will be reviewing code and auditing the code to ensure its meeting the standard (final eye)Migrating to a data warehouse. |
Senior Data Scientist GenAI NLP MLOps Kubernetes |
experience building enterprise level GenAI applications, designed and developed MLOps pipelines . The ideal candidate should have deep understanding of the NLP field, hands on experience in design and development of NLP models and experience in building LLM-based applications. Excellent written and verbal communication skills with the ability to collaborate effectively with domain experts and IT leadership team is key to be successful in this role. We are looking for candidates with expertise in Python, Pyspark, Pytorch, Langchain, GCP, Web development, Docker, Kubeflow etc. Key requirements and transition plan for the next generation of AI/ML enablement technology, tools, and processes to enable Walmart to efficiently improve performance with scale. Tools/Skills (hands-on experience is must):• Ability to transform designs ground up and lead innovation in system design• Deep understanding of GenAI applications and NLP field• Hands on experience in the design and development of NLP mode... |
skills, education, experience, and other qualifications. |
Energy policy analysis, regulatory impact modeling, distributed energy resource management. |
skills, modeling, energy data analysis, and critical thinking are required for a successful candidate. Knowledge of energy systems and distributed solar is required. |
Experience working in AWS environment (S3, Snowflake, EC2, APIs)Skilled in coding languages (Python, SQL, Spark)Ability to thrive in a fast-paced, evolving work environment Experience with BI tools like Tableau, QuicksightPrevious experience building and executing tools to monitor and report on data quality |
MultipleNegativesRankingLoss with these parameters:{
"scale": 20.0,
"similarity_fct": "cos_sim"
}
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16learning_rate: 2e-05num_train_epochs: 1warmup_ratio: 0.1batch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 2e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | ai-job-validation_cosine_accuracy | ai-job-test_cosine_accuracy |
|---|---|---|---|
| 0 | 0 | 0.9406 | - |
| 1.0 | 51 | 1.0 | 0.9804 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Base model
sentence-transformers/all-distilroberta-v1