ce-forkulous-large

This is the main reranker model for Forkulous, a free-as-in-freedom API for turning unstructured recipe data into nutrient information.

This model accepts an unparsed freetext ingredient line and a description from USDA FoodDataCentral, and outputs a logit that represents the similarity between the two inputs.

This is not a model for sequence/token classification of single ingredient lines.

This is a Cross Encoder model finetuned from cross-encoder/ettin-reranker-150m-v1 using the sentence-transformers library. It computes scores for pairs of texts, which can be used for text reranking and semantic search.

Model Details

Model Description

Model Sources

Full Model Architecture

CrossEncoder(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'ModernBertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
  (2): Dense({'in_features': 768, 'out_features': 768, 'bias': False, 'activation_function': 'torch.nn.modules.activation.GELU', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
  (3): LayerNorm({'dimension': 768})
  (4): Dense({'in_features': 768, 'out_features': 1, 'bias': True, 'activation_function': 'torch.nn.modules.linear.Identity', 'module_input_name': 'sentence_embedding', 'module_output_name': 'scores'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import CrossEncoder

# Download from the 🤗 Hub
model = CrossEncoder("Forkulous/ce-forkulous-large")
# Get scores for pairs of inputs
pairs = [
    ['query: 2/3 cup dark brown sugar packed soft', 'document: Cherries, dark red, sweet, raw'],
    ['query: 1/4 granulated sugar I used raw', 'document: C&H Granulated White Sugar (OK1) - NFY040Y38'],
    ['query: 1 cup yellow squash diced', 'document: Squash, yellow, raw'],
    ['query: 1/4 cup onion tops thinly sliced green', 'document: Asparagus, green, whole spear, raw'],
    ['query: 24 ounces mascarpone cheese chilled', 'document: Cheese, Mexican blend'],
]
scores = model.predict(pairs)
print(scores)
# [-5.9062  1.7188  7.0938 -0.9141  1.1953]

# Or rank different texts based on similarity to a single text
ranks = model.rank(
    'query: 2/3 cup dark brown sugar packed soft',
    [
        'document: Cherries, dark red, sweet, raw',
        'document: C&H Granulated White Sugar (OK1) - NFY040Y38',
        'document: Squash, yellow, raw',
        'document: Asparagus, green, whole spear, raw',
        'document: Cheese, Mexican blend',
    ]
)
# [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]

Evaluation

Metrics

Cross Encoder Correlation

Metric Value
pearson 0.8598
spearman 0.9014

Training Details

Training Dataset

Unnamed Dataset

  • Size: 7,845 training samples
  • Columns: query, document, and score
  • Approximate statistics based on the first 100 samples:
    query document score
    type string string float
    modality text text
    details
    • min: 7 tokens
    • mean: 13.04 tokens
    • max: 30 tokens
    • min: 6 tokens
    • mean: 13.44 tokens
    • max: 27 tokens
    • min: 0.0
    • mean: 0.56
    • max: 1.0
  • Samples:
    query document score
    query: 2/3 cup potato starch or corn starch document: FLOUR, CORN, YELLOW (FINE MEAL) (ENRICHED) 0.4
    query: 2 heads butter lettuce chopped document: Lettuce, for use on a sandwich 1.0
    query: whipped cream or ice cream, for serving document: Beef with cream or white sauce 0.0
  • Loss: BinaryCrossEntropyLoss with these parameters:
    {
        "activation_fn": "torch.nn.modules.linear.Identity",
        "pos_weight": null
    }
    

Evaluation Dataset

Unnamed Dataset

  • Size: 2,055 evaluation samples
  • Columns: query, document, and score
  • Approximate statistics based on the first 100 samples:
    query document score
    type string string float
    modality text text
    details
    • min: 8 tokens
    • mean: 14.54 tokens
    • max: 37 tokens
    • min: 6 tokens
    • mean: 14.19 tokens
    • max: 34 tokens
    • min: 0.0
    • mean: 0.49
    • max: 1.0
  • Samples:
    query document score
    query: 2/3 cup dark brown sugar packed soft document: Cherries, dark red, sweet, raw 0.0
    query: 1/4 granulated sugar I used raw document: C&H Granulated White Sugar (OK1) - NFY040Y38 0.9
    query: 1 cup yellow squash diced document: Squash, yellow, raw 1.0
  • Loss: BinaryCrossEntropyLoss with these parameters:
    {
        "activation_fn": "torch.nn.modules.linear.Identity",
        "pos_weight": null
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • num_train_epochs: 4
  • learning_rate: 2e-05
  • lr_scheduler_type: cosine
  • warmup_steps: 0.1
  • weight_decay: 0.01
  • bf16: True
  • per_device_eval_batch_size: 16
  • load_best_model_at_end: True
  • dataloader_pin_memory: False

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 16
  • num_train_epochs: 4
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: cosine
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.01
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: True
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 16
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: Forkulous/ce-forkulous-large
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: False
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Click to expand
Epoch Step Training Loss Validation Loss spearman
0.0204 10 1.6208 - -
0.0407 20 1.4429 - -
0.0611 30 1.0490 - -
0.0815 40 0.5679 - -
0.1018 50 0.5497 - -
0.1222 60 0.5334 - -
0.1426 70 0.5868 - -
0.1629 80 0.4883 - -
0.1833 90 0.4745 - -
0.2037 100 0.4767 0.4467 0.8302
0.2240 110 0.4124 - -
0.2444 120 0.4810 - -
0.2648 130 0.4300 - -
0.2851 140 0.4689 - -
0.3055 150 0.3960 - -
0.3259 160 0.4221 - -
0.3462 170 0.4342 - -
0.3666 180 0.4317 - -
0.3870 190 0.4269 - -
0.4073 200 0.4022 0.4376 0.8620
0.4277 210 0.4581 - -
0.4481 220 0.4255 - -
0.4684 230 0.4297 - -
0.4888 240 0.4564 - -
0.5092 250 0.4250 - -
0.5295 260 0.3977 - -
0.5499 270 0.4122 - -
0.5703 280 0.3910 - -
0.5906 290 0.3930 - -
0.6110 300 0.4239 0.4146 0.8745
0.6314 310 0.4190 - -
0.6517 320 0.4212 - -
0.6721 330 0.4102 - -
0.6925 340 0.4513 - -
0.7128 350 0.3958 - -
0.7332 360 0.4110 - -
0.7536 370 0.3697 - -
0.7739 380 0.4526 - -
0.7943 390 0.4417 - -
0.8147 400 0.3545 0.4165 0.8824
0.8350 410 0.3983 - -
0.8554 420 0.3533 - -
0.8758 430 0.3654 - -
0.8961 440 0.3874 - -
0.9165 450 0.3816 - -
0.9369 460 0.3923 - -
0.9572 470 0.3717 - -
0.9776 480 0.4023 - -
0.9980 490 0.4285 - -
1.0183 500 0.3794 0.4080 0.8872
1.0387 510 0.3357 - -
1.0591 520 0.2863 - -
1.0794 530 0.3658 - -
1.0998 540 0.3649 - -
1.1202 550 0.4076 - -
1.1405 560 0.3354 - -
1.1609 570 0.3882 - -
1.1813 580 0.3731 - -
1.2016 590 0.3672 - -
1.2220 600 0.3987 0.4114 0.8954
1.2424 610 0.4156 - -
1.2627 620 0.3384 - -
1.2831 630 0.3951 - -
1.3035 640 0.3695 - -
1.3238 650 0.3838 - -
1.3442 660 0.3622 - -
1.3646 670 0.3820 - -
1.3849 680 0.3633 - -
1.4053 690 0.3564 - -
1.4257 700 0.3516 0.4089 0.8893
1.4460 710 0.3841 - -
1.4664 720 0.3823 - -
1.4868 730 0.3567 - -
1.5071 740 0.3644 - -
1.5275 750 0.4058 - -
1.5479 760 0.3474 - -
1.5682 770 0.3692 - -
1.5886 780 0.3836 - -
1.6090 790 0.3581 - -
1.6293 800 0.3675 0.4039 0.8947
1.6497 810 0.3366 - -
1.6701 820 0.3354 - -
1.6904 830 0.4028 - -
1.7108 840 0.3859 - -
1.7312 850 0.3046 - -
1.7515 860 0.3429 - -
1.7719 870 0.3857 - -
1.7923 880 0.3485 - -
1.8126 890 0.3832 - -
1.8330 900 0.4025 0.3989 0.8942
1.8534 910 0.3480 - -
1.8737 920 0.3625 - -
1.8941 930 0.3900 - -
1.9145 940 0.3804 - -
1.9348 950 0.3413 - -
1.9552 960 0.3600 - -
1.9756 970 0.4013 - -
1.9959 980 0.3806 - -
2.0163 990 0.3380 - -
2.0367 1000 0.3269 0.4063 0.8977
2.0570 1010 0.3237 - -
2.0774 1020 0.3287 - -
2.0978 1030 0.3342 - -
2.1181 1040 0.3151 - -
2.1385 1050 0.3219 - -
2.1589 1060 0.3601 - -
2.1792 1070 0.3548 - -
2.1996 1080 0.3342 - -
2.2200 1090 0.3918 - -
2.2403 1100 0.3419 0.3969 0.8975
2.2607 1110 0.3515 - -
2.2811 1120 0.3274 - -
2.3014 1130 0.3447 - -
2.3218 1140 0.3262 - -
2.3422 1150 0.3223 - -
2.3625 1160 0.3541 - -
2.3829 1170 0.3219 - -
2.4033 1180 0.3526 - -
2.4236 1190 0.3106 - -
2.4440 1200 0.3169 0.4025 0.8995
2.4644 1210 0.3134 - -
2.4847 1220 0.3326 - -
2.5051 1230 0.3669 - -
2.5255 1240 0.3433 - -
2.5458 1250 0.3299 - -
2.5662 1260 0.3784 - -
2.5866 1270 0.3436 - -
2.6069 1280 0.3636 - -
2.6273 1290 0.2903 - -
2.6477 1300 0.3264 0.4039 0.8999
2.6680 1310 0.3589 - -
2.6884 1320 0.3355 - -
2.7088 1330 0.3434 - -
2.7291 1340 0.3430 - -
2.7495 1350 0.3399 - -
2.7699 1360 0.3572 - -
2.7902 1370 0.3102 - -
2.8106 1380 0.3460 - -
2.8310 1390 0.4116 - -
2.8513 1400 0.3252 0.3946 0.9005
2.8717 1410 0.3421 - -
2.8921 1420 0.3196 - -
2.9124 1430 0.3131 - -
2.9328 1440 0.4006 - -
2.9532 1450 0.3267 - -
2.9735 1460 0.3391 - -
2.9939 1470 0.3545 - -
3.0143 1480 0.3362 - -
3.0346 1490 0.3144 - -
3.0550 1500 0.3422 0.3990 0.9014

Training Time

  • Training: 10.5 minutes
  • Evaluation: 3.0 minutes
  • Total: 13.6 minutes

Framework Versions

  • Python: 3.13.11
  • Sentence Transformers: 5.5.1
  • Transformers: 5.15.0
  • PyTorch: 2.14.0.dev20260708
  • Accelerate: 1.14.0
  • Datasets: 5.0.1
  • Tokenizers: 0.22.2

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
Downloads last month
28
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Forkulous/ce-forkulous-large

Finetuned
(5)
this model

Dataset used to train Forkulous/ce-forkulous-large

Paper for Forkulous/ce-forkulous-large

Evaluation results