OLMo-2-1124-13B-TA8_1

This is a merge of pre-trained language models created using mergekit.

Merge Details

Merge Method

This model was merged using the Task Arithmetic merge method using /cluster/scratch/sharaj/OLMo-2-1124-13B as a base.

Models Merged

The following models were included in the merge:

  • /cluster/scratch/sharaj/OLMo-2-1124-13B-SFT
  • /cluster/scratch/sharaj/merges/OLMo-2-1124-13B-Reedsy-olmo-Linear

Configuration

The following YAML configuration was used to produce this model:

merge_method: task_arithmetic
base_model: /cluster/scratch/sharaj/OLMo-2-1124-13B
models:
  - model: /cluster/scratch/sharaj/OLMo-2-1124-13B-SFT
    parameters: {weight: 0.42}
  - model: /cluster/scratch/sharaj/merges/OLMo-2-1124-13B-Reedsy-olmo-Linear
    parameters: {weight: 1.00}     # sweep: 0.58 (≈TA1), 0.80, 1.00
parameters:
  normalize: false
tokenizer_source: /cluster/scratch/sharaj/merges/OLMo-2-1124-13B-Reedsy-olmo-Linear
dtype: float32
Downloads last month
10
Safetensors
Model size
14B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for sraj/OLMo-2-1124-13B-TA8_1