Columbia-NLP
/

gemma-2b-zephyr-dpo

@@ -1,201 +1,138 @@
 ---
-library_name: transformers
-tags: []
 ---
-# Model Card for Model ID
-<!-- Provide a quick summary of what the model is/does. -->
-## Model Details
-### Model Description
-<!-- Provide a longer summary of what this model is. -->
-This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
-- **Developed by:** [More Information Needed]
-- **Funded by [optional]:** [More Information Needed]
-- **Shared by [optional]:** [More Information Needed]
-- **Model type:** [More Information Needed]
-- **Language(s) (NLP):** [More Information Needed]
-- **License:** [More Information Needed]
-- **Finetuned from model [optional]:** [More Information Needed]
-### Model Sources [optional]
-<!-- Provide the basic links for the model. -->
-- **Repository:** [More Information Needed]
-- **Paper [optional]:** [More Information Needed]
-- **Demo [optional]:** [More Information Needed]
-## Uses
-<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
-### Direct Use
-<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
-[More Information Needed]
-### Downstream Use [optional]
-<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
-[More Information Needed]
-### Out-of-Scope Use
-<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
-[More Information Needed]
-## Bias, Risks, and Limitations
-<!-- This section is meant to convey both technical and sociotechnical limitations. -->
-[More Information Needed]
-### Recommendations
-<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
-Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
-## How to Get Started with the Model
-Use the code below to get started with the model.
-[More Information Needed]
-## Training Details
-### Training Data
-<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
-[More Information Needed]
-### Training Procedure
-<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
-#### Preprocessing [optional]
-[More Information Needed]
-#### Training Hyperparameters
-- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
-#### Speeds, Sizes, Times [optional]
-<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
-[More Information Needed]
-## Evaluation
-<!-- This section describes the evaluation protocols and provides the results. -->
-### Testing Data, Factors & Metrics
-#### Testing Data
-<!-- This should link to a Dataset Card if possible. -->
-[More Information Needed]
-#### Factors
-<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
-[More Information Needed]
-#### Metrics
-<!-- These are the evaluation metrics being used, ideally with a description of why. -->
-[More Information Needed]
-### Results
-[More Information Needed]
-#### Summary
-## Model Examination [optional]
-<!-- Relevant interpretability work for the model goes here -->
-[More Information Needed]
-## Environmental Impact
-<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
-Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
-- **Hardware Type:** [More Information Needed]
-- **Hours used:** [More Information Needed]
-- **Cloud Provider:** [More Information Needed]
-- **Compute Region:** [More Information Needed]
-- **Carbon Emitted:** [More Information Needed]
-## Technical Specifications [optional]
-### Model Architecture and Objective
-[More Information Needed]
-### Compute Infrastructure
-[More Information Needed]
-#### Hardware
-[More Information Needed]
-#### Software
-[More Information Needed]
-## Citation [optional]
-<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
-**BibTeX:**
-[More Information Needed]
-**APA:**
-[More Information Needed]
-## Glossary [optional]
-<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
-[More Information Needed]
-## More Information [optional]
-[More Information Needed]
-## Model Card Authors [optional]
-[More Information Needed]
-## Model Card Contact
-[More Information Needed]

 ---
+license: other
+license_name: gemma-terms-of-use
+license_link: https://ai.google.dev/gemma/terms
+base_model: google/gemma-2b
+tags:
+- alignment-handbook
+- trl
+- dpo
+- generated_from_trainer
+datasets:
+- argilla/dpo-mix-7k
+model-index:
+- name: gemma-2b-zephyr-dpo
+  results:
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: AI2 Reasoning Challenge (25-Shot)
+      type: ai2_arc
+      config: ARC-Challenge
+      split: test
+      args:
+        num_few_shot: 25
+    metrics:
+    - type: acc_norm
+      value: 52.22
+      name: normalized accuracy
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: HellaSwag (10-Shot)
+      type: hellaswag
+      split: validation
+      args:
+        num_few_shot: 10
+    metrics:
+    - type: acc_norm
+      value: 73.11
+      name: normalized accuracy
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: MMLU (5-Shot)
+      type: cais/mmlu
+      config: all
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 42.55
+      name: accuracy
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: TruthfulQA (0-shot)
+      type: truthful_qa
+      config: multiple_choice
+      split: validation
+      args:
+        num_few_shot: 0
+    metrics:
+    - type: mc2
+      value: 42.64
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: Winogrande (5-shot)
+      type: winogrande
+      config: winogrande_xl
+      split: validation
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 64.40
+      name: accuracy
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: GSM8k (5-shot)
+      type: gsm8k
+      config: main
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 19.94
+      name: accuracy
 ---
+# Model Card for Gemma 2B Zephyr SFT
+We trained the [google/gemma-2b](https://huggingface.co/google/gemma-2b) with DPO and data from `argilla/dpo-mix-7k`.
+We carefully selected the hyper-parameters to achieve the best DPO performance.
+## Model description
+- **Model type:** A 2.5B parameter GPT-like model fine-tuned on a mix of publicly available, synthetic datasets.
+- **Language(s) (NLP):** Primarily English
+- **License:** Gemma Terms of Use
+- **Finetuned from model:** [google/gemma-2b](https://huggingface.co/google/gemma-2b)
+## License
+This model has the same license as the [original Gemma model collection](https://ai.google.dev/gemma/terms)
+## OpenLLM Leaderboard Performance
+| Models                                  | Avg. | ARC   | HellaSwag | MMLU | TruthfulQA | Winogrande | GSM8k |
+|-----------------------------------------|------|-------|-----------|------|------------|------------|-------|
+| google/gemma-2b                         | 46.37| 48.38 | 71.77     | 41.77| 33.08      | 66.77      | 16.91 |
+| google/gemma-2b-it                      | 42.75| 43.94 | 62.70     | 37.65| 45.82      | 60.93      | 5.46 |
+| wandb/gemma-2b-zephyr-sft               | 47.18| 49.74 | 72.38     | 41.37| 34.42      | 66.93      | 18.27 |
+| wandb/gemma-2b-zephyr-dpo               | 46.92| 49.66 | 72.23     | 41.13| 34.47      | 66.54      | 17.51 |
+| Columbia-NLP/gemma-2b-zephyr-sft      | 48.75| 51.80 | 72.63     | 42.20| 41.96      | 63.85      | 20.09 |
+| **Columbia-NLP/gemma-2b-zephyr-dpo**        | 49.14| 52.22 | 73.11     | 42.55| 42.64      | 64.40      | 19.94 |
+## MT-Bench
+We evaluate our model with `GPT-4-0125-preview` as the judge.
+| Model                                    | Total | Coding | Extraction | Humanities | Math | Reasoning | Roleplay | STEM | Writing |
+|------------------------------------------|-------|--------|------------|------------|------|-----------|----------|------|---------|
+| google/gemma-2b-it                       | 4.71  | 2.95   | 4.35       | 6.15       | 2.90 | 3.50      | 5.60     | 5.50 | 6.70    |
+| wandb/gemma-2b-zephyr-sft                | 4.03  | 3.10   | 3.15       | 5.00       | 2.70 | 2.65      | 5.10     | 4.80 | 5.75    |
+| wandb/gemma-2b-zephyr-dpo                | 4.06  | 2.80   | 2.90       | 5.55       | 2.65 | 2.70      | 5.20     | 4.80 | 5.85    |
+| Columbia-NLP/gemma-2b-zephyr-sft     | 4.34  | 3.10   | 3.70       | 6.25       | 2.65 | 2.70      | 5.55     | 5.25 | 5.50    |
+| **Columbia-NLP/gemma-2b-zephyr-dpo**         | **4.75**  | 3.50   | 4.05       | 6.75       | 3.30 | 3.70      | 5.85     | 5.40 | 5.53    |