YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Best Model Checkpoint (Step 1000)

This repository contains the best performing model checkpoint from training step 1000, selected based on the highest overall evaluation performance.

Evaluation Results

All benchmark scores are reported to three decimal places:

Benchmark Category Score
Math Reasoning 0.550
Logical Reasoning 0.819
Code Generation 0.650
Question Answering 0.607
Reading Comprehension 0.700
Common Sense 0.736
Text Classification 0.828
Sentiment Analysis 0.792
Dialogue Generation 0.644
Summarization 0.767
Translation 0.804
Knowledge Retrieval 0.676
Creative Writing 0.610
Instruction Following 0.758
Safety Evaluation 0.739

Overall Weighted Score: 0.710

The overall score is calculated using a weighted average, with higher weights assigned to reasoning and specialized capability tasks:

  • 1.2x weight: Math Reasoning, Logical Reasoning
  • 1.1x weight: Code Generation, Question Answering, Instruction Following, Safety Evaluation
  • 1.0x weight: Reading Comprehension, Common Sense, Dialogue Generation, Summarization, Translation, Knowledge Retrieval
  • 0.9x weight: Text Classification, Sentiment Analysis, Creative Writing
Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support