MyAwesomeModel (Best Checkpoint)

This repository hosts the best-performing checkpoint of MyAwesomeModel, selected across all training checkpoints by the highest evaluation score.

  • Selected checkpoint: step_1000
  • Weighted overall evaluation score: 0.710
  • Selection method: highest weighted overall evaluation score across all 10 checkpoints (step_100 through step_1000).

Evaluation Results

Scores below are reported for the selected step_1000 checkpoint on all 15 benchmark categories, each rounded to 3 decimal places.

Category Group Benchmark Score
Core Reasoning Tasks Math Reasoning 0.550
Logical Reasoning 0.819
Common Sense 0.736
Language Understanding Reading Comprehension 0.700
Question Answering 0.607
Text Classification 0.828
Sentiment Analysis 0.792
Generation Tasks Code Generation 0.650
Creative Writing 0.610
Dialogue Generation 0.644
Summarization 0.767
Specialized Capabilities Translation 0.804
Knowledge Retrieval 0.676
Instruction Following 0.758
Safety Evaluation 0.739
Overall (weighted) — 0.710

Per-benchmark details

Benchmark Score
Math Reasoning 0.550
Logical Reasoning 0.819
Common Sense 0.736
Reading Comprehension 0.700
Question Answering 0.607
Text Classification 0.828
Sentiment Analysis 0.792
Code Generation 0.650
Creative Writing 0.610
Dialogue Generation 0.644
Summarization 0.767
Translation 0.804
Knowledge Retrieval 0.676
Instruction Following 0.758
Safety Evaluation 0.739

How these scores are computed

Each benchmark score is a deterministic function of the training step number, evaluated through the workspace evaluation/ pipeline. The overall score is the weighted average of the 15 benchmark scores, with reasoning/code tasks carrying slightly higher weights.

License

Released under the MIT License.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support