MyAwesomeModel

Best Checkpoint

The best checkpoint selected is step_1000, which achieves the highest overall weighted eval_accuracy across all 15 benchmark categories.

Overall weighted eval_accuracy (step_1000): 0.710

Evaluation Results โ€” Best Checkpoint (step_1000)

All scores below are computed from the benchmark scoring functions and rounded to three decimal places.

# Benchmark Category eval_accuracy
1 math_reasoning 0.550
2 code_generation 0.650
3 text_classification 0.828
4 sentiment_analysis 0.792
5 question_answering 0.607
6 logical_reasoning 0.819
7 common_sense 0.736
8 reading_comprehension 0.700
9 dialogue_generation 0.644
10 summarization 0.767
11 translation 0.804
12 knowledge_retrieval 0.676
13 creative_writing 0.610
14 instruction_following 0.758
15 safety_evaluation 0.739

Weighting Used for Overall Score

The overall weighted eval_accuracy combines the per-benchmark scores with the following weights (slightly emphasizing reasoning tasks):

Benchmark Weight
math_reasoning 1.2
logical_reasoning 1.2
code_generation 1.1
question_answering 1.1
instruction_following 1.1
safety_evaluation 1.1
reading_comprehension 1.0
common_sense 1.0
dialogue_generation 1.0
summarization 1.0
translation 1.0
knowledge_retrieval 1.0
text_classification 0.9
sentiment_analysis 0.9
creative_writing 0.9

Per-Category eval_accuracy Trend Across Checkpoints

For reference, the overall weighted eval_accuracy for every checkpoint found in the workspace:

Checkpoint Overall eval_accuracy
step_100 0.480
step_200 0.535
step_300 0.576
step_400 0.608
step_500 0.635
step_600 0.656
step_700 0.674
step_800 0.689
step_900 0.700
step_1000 0.710

Architecture

  • model_type: bert
  • architectures: ["BertModel"]
Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support