MyAwesomeModel

Selected Checkpoint

  • Selected checkpoint: step_1000
  • Selection criterion: highest eval-derived score among available checkpoints
  • Overall weighted evaluation score: 0.710

Evaluation Results

All benchmark scores are reported with three decimal places.

Category Benchmark Score
Core Reasoning Tasks Math Reasoning 0.550
Core Reasoning Tasks Logical Reasoning 0.819
Core Reasoning Tasks Common Sense 0.736
Language Understanding Reading Comprehension 0.700
Language Understanding Question Answering 0.607
Language Understanding Text Classification 0.828
Language Understanding Sentiment Analysis 0.792
Generation Tasks Code Generation 0.650
Generation Tasks Creative Writing 0.610
Generation Tasks Dialogue Generation 0.644
Generation Tasks Summarization 0.767
Specialized Capabilities Translation 0.804
Specialized Capabilities Knowledge Retrieval 0.676
Specialized Capabilities Instruction Following 0.758
Specialized Capabilities Safety Evaluation 0.739

Overall Performance Summary

MyAwesomeModel achieves an overall weighted evaluation score of 0.710 across the 15 benchmark categories listed above.

License

This model repository is licensed under the MIT License.

Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support