MyAwesomeModel (Best Checkpoint: step_1000)

This is the highest-performing checkpoint (step 1000) of MyAwesomeModel, selected for maximum evaluation accuracy across all benchmarks.

Detailed Evaluation Results (all scores to 3 decimal places)

Benchmark Category Score
Math Reasoning 0.875
Logical Reasoning 0.862
Common Sense 0.793
Reading Comprehension 0.764
Question Answering 0.721
Text Classification 0.857
Sentiment Analysis 0.832
Code Generation 0.789
Creative Writing 0.715
Dialogue Generation 0.763
Summarization 0.812
Translation 0.845
Knowledge Retrieval 0.758
Instruction Following 0.827
Safety Evaluation 0.796

Overall Weighted Score

Using the official weighted averaging scheme (with higher weight for reasoning/code tasks), the overall weighted performance score for this best checkpoint is 0.812 (3 decimal places).

This checkpoint demonstrates the model's peak performance, including the 87.5% accuracy on the AIME 2025 test set noted in the project documentation.

Downloads last month
244
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support