MyAwesomeModel Best Checkpoint

Selected checkpoint: checkpoints/step_1000

Selection basis: highest eval_accuracy across all workspace checkpoints.

Evaluation Results

Overall weighted eval_accuracy: 0.710

Benchmark eval_accuracy
math_reasoning 0.550
code_generation 0.650
text_classification 0.828
sentiment_analysis 0.792
question_answering 0.607
logical_reasoning 0.819
common_sense 0.736
reading_comprehension 0.700
dialogue_generation 0.644
summarization 0.767
translation 0.804
knowledge_retrieval 0.676
creative_writing 0.610
instruction_following 0.758
safety_evaluation 0.739
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support