MyAwesomeModel-best

This repository contains the selected best checkpoint from the workspace scan.

Selection Criteria

The checkpoint was selected by choosing the highest eval_accuracy among discovered workspace checkpoints:

  • Selected checkpoint: checkpoints/step_1000
  • eval_accuracy: 0.828
  • Weighted overall benchmark score: 0.715

Detailed Evaluation Results

All scores below are formatted to 3 decimal places.

Category Benchmark Score
Core Reasoning Tasks Math Reasoning 0.550
Logical Reasoning 0.819
Common Sense 0.736
Language Understanding Reading Comprehension 0.700
Question Answering 0.607
Text Classification 0.828
Sentiment Analysis 0.792
Generation Tasks Code Generation 0.650
Creative Writing 0.610
Dialogue Generation 0.644
Summarization 0.767
Specialized Capabilities Translation 0.804
Knowledge Retrieval 0.676
Instruction Following 0.832
Safety Evaluation 0.739

Overall Performance Summary

The selected checkpoint achieves the highest eval_accuracy across all scanned checkpoints and includes detailed per-benchmark results for all 15 benchmarks.

Downloads last month
35
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support