MyAwesomeModel

This repository contains the best checkpoint selected from the workspace evaluation sweep: checkpoints/step_1000.

Evaluation Results

The best checkpoint was selected by computing the weighted overall score across all 15 benchmark categories. The overall score for step_1000 was 0.710.

All 15 Benchmark Scores

Benchmark Score
Math Reasoning 0.550
Logical Reasoning 0.819
Common Sense 0.736
Reading Comprehension 0.700
Question Answering 0.607
Text Classification 0.828
Sentiment Analysis 0.792
Code Generation 0.650
Creative Writing 0.610
Dialogue Generation 0.644
Summarization 0.767
Translation 0.804
Knowledge Retrieval 0.676
Instruction Following 0.758
Safety Evaluation 0.739

Files

  • config.json: Hugging Face Transformers-compatible model configuration.
  • pytorch_model.bin: PyTorch checkpoint weights.
Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support