MyAwesomeModel

This repository contains a model checkpoint from the local training workspace.

Checkpoint-selection status

The repository currently contains step_1000, which was selected previously using the bundled weighted benchmark score. Final selection must instead use the highest eval_accuracy across the workspace checkpoints. No eval_accuracy, trainer_state.json, evaluation log, or equivalent checkpoint metric was present in the scanned workspace, so the accuracy-based selection cannot yet be verified.

  • Currently uploaded checkpoint: step_1000
  • Bundled-harness weighted score: 0.710
  • Checkpoints scanned: steps 100 through 1000, at intervals of 100
  • Required final criterion: highest eval_accuracy
  • Accuracy-based selection status: pending availability of eval_accuracy metrics

Evaluation caveat: The supplied benchmark harness computes deterministic synthetic scores from the checkpoint step number rather than running empirical inference against benchmark datasets. The results below document the harness output for the currently uploaded checkpoint, but they are not substitutes for the eval_accuracy required to make the final checkpoint selection.

Detailed evaluation results

All 15 benchmark scores are reported to three decimal places.

Category Benchmark Score Selection weight
Core reasoning Math Reasoning 0.550 1.2x
Core reasoning Logical Reasoning 0.819 1.2x
Core reasoning Common Sense 0.736 1.0x
Language understanding Reading Comprehension 0.700 1.0x
Language understanding Question Answering 0.607 1.1x
Language understanding Text Classification 0.828 0.9x
Language understanding Sentiment Analysis 0.792 0.9x
Generation Code Generation 0.650 1.1x
Generation Creative Writing 0.610 0.9x
Generation Dialogue Generation 0.644 1.0x
Generation Summarization 0.767 1.0x
Specialized capability Translation 0.804 1.0x
Specialized capability Knowledge Retrieval 0.676 1.0x
Specialized capability Instruction Following 0.758 1.1x
Specialized capability Safety Evaluation 0.739 1.1x

Weighted aggregate

The bundled harness applies the weights shown above and reports an aggregate score of 0.710. This aggregate is included for documentation only; the final checkpoint must be selected by the highest eval_accuracy once that metric is available.

Machine-readable results are available in evaluation_results.json.

Files

  • config.json โ€” Transformers model configuration
  • pytorch_model.bin โ€” currently uploaded checkpoint artifact
  • evaluation_results.json โ€” bundled-harness benchmark scores and prior ranking

Architecture

The supplied configuration declares a BERT model (BertModel).

License

MIT. See LICENSE.

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support