Evaluation Summary: Glimmer-1-Base
#2
by GODELEV - opened
MEMORANDUM
TO: Glint-Research Team / Model Owner
SUBJECT: Evaluation Summary: Glimmer-1-Base
A comprehensive evaluation suite has been completed for Glint-Research/Glimmer-1-Base. The data below reflects the baseline performance across standard language modeling, reasoning, knowledge, and linguistic benchmarks.
Benchmark Evaluation Metrics
| Category | Benchmark | Metric | Score / Value | Status |
|---|---|---|---|---|
| Linguistics & Grammar | COPA | Accuracy | 61.00% | Success |
| BLiMP | Accuracy | 54.34% | Success | |
| Commonsense & Reasoning | BoolQ | Accuracy | 62.08% | Success |
| WinoGrande | Accuracy | 49.01% | Success | |
| PIQA | Normalized Accuracy | 48.04% | Success | |
| TruthfulQA MC2 | Accuracy | 46.96% | Success | |
| HellaSwag | Normalized Accuracy | 25.08% | Success | |
| SWAG | Normalized Accuracy | 24.89% | Success | |
| RACE | Accuracy | 20.86% | Success | |
| CommonsenseQA | Accuracy | 19.57% | Success | |
| Academic & Knowledge | OpenBookQA | Normalized Accuracy | 28.20% | Success |
| ARC-Easy | Normalized Accuracy | 26.47% | Success | |
| ARC-Challenge | Normalized Accuracy | 25.34% | Success | |
| SciQ | Normalized Accuracy | 23.00% | Success | |
| MMLU | Accuracy | 22.95% | Success | |
| Language Modeling | WikiText-2 (Byte) | Byte Perplexity | 14.73 | Success |
| LAMBADA | Accuracy | 0.00% | Success | |
| WikiText-2 (Word) | Word Perplexity | 1,764,896.00 | Success |
Best regards,
Akshit
Your benchmark results have already been merged π₯
Great job, also you got here quick lol
CompactAI changed discussion status to closed