Experiment 27: meta-llama/Meta-Llama-3-8B-Instruct quality benchmark

  • Status: completed
  • Model: meta-llama/Meta-Llama-3-8B-Instruct
  • Revision: 8afb486c1db24fe5011ec46dfbe5b5dccdb575c2
  • Candidate run: /workspace/NanoQuant/evidence/027/027-compress-and-benchmark-meta-llama-3-8b-instruct
  • Backend: dense
  • Wall time: 141.53 seconds

completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.

Protocol

  • WikiText-2: 64 samples Γ— 128 tokens, batch 8
  • WikiText token hash: sha256:c7dc8810186996593a9a8a419db6bcae8e67e520c2412722425817fc5862a77d
  • Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
  • Tokenizer hash: sha256:8aa3159f5493d3660d5f6898b3b1a88d5b1626e4ac1f7dd5b60fed8f916080df

Quality results

Benchmark Metric BF16 NanoQuant Delta Ratio
WikiText-2 perplexity ↓ 24.956664 55.331093 +30.374429 (+121.71%) 2.2171x
piqa acc_norm ↑ 0.7700 0.6800 -0.0900 0.8831x
arc_easy acc_norm ↑ 0.7550 0.4850 -0.2700 0.6424x
arc_challenge acc_norm ↑ 0.5150 0.2900 -0.2250 0.5631x
hellaswag acc_norm ↑ 0.6550 0.5500 -0.1050 0.8397x
winogrande acc ↑ 0.7150 0.6350 -0.0800 0.8881x
boolq acc ↑ 0.8300 0.7400 -0.0900 0.8916x

Runtime and memory

Model Elapsed seconds Peak CUDA bytes Peak host bytes
BF16 29.73 18,194,890,752 33,399,840,768
NanoQuant 35.35 18,803,064,832 33,399,840,768

Provenance

  • Experiment config hash: sha256:801c828f06b50662d77573ed177f639142880d06b9d4f440261babe50ab94b96
  • Launcher: experiments/027-compress-and-benchmark-meta-llama-3-8b-instruct.py
  • Candidate identity: {"config_hash":"sha256:044785fa22c02ddeb31c43cffde7172becdc19475a8e342b0a8344613a6ef2fb","model_hash":"sha256:d590dbc8c4a1851df6feb003b377003e7e4ededacc99ed35abb96844d236322d","plan_hash":"sha256-57b3f751fb9b821e18abad776ffc22f4f90d5b6d73378e90037728ca2b1cdcbb"}
  • Global tuning: {"artifact_id":"sha256-68defbc90a8e75e2a59b51d40c054cff9bc393012039641fa3365b32787509e7","artifact_type":"global-tuning-result","schema_version":1}
Downloads last month
225
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for arelath/Meta-Llama-3-8B-Instruct-nanoquant-GGUF

Quantized
(277)
this model