YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Experiment 25: meta-llama/Llama-3.2-1B-Instruct quality benchmark

  • Status: completed
  • Model: meta-llama/Llama-3.2-1B-Instruct
  • Revision: 9213176726f574b556790deb65791e0c5aa438b6
  • Candidate run: D:\dev\research\NanoQuantRewrite\evidence\025\025-compress-and-benchmark-llama-3-2-1b-instruct
  • Backend: dense
  • Wall time: 50.07 seconds

completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.

Protocol

  • WikiText-2: 64 samples Γ— 128 tokens, batch 8
  • WikiText token hash: sha256:c7dc8810186996593a9a8a419db6bcae8e67e520c2412722425817fc5862a77d
  • Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
  • Tokenizer hash: sha256:5409af4b5ead403c8c413b60287460703373a37222ce25ce929035b49b81719c

Quality results

Benchmark Metric BF16 NanoQuant Delta Ratio
WikiText-2 perplexity ↓ 36.856393 116.980145 +80.123753 (+217.39%) 3.1739x
piqa acc_norm ↑ 0.7350 0.6650 -0.0700 0.9048x
arc_easy acc_norm ↑ 0.6200 0.4250 -0.1950 0.6855x
arc_challenge acc_norm ↑ 0.3300 0.2950 -0.0350 0.8939x
hellaswag acc_norm ↑ 0.6050 0.4550 -0.1500 0.7521x
winogrande acc ↑ 0.6150 0.5550 -0.0600 0.9024x
boolq acc ↑ 0.7500 0.6450 -0.1050 0.8600x

Runtime and memory

Model Elapsed seconds Peak CUDA bytes Peak host bytes
BF16 19.98 5,920,260,096 4,135,366,656
NanoQuant 18.02 6,192,889,856 4,869,566,464

Provenance

  • Experiment config hash: sha256:5be8cce6ef0ec17fcd90ecf10c7971523799070d4c61eeda245207a1ac69b319
  • Launcher: experiments/025-compress-and-benchmark-llama-3-2-1b-instruct.py
  • Candidate identity: {"config_hash":"sha256:a282bec0f20d082888d7322301dc95ef57f39a8b110317c9499aea8313acb3c4","model_hash":"sha256:0d4bbbbc32aa6ccf91325a258c47d5bf3839604149b6e081cb561c6a65e58f51","plan_hash":"sha256-3a9a46309740f9452a117627cb32f36f21f2fc9558d918aac3b601b5a0258da8"}
  • Global tuning: {"artifact_id":"sha256-5c4902d094ed2d025b63e3f24ba3b9b0ccca3e5bc0541da64bc6a3def1ca0b4a","artifact_type":"global-tuning-result","schema_version":1}
Downloads last month
42
GGUF
Model size
0.3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support