band-edge-detector
Detect whether benchmark scores are compressed at the tails of the model capability distribution.
What this does
A benchmark is usually calibrated around some median difficulty. Items that don't discriminate near the median are removed during construction. The result: scores have fine resolution near the median and compressed resolution at the tails. A model that is far from the median — much better or much worse — gets a score that is pulled toward the band edge.
This tool computes, for each benchmark in a (model × benchmark) matrix, a compression ratio for the low and high tails. Ratios below 0.5 indicate the benchmark cannot reliably distinguish models in that tail.
Install
pip install -e .
Quickstart
import pandas as pd
from band_edge_detector import BenchmarkMatrix, analyze
scores = pd.DataFrame({...}) # models × benchmarks
matrix = BenchmarkMatrix(scores=scores)
report = analyze(matrix, output_dir="report/")
print(report.summary())
CLI
band-edge-detector --input scores.csv --output report/
Method
For each benchmark:
- Sort models by a capability proxy (default: mean z-scored score).
- Split into quantile bins.
- Compute local discrimination
std(benchmark) / std(capability)in each bin. - Compression at a tail =
tail discrimination / median discrimination.
Capability proxies
The tool supports three proxy methods:
mean_z(default): z-score each benchmark across models, then take the mean. Simple, robust.pca: first principal component of the model × benchmark matrix. Better when benchmarks are correlated but have different scales.irt: Rasch model fit via logistic regression. Assumes a single ability parameter per model and a difficulty parameter per benchmark. Principled but requires more data.
Verdicts
tail_reliable: both tails have ratio ≥ threshold.high_tail_compressed: top models are scored with reduced resolution.low_tail_compressed: bottom models are scored with reduced resolution.both_tails_compressed: the benchmark is only reliable near the median.
Output
Running analyze() produces:
- A summary table with per-benchmark compression ratios and verdicts.
- A scatter grid showing capability proxy vs. score for each benchmark.
- A ratio heatmap highlighting low/high tail compression across benchmarks.
- A worst-case scatter focusing on the most compressed benchmark.
Limitations
- The capability proxy is itself a projection. Choose it carefully; the tool supports
mean_z,pca, andirtvariants. If the proxy is wrong, the compression ratios will be wrong. - Needs at least ~15 models with scores on the benchmark. Smaller samples produce noisy ratios.
- Does not tell you the true capability of any model. It tells you whether the benchmark is reliable at distinguishing models in the tails of the observed distribution.
- The tool does not detect compression modes other than band-edge. Other saturation patterns (e.g., mid-band compression, bimodal scoring) are not addressed.
Use cases
- Leaderboard auditing: identify which benchmarks have stopped discriminating between frontier models.
- Benchmark selection: choose benchmarks that remain reliable across the full capability range.
- Model comparison: weight models' scores by the reliability of the benchmarks that produced them.
- Research: quantify how benchmark construction choices affect measurement resolution.
Citation
If you use this tool in your research, please cite:
@software{band_edge_detector2026,
title = {band-edge-detector: Detecting Compression at the Tails of Benchmark Scores},
author = {zeechimp},
year = {2026},
url = {https://github.com/[user]/band-edge-detector}
}
License
MIT