band-edge-detector

Detect whether benchmark scores are compressed at the tails of the model capability distribution.

What this does

A benchmark is usually calibrated around some median difficulty. Items that don't discriminate near the median are removed during construction. The result: scores have fine resolution near the median and compressed resolution at the tails. A model that is far from the median — much better or much worse — gets a score that is pulled toward the band edge.

This tool computes, for each benchmark in a (model × benchmark) matrix, a compression ratio for the low and high tails. Ratios below 0.5 indicate the benchmark cannot reliably distinguish models in that tail.

Install

pip install -e .

Quickstart

import pandas as pd
from band_edge_detector import BenchmarkMatrix, analyze

scores = pd.DataFrame({...})  # models × benchmarks
matrix = BenchmarkMatrix(scores=scores)
report = analyze(matrix, output_dir="report/")
print(report.summary())

CLI

band-edge-detector --input scores.csv --output report/

Method

For each benchmark:

  1. Sort models by a capability proxy (default: mean z-scored score).
  2. Split into quantile bins.
  3. Compute local discrimination std(benchmark) / std(capability) in each bin.
  4. Compression at a tail = tail discrimination / median discrimination.

Capability proxies

The tool supports three proxy methods:

  • mean_z (default): z-score each benchmark across models, then take the mean. Simple, robust.
  • pca: first principal component of the model × benchmark matrix. Better when benchmarks are correlated but have different scales.
  • irt: Rasch model fit via logistic regression. Assumes a single ability parameter per model and a difficulty parameter per benchmark. Principled but requires more data.

Verdicts

  • tail_reliable: both tails have ratio ≥ threshold.
  • high_tail_compressed: top models are scored with reduced resolution.
  • low_tail_compressed: bottom models are scored with reduced resolution.
  • both_tails_compressed: the benchmark is only reliable near the median.

Output

Running analyze() produces:

  • A summary table with per-benchmark compression ratios and verdicts.
  • A scatter grid showing capability proxy vs. score for each benchmark.
  • A ratio heatmap highlighting low/high tail compression across benchmarks.
  • A worst-case scatter focusing on the most compressed benchmark.

Limitations

  • The capability proxy is itself a projection. Choose it carefully; the tool supports mean_z, pca, and irt variants. If the proxy is wrong, the compression ratios will be wrong.
  • Needs at least ~15 models with scores on the benchmark. Smaller samples produce noisy ratios.
  • Does not tell you the true capability of any model. It tells you whether the benchmark is reliable at distinguishing models in the tails of the observed distribution.
  • The tool does not detect compression modes other than band-edge. Other saturation patterns (e.g., mid-band compression, bimodal scoring) are not addressed.

Use cases

  • Leaderboard auditing: identify which benchmarks have stopped discriminating between frontier models.
  • Benchmark selection: choose benchmarks that remain reliable across the full capability range.
  • Model comparison: weight models' scores by the reliability of the benchmarks that produced them.
  • Research: quantify how benchmark construction choices affect measurement resolution.

Citation

If you use this tool in your research, please cite:

@software{band_edge_detector2026,
  title = {band-edge-detector: Detecting Compression at the Tails of Benchmark Scores},
  author = {zeechimp},
  year = {2026},
  url = {https://github.com/[user]/band-edge-detector}
}

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support