CodeBench

CodeBench is an evaluation framework for measuring the reliable code generation ability of AI coding agents, beyond inflated pass@k scores.

Key Findings

Finding Result
H1 β€” Category error Standard pass@k treats test-case counts as independent sample counts, producing different scores for agents with identical true reliability
H2 β€” Score inflation Current pass@5 β‰ˆ 0.96–0.97 collapses to reliability@5 β‰ˆ 0.00–0.12 when the category error is corrected
H3 β€” Proxy validity Single-rollout proxy score has low Spearman correlation with reliability@k; β‰₯5 rollouts needed for reliable ranking
H4 β€” Security gap Security-adjusted reliability@5 is meaningfully lower than reliability@5; leaderboard rank flips when security is accounted for

Metrics

  • reliability@k β€” correct operationalization of Chen et al. (2021) pass@k, using per-(task, agent) rollout counts and binary execution success
  • security_adjusted_reliability@k β€” reliability@k counting only rollouts that are both correct and produce code with no insecure patterns (eval, exec, os.system, yaml.load without Loader, pickle.loads)

Agents Evaluated

Agent Provider Model
anote-code Anthropic claude-sonnet-4-6 (Anote system prompt)
claude-code Anthropic claude-sonnet-4-6
codex OpenAI gpt-4o

Figures

Figure 1 β€” Baseline: pass@1 vs current pass@5

fig1

Figure 2 β€” H1: Category-error proof

fig2

Figure 3 β€” H2: Score inflation magnitude

fig3

Figure 4 β€” H3: Proxy vs reliability@k correlation

fig4

Figure 5 β€” H4: Security-adjusted reliability leaderboard

fig5

Datasets

Citation

@misc{codebench2026,
  title  = {CodeBench: Measuring Reliable Code Generation},
  author = {Anote AI},
  year   = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support