AIF Benchmarks
Evaluation datasets and results for benchmarking AIF agent performance against industry standards.
Benchmarks
| Benchmark | AIF Score | Industry Baseline |
|---|---|---|
| GAIA | 0.840 | 0.750 |
| SWE-bench | 0.782 | 0.650 |
| AgentBench | 0.820 | 0.700 |
| Protocol Conformance | 48/48 | Varies |
Contents
- Task definitions for each benchmark
- Ground truth answers and scoring rubrics
- AIF execution traces and results
- Historical performance trends
- Regression detection metadata
Usage
from datasets import load_dataset
ds = load_dataset("FrostyJay7813/aif-benchmarks")
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support