Crab-1 benchmark β 30 French companies, ground truth
The evaluation set used for the numbers in the Crab-1 model card: 30 French companies with hand-verified ground truth (official website, department, region, SIREN, registry data).
File
eval_ground_truth_v3_enriched.jsonβ array of 30 records, each with:nameβ company name as given to the modelexpectedβ ground truth: website, city, department (+ code), region, postal code, SIREN, headcount, sectordifficultyβ easy / medium / hardregistryβ SIREN/SIRET, NAF code, headcount code from the official French registry
How it's used
The eval harness (eval/run_eval.py in
github.com/gaidar0yegor/crab-1)
runs the model against these 30 companies with live web tools and scores each
profile with harness/reward.py (website at registrable-domain level,
location at department level, sector when ground truth exists).
β οΈ The benchmark is perishable β the tools hit the live web, so scores drift over time. Only compare models measured the same day.
License
CC0 β the data is public facts (French company registry + public websites).
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support