Crab-1 benchmark β€” 30 French companies, ground truth

The evaluation set used for the numbers in the Crab-1 model card: 30 French companies with hand-verified ground truth (official website, department, region, SIREN, registry data).

File

  • eval_ground_truth_v3_enriched.json β€” array of 30 records, each with:
    • name β€” company name as given to the model
    • expected β€” ground truth: website, city, department (+ code), region, postal code, SIREN, headcount, sector
    • difficulty β€” easy / medium / hard
    • registry β€” SIREN/SIRET, NAF code, headcount code from the official French registry

How it's used

The eval harness (eval/run_eval.py in github.com/gaidar0yegor/crab-1) runs the model against these 30 companies with live web tools and scores each profile with harness/reward.py (website at registrable-domain level, location at department level, sector when ground truth exists).

⚠️ The benchmark is perishable β€” the tools hit the live web, so scores drift over time. Only compare models measured the same day.

License

CC0 β€” the data is public facts (French company registry + public websites).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support