YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Reference DB System

A reusable data pipeline for building retailer-level reference datasets for Open Food Facts Canada.

Reference DB transforms retailer product data into standardized, validated, classified, grouped, and enriched product datasets while preserving source identifiers, provenance, and traceability.

The system is designed to support multiple Canadian retailers while keeping retailer-specific extraction and mapping logic isolated from the generic processing pipeline.


Overview

Retailer product data is often inconsistent across sources. Product names, identifiers, attributes, categories, nutrition information, and variants may use different structures and conventions.

Reference DB provides a structured pipeline that transforms this source data into a consistent retailer-level reference dataset.

The pipeline separates:

  • Source-specific ingestion
  • Data validation
  • Product identity
  • Classification and taxonomy
  • Product grouping
  • Variant modeling
  • Nutrition processing
  • Scores and enrichment
  • Production outputs
  • Quality validation and review

The resulting data can be used as a reference layer for Open Food Facts Canada workflows and downstream applications.


Architecture

Retailer Source
      ↓
Retailer Adapter
      ↓
Ingestion
      ↓
Phase 1 β€” Validation
      ↓
Phase 2 β€” Identity & Normalization
      ↓
Phase 3 β€” Classification + Taxonomy + Grouping
      ↓
Phase 4 β€” Variants
      ↓
Phase 5 β€” Nutrition
      ↓
Phase 6 β€” Scores & Enrichment
      ↓
Production Outputs
      ↓
Retailer Reference Dataset

The architecture is intentionally modular so that retailer-specific logic does not have to be duplicated across processing phases.


Processing Phases

Phase 1 β€” Validation

Validates the ingested product dataset against the ingestion and validation contracts.

Checks include:

  • Required fields
  • Data types
  • Product identifiers
  • Duplicate records
  • Basic structural consistency
  • Source data integrity

Phase 2 β€” Identity & Normalization

Builds a standardized representation of product identity.

This phase:

  • Normalizes identity-related attributes
  • Separates identity attributes from variant attributes
  • Preserves source identifiers
  • Generates deterministic identity information
  • Provides the foundation for product grouping

Identity and variant information are deliberately kept separate so that different variants of the same product can remain within the same product group.


Phase 3 β€” Classification, Taxonomy & Grouping

Phase 3 combines three related but distinct concepts:

Classification

Assigns product-level classifications based on product information and domain rules.

Reference Taxonomy

Maps products to the Reference DB taxonomy rather than relying entirely on retailer-specific categories.

Product Grouping

Groups records representing the same underlying product identity.

The grouping pipeline uses candidate generation, matching signals, identity information, and explicit decision logic.

Ambiguous cases can be isolated for review rather than being silently merged.


Phase 4 β€” Variants

Models product variants within established product groups.

Examples of variant-level differences include:

  • Size
  • Quantity
  • Flavor
  • Format
  • Packaging
  • Other product-specific attributes

Variants remain linked to their product group while preserving their own identifiers and attributes.


Phase 5 β€” Nutrition

Standardizes and validates available nutrition information and links nutrition data to products.

Missing nutrition information is treated as missing data, not as zero.

Nutrition processing also includes validation to prevent invalid or suspicious source values from silently becoming production data.


Phase 6 β€” Scores & Enrichment

Calculates available product scores and enrichment information.

Depending on the available source data, this phase can include:

  • Nutri-Score
  • Agribalyse references
  • Fruit and vegetable-related enrichment
  • Score eligibility
  • Other derived product attributes

Only products satisfying the required eligibility conditions receive the corresponding score or enrichment.


Production Outputs

The pipeline produces production-oriented datasets rather than exposing processing artifacts as the final interface.

Current production outputs include:

production_metadata
production_groups
production_variants
production_nutrition
production_scores

These outputs provide a stable representation of the processed Reference DB data while keeping intermediate processing artifacts separate.


Project Structure

reference-db-system/
β”‚
β”œβ”€β”€ contracts/
β”‚   β”œβ”€β”€ adapters/
β”‚   β”œβ”€β”€ ingestion/
β”‚   β”œβ”€β”€ phase_1/
β”‚   β”œβ”€β”€ phase_2/
β”‚   β”œβ”€β”€ phase_3/
β”‚   β”œβ”€β”€ phase_4/
β”‚   β”œβ”€β”€ phase_5/
β”‚   └── phase_6/
β”‚
β”œβ”€β”€ data/
β”‚   └── README.md
β”‚
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ phase1.py
β”‚   β”œβ”€β”€ phase5.py
β”‚   β”œβ”€β”€ phase6.py
β”‚   β”œβ”€β”€ publish_phase3.py
β”‚   β”œβ”€β”€ regression_after.py
β”‚   β”œβ”€β”€ regression_check.py
β”‚   └── qa/
β”‚       └── end_to_end_qa.py
β”‚
β”œβ”€β”€ src/
β”‚   └── reference_db/
β”‚       β”œβ”€β”€ adapters/
β”‚       β”œβ”€β”€ classification/
β”‚       β”œβ”€β”€ ingestion/
β”‚       β”œβ”€β”€ phase_1/
β”‚       β”œβ”€β”€ phase_2/
β”‚       β”œβ”€β”€ phase_3/
β”‚       β”‚   β”œβ”€β”€ classification/
β”‚       β”‚   β”œβ”€β”€ grouping/
β”‚       β”‚   └── taxonomy/
β”‚       β”œβ”€β”€ phase_4/
β”‚       β”œβ”€β”€ phase_5/
β”‚       β”œβ”€β”€ phase_6/
β”‚       β”œβ”€β”€ production/
β”‚       β”œβ”€β”€ publishing/
β”‚       β”œβ”€β”€ quality/
β”‚       β”œβ”€β”€ review/
β”‚       └── taxonomy/
β”‚
β”œβ”€β”€ tests/
β”‚
β”œβ”€β”€ pyproject.toml
β”œβ”€β”€ README.md
└── .gitignore

Retailer Adapters

The repository contains retailer-specific adapters for:

  • Compliments
  • Costco
  • Metro
  • VoilΓ 
  • Walmart

The adapter architecture isolates retailer-specific source mapping and extraction logic from the generic Reference DB processing phases.

This makes it possible to add or modify a retailer without rewriting the core processing pipeline.


Data Contracts

Each major pipeline stage has a versioned contract describing the expected behavior and data structure.

Contracts define aspects such as:

  • Expected input schema
  • Expected output schema
  • Required and optional fields
  • Primary identifiers
  • Processing responsibilities
  • Quality expectations
  • Data lineage

Contracts are stored under:

contracts/

The current contracts cover:

contracts/
β”œβ”€β”€ adapters/
β”œβ”€β”€ ingestion/
β”œβ”€β”€ phase_1/
β”œβ”€β”€ phase_2/
β”œβ”€β”€ phase_3/
β”œβ”€β”€ phase_4/
β”œβ”€β”€ phase_5/
└── phase_6/

Quality Gates & Review

Quality checks are performed at phase boundaries to prevent invalid or ambiguous data from silently progressing through the pipeline.

Possible quality states include:

Status Meaning
PASS The output satisfies the required checks.
WARNING An issue was detected but processing can continue.
AMBIGUOUS The decision cannot be made safely and the record can be isolated for review.
FAIL The affected pipeline path should not continue.

Ambiguity is detected by the phase responsible for the corresponding decision.

For example:

  • Phase 2 β†’ ambiguous identity or normalization
  • Phase 3 β†’ ambiguous classification, taxonomy, or grouping
  • Phase 4 β†’ ambiguous variant decisions

Review is not treated as a separate processing phase. Once a decision is resolved, the affected record can be reprocessed from the relevant phase.


Testing & Validation

The repository includes automated tests and an end-to-end QA workflow.

The current validation workflow includes:

ruff check src scripts tests
python -m compileall -q src scripts tests
pytest -q
python scripts/qa/end_to_end_qa.py

The end-to-end QA validates:

  • Phase outputs
  • External ID coverage
  • Group integrity
  • Variant integrity
  • Nutrition and score relationships
  • Production output consistency
  • Orphan detection
  • Basic nutrition quality
  • Score eligibility

The current test suite contains 74 passing tests, and the end-to-end QA workflow passes on the current implementation.


Current QA Snapshot

The current end-to-end QA run processes:

Output Rows Γ— Columns
Phase 1 4,440 Γ— 14
Phase 2 4,440 Γ— 34
Phase 3 Groups 4,440 Γ— 2
Phase 3 Classification 4,440 Γ— 12
Phase 3 Taxonomy 4,440 Γ— 7
Phase 4 Variants 4,320 Γ— 6
Phase 4 Mapping 4,440 Γ— 4
Phase 5 Nutrition 4,440 Γ— 77
Phase 6 Scores 4,440 Γ— 39
Production Metadata 4,440 Γ— 34
Production Groups 3,845 Γ— 2
Production Variants 4,320 Γ— 6
Production Nutrition 4,440 Γ— 77
Production Scores 4,440 Γ— 39

The QA checks currently report:

  • External ID missing/extra records: 0 / 0
  • Null production group_id: 0
  • Production group mismatches: 0
  • Orphan variants: 0
  • Orphan nutrition records: 0
  • Orphan score records: 0
  • Production metadata/nutrition/score mismatches: 0

These figures describe the current validated test dataset and are not claims about the size of all Reference DB source datasets.


Data

Raw source data and generated pipeline outputs are intentionally not committed to this code repository by default.

The data/ directory is reserved for local or generated pipeline data and currently contains documentation only.

Reference datasets and source collections are maintained separately on Hugging Face.

ReferenceDB Datasets

The ReferenceDB dataset collection contains Canadian retailer, brand, and source datasets used by the project.

ReferenceDB Hugging Face Collection


Technology Stack

  • Python β€” pipeline implementation
  • dlt β€” data ingestion and loading
  • DuckDB β€” local analytical processing and storage
  • pytest β€” automated testing
  • Ruff β€” linting and code quality checks

The implementation intentionally keeps the core processing logic lightweight and modular.


Design Principles

The system follows several core design principles:

  • Keep retailer-specific logic inside adapters.
  • Keep processing phases retailer-agnostic.
  • Preserve source identifiers and provenance.
  • Never invent missing source values.
  • Keep identity separate from variants.
  • Keep classification, taxonomy, and product grouping as separate concerns.
  • Preserve traceability from processed records back to source data.
  • Validate data at phase boundaries.
  • Use executable data contracts where appropriate.
  • Isolate ambiguous decisions rather than silently forcing them.
  • Keep processing artifacts separate from production outputs.
  • Prefer simple implementations before introducing additional infrastructure.

Development

Create and activate a virtual environment:

python -m venv .venv
source .venv/bin/activate

Install the project in editable mode:

pip install -e .

Run the test suite:

pytest -q

Run linting:

ruff check src scripts tests

Run the end-to-end QA:

python scripts/qa/end_to_end_qa.py

Repository Roles

The project uses different repositories and resources for different purposes.

GitHub β€” Source Code

The GitHub repository is the primary development repository containing the Reference DB implementation, contracts, tests, and pipeline scripts.

Reference DB System on GitHub

Hugging Face β€” Code Snapshot

A Hugging Face copy of the code is maintained as a convenient snapshot for review and reproducibility.

Reference DB System on Hugging Face

Hugging Face β€” Datasets

Reference datasets are maintained separately in the ReferenceDB collection.

ReferenceDB Collection

Notion β€” Architecture Documentation

The architectural decisions, processing model, data flow, and system design are documented separately.

Reference DB Architecture Documentation


Project Status

The core Reference DB processing pipeline is implemented and integrated across Phases 1–6.

The current repository includes:

  • Retailer adapters
  • Ingestion and validation
  • Identity normalization
  • Classification
  • Reference taxonomy
  • Product grouping
  • Variant modeling
  • Nutrition processing
  • Scores and enrichment
  • Production dataset construction
  • Quality gates
  • Review and quarantine components
  • Data contracts
  • Automated tests
  • End-to-end QA
  • Publishing utilities

The current implementation has been validated through automated tests and an end-to-end QA workflow.

The repository is intended to provide a reusable foundation for continuing Reference DB development and onboarding additional Canadian data sources.


Related Resources


GSoC 2026

This repository contains the implementation developed as part of the Open Food Facts Canada Reference DB project during Google Summer of Code 2026.

The repository should be read together with the architecture documentation and ReferenceDB datasets to understand the complete project.

The final project overview and work-product documentation provide the broader context, implementation summary, results, challenges, and project status.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support