YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- Reference DB System
Reference DB System
A reusable data pipeline for building retailer-level reference datasets for Open Food Facts Canada.
Reference DB transforms retailer product data into standardized, validated, classified, grouped, and enriched product datasets while preserving source identifiers, provenance, and traceability.
The system is designed to support multiple Canadian retailers while keeping retailer-specific extraction and mapping logic isolated from the generic processing pipeline.
Overview
Retailer product data is often inconsistent across sources. Product names, identifiers, attributes, categories, nutrition information, and variants may use different structures and conventions.
Reference DB provides a structured pipeline that transforms this source data into a consistent retailer-level reference dataset.
The pipeline separates:
- Source-specific ingestion
- Data validation
- Product identity
- Classification and taxonomy
- Product grouping
- Variant modeling
- Nutrition processing
- Scores and enrichment
- Production outputs
- Quality validation and review
The resulting data can be used as a reference layer for Open Food Facts Canada workflows and downstream applications.
Architecture
Retailer Source
β
Retailer Adapter
β
Ingestion
β
Phase 1 β Validation
β
Phase 2 β Identity & Normalization
β
Phase 3 β Classification + Taxonomy + Grouping
β
Phase 4 β Variants
β
Phase 5 β Nutrition
β
Phase 6 β Scores & Enrichment
β
Production Outputs
β
Retailer Reference Dataset
The architecture is intentionally modular so that retailer-specific logic does not have to be duplicated across processing phases.
Processing Phases
Phase 1 β Validation
Validates the ingested product dataset against the ingestion and validation contracts.
Checks include:
- Required fields
- Data types
- Product identifiers
- Duplicate records
- Basic structural consistency
- Source data integrity
Phase 2 β Identity & Normalization
Builds a standardized representation of product identity.
This phase:
- Normalizes identity-related attributes
- Separates identity attributes from variant attributes
- Preserves source identifiers
- Generates deterministic identity information
- Provides the foundation for product grouping
Identity and variant information are deliberately kept separate so that different variants of the same product can remain within the same product group.
Phase 3 β Classification, Taxonomy & Grouping
Phase 3 combines three related but distinct concepts:
Classification
Assigns product-level classifications based on product information and domain rules.
Reference Taxonomy
Maps products to the Reference DB taxonomy rather than relying entirely on retailer-specific categories.
Product Grouping
Groups records representing the same underlying product identity.
The grouping pipeline uses candidate generation, matching signals, identity information, and explicit decision logic.
Ambiguous cases can be isolated for review rather than being silently merged.
Phase 4 β Variants
Models product variants within established product groups.
Examples of variant-level differences include:
- Size
- Quantity
- Flavor
- Format
- Packaging
- Other product-specific attributes
Variants remain linked to their product group while preserving their own identifiers and attributes.
Phase 5 β Nutrition
Standardizes and validates available nutrition information and links nutrition data to products.
Missing nutrition information is treated as missing data, not as zero.
Nutrition processing also includes validation to prevent invalid or suspicious source values from silently becoming production data.
Phase 6 β Scores & Enrichment
Calculates available product scores and enrichment information.
Depending on the available source data, this phase can include:
- Nutri-Score
- Agribalyse references
- Fruit and vegetable-related enrichment
- Score eligibility
- Other derived product attributes
Only products satisfying the required eligibility conditions receive the corresponding score or enrichment.
Production Outputs
The pipeline produces production-oriented datasets rather than exposing processing artifacts as the final interface.
Current production outputs include:
production_metadata
production_groups
production_variants
production_nutrition
production_scores
These outputs provide a stable representation of the processed Reference DB data while keeping intermediate processing artifacts separate.
Project Structure
reference-db-system/
β
βββ contracts/
β βββ adapters/
β βββ ingestion/
β βββ phase_1/
β βββ phase_2/
β βββ phase_3/
β βββ phase_4/
β βββ phase_5/
β βββ phase_6/
β
βββ data/
β βββ README.md
β
βββ scripts/
β βββ phase1.py
β βββ phase5.py
β βββ phase6.py
β βββ publish_phase3.py
β βββ regression_after.py
β βββ regression_check.py
β βββ qa/
β βββ end_to_end_qa.py
β
βββ src/
β βββ reference_db/
β βββ adapters/
β βββ classification/
β βββ ingestion/
β βββ phase_1/
β βββ phase_2/
β βββ phase_3/
β β βββ classification/
β β βββ grouping/
β β βββ taxonomy/
β βββ phase_4/
β βββ phase_5/
β βββ phase_6/
β βββ production/
β βββ publishing/
β βββ quality/
β βββ review/
β βββ taxonomy/
β
βββ tests/
β
βββ pyproject.toml
βββ README.md
βββ .gitignore
Retailer Adapters
The repository contains retailer-specific adapters for:
- Compliments
- Costco
- Metro
- VoilΓ
- Walmart
The adapter architecture isolates retailer-specific source mapping and extraction logic from the generic Reference DB processing phases.
This makes it possible to add or modify a retailer without rewriting the core processing pipeline.
Data Contracts
Each major pipeline stage has a versioned contract describing the expected behavior and data structure.
Contracts define aspects such as:
- Expected input schema
- Expected output schema
- Required and optional fields
- Primary identifiers
- Processing responsibilities
- Quality expectations
- Data lineage
Contracts are stored under:
contracts/
The current contracts cover:
contracts/
βββ adapters/
βββ ingestion/
βββ phase_1/
βββ phase_2/
βββ phase_3/
βββ phase_4/
βββ phase_5/
βββ phase_6/
Quality Gates & Review
Quality checks are performed at phase boundaries to prevent invalid or ambiguous data from silently progressing through the pipeline.
Possible quality states include:
| Status | Meaning |
|---|---|
PASS |
The output satisfies the required checks. |
WARNING |
An issue was detected but processing can continue. |
AMBIGUOUS |
The decision cannot be made safely and the record can be isolated for review. |
FAIL |
The affected pipeline path should not continue. |
Ambiguity is detected by the phase responsible for the corresponding decision.
For example:
- Phase 2 β ambiguous identity or normalization
- Phase 3 β ambiguous classification, taxonomy, or grouping
- Phase 4 β ambiguous variant decisions
Review is not treated as a separate processing phase. Once a decision is resolved, the affected record can be reprocessed from the relevant phase.
Testing & Validation
The repository includes automated tests and an end-to-end QA workflow.
The current validation workflow includes:
ruff check src scripts tests
python -m compileall -q src scripts tests
pytest -q
python scripts/qa/end_to_end_qa.py
The end-to-end QA validates:
- Phase outputs
- External ID coverage
- Group integrity
- Variant integrity
- Nutrition and score relationships
- Production output consistency
- Orphan detection
- Basic nutrition quality
- Score eligibility
The current test suite contains 74 passing tests, and the end-to-end QA workflow passes on the current implementation.
Current QA Snapshot
The current end-to-end QA run processes:
| Output | Rows Γ Columns |
|---|---|
| Phase 1 | 4,440 Γ 14 |
| Phase 2 | 4,440 Γ 34 |
| Phase 3 Groups | 4,440 Γ 2 |
| Phase 3 Classification | 4,440 Γ 12 |
| Phase 3 Taxonomy | 4,440 Γ 7 |
| Phase 4 Variants | 4,320 Γ 6 |
| Phase 4 Mapping | 4,440 Γ 4 |
| Phase 5 Nutrition | 4,440 Γ 77 |
| Phase 6 Scores | 4,440 Γ 39 |
| Production Metadata | 4,440 Γ 34 |
| Production Groups | 3,845 Γ 2 |
| Production Variants | 4,320 Γ 6 |
| Production Nutrition | 4,440 Γ 77 |
| Production Scores | 4,440 Γ 39 |
The QA checks currently report:
- External ID missing/extra records: 0 / 0
- Null production
group_id: 0 - Production group mismatches: 0
- Orphan variants: 0
- Orphan nutrition records: 0
- Orphan score records: 0
- Production metadata/nutrition/score mismatches: 0
These figures describe the current validated test dataset and are not claims about the size of all Reference DB source datasets.
Data
Raw source data and generated pipeline outputs are intentionally not committed to this code repository by default.
The data/ directory is reserved for local or generated pipeline data and currently contains documentation only.
Reference datasets and source collections are maintained separately on Hugging Face.
ReferenceDB Datasets
The ReferenceDB dataset collection contains Canadian retailer, brand, and source datasets used by the project.
ReferenceDB Hugging Face Collection
Technology Stack
- Python β pipeline implementation
- dlt β data ingestion and loading
- DuckDB β local analytical processing and storage
- pytest β automated testing
- Ruff β linting and code quality checks
The implementation intentionally keeps the core processing logic lightweight and modular.
Design Principles
The system follows several core design principles:
- Keep retailer-specific logic inside adapters.
- Keep processing phases retailer-agnostic.
- Preserve source identifiers and provenance.
- Never invent missing source values.
- Keep identity separate from variants.
- Keep classification, taxonomy, and product grouping as separate concerns.
- Preserve traceability from processed records back to source data.
- Validate data at phase boundaries.
- Use executable data contracts where appropriate.
- Isolate ambiguous decisions rather than silently forcing them.
- Keep processing artifacts separate from production outputs.
- Prefer simple implementations before introducing additional infrastructure.
Development
Create and activate a virtual environment:
python -m venv .venv
source .venv/bin/activate
Install the project in editable mode:
pip install -e .
Run the test suite:
pytest -q
Run linting:
ruff check src scripts tests
Run the end-to-end QA:
python scripts/qa/end_to_end_qa.py
Repository Roles
The project uses different repositories and resources for different purposes.
GitHub β Source Code
The GitHub repository is the primary development repository containing the Reference DB implementation, contracts, tests, and pipeline scripts.
Hugging Face β Code Snapshot
A Hugging Face copy of the code is maintained as a convenient snapshot for review and reproducibility.
Reference DB System on Hugging Face
Hugging Face β Datasets
Reference datasets are maintained separately in the ReferenceDB collection.
Notion β Architecture Documentation
The architectural decisions, processing model, data flow, and system design are documented separately.
Reference DB Architecture Documentation
Project Status
The core Reference DB processing pipeline is implemented and integrated across Phases 1β6.
The current repository includes:
- Retailer adapters
- Ingestion and validation
- Identity normalization
- Classification
- Reference taxonomy
- Product grouping
- Variant modeling
- Nutrition processing
- Scores and enrichment
- Production dataset construction
- Quality gates
- Review and quarantine components
- Data contracts
- Automated tests
- End-to-end QA
- Publishing utilities
The current implementation has been validated through automated tests and an end-to-end QA workflow.
The repository is intended to provide a reusable foundation for continuing Reference DB development and onboarding additional Canadian data sources.
Related Resources
- Source code: GitHub Repository
- Code review snapshot: Hugging Face Code Repository
- Datasets: ReferenceDB Collection
- Architecture: Reference DB Architecture
GSoC 2026
This repository contains the implementation developed as part of the Open Food Facts Canada Reference DB project during Google Summer of Code 2026.
The repository should be read together with the architecture documentation and ReferenceDB datasets to understand the complete project.
The final project overview and work-product documentation provide the broader context, implementation summary, results, challenges, and project status.