YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Phase 2 β€” Semantic Normalization / Identity Extraction

Compliments Reference DB Pipeline

Version: 2.0.0 Date: 2026-07-30 Status: PASS


Purpose

Phase 2 reads the validated Phase 1 output and performs semantic normalization (brand β†’ brand_norm + product_line) and identity attribute extraction from the original title. Phase 2 is purely semantic β€” it does NOT duplicate basic string cleaning now handled by Phase 1.


Input

Property Value
Source phase1/outputs/phase1_output.parquet
Row Count 4,440
Column Count 20

What Phase 2 Does

1. Brand Semantic Normalization

  • Uses brand_clean from Phase 1 (already whitespace/case normalized)
  • Maps 8 cleaned brand variants to brand_norm (Compliments/Sensations) + product_line (Core/Organic/Balance/Naturally Simple/Green/Little Ones/Sensations)
  • Does NOT duplicate whitespace/case cleaning (Phase 1 handles this)

2. Identity Attribute Extraction (from ORIGINAL title)

  • 9 boolean flags from ORIGINAL title: is_organic, is_gluten_free, is_naturally_simple, is_sugar_free, is_unsalted, is_lactose_free, is_peanut_free, is_plant_based, is_reduced_sodium
  • Fat-level extraction: fat_free / reduced_fat / regular
  • Fat-percentage extraction: Numeric % from title
  • Flavour extraction: 32 keyword list
  • Formulation extraction: 14 keyword list

3. Title Processing

  • Core-title extraction: Strip brand prefix + size + bracket content (uses title_clean from Phase 1)
  • Variant attributes parsing: Size string β†’ structured JSON

4. Identity Hash

  • Deterministic 14-field hash (for reference; grouping in Phase 3)

Phase 2 Does NOT

  • Perform product grouping
  • Perform food/non-food classification
  • Use taxonomy
  • Use nutrition data
  • Use external data sources
  • Modify original title, upc, external_id, or other source identifiers
  • Duplicate string cleaning already done in Phase 1

Output Schema

The Phase 2 output (phase2_output.parquet) contains 38 columns:

Original Columns (14, preserved from Phase 1)

# Column Type
1 upc str
2 external_id str
3 brand str
4 title str
5 price float64
6 price_currency str
7 size str
8 size_amount float64
9 size_unit str
10 size_qty int64
11 size_unit_norm str
12 image_url str
13 source str
14 source_url str

Phase 1 Cleaned Columns (6, preserved from Phase 1)

# Column Type Description
15 brand_clean str Cleaned brand (from Phase 1)
16 brand_raw str Original brand (from Phase 1)
17 title_clean str Cleaned title (from Phase 1)
18 title_raw str Original title (from Phase 1)
19 upc_raw str Original UPC (from Phase 1)
20 external_id_raw str Original external_id (from Phase 1)

New Columns (18, added by Phase 2)

# Column Type Description
21 brand_norm str Normalized brand (Compliments/Sensations)
22 product_line str Product line (Core/Organic/Balance/Naturally Simple/Green/Little Ones/Sensations)
23 is_organic bool Organic flag
24 is_gluten_free bool Gluten-free flag
25 is_naturally_simple bool Naturally Simple flag
26 is_sugar_free bool Sugar-free flag
27 is_unsalted bool Unsalted flag
28 is_lactose_free bool Lactose-free flag
29 is_peanut_free bool Peanut-free flag
30 is_plant_based bool Plant-based flag
31 is_reduced_sodium bool Reduced sodium flag
32 fat_level str Fat level (fat_free/reduced_fat/regular)
33 fat_percentage float64 Fat percentage from title (nullable)
34 flavour list Extracted flavour keywords
35 formulation list Extracted formulation keywords
36 variant_attributes str Parsed size as JSON
37 core_title str Title with brand prefix + size stripped
38 identity_hash str Deterministic 14-field identity hash

Validation Results

Check Result
Row count preserved (4,440) PASS
All original columns preserved PASS
All 18 new columns added PASS
Brand normalization (2 values) PASS
Product line (7 values) PASS
Fat level values valid PASS
Identity flags are boolean PASS
Core titles not empty PASS
Identity hashes not empty PASS
No duplicate external_ids PASS
Overall PASS

Files

phase2/
    README.md
    src/phase2.py
    notebooks/phase2.ipynb
    outputs/
        phase2_output.parquet
        phase2_executed.ipynb
    validation/
        phase2_validation.json
    statistics/
        phase2_statistics.json

How to Run

# Using the Python script
cd /home/sara/gsoc/compliments-reference-db
python -m phase2.src.phase2

# Using the notebook
jupyter notebook phase2/notebooks/phase2.ipynb

Handoff to Phase 3

Phase 3 must read:

  • phase2/outputs/phase2_output.parquet

Phase 3 must NOT:

  • Re-download from HuggingFace
  • Read the original products.parquet directly
  • Read Phase 1 output directly
  • Use any other data source
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support