YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

production-raw-assets

Raw and processed visual assets for the synthetic YOLO FOD dataset pipeline, plus every data-preparation stage. Hugging Face dataset repository.

Scope

In scope: raw video β†’ background frames β†’ object cut-outs β†’ mm dimensions (sizes.json) β†’ dataset build (build_synthetic_yolo.py).

Out of scope: model training (production-yolo-synthetic/train/), weights & metrics (production-yolo-fod). System map: ../PIPELINE.md.

Dataflow

raw_videos/ ──linewatch──> backgrounds/ + backgrounds_timestamps/   (112 PNG each)
raw_objects/ ──rembg────> filtered_objects/   (28 RGBA cut-outs)
filtered_objects/ + manual sizes.json ─┐
backgrounds/ ─────────────────────────┴──build_synthetic_yolo──>
     ../production-yolo-synthetic/  (images/, label/, data.yaml)

Stage table, cross-repo contracts and invariants: ../PIPELINE.md. Rationale for every choice: DECISIONS.md. Agent-facing constraints: AGENTS.md.

Folder layout

production-raw-assets/
β”œβ”€β”€ raw_videos/            # source production-line footage
β”œβ”€β”€ backgrounds/           # ORIGINAL frames extracted by linewatch       <- output
β”œβ”€β”€ backgrounds_timestamps/ # same frames with border + burnt-in timestamp <- output
β”œβ”€β”€ raw_objects/           # raw object photos (background still included)
β”œβ”€β”€ filtered_objects/     # RGBA object cut-outs (transparent background)  <- output
β”‚   └── sizes.json         # real object dimensions in mm (fill in manually)
└── pipeline_scripts/
    β”œβ”€β”€ linewatch/                # stop detection + frame extraction (repo variant)
    β”œβ”€β”€ run_linewatch.sh         # one-command runner for linewatch (creates venv)
    β”œβ”€β”€ objects_cleanup.py       # background removal: raw_objects -> filtered_objects
    β”œβ”€β”€ run_objects_cleanup.sh   # one-command runner for objects_cleanup
    β”œβ”€β”€ build_synthetic_yolo.py  # dataset builder -> ../production-yolo-synthetic
    β”œβ”€β”€ run_build_synthetic.sh   # one-command runner for the dataset builder
    β”œβ”€β”€ requirements.txt
    └── extract.py                # (placeholder, superseded by linewatch/)

The final stage populates the sibling Hugging Face dataset repository production-yolo-synthetic/ (see the last section of this README).

linewatch β€” stop detection & frame extraction

pipeline_scripts/linewatch/ is the repo-local variant of the workspace's linewatch package (OpenCV). It watches the production-line video, detects when the line stops (frame differencing + optical-flow direction), and picks the sharpest frame of each stop interval (Sobel gradient energy).

Each selected frame is written TWICE:

Output Content
backgrounds/<video>-interval-XXXXXX.png the ORIGINAL frame, untouched
backgrounds_timestamps/<video>-interval-XXXXXX.png same frame with a uniform black border + burnt-in timestamp (provenance view)

Run artifacts (events.ndjson, intervals.json, scores.csv, run.json) go to pipeline_scripts/results/.

Run it

cd production-raw-assets/pipeline_scripts
./run_linewatch.sh ../raw_videos/production_line_10_min.mp4
# or manually inside the venv:
python -m linewatch --file ../raw_videos/production_line_10_min.mp4
# live camera instead of a file:
python -m linewatch --camera 0
# custom tuning (see --print-default-config for all knobs):
python -m linewatch --file <video> --config fod_config.json

The launcher creates .venv on the first run and installs requirements.txt (also used by objects_cleanup; adds opencv-python-headless).

Options worth knowing

Option Meaning
--backgrounds-dir / --backgrounds-timestamps-dir override the two frame destinations (default the repo folders shown above)
--image-prefix filename prefix; default is the video stem (or camera), so extracts from different sources never collide
--output run artifact directory (default pipeline_scripts/results/)
--config / --print-default-config JSON tuning (ROI, thresholds, settling time, border size, timestamp font, image format)

Notes:

  • The two destination folders are treated as persistent assets: the tool never deletes existing files there (unlike upstream linewatch, which wipes its per-run candidates/). Re-running the same video overwrites the same interval names.
  • backgrounds/ is the clean input for generate_yolo_dataset.py β€” pass --bg-border 0 there since these frames carry no annotation border. (The backgrounds_timestamps/ set is the annotated equivalent kept for traceability.)

objects_cleanup.py β€” background removal

Removes the background from every image in raw_objects/ and writes transparent RGBA cut-outs to filtered_objects/ using rembg (IS-Net/UΒ²-Net ONNX models).

1. Prerequisites (one-time, system packages)

A working python3 (3.9+) with the venv module and pip:

# Debian/Ubuntu (python3-venv is usually NOT preinstalled!)
sudo apt update && sudo apt install -y python3 python3-venv python3-pip
# Fedora:  sudo dnf install -y python3
# macOS:   brew install python      (or use the preinstalled one)
# Windows: https://www.python.org/downloads/  -> tick "Add python.exe to PATH"

2. Create the virtual environment (.venv)

The venv lives inside pipeline_scripts/.venv (git-ignored, see .gitignore), fully isolated from the system Python and from the parent workspace's other venvs:

cd production-raw-assets/pipeline_scripts
python3 -m venv .venv
source .venv/bin/activate              # Windows: .venv\Scripts\activate

After activation the shell prompt shows (.venv). Leave the venv with deactivate; re-enter later with source .venv/bin/activate.

3. Install the required libraries

With the venv active:

pip install --upgrade pip
pip install -r requirements.txt

requirements.txt pins rembg[cpu]==2.0.85 β€” that automatically pulls in onnxruntime, numpy, scipy, scikit-image and pymatting β€” plus pillow and tqdm. For an NVIDIA GPU, change rembg[cpu] to rembg[gpu] in requirements.txt before installing (optional; CPU is fine for a few dozen images).

4. Verify the installation

python -c "from rembg import new_session, remove; from PIL import Image; import tqdm; print('environment OK')"

5. Run objects_cleanup.py

Still inside the venv (defaults: ../raw_objects -> ../filtered_objects):

python objects_cleanup.py                        # default model: isnet-general-use
python objects_cleanup.py --model u2net         # alternative model
python objects_cleanup.py --alpha-matting       # finer soft edges, slower
python objects_cleanup.py --overwrite           # redo already-processed images

Already-processed images are skipped by default (idempotent), so interrupted runs can simply be restarted. The console summary lists processed / skipped / failed files and every cut-out produced.

Shortcut: the launcher does steps 2-5 automatically

run_objects_cleanup.sh creates .venv (if missing), installs requirements.txt into it, and runs the script β€” no manual setup needed:

cd production-raw-assets/pipeline_scripts
./run_objects_cleanup.sh [--model u2net] [--overwrite] ...

Reusing an existing venv that already has rembg

Instead of creating .venv, point at the parent workspace's src/.venv (rembg 2.0.85 already installed there β€” no downloads):

/workspace/src/.venv/bin/python objects_cleanup.py
# or with the launcher:
OBJECTS_CLEANUP_VENV=/workspace/src/.venv ./run_objects_cleanup.sh

Troubleshooting

Symptom Fix
cannot create venv / missing ensurepip (Ubuntu) sudo apt install python3-venv, delete the half-created .venv, retry
No module named rembg venv not active β€” source .venv/bin/activate, or use the launcher
model download fails behind a proxy export HTTPS_PROXY=http://proxy:port before running
no internet on the target machine copy the ~/.u2net model cache from a connected machine (or set U2NET_HOME)
Windows skip the .sh launcher; use python -m venv .venv, .venv\Scripts\activate, python objects_cleanup.py

Model cache

The first run downloads the ONNX model (170 MB for isnet-general-use) into rembg's cache (`/.u2netor$XDG_DATA_HOME/.u2net; override with U2NET_HOME`). Internet is needed once; subsequent runs are fully offline.

Options worth knowing

Option Meaning
--model isnet-general-use (default), u2net, u2netp, silueta, birefnet-general, birefnet-general-lite, birefnet-dis, bria-rmbg, u2net_human_seg
--alpha-matting refine soft edges with pymatting (slower)
--decontaminate remove colour fringing left on soft edges
--post-process-mask morphological cleanup of the raw mask
--no-trim / --pad N keep full canvas / set padding around the cut-out (default: trimmed with 8 px pad)
--overwrite redo images that already exist in filtered_objects/

build_synthetic_yolo.py β€” synthetic YOLO dataset builder

Final stage: build_synthetic_yolo.py reads the assets produced by the stages above and populates the production-yolo-synthetic Hugging Face dataset repository (sibling of this repo).

Three steps:

  1. Classify backgrounds β€” backgrounds/*.png are shuffled with a seeded RNG and split into a train and a test set (--test-split, default 20%). Deterministic: the same --seed always yields the same classification, and no background ever leaks into the other split.
  2. Combine with objects β€” for every background, --images-per-background composites are generated by pasting --objects-per-image cut-outs from filtered_objects/ at their physical size (px/mm = frame width / background_size_mm from sizes.json), with random rotation, optional defocus, soft drop shadow and IoU-aware placement. Compositing core is ported from the validated generate_yolo_dataset.py invariants (labels = tight alpha bbox of the rotated sprite, 6-decimal normalised YOLO coordinates, empty label files for pure negative samples).
  3. Populate the dataset β€” the provided repository structure is filled:
production-yolo-synthetic/
β”œβ”€β”€ images/train/synth_train_<bg-stem>_<k>.jpg
β”œβ”€β”€ images/test/synth_test_<bg-stem>_<k>.jpg
β”œβ”€β”€ label/train/*.txt, label/test/*.txt
β”œβ”€β”€ preview/                       # only with --preview N
└── data.yaml                       # Ultralytics-style train/val/test + names

File names embed the source background stem, so every synthetic image is traceable back to its linewatch interval.

Training and weights: the dataset repository owns the whole training stage β€” see production-yolo-synthetic/train/ and its README (prepare labels -> train -> evaluate -> export -> publish into the production-yolo-fod Hugging Face model repo).

Run it

cd production-raw-assets/pipeline_scripts
./run_build_synthetic.sh                                  # defaults
./run_build_synthetic.sh --images-per-background 6 --objects-per-image 2
./run_build_synthetic.sh --test-split 0.25 --seed 42
./run_build_synthetic.sh --preview 8 --clean             # regenerate + previews
# or manually inside the venv:
python build_synthetic_yolo.py --images-per-background 4

Only Pillow (plus optional tqdm) is needed β€” already part of requirements.txt and the shared .venv.

Options worth knowing

Option Meaning
--backgrounds-dir / --objects-dir / --sizes input folders (default the repo assets)
--dataset-dir dataset repo to populate (default ../production-yolo-synthetic)
--images-per-background composites generated per background (default 4)
--objects-per-image cut-outs pasted per composite (default 1)
--test-split / --seed background train/test classification
--max-iou / --max-blur / --no-shadow realism controls
--format / --jpeg-quality output encoding (default jpg 95)
--bg-border crop N px off each background first (default 0 β€” backgrounds are clean linewatch frames)
--preview N render N box-annotation previews into preview/
--clean remove previously generated synth_* files first (never touches data.yaml/README)

Ultralytics label-location note

data.yaml points Ultralytics at images/train and images/test. Ultralytics pairs each image with its label file by substituting images/ -> labels/ in the image path β€” this repository uses the provided label/ root instead. Before running yolo train, either rename label/ to labels/, or create a symlink ln -s label labels.

Downstream usage

The RGBA cut-outs in filtered_objects/ are consumed by pipeline_scripts/build_synthetic_yolo.py (alpha-keyed paste onto backgrounds/ frames, populating ../production-yolo-synthetic). The trim threshold used here (alpha >= 8) matches the builder's alpha-bbox logic, so no pixels are lost or hallucinated in between.

Object sizes: filtered_objects/sizes.json

sizes.json carries the real-world dimensions (millimetres) of every object cut-out and is passed to the dataset generator via --config:

{
  "background_size_mm": { "width": 800, "height": 600 },
  "objects": {
    "object_01.png": { "width_mm": 25, "height_mm": 60, "class": "part_a" }
  }
}
  • Fill in width_mm and height_mm for each entry (physical bounding box of the visible object). Entries left at null are skipped by the generator with a warning until measured.
  • class is optional β€” it defaults to the file stem (object_01). Map several PNGs to one YOLO class by giving them the same class value.
  • The generator warns when a configured mm aspect ratio differs > 10% from the PNG's pixel aspect β€” a handy typo/typo-in-measurement check.
  • background_size_mm must match the real area captured by the linewatch-extracted background photos (default 800 x 600 mm).

Then generate the dataset with:

cd pipeline_scripts
./run_build_synthetic.sh                 # or: python build_synthetic_yolo.py

(The old parent-workspace command is superseded by this stage:)

python generate_yolo_dataset.py --config ... --bg-border 0   # obsolete here

State

Item Status (2025-09-27)
Source video raw_videos/production_line_10_min.mp4 (517 MB, 640Γ—480 @ 30 fps, ~10 min)
Backgrounds 112 frames extracted (112/113 intervals completed; 1 tail interval incomplete); clean + timestamped copies in sync
Objects 28 raw photos β†’ 28 RGBA cut-outs; all 28 measured in sizes.json (scene 800Γ—600 mm)
Classes default file-stem names (object_01…); no merging yet
Dataset v1 generated β†’ ../production-yolo-synthetic/ (448 images, seed 42)
pipeline_scripts/results/ latest extraction artifacts, local-only (gitignored)
Known weak spot smallest objects (8–10 mm β†’ 6–8 px sprites) are the hardest classes downstream

Decisions

Load-bearing choices (full log: DECISIONS.md):

  • D-0001 three repos by concern, flow raw-assets β†’ synthetic β†’ fod
  • D-0002 dual linewatch output; asset folders persistent, never wiped
  • D-0004 rembg isnet-general-use (license over bria-rmbg)
  • D-0005 annotation removal is manual (automated removal rejected)
  • D-0006 alpha threshold 8 β€” must match the builder
  • D-0009 "classify backgrounds" = seeded train/test split by background (no leakage)
  • D-0010 JPG q95, native 640Γ—480, physical mm scale (0.8 px/mm)

Related repositories

  • Dataset: https://huggingface.co/datasets/simdengineer/production-yolo-synthetic
  • Model: https://huggingface.co/simdengineer/production-yolo-fod

AGENTS.md

Hard constraints for humans and AI coding agents: AGENTS.md β€” read before modifying anything in this repo.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support