YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- production-raw-assets
- Scope
- Dataflow
- Folder layout
- linewatch β stop detection & frame extraction
- objects_cleanup.py β background removal
- 1. Prerequisites (one-time, system packages)
- 2. Create the virtual environment (.venv)
- 3. Install the required libraries
- 4. Verify the installation
- 5. Run objects_cleanup.py
- Shortcut: the launcher does steps 2-5 automatically
- Reusing an existing venv that already has rembg
- Troubleshooting
- Model cache
- Options worth knowing
- build_synthetic_yolo.py β synthetic YOLO dataset builder
- Downstream usage
- State
- Decisions
- Related repositories
- AGENTS.md
- Scope
production-raw-assets
Raw and processed visual assets for the synthetic YOLO FOD dataset pipeline, plus every data-preparation stage. Hugging Face dataset repository.
Scope
In scope: raw video β background frames β object cut-outs β mm
dimensions (sizes.json) β dataset build (build_synthetic_yolo.py).
Out of scope: model training (production-yolo-synthetic/train/),
weights & metrics (production-yolo-fod). System map: ../PIPELINE.md.
Dataflow
raw_videos/ ββlinewatchββ> backgrounds/ + backgrounds_timestamps/ (112 PNG each)
raw_objects/ ββrembgββββ> filtered_objects/ (28 RGBA cut-outs)
filtered_objects/ + manual sizes.json ββ
backgrounds/ ββββββββββββββββββββββββββ΄ββbuild_synthetic_yoloββ>
../production-yolo-synthetic/ (images/, label/, data.yaml)
Stage table, cross-repo contracts and invariants: ../PIPELINE.md.
Rationale for every choice: DECISIONS.md. Agent-facing
constraints: AGENTS.md.
Folder layout
production-raw-assets/
βββ raw_videos/ # source production-line footage
βββ backgrounds/ # ORIGINAL frames extracted by linewatch <- output
βββ backgrounds_timestamps/ # same frames with border + burnt-in timestamp <- output
βββ raw_objects/ # raw object photos (background still included)
βββ filtered_objects/ # RGBA object cut-outs (transparent background) <- output
β βββ sizes.json # real object dimensions in mm (fill in manually)
βββ pipeline_scripts/
βββ linewatch/ # stop detection + frame extraction (repo variant)
βββ run_linewatch.sh # one-command runner for linewatch (creates venv)
βββ objects_cleanup.py # background removal: raw_objects -> filtered_objects
βββ run_objects_cleanup.sh # one-command runner for objects_cleanup
βββ build_synthetic_yolo.py # dataset builder -> ../production-yolo-synthetic
βββ run_build_synthetic.sh # one-command runner for the dataset builder
βββ requirements.txt
βββ extract.py # (placeholder, superseded by linewatch/)
The final stage populates the sibling Hugging Face dataset repository
production-yolo-synthetic/ (see the last section of this README).
linewatch β stop detection & frame extraction
pipeline_scripts/linewatch/ is the repo-local variant of the workspace's
linewatch package (OpenCV). It watches the production-line video, detects
when the line stops (frame differencing + optical-flow direction), and picks
the sharpest frame of each stop interval (Sobel gradient energy).
Each selected frame is written TWICE:
| Output | Content |
|---|---|
backgrounds/<video>-interval-XXXXXX.png |
the ORIGINAL frame, untouched |
backgrounds_timestamps/<video>-interval-XXXXXX.png |
same frame with a uniform black border + burnt-in timestamp (provenance view) |
Run artifacts (events.ndjson, intervals.json, scores.csv, run.json)
go to pipeline_scripts/results/.
Run it
cd production-raw-assets/pipeline_scripts
./run_linewatch.sh ../raw_videos/production_line_10_min.mp4
# or manually inside the venv:
python -m linewatch --file ../raw_videos/production_line_10_min.mp4
# live camera instead of a file:
python -m linewatch --camera 0
# custom tuning (see --print-default-config for all knobs):
python -m linewatch --file <video> --config fod_config.json
The launcher creates .venv on the first run and installs requirements.txt
(also used by objects_cleanup; adds opencv-python-headless).
Options worth knowing
| Option | Meaning |
|---|---|
--backgrounds-dir / --backgrounds-timestamps-dir |
override the two frame destinations (default the repo folders shown above) |
--image-prefix |
filename prefix; default is the video stem (or camera), so extracts from different sources never collide |
--output |
run artifact directory (default pipeline_scripts/results/) |
--config / --print-default-config |
JSON tuning (ROI, thresholds, settling time, border size, timestamp font, image format) |
Notes:
- The two destination folders are treated as persistent assets: the tool
never deletes existing files there (unlike upstream linewatch, which wipes
its per-run
candidates/). Re-running the same video overwrites the same interval names. backgrounds/is the clean input forgenerate_yolo_dataset.pyβ pass--bg-border 0there since these frames carry no annotation border. (Thebackgrounds_timestamps/set is the annotated equivalent kept for traceability.)
objects_cleanup.py β background removal
Removes the background from every image in raw_objects/ and writes
transparent RGBA cut-outs to filtered_objects/ using rembg
(IS-Net/UΒ²-Net ONNX models).
1. Prerequisites (one-time, system packages)
A working python3 (3.9+) with the venv module and pip:
# Debian/Ubuntu (python3-venv is usually NOT preinstalled!)
sudo apt update && sudo apt install -y python3 python3-venv python3-pip
# Fedora: sudo dnf install -y python3
# macOS: brew install python (or use the preinstalled one)
# Windows: https://www.python.org/downloads/ -> tick "Add python.exe to PATH"
2. Create the virtual environment (.venv)
The venv lives inside pipeline_scripts/.venv (git-ignored, see
.gitignore), fully isolated from the system Python and from the parent
workspace's other venvs:
cd production-raw-assets/pipeline_scripts
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
After activation the shell prompt shows (.venv). Leave the venv with
deactivate; re-enter later with source .venv/bin/activate.
3. Install the required libraries
With the venv active:
pip install --upgrade pip
pip install -r requirements.txt
requirements.txt pins rembg[cpu]==2.0.85 β that automatically pulls in
onnxruntime, numpy, scipy, scikit-image and pymatting β plus
pillow and tqdm. For an NVIDIA GPU, change rembg[cpu] to rembg[gpu]
in requirements.txt before installing (optional; CPU is fine for a few
dozen images).
4. Verify the installation
python -c "from rembg import new_session, remove; from PIL import Image; import tqdm; print('environment OK')"
5. Run objects_cleanup.py
Still inside the venv (defaults: ../raw_objects -> ../filtered_objects):
python objects_cleanup.py # default model: isnet-general-use
python objects_cleanup.py --model u2net # alternative model
python objects_cleanup.py --alpha-matting # finer soft edges, slower
python objects_cleanup.py --overwrite # redo already-processed images
Already-processed images are skipped by default (idempotent), so interrupted runs can simply be restarted. The console summary lists processed / skipped / failed files and every cut-out produced.
Shortcut: the launcher does steps 2-5 automatically
run_objects_cleanup.sh creates .venv (if missing), installs
requirements.txt into it, and runs the script β no manual setup needed:
cd production-raw-assets/pipeline_scripts
./run_objects_cleanup.sh [--model u2net] [--overwrite] ...
Reusing an existing venv that already has rembg
Instead of creating .venv, point at the parent workspace's src/.venv
(rembg 2.0.85 already installed there β no downloads):
/workspace/src/.venv/bin/python objects_cleanup.py
# or with the launcher:
OBJECTS_CLEANUP_VENV=/workspace/src/.venv ./run_objects_cleanup.sh
Troubleshooting
| Symptom | Fix |
|---|---|
cannot create venv / missing ensurepip (Ubuntu) |
sudo apt install python3-venv, delete the half-created .venv, retry |
No module named rembg |
venv not active β source .venv/bin/activate, or use the launcher |
| model download fails behind a proxy | export HTTPS_PROXY=http://proxy:port before running |
| no internet on the target machine | copy the ~/.u2net model cache from a connected machine (or set U2NET_HOME) |
| Windows | skip the .sh launcher; use python -m venv .venv, .venv\Scripts\activate, python objects_cleanup.py |
Model cache
The first run downloads the ONNX model (170 MB for isnet-general-use) into
rembg's cache (`/.u2netor$XDG_DATA_HOME/.u2net; override with U2NET_HOME`). Internet is needed once; subsequent runs are fully offline.
Options worth knowing
| Option | Meaning |
|---|---|
--model |
isnet-general-use (default), u2net, u2netp, silueta, birefnet-general, birefnet-general-lite, birefnet-dis, bria-rmbg, u2net_human_seg |
--alpha-matting |
refine soft edges with pymatting (slower) |
--decontaminate |
remove colour fringing left on soft edges |
--post-process-mask |
morphological cleanup of the raw mask |
--no-trim / --pad N |
keep full canvas / set padding around the cut-out (default: trimmed with 8 px pad) |
--overwrite |
redo images that already exist in filtered_objects/ |
build_synthetic_yolo.py β synthetic YOLO dataset builder
Final stage: build_synthetic_yolo.py reads the assets produced by the
stages above and populates the production-yolo-synthetic Hugging Face
dataset repository (sibling of this repo).
Three steps:
- Classify backgrounds β
backgrounds/*.pngare shuffled with a seeded RNG and split into a train and a test set (--test-split, default 20%). Deterministic: the same--seedalways yields the same classification, and no background ever leaks into the other split. - Combine with objects β for every background,
--images-per-backgroundcomposites are generated by pasting--objects-per-imagecut-outs fromfiltered_objects/at their physical size (px/mm = frame width / background_size_mmfromsizes.json), with random rotation, optional defocus, soft drop shadow and IoU-aware placement. Compositing core is ported from the validatedgenerate_yolo_dataset.pyinvariants (labels = tight alpha bbox of the rotated sprite, 6-decimal normalised YOLO coordinates, empty label files for pure negative samples). - Populate the dataset β the provided repository structure is filled:
production-yolo-synthetic/
βββ images/train/synth_train_<bg-stem>_<k>.jpg
βββ images/test/synth_test_<bg-stem>_<k>.jpg
βββ label/train/*.txt, label/test/*.txt
βββ preview/ # only with --preview N
βββ data.yaml # Ultralytics-style train/val/test + names
File names embed the source background stem, so every synthetic image is traceable back to its linewatch interval.
Training and weights: the dataset repository owns the whole training
stage β see production-yolo-synthetic/train/ and its README (prepare
labels -> train -> evaluate -> export -> publish into the
production-yolo-fod Hugging Face model repo).
Run it
cd production-raw-assets/pipeline_scripts
./run_build_synthetic.sh # defaults
./run_build_synthetic.sh --images-per-background 6 --objects-per-image 2
./run_build_synthetic.sh --test-split 0.25 --seed 42
./run_build_synthetic.sh --preview 8 --clean # regenerate + previews
# or manually inside the venv:
python build_synthetic_yolo.py --images-per-background 4
Only Pillow (plus optional tqdm) is needed β already part of
requirements.txt and the shared .venv.
Options worth knowing
| Option | Meaning |
|---|---|
--backgrounds-dir / --objects-dir / --sizes |
input folders (default the repo assets) |
--dataset-dir |
dataset repo to populate (default ../production-yolo-synthetic) |
--images-per-background |
composites generated per background (default 4) |
--objects-per-image |
cut-outs pasted per composite (default 1) |
--test-split / --seed |
background train/test classification |
--max-iou / --max-blur / --no-shadow |
realism controls |
--format / --jpeg-quality |
output encoding (default jpg 95) |
--bg-border |
crop N px off each background first (default 0 β backgrounds are clean linewatch frames) |
--preview N |
render N box-annotation previews into preview/ |
--clean |
remove previously generated synth_* files first (never touches data.yaml/README) |
Ultralytics label-location note
data.yaml points Ultralytics at images/train and images/test.
Ultralytics pairs each image with its label file by substituting
images/ -> labels/ in the image path β this repository uses the
provided label/ root instead. Before running yolo train, either rename
label/ to labels/, or create a symlink ln -s label labels.
Downstream usage
The RGBA cut-outs in filtered_objects/ are consumed by
pipeline_scripts/build_synthetic_yolo.py (alpha-keyed paste onto
backgrounds/ frames, populating ../production-yolo-synthetic). The trim
threshold used here (alpha >= 8) matches the builder's alpha-bbox logic, so
no pixels are lost or hallucinated in between.
Object sizes: filtered_objects/sizes.json
sizes.json carries the real-world dimensions (millimetres) of every object
cut-out and is passed to the dataset generator via --config:
{
"background_size_mm": { "width": 800, "height": 600 },
"objects": {
"object_01.png": { "width_mm": 25, "height_mm": 60, "class": "part_a" }
}
}
- Fill in
width_mmandheight_mmfor each entry (physical bounding box of the visible object). Entries left atnullare skipped by the generator with a warning until measured. classis optional β it defaults to the file stem (object_01). Map several PNGs to one YOLO class by giving them the sameclassvalue.- The generator warns when a configured mm aspect ratio differs > 10% from the PNG's pixel aspect β a handy typo/typo-in-measurement check.
background_size_mmmust match the real area captured by the linewatch-extracted background photos (default 800 x 600 mm).
Then generate the dataset with:
cd pipeline_scripts
./run_build_synthetic.sh # or: python build_synthetic_yolo.py
(The old parent-workspace command is superseded by this stage:)
python generate_yolo_dataset.py --config ... --bg-border 0 # obsolete here
State
| Item | Status (2025-09-27) |
|---|---|
| Source video | raw_videos/production_line_10_min.mp4 (517 MB, 640Γ480 @ 30 fps, ~10 min) |
| Backgrounds | 112 frames extracted (112/113 intervals completed; 1 tail interval incomplete); clean + timestamped copies in sync |
| Objects | 28 raw photos β 28 RGBA cut-outs; all 28 measured in sizes.json (scene 800Γ600 mm) |
| Classes | default file-stem names (object_01β¦); no merging yet |
| Dataset v1 | generated β ../production-yolo-synthetic/ (448 images, seed 42) |
pipeline_scripts/results/ |
latest extraction artifacts, local-only (gitignored) |
| Known weak spot | smallest objects (8β10 mm β 6β8 px sprites) are the hardest classes downstream |
Decisions
Load-bearing choices (full log: DECISIONS.md):
- D-0001 three repos by concern, flow raw-assets β synthetic β fod
- D-0002 dual linewatch output; asset folders persistent, never wiped
- D-0004 rembg
isnet-general-use(license overbria-rmbg) - D-0005 annotation removal is manual (automated removal rejected)
- D-0006 alpha threshold 8 β must match the builder
- D-0009 "classify backgrounds" = seeded train/test split by background (no leakage)
- D-0010 JPG q95, native 640Γ480, physical mm scale (0.8 px/mm)
Related repositories
- Dataset:
https://huggingface.co/datasets/simdengineer/production-yolo-synthetic - Model:
https://huggingface.co/simdengineer/production-yolo-fod
AGENTS.md
Hard constraints for humans and AI coding agents: AGENTS.md β
read before modifying anything in this repo.