Painting Vision Robotics Kit — robot-ready wall-paint coverage segmentation

A lite, edge-deployable research preview for autonomous painting and surface-finishing robots: from pixels to an executable paint plan.

The Painting Vision Robotics Kit reads a single camera frame of a wall and returns far more than a mask. It segments the paintable surface, the fixtures and openings that must not be painted, and the substrate — then hands a downstream planner a wall boundary, corners, a two-tool stroke plan, machine-readable waypoints, and a live paint-progress grid, independently cross-checked against depth/LiDAR.

It is built for the hard part of construction robotics: knowing exactly where the wall is, where the keep-outs are, and when the coat is complete.

Status: research preview. The segmentation architecture, dataset tooling, geometry planner and depth validator are implemented and unit-tested; the model weights are trained by the open train.py in this repo on session-separated, real capture data.

Lite model, frontier lineage

This kit is the lite tier of our models. The published painting-vision-robotics-kit is deliberately compact — a MobileNetV3/DeepLabV3-class backbone through SegFormer-B2 (models.py --arch) sized to run onboard a painting robot on CPU or a small accelerator, trading peak accuracy for latency, memory and offline operation (ONNX / TensorRT INT8 / CoreML targets; docs/ROADMAP.md, Phase 3).

Our frontier tier is the flagship. Constructelligence is developing frontier construction-AI models — larger backbones, higher-resolution and multimodal (RGB + depth/LiDAR) input, trained on session-separated real capture. Those are a separate, larger model class and are not what this card publishes. What is open here is the lite variant: the same task and the same honest, testable pipeline, at a fraction of the size, runnable on a laptop today.

Read this card as the floor of the range — a lite model shipped in the open by a frontier-model team — not the ceiling.


Why this model exists

General segmentation models tell you a wall is present. Painting robots need to know:

  • which pixels are paintable wall versus trim, skirting, openings and fixtures;
  • the exact boundary and corners of the wall, so a roller or spray head does not cross an edge;
  • a stroke path that respects a physical clearance margin and lifts over obstacles;
  • the substrate (e.g. drywall) so paint behaviour is not confused with surface type;
  • remaining coverage across repeated observations, so the robot knows when to stop.

The kit is designed end-to-end around those questions.

Capabilities

1. Ten-class semantic segmentation + substrate head

A configurable backbone (--arch: MobileNetV3-Large, ResNet-50/101 DeepLabV3, or a modern SegFormer) drives a ten-class semantic head and an independent drywall-material head, so paint state and substrate are never conflated:

other · wall_unpainted · wall_painted · wall_uncertain · skirting · switch/outlet · AC unit · door · window · wall_obstacle

Painted coverage is reported as wall_painted / (wall_painted + wall_unpainted) with lower/upper bounds that include uncertain wall area — a confidence-aware estimate, not a single false-precise number.

2. Window detection you can trust around

Windows are the most safety-critical keep-out in exterior painting. The window_postprocess stage merges mullion-split fragments, fits an oriented minimum-area rectangle with rotating calipers (so a window stays a window at any viewing angle), and filters specks and non-window shapes by fill ratio and aspect ratio — producing complete, regularised keep-out geometry.

3. Robot paint planner

robot_planner.py converts the masks into an executable plan:

  • ordered boundary polygon and corners (Moore-neighbour contour tracing + Douglas-Peucker — it follows real rooflines and voids, not a convex hull);
  • optional wall-plane rectification with edge lengths in millimetres;
  • a two-tool plan: roller passes for open areas plus trim/brush passes for edges and around fixtures the roller cannot reach;
  • machine-ready waypoints with paint on/off flags and travel distance;
  • paint-progress tracking (--state) that targets only what remains and tells the robot when the wall is done.

A documented real-facade run lifted planned coverage from 68% (roller only) to 97.5% (roller + trim) on the same wall — the residual is the physically unreachable remainder, reported explicitly.

4. Depth / LiDAR validation on top of vision

depth_validation.py independently checks the vision output in metric 3D: RANSAC wall-plane fitting, protrusion/recess detection (missed pipes or windows), recess-versus-window agreement, and a metric millimetres-per-pixel scale derived from depth instead of guessed — so the planner works in real units.

5. Balanced-data tooling

data_balance.py and audit_dataset.py score coverage and balance across indoor/exterior, substrate, lighting, weather, paint stage and classes, flag wall/session leakage and thin classes, and recommend exactly how many more samples each gap needs. train.py supports class-, wall- and hybrid-balanced sampling so the balanced collection survives into the optimizer.

6. Paint colour: named, matched, per face

paint_color.py places the painted pixels in CIELAB, trims the darkest and brightest 10% (shadow, glare, lap marks) and names the coat ("magnolia #F1EBD6"), reports a k-means palette and a colour per measured wall face — so a face painted in the wrong colour or a missed second coat is visible in the report. A specified colour (--reference-color) is compared with CIEDE2000 and banded (≤1 exact, ≤2 a touch-up match, ≤5 the same at a glance), and safety_gate.py stops the coat when the measured colour is not the specified one. The colour printed on a paint can's label is read and checked against the same reference, so the product on site is verified and not assumed.

7. The equipment around the wall

site_objects.py finds paint cans, trays, brushes, rollers, ladders and dust sheets around the wall, with no trained model (there is no annotated data for them): gradient outlines are cut above the frame's own noise, scene structure (skirting, ceiling, the wall/floor junction) is removed so a tin standing on the floor is not welded onto it, outlines are closed and filled back into bodies, and each body is scored per class from its own cues — aspect, fill, outline straightness, nap texture, the stripe a tin's rim or a brush's ferrule makes, the repeats tray ridges or ladder rungs make, frame position, colour. It answers what a site asks: is the gear the plan needs here, is anything standing where the paint is going (folded into the keep-out mask), and is the tin the right colour. Every detection carries the cues that fired and the scores of the classes it was not given; safety_gate.py warns by default on clutter and on missing tools, and stops on a colour mismatch.

Quickstart

python3 -m pip install -r requirements.txt

# 1. Segment a frame
python3 predict.py frame.jpg --checkpoint artifacts/best.pt \
    --paintable-mask paintable_wall_mask.png \
    --keepout-mask fixture_keepout_mask.png \
    --json result.json --tta

# 2. Plan the paint path (optional wall-plane rectification + mm scale)
python3 robot_planner.py --paintable paintable_wall_mask.png \
    --keepout fixture_keepout_mask.png \
    --corners "80,60 1180,55 1190,700 70,710" \
    --stroke-width 230 --trim-width 60 --overlap 0.2 --margin 40 --mm-per-unit 3.0 \
    --state wall_014.npz --plan plan.json --overlay plan.png

# 3. Validate against depth / LiDAR if available
python3 depth_validate.py --depth depth.npy --intrinsics "fx,fy,cx,cy" \
    --mask painting_mask.png --report validation.json

Repository layout

File Role
models.py architecture factory: DeepLabV3 (MobileNet/ResNet) + SegFormer, semantic + drywall heads
train.py training, balanced sampling, window loss boost, boundary/coverage metrics
predict.py inference: masks, coverage, geometry, regularised windows, paint colour, equipment on site
paint_color.py the coat's colour named, paletted, per face, and matched with CIEDE2000
site_objects.py model-free detection of paint cans, trays, brushes, rollers, ladders, dust sheets; the colour on a tin's label
robot_planner.py boundary/corners, two-tool stroke plan, waypoints, progress state
window_postprocess.py fragment merge + oriented rectangle fitting for windows
depth_validation.py LiDAR/depth plane fit, protrusion/recess checks, metric scale
data_balance.py coverage/balance analytics
audit_dataset.py dataset audit: schema, leakage, class and domain balance
fetch_real_data.py downloads and converts a real labeled dataset for testing
smoke_test.py end-to-end audit → train → predict → plan
benchmark.py head-to-head vs a generic baseline on the shared wall/window/other subset
provenance.py content-addressed dataset identity, integrity + licence ledger
safety_gate.py fail-closed approve/reject deployment gate
flywheel.py ingest corrected captures with content + perceptual dedupe
active_learning.py rank which frames to label next (novelty + domain gaps + uncertainty)
run_registry.py append-only, hash-linked ledger of training/eval runs
test_*.py geometry, window, depth and architecture regression tests
kaggle/ Kaggle training: notebook, dataset/kernel metadata, guarded push helper

Evaluation & status

Measured under a frozen protocol (benchmark/SPEC.md, v1.0) on a held-out test split of 136 real frames, pinned to a dataset_id and a checkpoint hash — so the numbers below are re-derivable from benchmark/records/:

Metric (test split) Value
mIoU, other / wall / window 0.802
wall / window / other IoU 0.872 / 0.717 / 0.819
window precision / recall 0.834 / 0.836
mIoU, full 10-class 0.698
painted-coverage MAE 0.162

The all-other floor is mIoU 0.110, so the model clears the floor by a wide margin. Two readings matter, and both are stated rather than hidden.

On its own validation data it looks strong: painted-coverage error is about 7.5 points, and it never marks a wall complete that is not (false-complete rate 0.0).

On 117 real photographs it was never tuned on, the substrate and condition questions are not solved — the honest half:

Question (117 wild photos) Result
Finding the wall 117 / 117
Wallpaper precision (22 positives) 0.50 — half its alarms are false
Painted wall called "bare board" (drywall) 50% raw · 65% calibrated
Damage (peeling / rust), AUROC 0.52 — close to a coin flip; it flags most sound walls
Sign-off on bare drywall 1 of 19 walls marked finished — a real robot would leave that wall unpainted

The quickest gain is fixing those real-photo numbers: hard negatives for painted vs drywall/wallpaper, or swapping in SAM 2 / Depth Anything (docs/ROADMAP.md, Phase 1–2). train.py reports the same metric families on any dataset you supply.

Intended use and limitations

Intended: assistive perception and path planning for painting/finishing robots; dataset curation and annotation QA for construction vision.

Not intended: unattended safety-critical control without calibrated geometry, clearance validation and a human-approved deployment; judging finish quality (streaking, thin coat, runs) — that needs dedicated defect labels and reference captures after the coat cures.

RGB alone cannot always reveal what is under a smooth surface; the model scores drywall only where visible evidence supports it and uses unknown otherwise. Always validate on held-out physical walls with the production camera.

Citation

@misc{constructelligence_painting_vision,
  title  = {Painting Vision Robotics Kit: robot-ready wall-paint coverage segmentation},
  author = {Constructelligence},
  year   = {2026},
  note   = {Research preview}
}

Keywords

wall painting robot · autonomous painting · paint coverage estimation · construction AI · building facade segmentation · drywall detection · skirting detection · window detection · semantic segmentation · LiDAR validation · depth sensing · robot path planning · paint progress tracking · construction robotics · BIM · surface finishing automation

Downloads last month
41
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results