YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Phase 1 β€” data pipeline and detector training

Detect cones, barriers and stop signs from a single RGB camera, and estimate the distance to each from focal length and pixel geometry. No depth sensor and no learned depth model at inference time β€” LiDAR appears only as scoring ground truth, never in the inference path.

This directory is phase 1: get every source into one schema with calibration and ground-truth distance attached, verify the geometry by eye and by assert, and train the detector.

Install

uv venv && source .venv/bin/activate
uv pip install -e .                      # manifest, merge, export, QA
uv pip install -e ".[nuscenes]"          # + the nuScenes extractor
uv pip install -e ".[hub]"               # + Hugging Face Hub push/pull
uv pip install -e ".[train]"             # + Ultralytics

The AV2, COCO and BDD extractors need no extra dependencies.

Everything runs as a module from the repo root, e.g. python -m src.data.merge.

Run order

The sample viewer goes immediately after the first extractor, before writing or running anything else. A sign error or an axis-order mistake in the 3D→2D projection produces plausible-looking garbage that stays invisible until the depth numbers make no sense a day later.

# 0. fetch nuScenes. Presigned URLs expire in minutes -- copy from the browser
#    and run immediately. See "Getting nuScenes" below.
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-mini/*' 'samples/CAM_FRONT/*'

# 1. nuScenes: 2D boxes projected from the 3D cuboids, plus gt distance and calibration
python -m src.data.extract_nuscenes --dataroot data/raw/nuscenes --version v1.0-mini

# 2. look at it, now
python -m src.data.view_samples \
    --manifest data/unified/parts/nuscenes.parquet --only-annotated --out qa/nusc.png

# 3. the other sources (any order; AV2 is the droppable one)
python -m src.data.extract_av2   --dataroot /data/av2/sensor/train --max-logs 20
python -m src.data.extract_coco  --dataroot /data/coco --split train2017
python -m src.data.extract_bdd_negatives \
    --dataroot /data/bdd100k --labels /data/bdd100k/labels/det_20/det_train.json --count 3000

# 4. merge: filter, split by scene, assert the geometry agrees across datasets
python -m src.data.merge

# 5. export
python -m src.data.to_yolo

# 6. publish, then train from the published copy
python -m src.hub.push --repo-id <user>/cone-distance-v1 --private
python -m src.hub.pull --repo-id <user>/cone-distance-v1 --dest data/hub
python -m src.train.train --data data/hub/data.yaml --models yolo11n.pt yolo11s.pt

UNIFIED_ROOT overrides data/unified for every step, or pass --unified-root.

To check the plumbing before any data has landed:

python -m tests.smoke_test

It builds a synthetic ground plane, projects cones onto it, and runs the real merge β†’ export β†’ QA chain over the result.

Layout

data/unified/
  manifest.parquet         # the source of truth: one row per object
  parts/<source>.parquet   # per-extractor fragments, concatenated by merge.py
  calib/<sensor_id>.yaml   # K, camera height, ego<-cam rotation
  class_priors.yaml        # per-class physical size, derived from the manifest
  images/<source>          # one symlink per dataset, never a copy
  yolo/                    # derived export: symlinked images + .txt + data.yaml

src/common/  schema.py geometry.py calib.py paths.py
src/data/    extract_*.py merge.py to_yolo.py view_samples.py check_ground_plane.py
src/bench/   pipeline.py eval_pt.py eval_onnx.py eval_ncnn.py compare.py
src/depth/   ground_plane.py known_size.py estimate.py evaluate.py
             accuracy.py view_detections.py
src/hub/     push.py pull.py
src/train/   train.py wandb_logger.py checkpoint_mirror.py
scripts/     fetch_nuscenes.sh fetch_coco_stopsigns.py fetch_av2.py make_evalpack.py make_evalpack.py
notebooks/   colab_train.ipynb local_train.ipynb

Getting nuScenes

nuScenes has no download API. Log in at https://www.nuscenes.org/nuscenes#download, right-click a file and copy the link address β€” it is a presigned CloudFront URL that expires within minutes.

Download with scripts/fetch_nuscenes.sh, which pipes curl straight into tar and keeps only the files we use. Each trainval blob is ~30 GB of six cameras, LiDAR and five radars; we use one camera. Unpacking one to disk needs ~60 GB of headroom for ~450 MB of useful data, so filtering in the stream is worth the one extra script.

# mini first -- build and verify the pipeline on it
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-mini/*' 'samples/CAM_FRONT/*'

# then annotations for all 850 scenes, and one blob of images
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-trainval/*'
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'samples/CAM_FRONT/*'

The metadata covers all 850 scenes while each blob holds a slice of the images, so the extractor skips samples whose blob you did not download and reports the count as image_not_downloaded.

Training

Extraction, merging and export are CLI scripts and run locally; there is no notebook for them. Training has two, and they train the same thing:

Notebook Runs on Dataset If it is interrupted
colab_train.ipynb Colab, T4 or better pulled from the Hub CheckpointMirror copies last.pt to Drive every 10 epochs, because the runtime's disk does not survive
local_train.ipynb any machine β€” it clones the repo and installs deps if they are not already there, and detects an existing checkout if they are the local export if present, else pulled Ultralytics' own last.pt is already on durable disk; section 7 resumes the same run from it

models, epochs, imgsz and batch are identical in both β€” yolo11n.pt, 200, 640, 16 β€” so runs from either are directly comparable. Only the data path, the device and the run directory differ.

Both the code and the dataset are published to the Hugging Face Hub, because a Colab runtime is ephemeral and re-downloading 18 GB of nuScenes every session is not worth it when the derived dataset is 2.4 GB:

Repo What
Aryan006/cone-distance this code, and trained weights under weights/
Aryan006/cone-distance-v1 the dataset: shards, manifest, calibration, priors

Both are private. The Colab notebook reads HF_TOKEN and WANDB_API_KEY from Colab Secrets; the local one reads them from the environment, then from a gitignored file of that name in the repo root, then by prompting. Neither needs editing.

What goes on the Hub

src/hub/push.py packs the YOLO export into ~1 GB tar shards β€” the Hub handles a few large files far better than 40,000 small ones, and a shard is a resumable unit if the upload drops β€” and uploads manifest.parquet, calib/, class_priors.yaml and a generated dataset card alongside them. The manifest stays previewable in the Hub viewer. src/hub/pull.py is the exact inverse, and rewrites data.yaml with an absolute path so training works from anywhere.

Only the local build uses symlinks. push.py dereferences them, so a shard holds real bytes and not a link into a nuScenes root that will not exist on the next runtime.

Why a manifest and not YOLO .txt files

.txt files cannot hold intrinsics or ground-truth distance, and phase 2 needs both. The manifest is the artifact; to_yolo.py is a pure function of it, so the YOLO directory can be deleted and regenerated at any time.

Decisions worth knowing

2D boxes are generated from the 3D cuboids, not taken from any pre-made 2D annotation file. That projection has to happen anyway to get gt_distance_m, so doing it for the training labels too means the boxes and the distances come from one convention, in the same domain the depth estimator is evaluated in.

gt_distance_m is forward depth along the camera optical axis β€” the Z component of the object centre in the camera frame β€” not Euclidean range. Phase 2 must compare like with like.

Barrel is not barrier. An AV2 CONSTRUCTION_BARREL is an orange drum; a nuScenes barrier is a jersey/fence barrier. Visually unrelated. They stay separate classes; merging them would tank the class and look like a training problem rather than a labelling problem.

Split by scene, per source. Adjacent video frames are near-duplicates, so a frame-level random split leaks the val set into training and inflates val mAP a long way. Splitting scenes independently within each source also guarantees every source reaches val β€” otherwise a random draw can leave val with no nuScenes cones, which is the one thing phase 2 needs.

BDD100K negatives exclude frames containing a traffic sign. BDD's ten classes have no cones and no barriers, so its frames are safe negatives for those β€” but its generic traffic sign bucket hides real stop signs, and pairing an image containing a stop sign with an empty label file trains the model to suppress stop signs. --labels enables that filter; without it the extractor warns.

Images that lose every object to the filters are kept as negatives, not dropped. Everything the filters remove is below the size we intend to detect at all, so training the model not to fire there matches the deployment target. --drop-filtered-frames picks the other behaviour.

Filter thresholds

All in src/common/schema.py, applied in merge.py:

Threshold Value Why
minimum box height 15 px a 6-pixel cone is an unlearnable label that only adds false-positive pressure
minimum visibility 0.4 drops the nuScenes v0-40 bucket only
maximum truncation 0.5 above this the box is mostly a guess about what is off-screen

What merge.py asserts

These are real asserts, not things to eyeball. A coordinate-convention mismatch between nuScenes and AV2 produces a manifest that trains YOLO perfectly well and makes the depth numbers nonsense a day later.

  • boxes are non-degenerate
  • gt_distance_m is always positive
  • within one camera, distance and box-bottom row are negatively correlated β€” objects lower in the frame must be nearer
  • median cone distance falls in a plausible range (3–80 m)

The projected 3D centre landing inside its own 2D box is asserted in the extractors, where the 3D box is still in hand.

Two convention bugs the cross-dataset checks caught

Both were found by scoring the ground-plane estimator against gt_distance_m before anything was built on it. Both would have produced a manifest that trains YOLO perfectly well and silently wrong distances a day later.

Argoverse 2 cameras are not pinhole

AV2 ships k1, k2, k3 alongside the intrinsics β€” around -0.24, -0.22, +0.34 for ring_front_center β€” and the images are not rectified. Projecting through K alone puts boxes tens of pixels off the objects they label. nuScenes ships rectified images and passes zeros, for which the model is the identity, so the same code path now serves both. Calibration.project() and Calibration.ray() own the forward and inverse; nothing else touches intrinsics directly.

The two datasets disagree about where the ego frame sits

nuScenes puts its ego origin on the ground, so the camera height is simply the sensor's z translation. Argoverse 2 puts its origin at the rear axle centre, about 0.26 m up. Taking tz_m as camera height for both makes every AV2 ground-plane distance read ~18% short, at every range:

Range MAE before MAE after Bias before Bias after
0–10 m 1.48 m 0.56 m βˆ’1.50 m +0.11 m
10–20 m 2.77 m 1.05 m βˆ’2.49 m +0.04 m
20–30 m 5.25 m 2.39 m βˆ’4.65 m +0.14 m
30–45 m 8.51 m 6.55 m βˆ’7.71 m +0.38 m
45 m+ 20.27 m 18.23 m βˆ’12.52 m +1.30 m

Rather than hardcode a vehicle constant, extract_av2.py measures it: cones and barrels rest on the road, so the median underside of their cuboids in the near field is the road surface. Applied per log, that also absorbs road grade. The same derivation is deliberately not applied to nuScenes β€” there it returns 1.412 m against a stated 1.511 m that the ray-plane fit independently confirms (implied/stated = 0.986), so the stated value is the better one.

AV2 also gets one sensor_id per log, because its calibration varies between logs: across these fifteen, focal length moves 1% and camera height 6%. A single shared av2_ring_front_center would hand phase 2 one arbitrary log's numbers for every image.

Argoverse 2 stop signs include the pole

AV2 annotates a stop sign as one cuboid from the ground to the top of the sign β€” median 3.22 m tall by 0.86 m wide β€” so projecting it whole gives a 2D box with a median aspect ratio of 3.95. COCO's boxes are tight on the octagon at 1.07. Training on both teaches the detector two incompatible definitions of one class, and a pole-inclusive box is the wrong input to a width-based distance estimator. A stop sign is a regular octagon, so extract_av2.py takes the cuboid's own width as the face height and keeps the top slice β€” no assumed sign standard. Aspect ratio after: 0.98, and the derived prior drops from 3.21 m to 0.90 m.

Measured on the assembled dataset

nuScenes v1.0-trainval (metadata for all 850 scenes, images from blob 01), Argoverse 2 (15 logs), COCO stop signs. 10,261 frames, 23,829 objects. nuScenes cam_height_m reads 1.511 m.

Source barrel barrier cone stop_sign
nuscenes β€” 3,224 3,112 β€”
av2 11,295 β€” 3,050 1,362
coco β€” β€” β€” 1,786

Val split: 287 barrels, 494 barriers, 284 cones, 60 stop signs with gt_distance_m.

Class priors: derive them, and derive them from enough data

Across all 850 trainval scenes, unfiltered, straight from sample_annotation.json:

Class Height (mean) Median 5–95% n
cone 1.07 m 1.11 m 0.70–1.42 97,959
barrier 0.98 m 0.98 m 0.78–1.22 152,087

Worth recording how this went, because the trap is a general one. Measured on v1.0-mini alone, cones came out at 0.78 m β€” and mini's ten scenes simply contain unusually short cones. A prior fitted there would have been 27% low, and it looked authoritative because it was derived from data rather than looked up. Deriving beats guessing only once the sample is representative; n_samples is in class_priors.yaml for exactly this reason, and it is worth reading before trusting the number above it.

Note also that the derived prior from our filtered manifest (cone 0.97 m) sits below the population figure, because the filters keep the closer, larger, less occluded instances. For phase 2 that is arguably the right bias β€” the estimator only ever sees objects the detector found, which are drawn from the same skewed distribution β€” but it is a choice, not an accident.

Ground-plane distance, dry-run against ground truth

Recovering distance from the box bottom row using only calib/ β€” phase 2's estimator run early, scored against gt_distance_m. Gate: visibility >= 0.7 and box bottom >= 20 px below the horizon row.

Range nuScenes MAE nuScenes bias AV2 MAE AV2 bias
0–10 m 0.76 m βˆ’0.42 m 0.56 m +0.11 m
10–20 m 2.62 m βˆ’0.58 m 1.05 m +0.04 m
20–30 m 5.87 m βˆ’0.54 m 2.39 m +0.14 m
30–45 m 10.33 m βˆ’4.26 m 6.55 m +0.38 m
45 m+ 17.67 m βˆ’16.18 m 18.23 m +1.30 m

n = 5,222 (nuScenes) and 12,271 (AV2).

Two things this settles. The geometry is correct: sub-metre error in the near field and near-zero bias out to 30 m on both datasets, derived independently from two different calibration sources, is not something a broken transform produces. Box bottoms sit at the ground contact point β€” measured directly against the cuboids' own bottom faces, the 2D box bottom lands within 1.9 px of the true contact row β€” so there is nothing to fix in the extractor.

The far field is a conditioning problem, not an accuracy problem. Sensitivity is Z^2/(fy*h): 0.03 m/px at 10 m, 1.8 m/px at 45 m. The worst estimates come from boxes bottoming out within a few pixels of the horizon row, where the ground ray runs nearly parallel to the plane and the intersection diverges β€” before the validity gate, single estimates of +3113 m and -2831 m. Phase 2 therefore needs a horizon-margin guard before it needs a better estimator, and the known-size cross-check earns its place above roughly 30 m. AV2 degrades more gracefully at range than nuScenes because its camera has higher angular resolution (fy 1778 vs 1266) in a portrait frame.

The training recipe is stock Ultralytics

src/train/train.py is plumbing around YOLO.train(), not a modified training loop. Verified against the reference YOLO11 fine-tuning notebook, which trains via the CLI:

yolo task=detect mode=train model=yolo11s.pt data=.../data.yaml epochs=10 imgsz=640 plots=True

Resolving both through ultralytics.cfg.get_cfg on the pinned 8.3.40 gives 105 config keys, of which four differ, all plumbing:

Key Reference Here What it controls
device None 0 which GPU
project None /content/runs where runs are written
name None yolo11n / yolo11s run directory name
exist_ok False True overwrite an existing run dir

All 55 recipe keys are identical: optimiser, lr0, lrf, momentum, weight decay, warmup, the box/cls/dfl loss gains, nbs, cos_lr, close_mosaic, patience, and every augmentation (hsv_*, degrees, translate, scale, shear, perspective, flipud, fliplr, mosaic, mixup, copy_paste, auto_augment, erasing). Validation matches too β€” only device differs, and Model.train() reloads best.pt on completion ([model.py:808]), so model.val() scores the same checkpoint the reference validates.

Why batch size is 16 and not -1

This one is not cosmetic. Ultralytics normalises the optimiser step to a nominal batch of 64 by accumulating gradients, and rescales weight decay to compensate:

accumulate   = max(round(nbs / batch), 1)
weight_decay = weight_decay * batch * accumulate / nbs

The compensation is exact only when the batch divides 64. batch=-1 hands the choice to AutoBatch, which computes b = int((f * fraction - p[1]) / p[0]) from free GPU memory β€” an arbitrary integer:

batch accumulate effective weight decay
16 / 32 / 64 4 / 2 / 1 1.000Γ—
37 2 1.156Γ— (+15.6%)
48 1 0.750Γ— (βˆ’25.0%)
100 1 1.562Γ— (+56.2%)

So --batch -1 would have quietly trained a different recipe depending on which GPU Colab handed out, and the run would not be comparable to the reference or to itself across sessions. The default is now 16 β€” Ultralytics' own default and the reference's β€” and train.py warns if a batch is passed that does not divide 64.

Ultralytics is pinned to <=8.3.40, the version this comparison was run against and the one the reference notebook pins, so a future default change cannot silently invalidate the table above.

Weights & Biases

--wandb logs per epoch, all on the same step so the curves overlay:

Series Source
train/box_loss, train/cls_loss, train/dfl_loss trainer.tloss
val/box_loss, val/cls_loss, val/dfl_loss trainer.metrics
metrics/* precision, recall, mAP50, mAP50-95
lr/pg0..2 per param group
grad_norm/mean, /max, /p95, /clipped_fraction see below

Authenticate with $WANDB_API_KEY. Each model gets its own run, so the YOLO11n and YOLO11s arms are two comparable runs rather than one interleaved mess.

Ultralytics ships its own W&B integration, but it is gated behind SETTINGS["wandb"], which defaults to False and is persisted to the user's global config β€” enabling it reconfigures their machine, not just the run. It also does not log gradient norm. Since a module was needed for that anyway, src/train/wandb_logger.py registers its own callbacks: one file states exactly which series are logged, and nothing outside the process is changed.

Getting at the gradient norm

This is the one part of the logging that is not a callback, and it is opt-in separately as --log-grad-norm. Losses and metrics come from callbacks that only read trainer.*; without this flag nothing in torch is touched.

BaseTrainer.optimizer_step calls clip_grad_norm_, whose return value is the total pre-clip gradient norm, and discards it. It then calls zero_grad() in the same method, and no callback fires between the two β€” the nearest, on_train_batch_end, runs after the gradients are already gone, so it cannot be recomputed either. Wrapping clip_grad_norm_ for the duration of training is the only hook available, and the most accurate one: the value is captured after scaler.unscale_(), so it is a true unscaled norm and the same number the optimiser acted on.

The wrapper only reads a return value. It never touches gradients, the optimiser or the loss, so it cannot change what the model learns. Three things it could still have cost you, all handled:

Risk Handling
on_train_end does not fire when training raises, so the patch survives and the second model's probe wraps the wrapper install() marks the function it creates and refuses to wrap a marked one; train.py calls close() from a finally
float() on a CUDA tensor synchronises the device, and clip_grad_norm_ does not otherwise sync (error_if_nonfinite is False), so converting per step would stall every step norms are kept as detached tensors and converted once per epoch
Coupling it to --wandb would force the patch on anyone wanting loss curves separate flag

tests/test_grad_norm_probe.py pins all of it against a stub torch β€” no GPU, no 2 GB install β€” covering restoration, refusal to nest, pass-through of the caller's value, and that constructing a probe without installing it leaves torch untouched.

grad_norm/clipped_fraction is the series to read first: the share of steps whose norm exceeded the clip of 10. Near 1.0 during warmup is normal. If it stays there after warmup, the effective step size is set by the clip rather than by lr0, and lowering the learning rate will do more than tuning anything else.

Verified on a real 3-epoch CPU run against a slice of the actual dataset, with W&B in offline mode so nothing reached the account. All twelve series arrived; gradient norm read mean 419 β†’ 1437 β†’ 1423 across the three epochs at clipped_fraction 1.0, which is warmup behaviour on 48 images.

Surviving a killed Colab runtime

Colab kills runtimes without warning, and Ultralytics rewrites weights/last.pt in place every epoch β€” so a run that dies at epoch 140 leaves one mutable file and no history. src/train/checkpoint_mirror.py copies last.pt out to Drive every N epochs under an epoch-stamped name.

It polls rather than hooking into training, so a failure in it cannot take the run down. It reads the epoch counter out of the checkpoint rather than counting its own ticks, so it stays correct across a restart and ignores the epoch: -1 that strip_optimizer writes when training finishes. And it verifies the copy, not the source, before os.replace-ing it into place: Ultralytics can begin rewriting last.pt between the epoch check and the copy, so the source having been loadable a moment ago says nothing about what actually landed. A name in the destination therefore only ever appears once the file behind it is whole.

It lives in the repo rather than in a notebook cell so that tests/test_checkpoint_mirror.py can cover it against a stub torch β€” no GPU, no 2 GB install β€” including the torn-copy race, the duplicate-epoch guard, and the stripped final checkpoint.

Phase 2 β€” distance from geometry

src/depth/ground_plane.py   ray-plane intersection. THE REFERENCE TO TRANSCRIBE.
src/depth/known_size.py     distance from apparent size, given a size prior
src/depth/estimate.py       box -> distance: sample, robust summary, class routing
src/depth/evaluate.py       score against LiDAR ground truth on nuScenes val

ground_plane.ground_depth is written to be read and rewritten β€” the maths is laid out a step at a time, with the two traps named (h is the height above the road, not a translation component; Argoverse 2 pixels need undistorting first).

Results, nuScenes val, ground-truth boxes

759 objects, 100% given an estimate.

Range n MAE bias p90 abs rel
0–10 m 103 1.05 m βˆ’1.22 m 1.48 m 14.1%
10–20 m 339 3.96 m βˆ’2.14 m 7.65 m 22.9%
20–30 m 147 4.42 m +0.72 m 8.64 m 17.9%
30–45 m 124 8.97 m βˆ’4.61 m 16.20 m 24.4%
45 m+ 46 12.13 m βˆ’9.77 m 20.66 m 22.7%

The dominant error is per scene, not per object

Reading the per-class row first is misleading: cones come out at βˆ’5.91 m bias and barriers at βˆ’1.23 m, which looks like a class problem. It is not. Broken down by scene, barriers and cones inside the same scene agree closely (βˆ’0.401 vs βˆ’0.351 in one, βˆ’0.059 vs βˆ’0.061 in another). The apparent cone bias is an artefact of one scene holding 120 of the 265 cones.

Median relative error by scene runs from βˆ’35% to +35%, standard deviation 24%. Within a scene it is close to a constant fraction, which is the signature of the flat-road assumption failing on a grade rather than of a noisy estimator.

Fitting a single constant row shift per scene β€” one pitch offset, nothing per-object β€” halves the error:

Scene n shift median abs error
c08e31a5 160 βˆ’50.0 px 6.56 β†’ 1.30 m
638bfb0a 97 βˆ’36.9 px 1.53 β†’ 0.42 m
cdb711af 52 βˆ’5.5 px 4.50 β†’ 3.23 m
fc01a4d7 295 βˆ’0.8 px 1.98 β†’ 1.91 m
79f1ca3a 75 +4.1 px 2.10 β†’ 1.70 m
mean 3.34 β†’ 1.71 m

βˆ’50 px at fy = 1266 is an effective pitch error of about 2.3Β°, which is an ordinary road grade or a loaded suspension. The ground-plane and known-size per-scene errors correlate only +0.25, so they fail independently β€” this is the geometry, not a ground-truth artefact.

The conclusion for the rest of phase 2 is that per-frame ground-plane estimation (horizon or vanishing-point) is worth more than any refinement of the per-object maths. The static extrinsic is the binding constraint.

Sampling the box

Distance comes from a band of candidate ground-contact rows centred on the bottom edge, not extended upward into the box. The interior of a box is not the ground β€” a pixel halfway up a cone is half a metre in the air β€” so an upward band biases every estimate outward: a synthetic cone at 20.0 m reads 21.5 m with a band over the bottom 20% of the box, and 20.0 m with the band centred.

Centring also makes the median exact. Distance is monotonic in row and the median commutes with monotonic maps, so the median of the sampled distances is exactly the distance at the median row. The band costs nothing in accuracy and buys two things: rays that escape over the horizon drop out as NaN instead of poisoning an average, and the spread of the survivors is a confidence signal.

Cross-check and routing

Class Estimator
cone, barrier, barrel ground plane, cross-checked against the height prior
stop_sign known size only β€” pole-mounted, so the box bottom is not a ground contact

Requiring the two estimators to agree within 40% keeps 87% of objects and moves MAE from 4.97 m to 4.11 m.

End to end with stock YOLO11n

Architecturally complete, numerically meaningless, and worth stating plainly. Stock YOLO11n knows COCO, which has no cone, barrel or barrier class β€” of the four classes here only stop_sign exists in COCO, and nuScenes contributes no stop signs. On nuScenes val it matched 2 of 759 ground-truth boxes at IoU β‰₯ 0.5. Relaxing to conf 0.05 and IoU 0.3 reaches 18 of 300, and those are coincidental overlaps with cars and people rather than detections of our classes: over 25 val frames the detector fires on car (54), person (24), traffic light (24), truck (1) and fire hydrant (1).

So --source gt is the number that means something today, and --source detect becomes meaningful the moment the fine-tuned weights exist. Keeping them separate is what makes it possible to tell the depth module's error from the detector's.

Phase 3 β€” what each export format costs

Three runners over one measured pipeline:

python -m src.bench.eval_pt    --threads 4
python -m src.bench.eval_onnx  --threads 4
python -m src.bench.eval_ncnn  --threads 4     # on the Pi
python -m src.bench.compare

They differ only in which file they hand to Ultralytics. That is deliberate: three independently written scripts would drift and make the timings incomparable. Each writes bench/<tag>_timing.json and a CSV of every detection with its estimated distance.

Stages, and why they are split

Stage What
read JPEG off disk into a numpy array
preprocess letterbox and normalise
inference the forward pass β€” the only stage the format changes
postprocess NMS and rescaling boxes
depth box β†’ distance, the phase 2 estimator

Images are read here rather than passing predict() a path, so disk I/O lands in read instead of inflating preprocess. Five warm-up frames are discarded: the first inference through any backend pays for lazy allocation and cold caches, and on NCNN it can be an order of magnitude slower than steady state. p50 and p95 rather than mean, because the tail is real and a mean reports the tail rather than the typical frame.

x86 CPU, 4 cores, 60 frames

stage (p50 ms) PyTorch ONNX
read 3.39 3.14
preprocess 1.33 1.78
inference 54.65 32.36
postprocess 0.52 1.03
depth 1.60 1.59
total 63.20 40.22
FPS 15.82 24.86
size 5.6 MB 10.6 MB

ONNX is 1.69Γ— on inference and 1.57Γ— end to end. Note inference is 86% of the PyTorch frame but only 80% of the ONNX one β€” the faster the model, the more the rest of the pipeline matters, which is the whole reason for splitting the stages.

depth is the control row

It is identical numpy over the same boxes whatever the backend, so it must be constant across columns. It came out at 1.60 and 1.59 ms, and that agreement is what makes the rest of the table trustworthy.

It did not start that way. Unpinned, it read 1.5 ms under PyTorch and 16.2 ms under ONNX β€” a 10Γ— swing in code that never changed. onnxruntime defaults to intra_op_num_threads = 0, meaning every core, while torch was pinned to --threads; worse, its pool spin-waits between inferences rather than sleeping, so it starved the single-threaded depth stage even after CPU affinity capped the process. Ultralytics builds its session with no SessionOptions, so pipeline.configure_onnxruntime wraps the constructor to inject intra_op_num_threads and session.intra_op.allow_spinning = 0.

Without that, the benchmark measures thread budget rather than export format. compare.py warns when the depth row varies by more than 25% between columns, because that is the signature of the same mistake returning.

--threads sets CPU affinity for the whole process, not just torch, since no single library-level setting binds them all. It also makes an x86 run a better proxy for a Pi 5, which has four cores.

NCNN does not run on this machine

It loads β€” both files parse, the graph builds, the input is accepted β€” and then segfaults in the forward pass. Every op in the export is standard, and a trivial two-layer graph runs fine in the same wheel, so this is a version mismatch between the pnnx that wrote the export and the installed ncnn reading it: the weight layout changed between them.

A segfault cannot be caught in-process, so eval_ncnn.py runs one forward pass in a child process first, where a crash is only a return code, and prints the diagnosis above instead of taking the benchmark down with no output. Re-export on the Pi and it should run there; that is the number that matters anyway, since NCNN's advantage is ARM-specific and largely absent on x86.

Accuracy, with stock weights

0 of the four classes detected, as in phase 2 β€” COCO has no cone, barrel or barrier. The detector fires on car, person, truck, traffic light. Every detection is still routed through the depth estimator so the depth stage is timed on a realistic detection count (~5.7 per frame), but the distances are meaningless until the fine-tuned weights land.

Shipping the benchmark to another machine

python scripts/make_evalpack.py --frames 100 --tar

Builds evalpack/ (40 MB, 37 MB tarred): the three model formats, a fixed set of frames as real files rather than symlinks, the calibration and priors those frames need, the depth and benchmark code, and four shell scripts. On the far end:

./setup.sh                  # once: venv + pinned deps. Uses uv if present.
TORCH_CPU=1 ./setup.sh      # same, without the 2.5 GB CUDA build of torch
./run_all.sh                # three formats, then the comparison table

Frames are taken in sorted order rather than sampled, so two builds of the same size contain the same frames and two machines really are comparable. Only the calibrations those frames use are copied. The tarball excludes .venv and results/, since the pack is usually archived after someone has already run setup in it and a venv built for the wrong architecture is worse than useless on the far end.

THREADS defaults to 4 β€” a Pi 5's core count β€” and should be the same on every machine being compared, for the reason in the depth row discussion above.

Verified end to end on a clean venv: setup resolved Python 3.11 with torch 2.14, ultralytics 8.3.40, onnxruntime 1.30 and ncnn; run_all.sh produced PyTorch 34.88 ms and ONNX 21.16 ms inference with the depth control row at 0.98 and 1.01 ms, and NCNN reported its version mismatch without taking the run down.

Shipping the benchmark to another machine

Numbers from two machines are only comparable if both ran the same code over the same frames with the same models. scripts/make_evalpack.py builds a self-contained package that carries all of it:

python scripts/make_evalpack.py --frames 100 --tar
# ship evalpack.tar.gz, then on the target machine:
./setup.sh            # or TORCH_CPU=1 ./setup.sh on a GPU-less x86 box
./run_all.sh

40 MB: three model formats, 100 fixed frames as real files, the calibration and priors those frames need, the depth and benchmark code, and one shell script per format.

Three decisions in the builder worth knowing:

  • Images are copied, not symlinked. The rest of the repo symlinks into the dataset roots to avoid duplicating 8 GB, which is right at home and useless in something being sent elsewhere.
  • Frames are taken in sorted order, never sampled. Two builds of the same size contain the same frames, so two machines can be compared directly.
  • Only the calibrations those frames use are copied β€” otherwise the pack carries fifteen Argoverse 2 calibrations nothing references.

run_all.sh lets NCNN fail without taking the run down, since on x86 it segfaults on a pnnx/ncnn mismatch and the other two numbers are still worth having. THREADS defaults to 4, matching a Pi 5, and should be the same on every machine being compared.

Scoring a model end to end

evaluate.py asks whether the depth maths works. accuracy.py asks how good a particular checkpoint is end to end, and saves the answer so checkpoints can be compared weeks apart without re-running the earlier one.

python -m src.depth.accuracy --weights trained/yolo11n/weights/best.pt --tag best

Writes results/depth_accuracy/<tag>_detections.csv (one row per detection: class, confidence, predicted distance, true distance, error) and <tag>_summary.json (the aggregates plus the metadata needed to know what produced them). Collect at a low --conf; every threshold in the report is applied afterwards, so one run covers all operating points.

Detections are matched to ground truth greedily, most confident first, and each ground-truth object can be claimed once. Without that rule two overlapping detections of one cone both count and "objects found" exceeds 100% β€” which it did before this was fixed. A detection is correct only if it overlaps at IoU β‰₯ 0.5 and carries the right class.

Results, 759 val objects, nuScenes

keep above detections % real objects found distance error
0.05 2011 27% 72% 3.7 m
0.25 863 54% 61% 3.5 m
0.50 447 73% 43% 2.1 m
0.60 300 91% 36% 1.6 m
0.75 104 96% 13% 1.1 m

Confidence predicts distance accuracy, which it has no obvious right to do β€” it is a statement about the box, not the geometry:

confidence n error within 10% within 25%
0.05–0.25 83 4.2 m 31% 80%
0.25–0.50 137 6.9 m 21% 50%
0.50–0.75 225 5.6 m 21% 51%
0.75+ 100 1.1 m 63% 98%

The link is the bottom edge. The model is confident when an object is close and unoccluded, which is exactly when its box bottom sits cleanly on the road β€” and that pixel is the entire distance estimate. So confidence is a usable proxy for "trust this distance", which is worth knowing when deciding what to act on.

Both checkpoints

mAP50 mAP50-95 detections at 0.5 % real error
best (epoch 106) 0.511 0.306 447 73% 2.1 m
last (epoch 120) 0.509 0.302 505 69% 2.5 m

A 0.004 mAP gap and identical timing β€” they are the same model. Use best.

Per class, barrel is the outlier: mAP50 0.155 against 0.74 for stop_sign, despite having 11,295 training instances, more than any other class. AV2 barrels sit at a median 73 m, so nearly all of them are small boxes near the horizon.

Seeing it

view_detections.py draws detections with predicted and true distance, coloured by relative error, one frame per scene:

python -m src.depth.view_detections --weights trained/yolo11n/weights/best.pt --only-matched

What phase 2 consumes

Needs Comes from
fx, fy, cx, cy calib/<sensor_id>.yaml
camera height, ground normal same file β€” Calibration.ground_normal_cam()
per-class height/width priors class_priors.yaml
GT distance to score against manifest.gt_distance_m
detections to estimate from runs/detect/yolo11n/weights/best.pt

Known gaps

  • extract_av2.py reads the AV2 feather files directly rather than through the devkit dataloader. The layout is documented and stable, but it has not yet been run against real AV2 data β€” verify the first log's output with view_samples.py before extracting the rest.
  • COCO rows carry sensor_id = coco_unknown and have no calibration file, on purpose: those images have no shared intrinsics and phase 2 must not try to estimate distance from them.
  • The nuScenes v1.0-trainval metadata covers all 850 scenes while each *_blobs.tgz holds only a slice of the images, so downloading one or two blobs is normal. The extractor skips samples whose image is not on disk and reports the count as image_not_downloaded.

Licence

Ultralytics YOLO11 weights are AGPL-3.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support