VIQ
VIQ extracts one or more people from video and packages the result as background-free media for compositing and AR prototyping. The active pipeline uses segmentation, matte cleanup, timing checks, and verified alpha-video export; it no longer pretends that training a clip classifier creates a cutout or a 3D person.
What it produces
masks/%06d.pngโ one grayscale alpha matte per source frame (the lossless source of truth)subject-alpha.movโ ProRes 4444 with alpha for editing and interchangesubject-alpha-hevc.movโ compact HEVC with alpha for AVFoundation and RealityKit (macOS only)subject-alpha.webmโ VP9 alpha for compatible web runtimespreview.mp4โ the cutout composited over a dark background for quick reviewmanifest.jsonโ source metadata, model settings, outputs, and mask-quality diagnostics
Multi-person split runs place the same deliverables under subjects/subject-<tracker-id>/; both
runs also create a combined/ scene containing all selected people.
The default export is only preview.mp4; this avoids unexpectedly creating several very large alpha
masters. Add the formats you need, for example --formats hevc,preview on a Mac or
--formats webm,preview for a compatible web target.
An alpha cutout is 2D/2.5D, not a complete 3D human. It can face the viewer as a billboard in AR. True side and rear views require a reconstruction stage and sufficient multi-view coverage. See the deep-research audit for current technology and licensing decisions.
Base setup
Requirements: Python 3.12+ and FFmpeg with ffprobe, alphamerge, and libx264. ProRes, VP9, true-360,
and Apple HEVC outputs have additional checks reported by viq doctor.
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e .
viq doctor
viq inspect 360.mp4
The media renderer is independent of the AI backend. If masks came from an editor, another model, or
another machine, name them 000000.png, 000001.png, โฆ at the exact source-video resolution and run:
viq render 360.mp4 \
--masks path/to/masks \
--formats hevc,preview \
--output outputs/360-person
Frame-numbered masks currently require constant-frame-rate input. viq inspect reports timing,
rotation, declared projection, and uncompressed working-set estimates. VIQ rejects detected
variable-frame-rate sources before expensive extraction; normalize those clips to CFR first.
Equirectangular / 360 extraction
VIQ can segment monoscopic equirectangular footage without asking the model to interpret the severe
distortion at the poles. It temporarily converts each panorama to a 3ร2 cubemap with FFmpeg's
v360 filter, runs the selected backend on that view,
then reprojects the completed masks to the original equirectangular resolution. Rendering always uses
the untouched source video.
viq extract panorama.mp4 \
--projection equirectangular \
--backend rembg \
--formats hevc,preview \
--output outputs/panorama-person
--projection auto is the default. It enables the cubemap path only when the file declares spherical
or equirectangular projection metadata; a 2:1 aspect ratio alone is reported as a hint, not treated as
proof. Use the explicit flag above for untagged panoramic exports, and use --projection flat to
override incorrect metadata.
The automatic cube-face size preserves the lower of the source's horizontal and vertical angular
sampling (4096ร2048 becomes a 3072ร2048 3ร2 cubemap). Use --cube-face-size 512 to trade detail
for faster inference. The chosen mode and working resolution are recorded in manifest.json.
This mode currently targets monoscopic 2:1 equirectangular input. Stereo over/under or side-by-side
360, partial/tiled equirectangular, dual-fisheye camera originals, and direct cubemap sources must be
converted or split first; auto rejects these layouts when their metadata declares them.
Cubemap face boundaries can still challenge a segmentation model when a subject crosses them, so
review the preview and exported masks around seams and poles.
The mode also requires an approximately 2:1 frame with no unresolved display-rotation metadata.
The masks and encoded frames retain the equirectangular pixel layout, but VIQ does not yet inject Spherical Video V2 projection boxes into filtered outputs. Many 360 players require those container atoms to activate spherical playback. The manifest records this explicitly; inject target-specific spatial metadata downstream when the result is meant to remain a navigable 360 video rather than an alpha/compositing asset.
Local CPU / Mac extraction
Install the optional rembg ONNX backend for a practical local fallback. Its human-segmentation model is downloaded on first use. It is lighter than SAM 3, although its frame-to-frame stability, crowded- scene accuracy, and fine hair edges are weaker.
pip install -e '.[local]'
viq extract 360.mp4 \
--backend rembg \
--formats hevc,preview \
--output outputs/360-person
The default --backend auto uses SAM 3 when a CUDA setup is ready and otherwise uses rembg when
installed. On supported Macs, rembg automatically uses Apple's Core ML execution provider; pass
--execution-provider cpu only for troubleshooting.
The rembg backend is a per-frame fallback, not a state-of-the-art video-matting model. Its legacy mask-averaging option now defaults to zero because averaging adjacent frames can leave motion ghosts. rembg itself is MIT-licensed, while downloaded model weights may have separate terms; review the exact model terms for production use.
MatAnyone 2 alpha refinement
MatAnyone 2 is now VIQ's preferred non-commercial refinement stage. It is the CVPR 2026 successor to MatAnyone, preserves fine boundary detail over time, and its official device selection supports CUDA, Apple MPS, and CPU. VIQ starts it from the first source frame where a selected subject has a non-empty segmentation mask, zero-pads earlier frames, then uses its temporally propagated alpha sequence for rendering.
Install the pinned runtime separately so its Torch dependencies cannot disturb rembg or SAM:
viq install-matanyone2
viq doctor
The installer works around an upstream wheel-packaging conflict, pins the compatible Torch 2.8 /
TorchVision 0.23 pair, and omits unrelated GUI/training packages. It installs under
.viq-tools/matanyone2, which --matting auto discovers automatically. The first refined extraction
downloads the 135 MiB MatAnyone 2 checkpoint and two pretrained ResNet backbones.
viq extract 360.mp4 \
--backend rembg \
--matting matanyone2 \
--formats hevc,preview \
--output outputs/360-person
--matting auto is the default: it uses MatAnyone 2 when installed and otherwise retains the coarse
segmentation masks. Use --matting none for a deliberate fast run. The target must be present in the
coarse mask on at least one frame. For large footage, --matting-max-size 1080 reduces memory and inference
time, then VIQ resizes the alpha back to the exact source resolution. The official runtime currently
loads the complete processing clip into memory, so split very long clips before refining them.
MatAnyone 2 uses the NTU S-Lab License 1.0 and is suitable here because this project is non-commercial. The selected model, license, parameters, executable, and cache path are recorded in manifest schema 5. SAM2Matting is newer as a generalized research system, but its current video release is CUDA-only and pins Torch-TensorRT, so it is not the default for this Mac-first pipeline.
SAM 3.1 extraction
Meta's official SAM 3.1 implementation currently requires an NVIDIA CUDA GPU, CUDA 12.6+, PyTorch 2.7+, and access to the gated checkpoint. Follow the official SAM 3 installation and authentication instructions, then install VIQ into that same environment.
hf auth login
viq doctor
viq extract 360.mp4 \
--backend sam3 \
--prompt person \
--formats hevc,preview \
--output outputs/360-person
Multiple people and scene splitting
SAM 3.1 assigns a run-scoped tracker ID to each person matched by the text prompt. VIQ can keep those people together, create a separate alpha scene for every person, or create both forms without running segmentation more than once:
# All detected people together. This remains the backward-compatible default.
viq extract clip.mp4 --backend sam3 --subject-output combined --output outputs/group
# One self-contained output scene per detected person.
viq extract clip.mp4 --backend sam3 --subject-output split --output outputs/people
# Individual scenes plus a combined scene.
viq extract clip.mp4 --backend sam3 --subject-output both --output outputs/people-and-group
After a split run, inspect the subject previews or the root manifest.json, then repeat
--object-id to make a scene containing only a chosen subset:
viq extract clip.mp4 \
--backend sam3 \
--object-id 2 \
--object-id 5 \
--subject-output combined \
--output outputs/people-2-and-5
--people-output is an alias for --subject-output. Tracker IDs belong only to that extraction run;
they are not biometric identities and should not be reused after a fresh run. Split and both modes
require the instance-aware SAM 3.1 backend. The portable rembg backend can keep everyone together but
returns one merged human mask, so it cannot reliably split or select people.
With MatAnyone 2 enabled, VIQ refines each person independently and constructs the combined scene by taking the pixelwise maximum of those individual alpha sequences. A person entering after frame zero starts matting at their first visible coarse mask and remains transparent on earlier frames. This workflow also applies to equirectangular input: subjects are split and refined in cubemap space, then each final matte is reprojected independently. Every subject directory has its own manifest, and the root schema-5 manifest catalogs the full scene collection.
Useful corrections:
# Prompt on a frame where the subject is clearly visible.
viq extract 360.mp4 --prompt person --prompt-frame 75 --output outputs/360-person
# After a first run, review segmentation.statistics.objects in manifest.json,
# then keep one SAM-tracked person on a replacement run.
viq extract 360.mp4 --prompt person --object-id 2 --overwrite --output outputs/360-person
# Preserve editing, web, preview, and every RGBA frame.
viq extract 360.mp4 --formats mov,webm,preview,png --output outputs/360-person
On macOS, --formats hevc first creates a temporary ProRes 4444 master and then uses Apple's
AVFoundation HEVC-with-alpha preset. VIQ verifies that the resulting track advertises an alpha channel
before accepting it. This is the intended compact asset for RealityKit.VideoMaterial; ProRes remains
the editing/interchange master.
VIQ refuses to overwrite existing deliverables by default. Add --overwrite only when replacement is
intentional. --compile can improve repeated SAM 3.1 inference on supported high-end GPUs but makes the
first run slower.
Development
pip install -e '.[dev]'
pytest
ruff check viq tests scripts main.py
ruff format --check viq tests scripts main.py
The older models/, modules/, and data_processing/ directories are retained as historical
prototype code. They are not imported by VIQ 0.6 and are intentionally excluded from the production
dependency set.