HOA7 Spatial Field Decoder (hoa64)
A formula-driven spatial field decoder for 7th-order Ambisonics — 64 channels (Ambix ACN + SN3D). It encodes/decodes spherical fields, rotates them with Wigner-D matrices, analyzes direction-of-arrival / energy, fuses vision detections onto the sphere, and emits spatial conditioning for generative pipelines. Pure NumPy; not an LLM and not a learned model — deterministic spherical-harmonic geometry, no weights.
| Property | Choice |
|---|---|
| Version | 0.5.0 |
| Geometry | Ambix ACN + SN3D |
| Channels | 64 = (7+1)², order 7 |
| Rotation | Wigner-D (~150× vs dense) |
| Audio | WAV / live mic (ffmpeg·arecord) |
| Vision | Boxes / YOLO labels / optional torchvision |
| Agents | HTTP :8765 + spatial-report CLI |
| Diffusion | Conditioning JSON + optional ComfyUI submit |
| Language | Python 3, NumPy only (optional torch/torchvision) |
What it does
- Encode point sources / plane waves / scene mixes → HOA-7 coefficients (
encode_points,encode_plane_waves,encode_scene). - Decode coefficients → samples on the sphere, product-grid render, or a single beamform readout (
decode_directions,decode_grid,beamform). - Rotate the field in the listener frame via Wigner-D / zyz (
hoa_rotation_matrix,apply_hoa_rotation,rotate_yaw_pitch_roll). - Analyze DOA (intensity vector + peak search), directional power, field energy, per-STFT-band and per-frame reports (
analysis.py,report.py). - Fuse vision — object boxes / rays / YOLO detections projected onto the sphere alongside audio (
vision.py,detector.py,fuse_reports). - Condition — spatial reports → plain-text control lines for T2I/T2V prompts, structured JSON for ControlNet-style nodes, or a ComfyUI API payload (
conditioning.py). - Serve — HTTP API and iterative agent state (
server.py,rnn_stub.py).
Coordinates (always)
- +X front, +Y left, +Z up (Ambix listener frame).
- Azimuth 0° = front, +90° = left, −90° = right.
- Elevation 0° = horizon, +90° = zenith.
- W (omnidirectional HOA channel) maps to field size / POV: low W → tight / subject-focused / narrow FOV; high W → wide / environmental / immersive FOV.
Install / run
cd spatial-hoa # this repo root
export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}"
python3 -m hoa64 --help
# or install the CLI on PATH:
ln -s "$PWD/scripts/spatial-report" ~/.local/bin/spatial-report
Deps: NumPy (system ffmpeg/arecord for live capture; optional torch/torchvision for the real object detector).
CLI map (spatial-report)
| Command | Purpose |
|---|---|
analyze |
Ambix / mono WAV → spatial JSON |
demo-scene |
Synthetic multi-source audio |
vision |
Raw sphere boxes → report |
detect |
Image / YOLO / demo → report |
live |
Mic capture → report |
condition |
Report → diffusion prompt + control vector |
serve |
HTTP API :8765 |
HTTP API (POST /v1/spatial/analyze)
Modes: demo_scene · ambix_file · mono_file · vision · fuse · detect · live · condition
curl -s -X POST http://127.0.0.1:8765/v1/spatial/analyze \
-H 'Content-Type: application/json' \
-d '{"mode":"demo_scene","order":3}'
An OpenAI-compatible function schema is provided at tools/spatial_analyze.openai.json.
Agent integration (Qwythos / Pi)
The HOA-7 calculator is not part of any LLM's weights. Agents should call it as a tool and treat the returned one_liner / doa_* / fuse fields as ground-truth geometry:
- Run the CLI:
spatial-report analyze /path/to.wav --ambix -o /tmp/spatial.json && cat /tmp/spatial.json - Or HTTP:
POST http://127.0.0.1:8765/v1/spatial/analyze - See
integrations/qwythos_system_snippet.mdfor the exact agent prompt block.
Quick test pack
python3 examples/demo_e2e_testpack.py # artifacts in /tmp/spatial_hoa_e2e/
spatial-report analyze /tmp/spatial_hoa_e2e/scene_ambix4.wav --ambix -o /tmp/a.json
spatial-report detect --demo-image /tmp/frame.png -o /tmp/v.json
spatial-report condition /tmp/spatial_hoa_e2e/fuse_report.json --prompt 'cinematic interior' -o /tmp/c.json
Tests
python3 tests/test_basis.py
python3 tests/test_encode_decode.py
python3 tests/test_rotate_rnn.py
python3 tests/test_phase1_audio.py
python3 tests/test_phase2_wigner.py
python3 tests/test_phase3_vision.py
python3 tests/test_integration_extras.py
Limitations
- Deterministic geometry, not a learned model. There are no trained weights; accuracy is bounded by the exact spherical-harmonic basis (Farina / Ambix), not by data. The
rnn_stub.pyintegrator is an explicit Euler pose/rotation loop — learned field dynamics are not implemented yet. - No audio synthesis. This decodes/analyzes spatial fields and produces conditioning; it does not generate audio content.
- Live capture requires system
ffmpeg/arecord. - Object detection is optional: the demo path uses synthetic boxes; the real detector downloads torchvision weights on first use.
- Coordinate convention is Ambix ACN/SN3D. Interop with other conventions (FuMa, N3D) requires explicit conversion.
License
This repository carries no explicit license file. The code is provided as-is by the author; contact woodfireind for usage terms. Third-party optional deps (torch/torchvision) keep their own licenses.