Variable-Array Reconstruction for Latent Acoustic Mapping
VAR-LAM is a research codebase for Variable-Array Reconstruction for Latent Acoustic Mapping. It predicts complex cross-spectral matrices (CSMs) from multi-channel audio by combining a learned STFT encoder, residual vector quantisation, microphone-pair reasoning, a factorised transformer, and a LAM-style steering-matrix decoder.
What This Repository Contains
- Checkpoints: published VAR-LAM baselines, ablations, direction-grid variants, and 4-to-32-channel upsampling models.
- Source: the complete implementation is maintained on GitHub.
- Documentation: setup, training, inference, and evaluation documentation is available at philippxxy.github.io/var-lam.
Setup
Install the source repository and its dependencies:
git clone https://github.com/PhilippXXY/var-lam.git
cd var-lam
uv sync --all-groups
The source repository contains small JSON pointers for published models. Its loader downloads checkpoints from this model repository and reuses the Hugging Face cache.
To download a standalone checkpoint explicitly:
hf download PhilippXXY/var-lam \
checkpoints/32ch-to-32ch/ablations/components/transformer-off.pt \
--revision main \
--local-dir .
Checkpoints
Checkpoint paths are organised by channel mapping and experiment:
checkpoints/
βββ 32ch-to-32ch/
β βββ baseline.pt
β βββ direction-grid/{64,128,256,512,1024,2048,4096,8192}.pt
β βββ frequency/{linear-0hz-1khz,linear-0hz-4khz,mel-50hz-12khz}.pt
β βββ input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
β βββ ablations/
β βββ components/{cross-channel-attention-on,lstm-off-temporal-attention-off,lstm-off-token-comparison-off-transformer-off,token-comparison-no-diagonal,token-comparison-no-differences,token-comparison-no-geometry,token-comparison-no-products,token-comparison-off-transformer-off,token-comparison-off,transformer-off}.pt
β βββ transformer-attention/{band-only,pair-only}.pt
β βββ rvq/{q2-k1024,q4-k1024,q4-k4096,q8-k64,q8-k256,q8-k512,q8-k2048,q12-k16,q16-k1024}.pt
βββ 4ch-to-32ch/
β βββ baseline.pt
β βββ direction-grid/{64,128,256,512,1024,2048,4096,8192}.pt
β βββ input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
β βββ ablations/
β βββ components/{cross-channel-attention-on,lstm-off-token-comparison-off-transformer-off,rvq-off,token-comparison-off-transformer-off}.pt
β βββ microphone-layout/{horizontal,random,vertical-front}.pt
βββ 4ch-to-4ch/
βββ baseline.pt
βββ input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
βββ ablations/components/{cross-channel-attention-on,lstm-off-temporal-attention-off,lstm-off-token-comparison-off-transformer-off,token-comparison-off-transformer-off,token-comparison-off,transformer-off}.pt
transformer-attention/ contains ablations isolated to the factorised transformer's attention axes. Optional modules and combined component ablations remain under components/.
The two direction-grid/128.pt files are aliases of their corresponding baselines, whose decoder direction grid contains 128 points.
Training
Training code, configuration, and instructions live in the source repository. Training writes timestamped latest.pt and best.pt files locally; publish selected outputs here only when they are ready for reuse.
Inference
Run a pointer-backed checkpoint through the source repository:
uv run src/infer.py \
--config configs/inference_config.yaml \
--device cuda \
--checkpoint checkpoints/32ch-to-32ch/ablations/components/transformer-off.pt \
--checkpoint-revision main \
--dataset locata
The loader reconstructs the architecture and channel-input mode from checkpoint metadata before loading the state dictionary.
Documentation
See the rendered VAR-LAM documentation for dataset setup, configuration, training, inference, metrics, and visualisation workflows.