Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
Paper • 2211.06687 • Published • 7
Core ML conversion of laion/larger_clap_general, used by Gridshift's smart sample search.
GridshiftCLAP.mlpackage.zip: audio encoder. Input audio: fp32 [1, 480000], mono 48 kHz, 10 s (HF repeatpad for shorter files). Output embedding: fp32 [1, 512], L2-normalized. Log-mel front-end in the graph (matches ClapFeatureExtractor), front-end ops fp32, int8 weights (convolutions excepted).GridshiftCLAPText.mlpackage.zip: text encoder. Inputs input_ids, attention_mask: int32 [1, 64] (RoBERTa tokenizer in tokenizer.tokenizer.zip). Output embedding: fp32 [1, 512], same space.Validated against the PyTorch reference (transformers ClapModel): audio cosine ≥ 0.9995 per sample, pairwise structure Δ ≤ 0.012; text cosine ≥ 0.999. SHA-256s are in manifest.json and text_manifest.json.
Upstream: LAION-CLAP, Apache-2.0. Wu et al., 2022, Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation (arXiv:2211.06687).
Base model
laion/larger_clap_general