Core ML
clap
audio
embeddings

CLAP (larger_clap_general) for Core ML — Gridshift

Core ML conversion of laion/larger_clap_general, used by Gridshift's smart sample search.

  • GridshiftCLAP.mlpackage.zip: audio encoder. Input audio: fp32 [1, 480000], mono 48 kHz, 10 s (HF repeatpad for shorter files). Output embedding: fp32 [1, 512], L2-normalized. Log-mel front-end in the graph (matches ClapFeatureExtractor), front-end ops fp32, int8 weights (convolutions excepted).
  • GridshiftCLAPText.mlpackage.zip: text encoder. Inputs input_ids, attention_mask: int32 [1, 64] (RoBERTa tokenizer in tokenizer.tokenizer.zip). Output embedding: fp32 [1, 512], same space.

Validated against the PyTorch reference (transformers ClapModel): audio cosine ≥ 0.9995 per sample, pairwise structure Δ ≤ 0.012; text cosine ≥ 0.999. SHA-256s are in manifest.json and text_manifest.json.

Upstream: LAION-CLAP, Apache-2.0. Wu et al., 2022, Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation (arXiv:2211.06687).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gridshiftstudio/clap-general-coreml

Finetuned
(3)
this model

Paper for gridshiftstudio/clap-general-coreml