Silero VAD for Core AI

The official Silero 16 kHz VAD converted for public Apple Core AI APIs on macOS 27 and iOS 27. The bundle batches eight consecutive 32 ms frames (256 ms), while preserving the original 128-dimensional LSTM state and 64-sample audio context across calls.

Use CPU-only specialization for this small recurrent model. Each input audio tensor has shape [8, 1, 1, 576], containing the preceding 64 samples and current 512 samples for each frame. h and c are [1, 128, 1, 1]. Outputs are probability, next_h and next_c. All tensors use FP16. Begin a new recording with zero states and context; do not reset between updates. Buffer incomplete batches until more audio arrives; zero-pad only the final batch and discard padded probabilities.

On an M3 MacBook Air, CPU-only specialization took 0.10 seconds. Processing 688 frames (22.0 seconds including validation silence and noise) took 0.024 seconds, excluding preparation and file I/O. Compared with the official FP32 JIT checkpoint, mean probability error was 0.00081, maximum error 0.01681, and no decisions differed at the 0.5 threshold in this test. These are conversion checks, not a speech-detection accuracy benchmark.

metadata.json defines the batch and state conventions. Hugging Face provides checksums for downloadable files in each repository snapshot. No ahead-of-time compiled assets are included.

Source: snakers4/silero-vad, revision 5cd7945676eb32225748052e2e6a0580e4686a08. Weights originate from the official silero_vad.jit, not a third-party conversion. The original MIT license is included.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including coder543/silero-vad-coreai