rnnoise-coreai

Original RNNoise model and pretrained weights: Jean-Marc Valin and the RNNoise/Xiph.Org contributors. Original copyright notices also credit Amazon, Mozilla, Xiph.Org Foundation, Mark Borgerding, and contributors named in individual source files.

Official upstream: Xiph RNNoise (canonical repository). Original paper: Jean-Marc Valin, A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement. The published checkpoints use the current RNNoise architecture.

License: upstream BSD-style notices, preserved in vendor/rnnoise/LICENSE and source files.

Core AI conversion, packaging, and validation: Max Farrell. Original architectures and pretrained parameters are credited to the upstream projects; no model training was performed for these conversions. This is an independent community conversion, without upstream or Apple endorsement.

Pretrained float32 Core AI assets for RNNoise (10Ga and 10Gb). CPU parity validated on macOS 27.0.1.

This Hub repository includes the named model assets and original checkpoint(s). Clone the complete source project for all model variants and the full documented workflow. Model-specific licenses are in LICENSE and vendor; new conversion code is MIT as recorded in the source project.

Speech enhancement models for Apple Core AI

Pretrained RNNoise and UL-UNAS streaming neural networks exported as self-contained .aimodel directories. Conversion and validation by Max Farrell; original models and weights by the upstream authors linked below. This is an independent community conversion project.

Model Checkpoint Asset License
RNNoise official rnnoise10Ga_12.pth exports/rnnoise10Ga_12_float32_streaming.aimodel upstream BSD-style notices
RNNoise official rnnoise10Gb_15.pth exports/rnnoise10Gb_15_float32_streaming.aimodel upstream BSD-style notices
UL-UNAS DNS3 model_trained_on_dns3.tar exports/ulunas_dns3_float32_streaming.aimodel MIT

Download complete model directories from this repository, the GitHub release, or the dedicated model repositories on Hugging Face: RNNoise Core AI and UL-UNAS Core AI. Integrity hashes are in SHA256SUMS; upstream revisions and checkpoint hashes are in provenance.json.

Reproduce

An Apple Silicon Mac with the Core AI runtime is required to run parity checks. Validated on macOS 27.0.1 (26A434), with CPU specialization. Export uses PyTorch 2.11.0, coreai-torch 0.4.3 and coreai-core 1.0.0b3. GPU, Neural Engine, iOS, and earlier OS releases have not been validated. No accelerator latency or energy-efficiency claim is made.

git clone https://github.com/maxffarrell/speech-enhancement-coreai.git
cd speech-enhancement-coreai
uv sync --frozen
# Check the published assets:
uv run python convert.py rnnoise10Ga_12 --validate-only
uv run python convert.py rnnoise10Gb_15 --validate-only
uv run python convert.py ulunas_dns3 --validate-only
# Optional real-audio check, using your own 16 kHz mono file:
uv run python convert.py ulunas_dns3 --validate-only --wav noisy.wav
# To regenerate, move the matching exports/*.aimodel directory elsewhere first:
uv run python convert.py ulunas_dns3 --dtype float32

The exporter refuses to overwrite existing assets. It loads checkpoint dictionaries using weights_only=True and strict state-dict matching, exports with torch.export.export, runs Apple's decomposition table, then saves the Core AI AIProgram with provenance/license metadata. No retraining, checkpoint edits, quantization, or third-party Metal kernels are involved.

UL-UNAS audio integration

Upstream code, checkpoints and paper. The checkpoint is trained on DNS3. The upstream frame-streaming implementation is used, with input caches cloned at the export boundary to preserve functional state semantics.

Audio is 16 kHz mono, with a 512-point periodic Hann-window STFT, 256-sample hop (16 ms), centered reflect padding as in upstream PyTorch, and matching inverse STFT. The exported graph accepts one real/imaginary STFT frame, rather than raw audio.

Input float32 shape Output
mix [1,257,1,2] enh (same shape)
conv_cache [1,5358] conv_cache_out
tfa_cache [1,402] tfa_cache_out
inter_cache [1,1056] inter_cache_out

Initialize caches to zero per stream. Feed each returned cache into the matching input for the next frame. Reset all caches when starting a new stream. Centered STFT requires buffering; the 16 ms hop is not a claim of zero-lookahead end-to-end latency.

uv run python enhance.py noisy.wav enhanced.wav

This complete WAV example runs the neural network in Core AI on CPU; PyTorch is used only for STFT/ISTFT. It preserves upstream handling of the trailing audio samples. It does not resample or downmix automatically.

RNNoise integration

Upstream RNNoise (canonical repository). Both checkpoints come from the exact official model archive in provenance.json; its SHA-256 matches upstream model_version. They use the current 65-feature, 32-gain architecture, with 128 convolution conditioning channels and 384 recurrent units. No claim is made that either checkpoint exactly matches upstream's quantized default C model.

RNNoise operates on 48 kHz mono, 480-sample (10 ms) audio frames. Its native DSP produces 65 normalized features. The Core AI graph converts these to 32 band gains and a voice-activity score. Native pitch analysis/filtering, feature extraction, gain smoothing/application, and waveform synthesis remain external. This asset alone does not accept PCM or generate cleaned audio. Preserve upstream feature conventions and DSP when integrating; do not pass raw PCM, GTCRN features, or arbitrary STFT bins as features.

Input float32 shape Output
features [1,1,65] gains [1,1,32], vad [1,1,1]
conv1_cache [1,65,2] conv1_cache_out
conv2_cache [1,128,2] conv2_cache_out
gru1_state [1,1,384] gru1_state_out
gru2_state [1,1,384] gru2_state_out
gru3_state [1,1,384] gru3_state_out

Initialize all five states to zero and carry all outputs forward on every frame. The convolution caches make the valid-convolution training graph causal for frame-wise inference. Cold-start caches match the native streaming convention. The source check compares against upstream PyTorch sliding five-frame windows after the four-frame warmup, with identical incoming GRU states.

Validation and limits

Every output is compared over 32 consecutive frames with independent PyTorch and Core AI state trajectories, including convolution caches and GRU states. Reports are in docs/.

  • RNNoise 10Ga float32: minimum 104.44 dB across all outputs; source graph agreement 119.08 dB.
  • RNNoise 10Gb float32: minimum 108.11 dB; source graph agreement 115.74 dB.
  • UL-UNAS float32 on upstream real-audio frames: minimum 127.83 dB; streaming versus upstream offline spectral graph 160.58 dB. A separate deterministic synthetic-input report is also included.
  • Full 10-second UL-UNAS WAV example versus upstream waveform output: 144.11 dB PSNR, 160,000 output samples, all finite.
  • RNNoise float16 was rejected: VAD minimum PSNR was 35.13 dB, below the 40 dB gate. UL-UNAS upstream float16 export hits a recurrent input/weight dtype mismatch. Only validated float32 assets are distributed.

These are conversion-parity checks, not speech-quality benchmark scores. RNNoise testing uses synthetic normalized-feature-shaped inputs and does not establish parity with native quantized C/DSP, end-to-end audio quality, or latency. UL-UNAS real-audio testing uses the upstream audio/noisy/0174.wav sample locally; that recording is not redistributed. A conversion can reproduce upstream outputs without establishing quality on your microphone or noise environment.

Licensing and attribution

New conversion/example code: MIT. UL-UNAS source and weights retain upstream MIT. RNNoise source and weights retain upstream BSD-style license and individual source-file notices. The root license does not replace those terms. See third-party notices for authors, exact revisions, modifications, and citations. Model weights use the upstream repository licenses; no separate checkpoint-specific license was supplied upstream.

Related: GTCRN Core AI, including DNS3 and VCTK streaming assets. GTCRN is maintained separately and is not redistributed by this project.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for maxffarrell/rnnoise-coreai