Vanilla-Seq2seq Baseline (No Side Input) β€” Single Interfering Voice (NoSideInputSingleInterfering)

Single-source sequence-to-sequence speech enhancement baseline (Vanilla-Seq2seq / NoSideInput) trained on 24 kHz LibriTTS user speech mixed at 0 dB SNR with reverberant LJ Speech without any side input.

Open-Source Reproduction Notice: This pretrained model was trained using the standalone open-source reproduction library (wq2012/tec) built on lingvo and tensorflow. It does not use Google's proprietary internal codebase or internal training data infrastructure.


Model Performance

1. Open-Source Reproduction Evaluation (NoSideInputSingleInterfering)

Evaluated on 24 kHz LibriTTS (test-clean, test-other) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the Single interfering voice (LibriTTS + LJ Speech) condition, using Qwen3-ASR-0.6B-F16 via audio.cpp for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD:

Metric test-clean test-other Paper Reference (Vanilla-Seq2seq)
WER (%) ↓ 45.23% (90/199) 91.53% (162/177) 25.4% / 54.0%
MCD (dB) ↓ 9.58 dB 11.34 dB 7.85 dB / 8.84 dB
Side Input Size (KB) ↓ 0.000 KB 0.000 KB 0 KB
Computational Complexity ↓ 6.32 GFLOPS 6.32 GFLOPS 6.32 GFLOPS

2. Full Comparison Across All Pretrained Models & Baselines

Condition Method Hugging Face Model WER (%) test-clean ↓ WER (%) test-other ↓ MCD (dB) test-clean ↓ MCD (dB) test-other ↓ Side Input test-clean (KB) ↓
Single interfering voice (LibriTTS + LJSpeech) GroundTruth β€” 3.52 6.78 0.00 0.00 0.000
MicrophoneSignal β€” 90.45 114.12 12.86 14.61 0.000
NlmsAec (AEC-NLMS) β€” 88.44 107.34 12.80 14.48 243.465
NoSideInputSingleInterfering wq2012/vanilla_seq2seq_single_interfering 45.23 91.53 9.58 11.34 0.000
AecSingleInterfering wq2012/aec_single_interfering 12.06 23.16 8.85 9.86 243.465
TecSingleInterfering wq2012/tec_single_interfering 21.61 46.89 8.24 9.28 0.076
Multiple interfering voices (LibriTTS + VCTK) GroundTruth β€” 5.03 7.82 0.00 0.00 0.000
MicrophoneSignal β€” 34.17 48.97 7.67 7.70 0.000
NlmsAec (AEC-NLMS) β€” 28.64 34.98 7.92 8.40 186.922
NoSideInputMultiInterfering wq2012/vanilla_seq2seq_multi_interfering 31.16 42.39 7.93 8.72 0.000
AecMultiInterfering wq2012/aec_multi_interfering 8.54 22.22 7.80 7.88 186.922
TecMultiInterfering wq2012/tec_multi_interfering 26.63 45.27 7.96 8.40 0.037

Files in This Repository

  • best.ckpt.data-00000-of-00001, best.ckpt.index, best.ckpt.meta, checkpoint: TensorFlow / Lingvo checkpoint for NoSideInputSingleInterfering.
  • model.tflite: Dynamic-range quantized TensorFlow Lite (.tflite) FlatBuffer model for on-device inference.
  • evaluation_metrics.json: Verified evaluation results (test-clean and test-other) for NoSideInputSingleInterfering.

How to Use

1. Install textual-echo-cancellation

pip3 install textual-echo-cancellation huggingface_hub

2. Download the Model from Hugging Face

from huggingface_hub import snapshot_download

model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_single_interfering")
print("Downloaded model to:", model_dir)

3. Run Inference via CLI (scripts/inference.py)

python3 -m scripts.inference \
  --model NoSideInputSingleInterfering \
  --checkpoint_path "${MODEL_DIR}/best.ckpt" \
  --mixed_wav /path/to/mixed_input.wav \
  # No side input required for Vanilla-Seq2seq \
  --output_wav /tmp/enhanced_clean.wav

4. Run Inference via Python API

import os
from huggingface_hub import snapshot_download
from tec import inference

model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_single_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")

result = inference.run_inference_on_wav(
    model_name="NoSideInputSingleInterfering",
    mixed_wav_path="/path/to/mixed_input.wav",
    interfering_text="currently in mountain view it is 72 degrees",
    checkpoint_path=ckpt_path,
    output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)

5. On-Device Inference with Quantized TFLite (model.tflite)

import os
import numpy as np
import tensorflow as tf
from huggingface_hub import snapshot_download

model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_single_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")

interpreter = tf.lite.Interpreter(model_path=tflite_path)
interpreter.allocate_tensors()

for detail in interpreter.get_input_details():
    interpreter.set_tensor(
        detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"])
    )

interpreter.invoke()
output_details = interpreter.get_output_details()
enhanced_mel = interpreter.get_tensor(output_details[0]["index"])
print("TFLite predicted log-Mel shape:", enhanced_mel.shape)

Citation

If you use this model or the textual-echo-cancellation library in your research, please cite the original paper:

@inproceedings{ding2021textual,
  title={Textual Echo Cancellation},
  author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
  booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
  pages={666--673},
  year={2021},
  organization={IEEE}
}
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using wq2012/vanilla_seq2seq_single_interfering 1

Collection including wq2012/vanilla_seq2seq_single_interfering

Paper for wq2012/vanilla_seq2seq_single_interfering