AEC-Seq2seq Neural Baseline β Single Interfering Voice (AecSingleInterfering)
Multi-source attention sequence-to-sequence Acoustic Echo Cancellation (AEC-Seq2seq) neural baseline model trained on 24 kHz LibriTTS user speech mixed at 0 dB SNR with reverberant LJ Speech (single-speaker TTS interference). Uses dual audio encoders (SpeechEncoderV1) on both the microphone mixture and the full reference TTS playback audio (~240β310 KB side input).
- GitHub Repository: https://github.com/wq2012/tec
- PyPI Package:
textual-echo-cancellation - Paper: Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)
- Audio Demo Page: https://google.github.io/speaker-id/publications/TEC/
Open-Source Reproduction Notice: This pretrained model was trained using the standalone open-source reproduction library (
wq2012/tec) built onlingvoandtensorflow. It does not use Google's proprietary internal codebase or internal training data infrastructure.
Model Performance
1. Open-Source Reproduction Evaluation (AecSingleInterfering)
Evaluated on 24 kHz LibriTTS (test-clean, test-other) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the Single interfering voice (LibriTTS + LJ Speech) condition, using Qwen3-ASR-0.6B-F16 via audio.cpp for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD:
| Metric | test-clean |
test-other |
Paper Reference (AEC-Seq2seq) |
|---|---|---|---|
| WER (%) β | 12.06% (24/199) | 23.16% (41/177) | 8.30% / 24.3% |
| MCD (dB) β | 8.85 dB | 9.86 dB | 6.38 dB / 7.07 dB |
| Side Input Size (KB) β | 243.465 KB | 209.085 KB | 310 KB |
| Computational Complexity β | 9.51 GFLOPS | 9.51 GFLOPS | 9.51 GFLOPS |
2. Full Comparison Across All Pretrained Models & Baselines
| Condition | Method | Hugging Face Model | WER (%) test-clean β | WER (%) test-other β | MCD (dB) test-clean β | MCD (dB) test-other β | Side Input test-clean (KB) β |
|---|---|---|---|---|---|---|---|
| Single interfering voice (LibriTTS + LJSpeech) | GroundTruth |
β | 3.52 | 6.78 | 0.00 | 0.00 | 0.000 |
MicrophoneSignal |
β | 90.45 | 114.12 | 12.86 | 14.61 | 0.000 | |
NlmsAec (AEC-NLMS) |
β | 88.44 | 107.34 | 12.80 | 14.48 | 243.465 | |
NoSideInputSingleInterfering |
wq2012/vanilla_seq2seq_single_interfering |
45.23 | 91.53 | 9.58 | 11.34 | 0.000 | |
AecSingleInterfering |
wq2012/aec_single_interfering |
12.06 | 23.16 | 8.85 | 9.86 | 243.465 | |
TecSingleInterfering |
wq2012/tec_single_interfering |
21.61 | 46.89 | 8.24 | 9.28 | 0.076 | |
| Multiple interfering voices (LibriTTS + VCTK) | GroundTruth |
β | 5.03 | 7.82 | 0.00 | 0.00 | 0.000 |
MicrophoneSignal |
β | 34.17 | 48.97 | 7.67 | 7.70 | 0.000 | |
NlmsAec (AEC-NLMS) |
β | 28.64 | 34.98 | 7.92 | 8.40 | 186.922 | |
NoSideInputMultiInterfering |
wq2012/vanilla_seq2seq_multi_interfering |
31.16 | 42.39 | 7.93 | 8.72 | 0.000 | |
AecMultiInterfering |
wq2012/aec_multi_interfering |
8.54 | 22.22 | 7.80 | 7.88 | 186.922 | |
TecMultiInterfering |
wq2012/tec_multi_interfering |
26.63 | 45.27 | 7.96 | 8.40 | 0.037 |
Files in This Repository
best.ckpt.data-00000-of-00001,best.ckpt.index,best.ckpt.meta,checkpoint: TensorFlow / Lingvo checkpoint forAecSingleInterfering.model.tflite: Dynamic-range quantized TensorFlow Lite (.tflite) FlatBuffer model for on-device inference.evaluation_metrics.json: Verified evaluation results (test-cleanandtest-other) forAecSingleInterfering.
How to Use
1. Install textual-echo-cancellation
pip3 install textual-echo-cancellation huggingface_hub
2. Download the Model from Hugging Face
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering")
print("Downloaded model to:", model_dir)
3. Run Inference via CLI (scripts/inference.py)
python3 -m scripts.inference \
--model AecSingleInterfering \
--checkpoint_path "${MODEL_DIR}/best.ckpt" \
--mixed_wav /path/to/mixed_input.wav \
--interfering_wav /path/to/reference_tts.wav \
--output_wav /tmp/enhanced_clean.wav
4. Run Inference via Python API
import os
from huggingface_hub import snapshot_download
from tec import inference
model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")
result = inference.run_inference_on_wav(
model_name="AecSingleInterfering",
mixed_wav_path="/path/to/mixed_input.wav",
interfering_text="currently in mountain view it is 72 degrees",
checkpoint_path=ckpt_path,
output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)
5. On-Device Inference with Quantized TFLite (model.tflite)
import os
import numpy as np
import tensorflow as tf
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")
interpreter = tf.lite.Interpreter(model_path=tflite_path)
interpreter.allocate_tensors()
for detail in interpreter.get_input_details():
interpreter.set_tensor(
detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"])
)
interpreter.invoke()
output_details = interpreter.get_output_details()
enhanced_mel = interpreter.get_tensor(output_details[0]["index"])
print("TFLite predicted log-Mel shape:", enhanced_mel.shape)
Citation
If you use this model or the textual-echo-cancellation library in your research, please cite the original paper:
@inproceedings{ding2021textual,
title={Textual Echo Cancellation},
author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
pages={666--673},
year={2021},
organization={IEEE}
}
- Downloads last month
- 7