YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
VOXLESS: Real-Time Camera-Based Silent Speech-to-Personalized Voice Synthesis
An End-to-End Visual Silent Speech Interface (VSSI) for Non-Vocal Speakers
๐ฌ Central Research Hypothesis
A multi-scale visual articulatory representation combining lip appearance, temporal lip motion, jaw dynamics and lower-face motion, coupled with phoneme-viseme multi-task supervision and speaker-adaptive neural speech synthesis, can produce lower-latency and more intelligible personalized speech than conventional lip-to-text-to-TTS pipelines for non-vocal speakers.
๐๏ธ System Architecture
RGB CAMERA (30โ60 FPS)
โ
โผ
Face Tracking & Alignment
(MediaPipe Face Mesh + Affine Warp)
โ
โผ
Multi-Scale Articulatory ROI
โโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ
Lip Region Jaw & Chin Cheeks & Contour
(96 ร 96) (112 ร 112) (64 ร 64)
โ โ โ
โโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโ
โผ
Multi-Stream Spatial Encoder
(ConvNeXt-VSSI / 2D-3D Hybrid Frontends)
โ
โผ
Cross-Articulatory Attention Fusion
(Lip Appearance + Dynamics + Jaw + Pose)
โ
โผ
Temporal Articulatory Encoder
(Conformer / Transformer / BiLSTM)
โ
โผ
Visual Speech Representation
โ
โโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโ
โผ โผ
Viseme Decoder (Auxiliary) Phoneme Decoder (CTC)
(Fisher/Jeffers Visemes) (ARPAbet / IPA Tokens)
โ โ
โโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ
โผ
Multi-Task Articulatory Loss
L = ฮปโL_ctc + ฮปโL_phoneme + ฮปโL_viseme + ฮปโL_lang
โ
โผ
Streaming Prefix Beam Search + LM
+ Confidence & Uncertainty Estimation
โ
โโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโ
โผ โผ
[Pipeline A: Cascaded] [Pipeline B: Direct]
Reconstructed Text/Phonemes Articulatory Mel-Gen
โ โ
โผ โ
Speaker Adapter โ
(LoRA / Emb Conditioning) โ
โ โ
โผ โผ
Personalized Neural TTS Acoustic Vocoder
(Acoustic Model + HiFi-GAN) (HiFi-GAN / Mel)
โ โ
โโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ
โผ
Real-Time Audio Buffer
(<300 ms End-to-End Latency)
๐ Key Innovations
- Multi-Scale Articulatory Decomposition: Processes not just lips ($96 \times 96$), but mandibular excursion ($112 \times 112$), cheek dynamics ($64 \times 64$), and 3D head pose simultaneously.
- Multi-Task Phoneme-Viseme Supervision: Resolves homophene visual ambiguities by co-learning viseme clusters (Fisher & Neti) alongside CTC phonemes and text.
- Dual-Track Personalized Synthesis:
- Pipeline A (Cascaded): Visual $\to$ Reconstructed Text $\to$ Speaker-conditioned TTS.
- Pipeline B (Direct Visual-to-Speech): Visual Embeddings $\to$ Direct Mel-Spectrogram $\to$ Vocoder.
- Rapid Speaker Personalization: Low-Rank Adaptation (LoRA) and few-shot calibration calibrate new speakers in under 3 minutes.
- Confidence-Aware Assistive Gating: Evaluates CTC posterior probabilities and entropy to prevent hallucinations, triggering polite fallback repetitions ("Low confidence. Please repeat.") when uncertain.
- Sub-300ms Real-Time Streaming: Asynchronous frame grabber, visual VAD, and chunk-level streaming prefix beam decoder.
๐ฆ Installation
# 1. Clone repository and navigate
git clone https://github.com/vocalis/voxless.git
cd voxless
# 2. Create isolated virtual environment
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/macOS
# 3. Install PyTorch with CUDA 12.x support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124
# 4. Install dependencies
pip install -r requirements.txt
pip install -e .
๐ฎ Usage & CLI
1. Run Real-Time Camera Demonstration
python -m voxless.cli demo --camera 0 --pipeline pipeline_a
Keyboard controls:
q: Quit applicationr: Reset beam decoder historys: Toggle between Pipeline A (Cascaded) and Pipeline B (Direct)
2. Calibrate & Enroll a New Speaker
python -m voxless.cli enroll --speaker-id spk_01 --speaker-name "Participant"
3. Record Custom Dataset (VOXLESS-INDIA)
python -m voxless.cli collect --speaker-id spk_01 --num-prompts 5
4. Run Scientific Ablation Study Matrix
python -m voxless.cli ablate
5. Compare with Mandatory Baselines
python -m voxless.cli benchmark
6. Launch Interactive Web Dashboard
streamlit run voxless/app/web_dashboard.py
๐งช Testing
Run the automated test suite covering vision, models, decoders, adaptation, synthesis, and streaming latency:
pytest tests/ -v
๐ก๏ธ Ethical Safeguards & Consent Gate
Voice cloning and personalized synthesis require cryptographic token verification (EthicalConsentGate). Voice profiles cannot be synthesized without participant-signed consent.