YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

VOXLESS: Real-Time Camera-Based Silent Speech-to-Personalized Voice Synthesis

An End-to-End Visual Silent Speech Interface (VSSI) for Non-Vocal Speakers

Python 3.11+ PyTorch 2.4+ CUDA Acceleration License: MIT


๐Ÿ”ฌ Central Research Hypothesis

A multi-scale visual articulatory representation combining lip appearance, temporal lip motion, jaw dynamics and lower-face motion, coupled with phoneme-viseme multi-task supervision and speaker-adaptive neural speech synthesis, can produce lower-latency and more intelligible personalized speech than conventional lip-to-text-to-TTS pipelines for non-vocal speakers.


๐Ÿ›๏ธ System Architecture

                       RGB CAMERA (30โ€“60 FPS)
                                 โ”‚
                                 โ–ผ
                     Face Tracking & Alignment
                (MediaPipe Face Mesh + Affine Warp)
                                 โ”‚
                                 โ–ผ
                    Multi-Scale Articulatory ROI
          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
          โ–ผ                      โ–ผ                      โ–ผ
     Lip Region              Jaw & Chin             Cheeks & Contour
     (96 ร— 96)              (112 ร— 112)                 (64 ร— 64)
          โ”‚                      โ”‚                      โ”‚
          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ–ผ
                   Multi-Stream Spatial Encoder
               (ConvNeXt-VSSI / 2D-3D Hybrid Frontends)
                                 โ”‚
                                 โ–ผ
               Cross-Articulatory Attention Fusion
              (Lip Appearance + Dynamics + Jaw + Pose)
                                 โ”‚
                                 โ–ผ
                   Temporal Articulatory Encoder
               (Conformer / Transformer / BiLSTM)
                                 โ”‚
                                 โ–ผ
                     Visual Speech Representation
                                 โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ                               โ–ผ
       Viseme Decoder (Auxiliary)     Phoneme Decoder (CTC)
       (Fisher/Jeffers Visemes)       (ARPAbet / IPA Tokens)
                 โ”‚                               โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ–ผ
                  Multi-Task Articulatory Loss
           L = ฮปโ‚L_ctc + ฮปโ‚‚L_phoneme + ฮปโ‚ƒL_viseme + ฮปโ‚„L_lang
                                 โ”‚
                                 โ–ผ
                 Streaming Prefix Beam Search + LM
               + Confidence & Uncertainty Estimation
                                 โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ                               โ–ผ
           [Pipeline A: Cascaded]         [Pipeline B: Direct]
           Reconstructed Text/Phonemes    Articulatory Mel-Gen
                 โ”‚                               โ”‚
                 โ–ผ                               โ”‚
           Speaker Adapter                       โ”‚
          (LoRA / Emb Conditioning)              โ”‚
                 โ”‚                               โ”‚
                 โ–ผ                               โ–ผ
        Personalized Neural TTS          Acoustic Vocoder
        (Acoustic Model + HiFi-GAN)       (HiFi-GAN / Mel)
                 โ”‚                               โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ–ผ
                     Real-Time Audio Buffer
                   (<300 ms End-to-End Latency)

๐Ÿš€ Key Innovations

  1. Multi-Scale Articulatory Decomposition: Processes not just lips ($96 \times 96$), but mandibular excursion ($112 \times 112$), cheek dynamics ($64 \times 64$), and 3D head pose simultaneously.
  2. Multi-Task Phoneme-Viseme Supervision: Resolves homophene visual ambiguities by co-learning viseme clusters (Fisher & Neti) alongside CTC phonemes and text.
  3. Dual-Track Personalized Synthesis:
    • Pipeline A (Cascaded): Visual $\to$ Reconstructed Text $\to$ Speaker-conditioned TTS.
    • Pipeline B (Direct Visual-to-Speech): Visual Embeddings $\to$ Direct Mel-Spectrogram $\to$ Vocoder.
  4. Rapid Speaker Personalization: Low-Rank Adaptation (LoRA) and few-shot calibration calibrate new speakers in under 3 minutes.
  5. Confidence-Aware Assistive Gating: Evaluates CTC posterior probabilities and entropy to prevent hallucinations, triggering polite fallback repetitions ("Low confidence. Please repeat.") when uncertain.
  6. Sub-300ms Real-Time Streaming: Asynchronous frame grabber, visual VAD, and chunk-level streaming prefix beam decoder.

๐Ÿ“ฆ Installation

# 1. Clone repository and navigate
git clone https://github.com/vocalis/voxless.git
cd voxless

# 2. Create isolated virtual environment
python -m venv .venv
.venv\Scripts\activate  # Windows
# source .venv/bin/activate  # Linux/macOS

# 3. Install PyTorch with CUDA 12.x support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124

# 4. Install dependencies
pip install -r requirements.txt
pip install -e .

๐ŸŽฎ Usage & CLI

1. Run Real-Time Camera Demonstration

python -m voxless.cli demo --camera 0 --pipeline pipeline_a

Keyboard controls:

  • q: Quit application
  • r: Reset beam decoder history
  • s: Toggle between Pipeline A (Cascaded) and Pipeline B (Direct)

2. Calibrate & Enroll a New Speaker

python -m voxless.cli enroll --speaker-id spk_01 --speaker-name "Participant"

3. Record Custom Dataset (VOXLESS-INDIA)

python -m voxless.cli collect --speaker-id spk_01 --num-prompts 5

4. Run Scientific Ablation Study Matrix

python -m voxless.cli ablate

5. Compare with Mandatory Baselines

python -m voxless.cli benchmark

6. Launch Interactive Web Dashboard

streamlit run voxless/app/web_dashboard.py

๐Ÿงช Testing

Run the automated test suite covering vision, models, decoders, adaptation, synthesis, and streaming latency:

pytest tests/ -v

๐Ÿ›ก๏ธ Ethical Safeguards & Consent Gate

Voice cloning and personalized synthesis require cryptographic token verification (EthicalConsentGate). Voice profiles cannot be synthesized without participant-signed consent.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Jegan022/Vocalis 1