DecodeX Base: Real-Time Neural Audio Engine

This is the primary foundational model for the DecodeX audio verification platform, running in parallel with V2 to provide highly accurate, real-time deepfake and synthetic voice detection. It achieves the following results on the evaluation set:

  • Loss: 0.0829
  • Accuracy: 0.9882

Model description

DecodeX Base operates as the core Wav2Vec 2.0 layer in our dual-engine cross-verification architecture. It is optimized for high-speed inference on raw PCM audio streams, specifically designed to identify zero-day synthetic voices and voice cloning attempts before they breach communication networks.

Intended uses & limitations

Intended for deployment in enterprise PBX systems, SIP trunks, and local edge hardware to act as a first line of defense against AI-generated social engineering and spoofing attacks.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 3e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 32
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 5

Training results

Training Loss Epoch Step Accuracy Validation Loss
0.1448 1.0 1900 0.9601 0.1447
0.0673 2.0 3800 0.9824 0.0817
0.0178 3.0 5700 0.9796 0.1054
0.0002 4.0 7600 0.9824 0.1074
0.0108 5.0 9500 0.9882 0.0829

Framework versions

  • Transformers 4.39.3
  • Pytorch 2.1.2
  • Datasets 2.18.0
  • Tokenizers 0.15.2
Downloads last month
-
Safetensors
Model size
94.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support