Tiny Turn Detector
A lightweight real-time audio turn detection model that predicts whether a speaker is DONE speaking or PAUSING/CONTINUING in conversational audio.
Model Description
This model addresses a critical challenge in building responsive voice assistants and conversation systems: determining when a speaker has actually finished their turn versus just pausing mid-sentence.
Key Features:
- โก Fast inference (~100ms on CPU)
- ๐ฏ High accuracy (95.5% on training set, 73% on validation)
- ๐ Works with 8-second audio clips
- ๐ Easy to integrate with existing systems
- ๐ฆ Small model size (~150MB with Whisper Tiny)
Architecture
Audio (8 sec, 16kHz)
โ
Whisper Tiny Encoder (frozen/fine-tuned)
โ
Mean Pooling
โ
MLP Head (384 โ 64 โ 1)
โ
Sigmoid โ P(end_turn)
โ
Binary Decision: END or CONTINUE
Components:
- Encoder: OpenAI Whisper Tiny (pretrained)
- Classifier: 2-layer MLP with ReLU activation
- Input: 8-second audio clips at 16kHz
- Output: Binary classification (0=CONTINUE, 1=END)
Training Results
The model was trained on the pipecat-ai/smart-turn-data-v3.2-train dataset.
Final Metrics (Best Model)
| Split | Loss | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|---|
| Train | 0.1391 | 95.50% | 95.90% | 95.34% | 95.62% |
| Val | 4.9075 | 73.00% | 70.00% | 89.09% | 78.40% |
Training Configuration:
- Epochs: Multiple epochs with early stopping
- Optimizer: AdamW
- Loss Function: Binary Cross-Entropy with Logits
- Best model selected based on validation F1 score (0.7840)
Note: The validation loss is higher due to the model being optimized for F1 score rather than loss. The high recall (89%) indicates the model is conservative about marking turn endings, which is desirable for real-time applications to avoid premature interruptions.
Usage
Download the Model
from huggingface_hub import hf_hub_download
import torch
# Download model
model_path = hf_hub_download(
repo_id="YOUR_USERNAME/tiny-turn-detector",
filename="best_model.pt"
)
# Load model
model = torch.load(model_path, map_location='cpu')
model.eval()
Run Inference
import torch
import librosa
from transformers import WhisperProcessor
# Load processor
processor = WhisperProcessor.from_pretrained("openai/whisper-tiny")
# Load audio (8 seconds at 16kHz)
audio, sr = librosa.load("your_audio.wav", sr=16000, duration=8.0)
# Process audio
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
# Predict
with torch.no_grad():
outputs = model(inputs.input_features)
probability = torch.sigmoid(outputs).item()
# Decision
threshold = 0.5
decision = "END" if probability > threshold else "CONTINUE"
print(f"Probability: {probability:.4f}")
print(f"Decision: {decision}")
Full Inference Script
For complete inference code with audio loading, preprocessing, and visualization, see the GitHub repository.
Installation
pip install torch torchaudio transformers librosa huggingface_hub
Use Cases
- ๐๏ธ Voice assistants and chatbots
- ๐ Real-time conversation systems
- ๐ง Meeting transcription tools
- ๐ค Interactive voice response (IVR) systems
- ๐ฌ Voice-based interfaces
- ๐ฎ Voice-controlled applications
Model Details
- Model Type: Audio Classification (Binary)
- Base Model: OpenAI Whisper Tiny
- Language: English (primarily)
- Sampling Rate: 16kHz
- Input Duration: 8 seconds
- Framework: PyTorch
- Parameters: ~39M (Whisper) + ~25K (Classifier)
Training Data
Dataset: pipecat-ai/smart-turn-data-v3.2-train
The dataset contains conversational audio clips labeled with turn-taking information:
endpoint_bool: Binary label (0=continue, 1=end)- Audio clips of varying lengths (processed to 8 seconds)
- Real-world conversational scenarios
Limitations
Validation Gap: The model shows some overfitting (95.5% train vs 73% val accuracy). This could be improved with:
- Data augmentation
- Regularization techniques
- More diverse training data
8-Second Window: Requires exactly 8 seconds of audio context
English Focus: Primarily trained on English conversations
VAD Dependency: Works best when combined with Voice Activity Detection (VAD) for silence removal
Future Improvements
- Add multi-language support
- Reduce validation gap through regularization
- Variable-length audio support
- Real-time streaming inference
- Integration with VAD systems
- Ensemble with acoustic features (pause duration, pitch)
GitHub Repository
Full training code, evaluation scripts, and inference examples:
๐ https://github.com/Nitin1613/Turn_detector/tree/main
The repository includes:
- Complete training pipeline
- Dataset preparation scripts
- Evaluation and benchmarking tools
- Inference examples
- Google Colab notebook for easy experimentation
Citation
If you use this model in your research or application, please cite:
@misc{tiny-turn-detector-2026,
title={Tiny Turn Detector: Real-time Audio Turn Detection with Whisper},
author=Nitinbudania,
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/Nitinbudania/tiny-turn-detector}}
}
License
MIT License - See LICENSE file for details
Acknowledgments
- OpenAI for the Whisper model
- Pipecat.ai for the training dataset
- Hugging Face for hosting and tools
Model Card Authors: YOUR_NAME
Contact: YOUR_EMAIL or GitHub
Last Updated: August 2026
Model tree for Nitinbudania/tiny-turn-detector
Base model
openai/whisper-tiny