YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- π€ Vocallet AI Service
- π Architecture Overview
- π οΈ Prerequisites
- π Quick Start (5 Steps)
- π Environment Variables
- π‘ API Documentation
- π§ Model Information
- π΅ Audio Format Support
- π§ Troubleshooting
- π³ Docker Deployment
- π Integration with Node.js / Express
- π§ͺ Running Tests
- β
First Run Checklist
- π Available Endpoints Summary
- π Architecture Overview
π€ Vocallet AI Service
A production-grade Python FastAPI microservice that processes voice commands for the Vocallet accessible finance application. Built for Indonesian users including visually impaired individuals, UMKM small businesses, and personal finance management.
π Architecture Overview
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β CLIENT (React 19 / Node.js) β
β POST /api/voice-command (multipart audio) β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β VOCALLET AI SERVICE (FastAPI) β
β βββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββββββ β
β β AudioServiceβ β STT Service β β Intent Service β β
β β validate ββββΆβ (Whisper STT) ββββΆβ (XLM-RoBERTa XNLI) β β
β β save temp β β transcribe() β β classify_intent() β β
β β load 16kHz β β β β zero-shot labels β β
β βββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββββββ β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β ModelRegistry (Singleton) β β
β β _stt_pipeline (Whisper) β _intent_pipeline (RoBERTa) β β
β β Loaded ONCE at startup via asyncio.gather() β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π οΈ Prerequisites
| Requirement | Version | Notes |
|---|---|---|
| Python | 3.10+ | Required for match statements |
| pip | 23.0+ | For dependency resolution |
| ffmpeg | Any recent | Required for MP3/M4A/WebM processing |
| RAM | 4GB minimum | For Whisper Small + XLM-RoBERTa Large |
| RAM | 8GB recommended | For smooth concurrent inference |
| NVIDIA GPU | Optional | Enables faster inference; CUDA 11.8+ |
Installing ffmpeg
# Ubuntu/Debian
sudo apt-get install ffmpeg
# macOS (Homebrew)
brew install ffmpeg
# Windows (Chocolatey)
choco install ffmpeg
# Windows (Scoop)
scoop install ffmpeg
π Quick Start (5 Steps)
1. Clone and enter the project
cd vocallet-ai-service
2. Create virtual environment
python -m venv venv
# Activate (Linux/macOS)
source venv/bin/activate
# Activate (Windows)
venv\Scripts\activate
3. Install dependencies
# CPU-only (recommended for most users)
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
# GPU (CUDA 12.1) β skip if using CPU
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
4. Configure environment
cp .env.example .env
# Edit .env as needed (optional β defaults work out of the box)
5. Run the service
uvicorn app.main:app --host 0.0.0.0 --port 8001 --reload
The service will be available at:
- API Docs (Swagger):
http://localhost:8001/docs - API Docs (ReDoc):
http://localhost:8001/redoc - Test UI:
http://localhost:8001/ui/index.html - Health:
http://localhost:8001/health
β οΈ First run note: On first startup, models (~1.3GB total) will download from Hugging Face. This can take 5β20 minutes depending on your internet connection. The service will not accept requests until downloads complete.
π Environment Variables
| Variable | Default | Description |
|---|---|---|
APP_NAME |
Vocallet AI Service |
Service display name |
APP_VERSION |
1.0.0 |
Service version string |
DEBUG |
false |
Enable debug mode + verbose logging |
HOST |
0.0.0.0 |
Bind host address |
PORT |
8001 |
Bind port |
ALLOWED_ORIGINS |
http://localhost:3000,... |
Comma-separated CORS allowed origins |
HF_TOKEN |
(empty) | Hugging Face access token (optional for public models) |
HF_CACHE_DIR |
./models_cache |
Local directory for model weight cache |
STT_MODEL_NAME |
openai/whisper-small |
Hugging Face model ID for speech-to-text |
INTENT_MODEL_NAME |
joeddav/xlm-roberta-large-xnli |
Hugging Face model ID for intent classification |
DEVICE |
auto |
Inference device: auto, cpu, cuda, cuda:0, mps |
TORCH_DTYPE |
float32 |
float32 (default) or float16 (GPU, faster + less VRAM) |
MAX_AUDIO_SIZE_MB |
25 |
Maximum allowed upload file size in MB |
MAX_AUDIO_DURATION_SECONDS |
60 |
Maximum allowed audio duration in seconds |
INTENT_CONFIDENCE_THRESHOLD |
0.45 |
Scores below this fall back to "tidak diketahui" |
INTENT_LABELS |
catat pengeluaran,... |
Comma-separated intent label strings |
TEMP_DIR |
./temp_audio |
Directory for temporary audio file storage |
π‘ API Documentation
POST /api/voice-command
Process a voice audio file and return the transcript and intent.
Request:
Content-Type: multipart/form-data
Body: file=<audio_file>
curl -X POST http://localhost:8001/api/voice-command \
-F "file=@recording.wav;type=audio/wav"
Successful Response (200):
{
"status": "success",
"transcript": "tolong catat jualan basreng 50 ribu",
"transcript_language": "id",
"intent": "catat penjualan umkm",
"confidence_score": 0.9241,
"all_intent_scores": {
"catat penjualan umkm": 0.9241,
"catat pengeluaran": 0.0412,
"hitung zakat": 0.0214,
"baca laporan": 0.0091,
"tidak diketahui": 0.0042
},
"processing_time_ms": 1234.56,
"audio_duration_seconds": 3.2,
"model_info": {
"stt_model": "openai/whisper-small",
"intent_model": "joeddav/xlm-roberta-large-xnli",
"device": "cpu"
}
}
Error Response (4xx/5xx):
{
"status": "error",
"error_code": "AUDIO_TOO_LONG",
"message": "Audio duration (75.0s) exceeds maximum allowed (60s).",
"detail": "...",
"timestamp": "2024-08-10T12:34:56.789Z"
}
Error Codes:
| Code | HTTP | Description |
|---|---|---|
FILE_TOO_LARGE |
400 | File size exceeds MAX_AUDIO_SIZE_MB |
UNSUPPORTED_FORMAT |
400 | File format not in allowed list |
AUDIO_TOO_LONG |
400 | Audio duration exceeds MAX_AUDIO_DURATION_SECONDS |
SERVICE_UNAVAILABLE |
503 | Models not loaded yet |
TRANSCRIPTION_FAILED |
422 | Whisper could not transcribe the audio |
CLASSIFICATION_FAILED |
422 | Intent classification failed |
INTERNAL_ERROR |
500 | Unexpected server error |
GET /health
Returns service health and model loading status.
curl http://localhost:8001/health
{
"status": "healthy",
"service": "Vocallet AI Service",
"version": "1.0.0",
"uptime_seconds": 3600.5,
"models": {
"stt": {
"name": "openai/whisper-small",
"loaded": true,
"load_time_seconds": 12.3,
"device": "cpu",
"error": null
},
"intent": {
"name": "joeddav/xlm-roberta-large-xnli",
"loaded": true,
"load_time_seconds": 25.7,
"device": "cpu",
"error": null
}
}
}
Status values: healthy (both models loaded) | degraded (one model failed) | unhealthy (both failed)
GET /health/models
Detailed model information including GPU memory usage (if applicable).
curl http://localhost:8001/health/models
GET /health/ping
Simple liveness probe for Docker/Kubernetes healthchecks.
curl http://localhost:8001/health/ping
# Response: {"pong": true}
π§ Model Information
Whisper STT Models
| Model | Size | Speed | Quality | Recommended Use |
|---|---|---|---|---|
whisper-tiny |
~75 MB | β‘β‘β‘ Fastest | β Basic | Development / testing |
whisper-base |
~145 MB | β‘β‘ Fast | ββ Decent | Low-resource deployment |
whisper-small |
~242 MB | β‘ Good | βββ Good | Recommended β |
whisper-medium |
~769 MB | π’ Slow | ββββ High | High-accuracy use cases |
whisper-large-v3 |
~1.5 GB | π’π’ Slowest | βββββ Best | Maximum accuracy |
Intent Classification Models
| Model | Size | Accuracy | Languages |
|---|---|---|---|
joeddav/xlm-roberta-large-xnli |
~1.1 GB | βββββ Best | 100+ (incl. Indonesian) β |
cross-encoder/nli-MiniLM2-L6-H768 |
~120 MB | βββ Good | Primarily English |
π΅ Audio Format Support
| Format | Extension | MIME Type | Notes |
|---|---|---|---|
| WAV | .wav |
audio/wav |
Recommended, no conversion needed |
| MP3 | .mp3 |
audio/mp3, audio/mpeg |
Requires ffmpeg |
| OGG Vorbis | .ogg |
audio/ogg |
Web standard |
| WebM | .webm |
audio/webm |
Browser MediaRecorder default |
| FLAC | .flac |
audio/flac |
Lossless |
| M4A/AAC | .m4a |
audio/m4a |
iOS recording format |
π§ Troubleshooting
"Model download is slow"
Models are downloaded from Hugging Face Hub on first run:
whisper-smallβ ~242MBxlm-roberta-large-xnliβ ~1.1GB
After the first download, they are cached in ./models_cache/ and reused on subsequent restarts. To pre-warm the cache:
python -c "
from transformers import pipeline
pipeline('automatic-speech-recognition', 'openai/whisper-small', model_kwargs={'cache_dir': './models_cache'})
pipeline('zero-shot-classification', 'joeddav/xlm-roberta-large-xnli', model_kwargs={'cache_dir': './models_cache'})
"
"CUDA out of memory"
Switch to CPU or a smaller model:
# In .env:
DEVICE=cpu
STT_MODEL_NAME=openai/whisper-tiny
Or use float16 for less VRAM usage (GPU only):
DEVICE=cuda
TORCH_DTYPE=float16
"Audio processing failed" / librosa error
Ensure ffmpeg is installed and available in your PATH:
ffmpeg -version
If missing, install it:
# Ubuntu
sudo apt-get install ffmpeg
# macOS
brew install ffmpeg
"Import errors" / ModuleNotFoundError
Make sure you are inside the virtual environment:
source venv/bin/activate # Linux/macOS
venv\Scripts\activate # Windows
# Then reinstall:
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
"TranscriptionError: Whisper returned an empty transcript"
The audio may be:
- Completely silent
- Contains only noise without recognizable speech
- Too short (< 0.5 seconds)
Try recording a clear, audible voice command.
π³ Docker Deployment
Build and run with Docker Compose
# Build the image
docker-compose build
# Start service
docker-compose up -d
# View logs
docker-compose logs -f vocallet-ai
# Stop service
docker-compose down
Environment variables for Docker
# Pass HF token securely via host environment
export HF_TOKEN=hf_your_token_here
docker-compose up -d
Check service health
docker-compose ps
curl http://localhost:8001/health/ping
π Integration with Node.js / Express
Example: Forwarding audio from Express to the Python microservice.
const axios = require('axios');
const FormData = require('form-data');
const multer = require('multer');
const upload = multer({ storage: multer.memoryStorage() });
// Express route: POST /voice-command
router.post('/voice-command', upload.single('audio'), async (req, res) => {
try {
const form = new FormData();
form.append('file', req.file.buffer, {
filename: req.file.originalname,
contentType: req.file.mimetype,
});
const response = await axios.post(
'http://localhost:8001/api/voice-command',
form,
{
headers: form.getHeaders(),
timeout: 60000 // 60s timeout for model inference
}
);
res.json(response.data);
} catch (err) {
const detail = err.response?.data?.detail || err.message;
res.status(err.response?.status || 500).json({
error: 'voice_processing_failed',
detail
});
}
});
π§ͺ Running Tests
# Install dev dependencies
pip install -r requirements-dev.txt
# Run all tests
pytest tests/ -v
# Run with coverage report
pytest tests/ -v --cov=app --cov-report=html
# Open coverage report
open htmlcov/index.html # macOS
xdg-open htmlcov/index.html # Linux
β First Run Checklist
When you start the service for the first time, here is what happens:
- FastAPI boots up β Uvicorn starts, lifespan context begins
detect_device()runs β Auto-detects CUDA β MPS β CPUasyncio.gather()kicks off β Both models begin downloading/loading concurrently- Whisper downloads β ~242MB for
whisper-small; goes to./models_cache/ - XLM-RoBERTa downloads β ~1.1GB for
xlm-roberta-large-xnli; goes to./models_cache/ - Models load into memory β Weights deserialized into PyTorch; moved to target device
- "Vocallet AI Service ready!" is logged β Now accepting requests
GET /healthreturns"status": "healthy"β Both models are live
β±οΈ Total startup time on first run (CPU): 5β25 minutes (mostly download) β±οΈ Total startup time on subsequent runs (cache warm): 30β90 seconds (model load only)
π Available Endpoints Summary
| Method | Path | Description |
|---|---|---|
GET |
/ |
Service info |
POST |
/api/voice-command |
Main endpoint β process audio |
GET |
/health |
Readiness probe with model status |
GET |
/health/models |
Detailed model metadata |
GET |
/health/ping |
Liveness probe |
GET |
/docs |
Swagger UI |
GET |
/redoc |
ReDoc UI |
GET |
/ui/index.html |
Standalone HTML test UI |