YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

🎀 Vocallet AI Service

A production-grade Python FastAPI microservice that processes voice commands for the Vocallet accessible finance application. Built for Indonesian users including visually impaired individuals, UMKM small businesses, and personal finance management.


πŸ“ Architecture Overview

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                         CLIENT (React 19 / Node.js)                  β”‚
β”‚              POST /api/voice-command  (multipart audio)              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚
                            β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     VOCALLET AI SERVICE (FastAPI)                     β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚ AudioServiceβ”‚   β”‚   STT Service    β”‚   β”‚  Intent Service      β”‚  β”‚
β”‚  β”‚  validate   │──▢│  (Whisper STT)   │──▢│ (XLM-RoBERTa XNLI)  β”‚  β”‚
β”‚  β”‚  save temp  β”‚   β”‚  transcribe()    β”‚   β”‚  classify_intent()   β”‚  β”‚
β”‚  β”‚  load 16kHz β”‚   β”‚                  β”‚   β”‚  zero-shot labels    β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                                       β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚                     ModelRegistry (Singleton)                 β”‚   β”‚
β”‚  β”‚  _stt_pipeline (Whisper)    β”‚    _intent_pipeline (RoBERTa)  β”‚   β”‚
β”‚  β”‚  Loaded ONCE at startup via asyncio.gather()                  β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ› οΈ Prerequisites

Requirement Version Notes
Python 3.10+ Required for match statements
pip 23.0+ For dependency resolution
ffmpeg Any recent Required for MP3/M4A/WebM processing
RAM 4GB minimum For Whisper Small + XLM-RoBERTa Large
RAM 8GB recommended For smooth concurrent inference
NVIDIA GPU Optional Enables faster inference; CUDA 11.8+

Installing ffmpeg

# Ubuntu/Debian
sudo apt-get install ffmpeg

# macOS (Homebrew)
brew install ffmpeg

# Windows (Chocolatey)
choco install ffmpeg

# Windows (Scoop)
scoop install ffmpeg

πŸš€ Quick Start (5 Steps)

1. Clone and enter the project

cd vocallet-ai-service

2. Create virtual environment

python -m venv venv

# Activate (Linux/macOS)
source venv/bin/activate

# Activate (Windows)
venv\Scripts\activate

3. Install dependencies

# CPU-only (recommended for most users)
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt

# GPU (CUDA 12.1) β€” skip if using CPU
pip install torch==2.3.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt

4. Configure environment

cp .env.example .env
# Edit .env as needed (optional β€” defaults work out of the box)

5. Run the service

uvicorn app.main:app --host 0.0.0.0 --port 8001 --reload

The service will be available at:

  • API Docs (Swagger): http://localhost:8001/docs
  • API Docs (ReDoc): http://localhost:8001/redoc
  • Test UI: http://localhost:8001/ui/index.html
  • Health: http://localhost:8001/health

⚠️ First run note: On first startup, models (~1.3GB total) will download from Hugging Face. This can take 5–20 minutes depending on your internet connection. The service will not accept requests until downloads complete.


🌍 Environment Variables

Variable Default Description
APP_NAME Vocallet AI Service Service display name
APP_VERSION 1.0.0 Service version string
DEBUG false Enable debug mode + verbose logging
HOST 0.0.0.0 Bind host address
PORT 8001 Bind port
ALLOWED_ORIGINS http://localhost:3000,... Comma-separated CORS allowed origins
HF_TOKEN (empty) Hugging Face access token (optional for public models)
HF_CACHE_DIR ./models_cache Local directory for model weight cache
STT_MODEL_NAME openai/whisper-small Hugging Face model ID for speech-to-text
INTENT_MODEL_NAME joeddav/xlm-roberta-large-xnli Hugging Face model ID for intent classification
DEVICE auto Inference device: auto, cpu, cuda, cuda:0, mps
TORCH_DTYPE float32 float32 (default) or float16 (GPU, faster + less VRAM)
MAX_AUDIO_SIZE_MB 25 Maximum allowed upload file size in MB
MAX_AUDIO_DURATION_SECONDS 60 Maximum allowed audio duration in seconds
INTENT_CONFIDENCE_THRESHOLD 0.45 Scores below this fall back to "tidak diketahui"
INTENT_LABELS catat pengeluaran,... Comma-separated intent label strings
TEMP_DIR ./temp_audio Directory for temporary audio file storage

πŸ“‘ API Documentation

POST /api/voice-command

Process a voice audio file and return the transcript and intent.

Request:

Content-Type: multipart/form-data
Body: file=<audio_file>
curl -X POST http://localhost:8001/api/voice-command \
  -F "file=@recording.wav;type=audio/wav"

Successful Response (200):

{
  "status": "success",
  "transcript": "tolong catat jualan basreng 50 ribu",
  "transcript_language": "id",
  "intent": "catat penjualan umkm",
  "confidence_score": 0.9241,
  "all_intent_scores": {
    "catat penjualan umkm": 0.9241,
    "catat pengeluaran": 0.0412,
    "hitung zakat": 0.0214,
    "baca laporan": 0.0091,
    "tidak diketahui": 0.0042
  },
  "processing_time_ms": 1234.56,
  "audio_duration_seconds": 3.2,
  "model_info": {
    "stt_model": "openai/whisper-small",
    "intent_model": "joeddav/xlm-roberta-large-xnli",
    "device": "cpu"
  }
}

Error Response (4xx/5xx):

{
  "status": "error",
  "error_code": "AUDIO_TOO_LONG",
  "message": "Audio duration (75.0s) exceeds maximum allowed (60s).",
  "detail": "...",
  "timestamp": "2024-08-10T12:34:56.789Z"
}

Error Codes:

Code HTTP Description
FILE_TOO_LARGE 400 File size exceeds MAX_AUDIO_SIZE_MB
UNSUPPORTED_FORMAT 400 File format not in allowed list
AUDIO_TOO_LONG 400 Audio duration exceeds MAX_AUDIO_DURATION_SECONDS
SERVICE_UNAVAILABLE 503 Models not loaded yet
TRANSCRIPTION_FAILED 422 Whisper could not transcribe the audio
CLASSIFICATION_FAILED 422 Intent classification failed
INTERNAL_ERROR 500 Unexpected server error

GET /health

Returns service health and model loading status.

curl http://localhost:8001/health
{
  "status": "healthy",
  "service": "Vocallet AI Service",
  "version": "1.0.0",
  "uptime_seconds": 3600.5,
  "models": {
    "stt": {
      "name": "openai/whisper-small",
      "loaded": true,
      "load_time_seconds": 12.3,
      "device": "cpu",
      "error": null
    },
    "intent": {
      "name": "joeddav/xlm-roberta-large-xnli",
      "loaded": true,
      "load_time_seconds": 25.7,
      "device": "cpu",
      "error": null
    }
  }
}

Status values: healthy (both models loaded) | degraded (one model failed) | unhealthy (both failed)


GET /health/models

Detailed model information including GPU memory usage (if applicable).

curl http://localhost:8001/health/models

GET /health/ping

Simple liveness probe for Docker/Kubernetes healthchecks.

curl http://localhost:8001/health/ping
# Response: {"pong": true}

🧠 Model Information

Whisper STT Models

Model Size Speed Quality Recommended Use
whisper-tiny ~75 MB ⚑⚑⚑ Fastest ⭐ Basic Development / testing
whisper-base ~145 MB ⚑⚑ Fast ⭐⭐ Decent Low-resource deployment
whisper-small ~242 MB ⚑ Good ⭐⭐⭐ Good Recommended βœ…
whisper-medium ~769 MB 🐒 Slow ⭐⭐⭐⭐ High High-accuracy use cases
whisper-large-v3 ~1.5 GB 🐒🐒 Slowest ⭐⭐⭐⭐⭐ Best Maximum accuracy

Intent Classification Models

Model Size Accuracy Languages
joeddav/xlm-roberta-large-xnli ~1.1 GB ⭐⭐⭐⭐⭐ Best 100+ (incl. Indonesian) βœ…
cross-encoder/nli-MiniLM2-L6-H768 ~120 MB ⭐⭐⭐ Good Primarily English

🎡 Audio Format Support

Format Extension MIME Type Notes
WAV .wav audio/wav Recommended, no conversion needed
MP3 .mp3 audio/mp3, audio/mpeg Requires ffmpeg
OGG Vorbis .ogg audio/ogg Web standard
WebM .webm audio/webm Browser MediaRecorder default
FLAC .flac audio/flac Lossless
M4A/AAC .m4a audio/m4a iOS recording format

πŸ”§ Troubleshooting

"Model download is slow"

Models are downloaded from Hugging Face Hub on first run:

  • whisper-small β†’ ~242MB
  • xlm-roberta-large-xnli β†’ ~1.1GB

After the first download, they are cached in ./models_cache/ and reused on subsequent restarts. To pre-warm the cache:

python -c "
from transformers import pipeline
pipeline('automatic-speech-recognition', 'openai/whisper-small', model_kwargs={'cache_dir': './models_cache'})
pipeline('zero-shot-classification', 'joeddav/xlm-roberta-large-xnli', model_kwargs={'cache_dir': './models_cache'})
"

"CUDA out of memory"

Switch to CPU or a smaller model:

# In .env:
DEVICE=cpu
STT_MODEL_NAME=openai/whisper-tiny

Or use float16 for less VRAM usage (GPU only):

DEVICE=cuda
TORCH_DTYPE=float16

"Audio processing failed" / librosa error

Ensure ffmpeg is installed and available in your PATH:

ffmpeg -version

If missing, install it:

# Ubuntu
sudo apt-get install ffmpeg

# macOS
brew install ffmpeg

"Import errors" / ModuleNotFoundError

Make sure you are inside the virtual environment:

source venv/bin/activate  # Linux/macOS
venv\Scripts\activate     # Windows

# Then reinstall:
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt

"TranscriptionError: Whisper returned an empty transcript"

The audio may be:

  • Completely silent
  • Contains only noise without recognizable speech
  • Too short (< 0.5 seconds)

Try recording a clear, audible voice command.


🐳 Docker Deployment

Build and run with Docker Compose

# Build the image
docker-compose build

# Start service
docker-compose up -d

# View logs
docker-compose logs -f vocallet-ai

# Stop service
docker-compose down

Environment variables for Docker

# Pass HF token securely via host environment
export HF_TOKEN=hf_your_token_here
docker-compose up -d

Check service health

docker-compose ps
curl http://localhost:8001/health/ping

πŸ”— Integration with Node.js / Express

Example: Forwarding audio from Express to the Python microservice.

const axios = require('axios');
const FormData = require('form-data');
const multer = require('multer');

const upload = multer({ storage: multer.memoryStorage() });

// Express route: POST /voice-command
router.post('/voice-command', upload.single('audio'), async (req, res) => {
  try {
    const form = new FormData();
    form.append('file', req.file.buffer, {
      filename: req.file.originalname,
      contentType: req.file.mimetype,
    });

    const response = await axios.post(
      'http://localhost:8001/api/voice-command',
      form,
      {
        headers: form.getHeaders(),
        timeout: 60000  // 60s timeout for model inference
      }
    );

    res.json(response.data);
  } catch (err) {
    const detail = err.response?.data?.detail || err.message;
    res.status(err.response?.status || 500).json({
      error: 'voice_processing_failed',
      detail
    });
  }
});

πŸ§ͺ Running Tests

# Install dev dependencies
pip install -r requirements-dev.txt

# Run all tests
pytest tests/ -v

# Run with coverage report
pytest tests/ -v --cov=app --cov-report=html

# Open coverage report
open htmlcov/index.html  # macOS
xdg-open htmlcov/index.html  # Linux

βœ… First Run Checklist

When you start the service for the first time, here is what happens:

  1. FastAPI boots up β€” Uvicorn starts, lifespan context begins
  2. detect_device() runs β€” Auto-detects CUDA β†’ MPS β†’ CPU
  3. asyncio.gather() kicks off β€” Both models begin downloading/loading concurrently
  4. Whisper downloads β€” ~242MB for whisper-small; goes to ./models_cache/
  5. XLM-RoBERTa downloads β€” ~1.1GB for xlm-roberta-large-xnli; goes to ./models_cache/
  6. Models load into memory β€” Weights deserialized into PyTorch; moved to target device
  7. "Vocallet AI Service ready!" is logged β€” Now accepting requests
  8. GET /health returns "status": "healthy" β€” Both models are live

⏱️ Total startup time on first run (CPU): 5–25 minutes (mostly download) ⏱️ Total startup time on subsequent runs (cache warm): 30–90 seconds (model load only)


🌐 Available Endpoints Summary

Method Path Description
GET / Service info
POST /api/voice-command Main endpoint β€” process audio
GET /health Readiness probe with model status
GET /health/models Detailed model metadata
GET /health/ping Liveness probe
GET /docs Swagger UI
GET /redoc ReDoc UI
GET /ui/index.html Standalone HTML test UI
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support