V2C1: INT8 acoustic drone detection
V2C1 classifies drone versus background from a 1-second audio window. This release contains the original frozen model: 40,200 bytes, fully integer INT8, with a 32 × 32 log-mel input and a compact CNN containing 30,066 parameters.
It is a research prototype with measured errors. It has not been field validated. Its benchmark results do not establish outdoor detection range, performance on a new microphone or airframe, or ESP32 latency, RAM use, or power consumption.
| Released artifact | Value |
|---|---|
| Deployment file | v2_deploy/drone_model_int8.tflite |
| Size | 40,200 bytes, approximately 39.26 KiB |
| SHA256 | e2e4e4b0e49d50f0d3c0b2129240564ec84b93b4c252651ef10b82560b9d72ae |
| Input | INT8, allocated shape (1, 32, 32, 1), signature (-1, 32, 32, 1) |
| Output | INT8, allocated shape (1, 2), signature (-1, 2); index 0 background, index 1 drone |
| Operating threshold | 0.4684000015258789, selected on validation |
Download and run
Download the repository with the Hugging Face CLI:
python -m pip install -U huggingface_hub
hf download Umranz/drone-detector-tinyml-v2c1 --local-dir V2C1
cd V2C1
python -m pip install -r requirements.txt
python -B verify_baseline.py
python run_v2c1.py path/to/audio.wav --json
Replace path/to/audio.wav with your 16 kHz WAV. The runner averages stereo channels into mono, rejects empty files and other sample rates, and does not resample. For a directory of recordings, pass the directory instead of a filename. JSON output includes results and errors; any failed input produces a nonzero exit status. Keep run_v2c1.py beside v2_deploy/ so it can find the model, filterbank and metadata.
The reference runner processes consecutive, non-overlapping 1-second windows using allocated batch size 1. The serialized tensor signatures allow a dynamic batch dimension; this release's reference verification uses batch size 1. Files shorter than one second are right-zero-padded and produce one window. For longer files, a trailing incomplete window is discarded. It returns each window's drone score and decision; the file verdict is DRONE if any window reaches the threshold. That file rule is a demonstration behavior, not a validated event-level alarm policy.
The release is run through the supplied TFLite reference script. The audio-classification task label does not imply a Transformers pipeline() interface or hosted inference support.
The repository also includes the original V2C1.keras checkpoint, float32 TFLite model, embedded C headers, and historical evaluation records. The INT8 file above is the deployment model. MANIFEST.sha256 lists release file hashes; RELEASE_VERIFICATION.json records integrity, batch-1 inference, seven reference-vector matches and audio-input checks. The dependency versions in requirements.txt were verified with Python 3.10.11.
The serialized INT8 graph can also be checked without TensorFlow or any source audio:
python -B same_size_verify.py --candidate v2_deploy/drone_model_int8.tflite
Measured INT8 performance
These measurements concern the released .tflite artifact at its frozen validation threshold, on 1,621 historical TEST windows: 198 drone and 1,423 background windows from 296 recording groups. The original model did not train on this TEST split. The project has subsequently consulted this benchmark, so it should not be treated as a new independent test for future changes.
| Metric | Result |
|---|---|
| Drone recall | 95.45%, 189/198 |
| Background specificity | 98.8756%, 1,407/1,423 |
| False-positive rate | 1.1244%, 16/1,423 |
| Precision | 92.20% |
| F1 | 0.937965 |
| Accuracy | 98.46% |
| ROC AUC | 0.997083 |
| Average precision / PR AUC | 0.983962 |
| TP / FN / FP / TN | 189 / 9 / 16 / 1,407 |
Three original operating-performance targets were missed:
| Target | Requirement | Observed | Status |
|---|---|---|---|
| Drone recall | ≥97.5% | 95.45% | FAIL |
| Specificity | ≥99.2% | 98.8756% | FAIL |
| False-positive rate | ≤0.8% | 1.1244% | FAIL |
Mechanical sounds can be confused with drones. On the historical TEST chainsaw subset, 7 of 35 windows were false positives; those windows came from only 3 recording groups. Small category subsets do not establish population-level rates. Continuous event-level false alarms per hour have not been measured.
See FINAL_V2C1_READINESS.md and AUDIT_V2C1_BINARY.md for the measured evidence and its limits.
Preprocessing and quantization
The frontend is part of the model contract. Use the supplied mel_filterbank.npy and model_metadata.json, and reproduce the following steps in order:
- Convert 16 kHz audio to mono and take 16,000 samples per window, padding a short file on the right with zeros.
- Apply a causal fourth-order Butterworth bandpass from 100 to 6,000 Hz, implemented with second-order sections and
sosfilt. - Reflect-pad the filtered window by 256 samples on each side.
- Take 32 frames with FFT length 512 and frame hop 512. Apply the periodic Hann window
0.5 - 0.5*cos(2*pi*n/512). - Compute the real FFT and power spectrum, then 32 mel bands spanning 100–6,000 Hz, using the HTK mel scale and Slaney bandwidth normalization.
- Compute
log(mel_power + 1e-6)and normalize with the fixed TRAIN-derived constants below. - Quantize to INT8, invoke TFLite, dequantize output index 1 and compare it with the frozen threshold.
x_norm = (logmel - (-7.041979789733887)) / (4.118093490600586 + 1e-6)
q = clip(round(x_norm / 0.018346942961215973) - 38, -128, 127).astype(int8)
p_drone = (int(output_int8[1]) + 128) * 0.00390625
is_drone = p_drone >= 0.4684000015258789
This snippet specifies the arithmetic; the supplied runner provides the complete implementation. The model expects normalized features, not a raw waveform. Input scale is 0.018346942961215973, input zero point is -38; output scale is 1/256, output zero point is -128. Normalization is global with constants frozen from TRAIN. Do not replace it with per-window normalization.
For embedded integration, model_data.h contains the same model bytes, and mel_filterbank.h contains the filterbank. Follow HANDOFF_V2C1.md and the golden vectors when available. Serialized model size is not a measurement of the runtime tensor arena, SRAM footprint or total firmware flash.
Training data and evaluation scope
The corpus comes from DroneAudioDataset, containing indoor propeller recordings and augmented audio for Bebop and Mambo; repository metadata spells the second airframe membo. Its background sources include ESC-50, Speech Commands white noise and synthetic silence, as recorded in the upstream dataset README.
The retained feature corpus contains 10,817 windows in 1,912 recording groups, after 12 documented mixed-audio exclusions:
| Split | Windows | Drone | Background | Recording groups |
|---|---|---|---|---|
| TRAIN | 7,573 | 920 | 6,653 | 1,328 |
| Validation | 1,623 | 202 | 1,421 | 288 |
| TEST | 1,621 | 198 | 1,423 | 296 |
Recording groups do not span splits. Global normalization and the post-training INT8 representative dataset use TRAIN only; the deployed threshold was selected on validation. Quantization is post-training quantization, not verified quantization-aware training.
Recording groups are the available evidence unit, not proof of independent sessions. The corpus has no session metadata. It retains 26 background-dominated mixed files in the 5–20% residual-energy band; some share background source components across splits. This correlation and the two-airframe scope limit generalization claims.
The release is the original V2C1 baseline. A separate same-size experiment reached 96.97% recall but increased FPR to 1.69% on the historical benchmark and was rejected as a better overall replacement. Those experimental weights are not the default model in this release.
Intended research use and limitations
Use the model to reproduce the documented acoustic-classification benchmark and study TinyML frontend or inference implementations, subject to the unresolved permissions below. Further development needs recordings from independent sessions, microphones, environments and airframes, followed by validation chosen before a new independent evaluation.
The evidence does not establish Seshachalam or other outdoor field performance; detection at 50–100 metres or any other distance; performance on unseen airframes; an operational surveillance or safety system; ESP32 latency, RAM, power or battery life; or real-world recall and false-alarm guarantees. A high score is a classification output, not proof that a physical drone is present.
Source data permissions and model licensing
The current DroneAudioDataset license, checked 2026-10-07, limits dataset use to education and noncommercial academic research, imposes military/defence/intelligence/security restrictions, and restricts dataset redistribution. Third-party materials retain their own terms. The current ESC-50 license labels the dataset as CC-BY-NC 3.0 and lists individual-clip attribution and license information.
These are source-data terms. This repository has no established model-weight license or verified public-redistribution permission. The metadata uses license: unknown to record that gap. The card does not determine how dataset terms apply to learned weights, what terms governed historical acquisition, or whether later terms apply retroactively. It grants no MIT, Apache, unrestricted commercial-use or other broad rights. The availability of these files does not certify that downstream reuse or redistribution has been cleared; resolve the required permissions with the relevant rights holders before such use.
Attribution
Source dataset publication: Sara Al-Emadi, Abdulla Al-Ali, Abdulaziz Al-Ali and Amr Mohamed, “Audio Based Drone Detection and Identification using Deep Learning,” IWCMC, 2019.
Background dataset: Karol J. Piczak, “ESC: Dataset for Environmental Sound Classification,” ACM Multimedia, 2015. The ESC-50 repository provides dataset and clip attribution information.
The upstream DroneAudioDataset README also credits Pete Warden's “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.” Source authors are credited for the data; their work does not imply endorsement of this model or its release.
- Downloads last month
- -