Parakeet-TDT-0.6B-v3 Fine-Tuned for ATC — GGUF
This is for the ATC Audio Analyzer project
This repository contains a GGUF conversion of qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC, a Parakeet-TDT-0.6B-v3 model fine-tuned for air traffic control (ATC) speech recognition.
The original model was fine-tuned on the jacktol/ATC-ASR-Dataset using NVIDIA NeMo. The GGUF conversion is intended to make the model easier to run with local inference runtimes that support the Parakeet GGUF format, including CrispASR.
This repository contains a format conversion, not an additional fine-tuning or training run.
Model Details
Model Description
- Model type: Automatic Speech Recognition (ASR)
- Architecture: Parakeet-TDT / Transformer-Transducer (TDT)
- Base model:
nvidia/parakeet-tdt-0.6b-v3 - Fine-tuned model:
qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC - Domain: Air Traffic Control communications
- Language: English
- Sampling rate: 16 kHz
- Model format: GGUF
- Conversion precision: F16 by default
- License: MIT
- Conversion framework/runtime: CrispASR
The upstream fine-tuned model reports a validation WER of 0.0558 and a test WER of 0.0599 on the ATC-ASR dataset's evaluation splits.
Important: These WER results are results reported for the original fine-tuned model. They should not be interpreted as an independently reproduced evaluation of this GGUF conversion.
Model Sources
- Fine-tuned model: https://huggingface.co/qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC
- Base model: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
- Conversion script: https://github.com/CrispStrobe/CrispASR/blob/main/models/convert-parakeet-to-gguf.py
- CrispASR: https://github.com/CrispStrobe/CrispASR
- Training dataset: https://huggingface.co/datasets/jacktol/ATC-ASR-Dataset
Uses
Direct Use
This model is intended for automatic transcription of English ATC communications, particularly speech containing aviation and air-traffic-control terminology.
The GGUF format is intended for local/offline inference in compatible runtimes.
Potential uses include:
- ATC speech transcription
- Aviation speech recognition experiments
- Local/offline ASR applications
- Benchmarking domain-specific speech recognition
- Research into aviation communication ASR
- Integration into local speech-processing pipelines
Downstream Use
The model can be integrated into larger speech-processing systems, including applications that combine:
- Audio capture
- Voice activity detection
- Streaming or batch ASR
- Post-processing
- Domain-specific information extraction
For example, the model can be used as the ASR component of an offline aviation speech-processing application.
Applications using this model should perform their own validation on the target audio conditions before relying on transcriptions operationally.
Out-of-Scope Use
This model is not intended to:
- Serve as a general-purpose multilingual ASR model
- Be treated as a general conversational speech recognizer
- Replace certified aviation communication or safety systems
- Make autonomous air-traffic-control decisions
- Provide safety-critical instructions without human verification
- Be assumed to have the same accuracy on domains outside ATC communications
- Be used as the sole source of truth for safety-critical aviation decisions
Bias, Risks, and Limitations
The model was fine-tuned for the ATC domain and therefore may perform differently on speech outside this domain.
Potential limitations include:
- Accented English may result in higher transcription error rates.
- Strong background noise can reduce transcription accuracy.
- Radio communication artifacts may affect recognition.
- Unusual speaker characteristics may affect performance.
- Aviation terminology not represented in the training data may be transcribed incorrectly.
- Poor audio quality, clipping, overlapping speech, or very short utterances may reduce accuracy.
- Results reported on the ATC-ASR dataset may not generalize to other ATC environments, countries, accents, equipment, or communication channels.
- GGUF conversion does not guarantee bit-for-bit equivalence with the original inference implementation.
The model should therefore be evaluated on representative target audio before deployment.
Recommendations
For applications where transcription errors could have significant consequences:
- Keep a human in the loop.
- Validate the transcription against the original audio.
- Evaluate performance on representative local ATC audio.
- Measure WER/CER separately for different accents, noise conditions, speakers, and communication environments.
- Do not use the model as the sole component of a safety-critical aviation system.
How to Get Started with the Model
Using CrispASR
The recommended way to use this GGUF conversion is with a runtime that supports the Parakeet GGUF architecture, such as CrispASR.
The conversion was generated using the following script:
https://github.com/CrispStrobe/CrispASR/blob/main/models/convert-parakeet-to-gguf.py
The converter accepts either a NeMo checkpoint or a Hugging Face Parakeet model:
python models/convert-parakeet-to-gguf.py \
--hf qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC \
--output parakeet-tdt-0.6b-v3-atc-f16.gguf
The converter defaults to F16.
For supported quantized variants, the conversion script also accepts:
--quant q4_k
or:
--quant q8_0
For example:
python models/convert-parakeet-to-gguf.py \
--hf qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC \
--output parakeet-tdt-0.6b-v3-atc-q4_k.gguf \
--quant q4_k
The exact quantization used should correspond to the GGUF file provided in this repository.
Training Details
Training Data
The original fine-tuned model was trained on the Jacktol ATC-ASR Dataset.
The dataset contains segmented audio/transcript pairs for ASR in the ATC domain. It contains 16 kHz mono WAV audio and aligned transcripts, with training, validation, and test splits.
The dataset combines ATC speech resources including:
- UWB ATC Corpus
- ATCO2 1-Hour Test Subset
The dataset is intended for aviation communication ASR and contains accented and noisy English speech.
Training Procedure
The GGUF model itself was not trained. It was converted from the upstream fine-tuned model.
The original fine-tuning was performed using NVIDIA NeMo.
According to the upstream model card:
- Base model:
nvidia/parakeet-tdt-0.6b-v3 - Dataset:
jacktol/ATC-ASR-Dataset - Epochs: 16
- Batch size: 16
- Learning rate: 1e-4
- Optimizer: AdamW
- Weight decay: 1e-3
- Scheduler: CosineAnnealing
- Warmup steps: 5000
- Minimum learning rate: 5e-6
- Precision: Mixed precision FP16
- Tokenizer: Parakeet default subword tokenizer
- Training hardware: NVIDIA H200
- Reported training time: less than 1 hour
Conversion Procedure
The GGUF conversion was performed using the convert-parakeet-to-gguf.py script from CrispASR.
The converter:
- Loads the Parakeet model from a Hugging Face model repository or NeMo checkpoint.
- Loads the model configuration and tokenizer.
- Maps the Parakeet model tensors to the GGUF tensor layout expected by the Parakeet runtime.
- Stores the SentencePiece vocabulary in GGUF metadata.
- Stores Parakeet-specific model configuration in GGUF metadata.
- Writes the resulting model in GGUF format.
The converter supports F16 output by default and optional q4_k and q8_0 linear-weight quantization.
Evaluation
Testing Data, Factors & Metrics
Testing Data
The upstream model was evaluated on the official test split of the ATC-ASR-Dataset.
The dataset contains English ATC communications with varying speakers, accents, and acoustic conditions.
Factors
Relevant factors for evaluation include:
- Speaker accent
- Background noise
- Radio communication quality
- Aviation terminology
- Speaking rate
- Speaker characteristics
- Audio quality
- Domain-specific vocabulary
Metrics
Word Error Rate (WER) is used to measure ASR transcription accuracy.
WER is commonly calculated as:
WER = (Substitutions + Deletions + Insertions) / Number of reference words
Lower WER indicates fewer word-level transcription errors.
Results
The upstream fine-tuned model reports:
| Metric | Result |
|---|---|
| Validation WER | 0.0558 |
| Test WER | 0.0599 |
| Training time | < 1 hour |
| Training hardware | NVIDIA H200 |
These figures are taken from the upstream fine-tuned model card and have not been independently reproduced here for the GGUF file.
Summary
The GGUF conversion preserves the trained Parakeet-TDT model parameters and tokenizer in a format intended for local inference. The conversion itself does not introduce additional training or fine-tuning.
Users should independently benchmark the GGUF model if they require verified accuracy for a particular deployment environment.
Model Examination
No additional interpretability or model examination was performed as part of the GGUF conversion.
The GGUF file should be considered a converted representation of the upstream fine-tuned model rather than a separately trained model.
Environmental Impact
The GGUF conversion is a model-format conversion and is substantially different from the original model's training process.
No formal carbon-emissions measurement was recorded for this conversion.
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
- Hardware Type: Not recorded
- Hours used: Not recorded
- Cloud Provider: None recorded
- Compute Region: Not recorded
- Carbon Emitted: Not measured
Technical Specifications
Model Architecture and Objective
The model uses the Parakeet-TDT architecture for automatic speech recognition.
The GGUF converter records the following architecture parameters:
- Sample rate: 16,000 Hz
- Mel features: 128
- FFT size: 512
- Window length: 400 samples
- Hop length: 160 samples
- Encoder dimension: 1024
- Encoder layers: 24
- Attention heads: 8
- Attention head dimension: 128
- Feed-forward dimension: 4096
- Subsampling factor: 8
- Subsampling channels: 256
- Convolution kernel: 9
- Prediction network hidden dimension: 640
- Prediction network layers: 2
- Joint network hidden dimension: 640
- Vocabulary size: 8192
- Blank token ID: 8192
- TDT durations: 0, 1, 2, 3, 4
The encoder is based on a 24-layer FastConformer architecture with a transducer prediction network and TDT joint network.
Compute Infrastructure
The GGUF format is designed for local inference and can be used on compatible CPU and/or GPU configurations depending on the inference runtime.
Actual inference speed and memory usage depend on:
- CPU architecture
- GPU architecture
- Number of GPU layers/offloading configuration
- GGUF precision/quantization
- Audio duration
- Runtime implementation
- Streaming configuration
Hardware
No specific hardware requirement is imposed by the GGUF format itself.
For production or real-time applications, users should benchmark the model on their target hardware.
Software
The conversion was performed using:
- Python
- PyTorch
gguf- SentencePiece
- CrispASR conversion tooling
Conversion script:
https://github.com/CrispStrobe/CrispASR/blob/main/models/convert-parakeet-to-gguf.py
Runtime:
https://github.com/CrispStrobe/CrispASR
Citation
If you use this model, please cite the original fine-tuned model, the Parakeet base model, the dataset, and the GGUF conversion tooling.
Fine-tuned Model
@misc{qenneth_parakeet_atc,
title={Parakeet-TDT-0.6B-v3 Fine-Tuned on ATC-ASR Dataset},
author={qenneth},
year={2025},
publisher={Hugging Face},
howpublished={https://huggingface.co/qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC}
}
Base Model
@misc{nvidia2024parakeet,
title={Parakeet-TDT-0.6B-v3},
author={NVIDIA},
year={2024},
publisher={Hugging Face},
howpublished={https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3}
}
Dataset
@dataset{jacktol_atc_asr,
title={ATC-ASR Dataset},
author={Jacktol},
year={2023},
howpublished={https://huggingface.co/datasets/jacktol/ATC-ASR-Dataset}
}
GGUF Conversion
CrispASR Parakeet-to-GGUF conversion script:
https://github.com/CrispStrobe/CrispASR/blob/main/models/convert-parakeet-to-gguf.py
Glossary
- ASR: Automatic Speech Recognition
- ATC: Air Traffic Control
- GGUF: A model file format designed for efficient local inference
- TDT: Token-and-Duration Transducer
- WER: Word Error Rate
- F16: 16-bit floating-point representation
- Q4_K: 4-bit GGUF quantization format
- Q8_0: 8-bit GGUF quantization format
More Information
Original Fine-Tuned Model
https://huggingface.co/qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC
NVIDIA Parakeet-TDT-0.6B-v3
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
ATC-ASR Dataset
https://huggingface.co/datasets/jacktol/ATC-ASR-Dataset
CrispASR
https://github.com/CrispStrobe/CrispASR
Conversion Script
https://github.com/CrispStrobe/CrispASR/blob/main/models/convert-parakeet-to-gguf.py
Model Card Authors
GGUF conversion: CrispASR community / contributors.
The underlying model was developed and fine-tuned by the upstream model authors. Please refer to the original model repository for attribution and training details.
Model Card Contact
For issues specific to the GGUF conversion or CrispASR runtime, please use the CrispASR GitHub repository.
For issues concerning the underlying fine-tuned model, please contact the authors of qenneth/parakeet-tdt-0.6b-v3-finetuned-for-ATC.
- Downloads last month
- 20
We're not able to determine the quantization variants.
Model tree for pronoobie/Parakeet-v3-For_ATC
Base model
nvidia/parakeet-tdt-0.6b-v3