EdgeCard Model Weights
This repository contains the pretrained and fine-tuned model weights used in EdgeCard: Real-Time Student Card Detection and Recognition on Low-Power Devices.
Source code: https://github.com/HuyHoang172004/EdgeCard-OCR
Pipeline
EdgeCard consists of four stages:
- Card detection and perspective alignment using YOLO26n-Pose.
- Text-region detection and spatial assignment using YOLO26n-OBB.
- Text recognition using PP-OCRv6 Small Recognition.
- Rule-based parsing and post-processing to produce structured student-card information.
The final end-to-end pipeline uses the models marked as Selected below.
Available Models
| Stage | Model | Role | Status |
|---|---|---|---|
| Stage 1 | YOLO26n-Pose | Student-card detection and four semantic corner keypoints | Selected |
| Stage 2 | YOLO26n-OBB | Oriented text-region detection | Selected |
| Stage 2 | DBNet | Text-detection benchmark | Benchmark |
| Stage 2 | PP-OCRv6 Small Detection | Text-detection benchmark | Benchmark |
| Stage 3 | PP-OCRv6 Small Recognition | Text recognition | Selected |
| Stage 3 | MobileNetV3-CRNN | Text-recognition benchmark | Benchmark |
| Stage 3 | RepSVTR | Text-recognition benchmark | Benchmark |
Repository Structure
stage-1/
βββ best-stage1.pt
βββ best-stage1_ncnn_model/
βββ metadata.yaml
βββ model.ncnn.bin
βββ model.ncnn.param
stage-2/
βββ dbnet/
βββ pp-ocrv6/
βββ yolo26n-obb/
βββ best-stage2.pt
βββ best-stage2_ncnn_model/
stage-3/
βββ mobilenetv3-crnn/
β βββ best.pth
β βββ mobilenetv3_crnn.onnx
βββ pp-ocrv6/
β βββ ppocrv6_small_rec.onnx
βββ repsvtr/
βββ repsvtr.onnx
Selected Models and Results
Stage 1 β YOLO26n-Pose
YOLO26n-Pose predicts four semantic corner keypoints in the fixed order: top-left, top-right, bottom-right, bottom-left.
Stage 2 β YOLO26n-OBB
| Model | Precision (%) | Recall (%) | F1-score (%) | Mean matched IoU (%) |
|---|---|---|---|---|
| YOLO26n-OBB | 96.31 | 98.28 | 97.28 | 79.04 |
| DBNet | 88.69 | 95.42 | 91.93 | 76.41 |
| PP-OCRv6 Small Detection | 92.18 | 93.91 | 93.04 | 73.07 |
Stage 3 β PP-OCRv6 Small Recognition
| Model | Exact Accuracy (%) | CER (%) | NES (%) | Latency (ms/crop) |
|---|---|---|---|---|
| MobileNetV3-CRNN | 98.22 | 0.26 | 99.73 | 11.18 |
| PP-OCRv6 Small Recognition | 99.03 | 0.11 | 99.91 | 10.24 |
| RepSVTR | 98.06 | 0.20 | 99.83 | 12.57 |
NES denotes Normalized Edit Similarity, where higher values are better.
Deployment Formats
Depending on the model and deployment target, the repository provides one or more of the following formats:
- PyTorch:
.pt,.pth - ONNX:
.onnx - NCNN:
.param,.bin - Paddle inference:
.pdmodel,.pdiparams,.yml,.json
End-to-End Performance
The complete EdgeCard pipeline achieved an Overall Exact Accuracy of 94.25% on both the PC and Raspberry Pi 4 evaluation setups.
On Raspberry Pi 4, the mean end-to-end processing time was approximately 1105.82 ms per card, corresponding to about 0.90 card/s, with a peak RSS of approximately 795 MiB.
Dataset Availability
The models were trained and evaluated using student-card data from the Academy of Cryptography Techniques (ACTVN).
Because the dataset contains personally identifiable information, including student names, student identifiers, and class information, the raw images and annotations are not publicly released.
For reproducibility purposes, qualified researchers may contact the authors regarding possible dataset access, subject to applicable institutional and privacy requirements.
Reproducibility
Training, benchmarking, and end-to-end evaluation code are available at:
https://github.com/HuyHoang172004/EdgeCard-OCR
Citation
If you use these model weights or the EdgeCard implementation in your research, please cite the corresponding EdgeCard paper.
A complete BibTeX entry will be added after publication.
License
Please refer to the licenses and terms of the original model frameworks and pretrained models used by each component. Dataset access and redistribution are subject to separate privacy and institutional restrictions.