⚡️ PrunaSuperPoint
TensorRT SuperPoint on Jetson Orin Nano
~1.9× faster at 640×480 · ~2.1× faster at 1080p · Optimized from SuperPoint
PrunaSuperPoint is an optimization of the SuperPoint (SP) keypoint detection model. Based on rpautrat/SuperPoint by Rémi Pautrat and Paul-Edouard Sarlin. Optimized for the TensorRT runtime on Jetson Orin Nano devices, but also applicable to other runtime backends. The ideas used can be applied to other similar network architectures.
This work was done in collaboration with CEA (Aidge) for the DeepGreen project. You can also find a pruning tutorial for SuperPoint using Aidge operators in the Aidge codebase.
Example
Optimization methods
To achieve better performance, we use the following methods:
- Structural pruning: It's the main optimization method used. The SuperPoint model architecture consists of eight sequential convolutional layers, and after every two convolutional layers, the model has a pooling layer, which halves the width and height of the feature map. So, the model's biggest bottlenecks are the first convolutional layers, since they have to process the image at full input resolution. By reducing the input/output channel sizes of desired convolutional layers, the runtime can be improved. In the case of channel pruning, we keep the channels corresponding to kernels with the highest mean absolute weight values, which provides a good initialization point for potential distillation training.
- Hierarchical top-k keypoint selection: Instead of selecting the global top-k keypoints directly, this can be done hierarchically by splitting the image into chunks, selecting the top k keypoints from each chunk, and then aggregating the candidates to obtain the final result. This guarantees the original result, using the same k and assuming no scoring ties, while providing a noticeable speedup on constrained devices.
- Skipping keypoint refinement: The base model performs two refinement steps after non-maximum suppression to recover additional high-scoring keypoints. Skipping them, even if not a direct optimization, avoids four extra full-resolution passes, which is very expensive.
Benchmark
Metrics compare five different model variations.
original- base SP with original weightsoriginal-topk- base SP with hierarchical top-k enabled. Each variant uses the original k for sub-chunks, which guarantees the original results. The resolution 640x480 uses 32 chunks, and 1920x1080 uses 36 chunks.pruned-light- base SP where input channels of thebackbone_0_1layer have been pruned from 64 to 32 with no additional training. This demonstrates the effect of the most expensive layer and how the kernel selection heuristic can retain performance. Previous top-k optimization is also enabled.pruned- uses the16_16_24_32_64pruning config that affects the first six layers of the model and uses our distillation checkpoint weights, which were trained on 250 examples. We use a combination of KL divergence and cross-entropy loss to recover keypoint locations, and a cosine similarity-based loss for descriptors. Previous top-k optimization is also enabled.pruned-noref- previousprunedmodel with no NMS refinement steps.
Latency
| Setting | Description |
|---|---|
| Model format | ONNX models serialized as TensorRT engines using the trtexec CLI utility. |
| Precision | FP16 |
| Device | Jetson Orin Nano 8GB. |
| Metrics | 1) Mean inference time over 1000 images from the indoor trace using our benchmarking script 2) Mean inference time reported by the TensorRT trtexec benchmarking utility. |
| Units | Milliseconds (ms). |
Standard comparison:
| Input | original | pruned | Speedup |
|---|---|---|---|
| 640×480 | 17.8 / 16.8 | 9.6 / 8.5 | 1.85× / 1.98× |
| 1920×1080 | 111.1 / 109.6 | 54.2 / 52.9 | 2.05× / 2.07× |
All variants:
| Input | original | original-topk | pruned-light | pruned | pruned-noref |
|---|---|---|---|---|---|
| 640 × 480 | 17,8 / 16,8 | 17,0 / 16,0 | 14,4 / 13,5 | 9,6 / 8,5 | 8,3 / 7,0 |
| 1920 × 1080 | 111,1 / 109,6 | 102,5 / 102,1 | 86,9 / 84,6 | 54,2 / 52,9 | 45,1 / 43,5 |
Quality
| Setting | Description |
|---|---|
| Model format | PyTorch models |
| Input resolution | Native dataset resolution of 640×480 |
| Keypoints | 1024 keypoints per image |
| Datasets | 1) Indoor dataset with 767 images from UZH FPV Indoor forward facing trace #3 2) Outdoor dataset with 500 images from UZH FPV Outdoor forward facing trace #1 |
| Metrics | 1) Average keypoints covered, the keypoints coverage compares the spatial distribution of predictions against the original model over 8×8 image regions, with 1024 as the maximum score 2) Mean descriptor difference L2 norm, the L2 norm is calculated by comparing the descriptor vectors at shared keypoint locations and calculating the norm of their difference. |
| Trace | original | original-topk | pruned-light | pruned | pruned-noref |
|---|---|---|---|---|---|
| Inside | 1024 / 0 | 1024 / 0 | 637 / 0,92 | 817 / 0,24 | 700 / - |
| Outside | 1024 / 0 | 1024 / 0 | 485 / 0,89 | 754 / 0,27 | 671 / - |
Keypoint visualization examples
A few examples of the predicted keypoints are shown below. The first row uses image_1_1200.png from the Indoor trace, while the second uses image_0_1300.png from the Outdoor trace. Each target model is compared against the original model predictions. Green points indicate keypoints predicted at the exact same location by both models, while red and blue indicate keypoints predicted by only one of the two models.
Although there are relatively few exact overlaps, many predictions lie very close to each other. In addition, even when target model predictions do not match the original exactly, they often still correspond to sensible keypoint locations.
Quickstart
Install the required dependencies and the package in editable mode:
pip install -r requirements.txt
pip install -e .
Download the default dataset, UZH FPV Indoor forward-facing trace #3, to src/superpoint_pruning/distillation/data:
superpoint-pruning setup
Compare the keypoint predictions from the pruned model against the original SuperPoint model on an image:
superpoint-pruning plot \
--image-name image_1_1200.png \
--output ./test.png \
--num-keypoints 512 \
--pruning-config ./src/superpoint_pruning/weights/16_16_24_32_64.ckpt \
--overlay
The resulting comparison is written to ./test.png. To run the distillation and evaluation scripts, see the Annex: Run your own distillation and compare inference time section.
Limitations
- Other popular SuperPoint implementations exist, especially versions paired with matchers such as LightGlue. This repository does not use them because of their restrictive licenses. The work is meant to show that SuperPoint-like models can be pruned and distilled. You can apply the same methods to those implementations after the necessary structural changes and a distillation run with your configs.
- The right way to judge a pruned SuperPoint is to test it in the intended downstream task and see if it still retains performance. We don't explore it here since we are not currently looking at a specific task.
- Distillation used 250 frames from UZH-FPV Indoor forward facing trace #3 (767 eval images). Direct training images were excluded from evaluation, but results may still be biased because the frames are sequential video frames. Outdoor forward-facing trace #1 was not used for recovery training.
- Skipping NMS refinement usually still produces sensible keypoints, but that depends on the task.
License
This repository is licensed under the Apache License 2.0.
Third-party Attributions
- SuperPoint implementation and weights — Rémi Pautrat and Paul-Edouard Sarlin, rpautrat/SuperPoint.
- SuperPoint method — Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich, SuperPoint: Self-Supervised Interest Point Detection and Description, CVPR Workshops, 2018.
- UZH-FPV dataset — Jeffrey Delmerico, Titus Cieslewski, Henri Rebecq, Matthias Faessler, and Davide Scaramuzza, Are We Ready for Autonomous Drone Racing? The UZH-FPV Drone Racing Dataset, ICRA, 2019. Dataset
What's next?
- Experiment with PrunaSuperPoint for computer vision, drone, and edge-device workloads
- Compress your own models with Pruna and give us a ⭐️
Annex: Run your own distillation and compare inference time
For installing the package with dependencies for running the basic distillation, evaluation, and export scripts, run:
pip install -r requirements.txt
pip install -e .
To install the base dataset images (UZH FPV Indoor forward facing trace #3) used for training to the default path src/superpoint_pruning/distillation/data, run:
superpoint-pruning setup
Alternatively, specify the desired trace download link with the --download-url flag, e.g., the one for the UZH FPV Outdoor forward facing trace #1, and the local dataset directory name with --output-dir-name.
Run distillation training
The base PyTorch Lightning training config that we used to produce the pruned model can be found under src/superpoint_pruning/distillation/base_config.yaml.
Since this includes the hyperparameters and the data setup used for our training, then this should allow to reproduce the trained checkpoint that is provided under src/superpoint_pruning/weights/16_16_24_32_64.ckpt. The training script itself is very basic but can serve as a starting point for further experiments.
To replicate the training, assuming that the main Indoor data trace is installed in the previous, run:
superpoint-pruning setup --generate-gt
superpoint-pruning train
The first command computes and presaves the target keypoint and descriptor feature maps and the other initiate the Lightning training script.
Run evaluation
The first row of quality evaluation results can be reproduced using the following commands corresponding to the different target models:
superpoint-pruning evaluate
superpoint-pruning evaluate --hierarchical
superpoint-pruning evaluate --backbone_0_1 32 --hierarchical
superpoint-pruning evaluate --pruning-config ./src/superpoint_pruning/weights/16_16_24_32_64.ckpt --hierarchical
superpoint-pruning evaluate --pruning-config ./src/superpoint_pruning/weights/16_16_24_32_64.ckpt --hierarchical --skip-refinement
Note that in this script and other similar ones, we can define the pruning either with format --backbone_x_y new_input_channel_size or by pointing it to a ckpt file named in the format a_b_...ckpt, which is translated to --backbone_0_1 a, --backbone_1_0 b and so on.
For the evaluation using the outdoor trace (or some other trace), the --image-dir parameter has to be specified including the other specific options to use the same image range (--no-skip --start-idx 1000 --end-idx 1500).
Plot keypoint comparisons
For a quick keypoint prediction comparison plot between a target model and the original on a specific image, one can for example use:
superpoint-pruning plot --image-name image_1_1200.png --output ./test.png --num-keypoints 512 --pruning-config ./src/superpoint_pruning/weights/16_16_24_32_64.ckpt --overlay
This once again defaults to the UZH FPV Indoor forward facing trace #3 and accepts file names from the dataset.
Export model to ONNX
To export the desired PyTorch model to ONNX:
superpoint-pruning export --backbone_0_1 32
The model pruning configuration can once again be controlled with same flags as shown in the evaluation example.
As an example, we have published the ONNX export of the original model to https://huggingface.co/EclipseAidge/SuperPoint.
Run TRT inference benchmark
In case the device supports TensorRT (with the necessary packages installed) and there is a serialized TRT model (here a dummy example of superpoint.engine), then the local python benchmark with a TRT wrapper can be run as:
superpoint-pruning benchmark --model-path superpoint.engine






