Instructions to use davidpblcrd/ramp-susy-tinyvit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use davidpblcrd/ramp-susy-tinyvit with timm:
import timm model = timm.create_model("hf-hub:davidpblcrd/ramp-susy-tinyvit", pretrained=True) - Notebooks
- Google Colab
- Kaggle
RAMP SuSy TinyViT
ONNX checkpoint of TinyViT for synthetic image detection on the SuSy Dataset and the per-layer mixed-precision quantization policies that RAMP selected for two edge CPUs, from the paper RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models (BMVC 2026). Applying a policy to the FP32 model rebuilds the quantized models. RAMP matches the FP32 accuracy (97.4% against 97.5%) on a Raspberry Pi 5 while cutting latency from 167.8 ms to 124.8 ms per image (1.34×); uniform INT8 is faster but its accuracy collapses to 69.0%.
Model Details
Model Description
This repository contains the FP32 model and, for each tested CPU, the quantization policies of every configuration RAMP evaluated on it.
| File | Description |
|---|---|
tinyvit_fp32.onnx |
Full-precision baseline, fine-tuned on the SuSy Dataset. |
quantization_policies_applem1.json |
Policies selected on the Apple M1. |
quantization_policies_cortexa76.json |
Policies selected on the Raspberry Pi 5 (Cortex-A76). |
Each JSON lists FP32, uniform INT8 and the configuration of the RAMP (JSD) and JSD + Latency sweeps, one per K-Means threshold, with the precision (int8/fp32) of each of the 88 quantizable nodes. The configurations reported in the paper are marked in selected_for:
selected_for |
Apple M1 | Cortex-A76 | INT8 nodes (of 88) |
|---|---|---|---|
INT8 unif. |
all_int8 |
all_int8 |
88 |
RAMP (JSD) |
jsd_2 |
jsd_2 |
62 / 62 |
JSD + Latency |
hw_jsd_3 |
hw_jsd_2 |
61 / 59 |
Uniform INT8 quantization collapses this architecture (63.2% accuracy on Apple M1, 69.0% on Cortex-A76), whereas the RAMP policy keeps the most sensitive layers in FP32 and stays within 0.2 accuracy points of the FP32 baseline on both tested platforms.
- Developed by: David Población-Criado, Dario Garcia-Gasulla, Eduardo Quinones (Barcelona Supercomputing Center (BSC))
- Funded by: DARE SGA1 (EuroHPC JU, grant agreement No 101202459) and ODISSEE (Horizon Europe, grant agreement No 101188332)
- Shared by: David Población-Criado
- Model type: Image classifier (TinyViT-21M hybrid convolution-transformer) in ONNX format (FP32), with the policies to quantize it to uniform INT8 or mixed-precision INT8/FP32 (static post-training quantization, QDQ)
- License: Apache 2.0
- Finetuned from model: timm/tiny_vit_21m_224.in1k
Model Sources
Uses
Direct Use
Classifying 224×224 RGB images into one of the six SuSy classes (one authentic source and five generators) with ONNX Runtime:
| Index | Label |
|---|---|
| 0 | coco (authentic) |
| 1 | dalle-3-images |
| 2 | diffusiondb |
| 3 | midjourney-images |
| 4 | midjourney_tti |
| 5 | realisticSDXL |
Downstream Use
The model can serve as a lightweight synthetic-image pre-filter in an on-device pipeline, flagging images for further review. The quantization procedure itself can be applied to other vision models with the code in the repository.
Out-of-Scope Use
- Sole evidence in forensic, legal, journalistic or moderation decisions. The model can be wrong and must not be used on its own to label content or people.
- Generators outside the training set. Images from newer or unseen generators, or heavily post-processed images (compression, resizing, screenshots), are likely to be misclassified.
Bias, Risks, and Limitations
- Hardware dependence. Latency gains were measured on ARM64 (Apple M1, Raspberry Pi 5 Cortex-A76). On x86-64 the default ONNX Runtime CPU provider degraded latency for all tested models, another execution provider should be used.
- Calibration dependence. As a data-driven PTQ method, the JSD ranking and the activation ranges depend on the 1000 calibration images from the SuSy training split. Out-of-distribution inputs may behave differently from the reported results.
- One-At-a-Time sensitivity. Layer sensitivity is measured in isolation and ignores inter-layer interactions so the selected policy is robust but not guaranteed to be optimal.
- Task scope. RAMP has only been evaluated on image classification.
- Dataset biases. The SuSy Dataset covers a limited set of generators and uses COCO as the only source of authentic images, so the model inherits its content and style distribution.
Recommendations
Benchmark the model on your target device before deploying it, since operator fusion and INT8 kernel availability change between platforms. Validate accuracy on data representative of your deployment, and treat predictions as probabilistic signals rather than definitive proof that an image is synthetic.
How to Get Started with the Model
1. Rebuild the quantized model from the FP32 model and the policy of your CPU, with ramp-mpq. It applies the same pipeline as the paper: ONNX Runtime preprocessing, MinMax calibration on 1000 images of the SuSy train split and QDQ quantization of the policy's INT8 nodes.
git clone https://github.com/davidpob99/ramp-mpq && cd ramp-mpq && uv sync
hf download davidpblcrd/ramp-susy-tinyvit --local-dir hf-ramp-susy-tinyvit
# What each policy file holds, and which configuration the paper selected
uv run python scripts/build_quantized.py hf-ramp-susy-tinyvit/quantization_policies_cortexa76.json --list
uv run python scripts/build_quantized.py hf-ramp-susy-tinyvit/quantization_policies_cortexa76.json \
--policy "RAMP (JSD)" \
--onnx hf-ramp-susy-tinyvit/tinyvit_fp32.onnx \
--calibration-dir <path>/SuSy-Dataset/data/train \
--output model_ramp.onnx
--config <name> rebuilds any other configuration of the sweep (e.g. all_int8, jsd_1).
2. Run it with ONNX Runtime:
import numpy as np
import onnxruntime as ort
from PIL import Image
from torchvision.transforms import v2 as T
import torch
LABELS = ["coco", "dalle-3-images", "diffusiondb", "midjourney-images", "midjourney_tti", "realisticSDXL"]
model_path = "model_ramp.onnx" # built in step 1
opts = ort.SessionOptions()
opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
opts.intra_op_num_threads = 8 # the number of CPU cores: 8 on the Apple M1, 4 on the Raspberry Pi 5
opts.inter_op_num_threads = 1
session = ort.InferenceSession(model_path, opts, providers=["CPUExecutionProvider"])
transform = T.Compose([T.CenterCrop(224), T.ToImage(), T.ToDtype(torch.float32, scale=True)])
x = transform(Image.open("image.jpg").convert("RGB")).unsqueeze(0).numpy()
logits = session.run(None, {"input": x})[0]
print(LABELS[int(np.argmax(logits, axis=1)[0])])
Training Details
Training Data
The FP32 model was fine-tuned on the train split of the SuSy Dataset, a collection of authentic (COCO) and synthetic images from five generator sources. The same split provided the 1000 images used to calibrate the static quantization and to profile layer sensitivity.
Training Procedure
Two stages:
- FP32 fine-tuning of the ImageNet-pretrained timm TinyViT-21M with a new 6-class head
- RAMP post-training quantization with ONNX Runtime (no retraining)
Preprocessing
- Training:
RandomCrop(224),RandomHorizontalFlip(p=0.5), scaling to[0, 1]. - Evaluation and calibration:
CenterCrop(224), scaling to[0, 1].
Training Hyperparameters
- Training regime: fp32
- Optimizer: AdamW (lr = 1e-3, weight decay = 0.05)
- Scheduler: cosine annealing with 10-epoch linear warmup (start factor 0.01)
- Loss: cross-entropy
- Quantization: ONNX Runtime static PTQ, QDQ format, symmetric INT8 weights (
QInt8), MinMax calibration on 1000 images; quantizable ops: Conv, Gemm, MatMul - RAMP: JSD sensitivity metric, 1D K-Means with K = 5, knee-point selection on the accuracy-latency Pareto front
Speeds, Sizes, Times
Median per-image latency reported in the paper, batch size 1, ONNX Runtime default CPU execution provider:
| Platform | FP32 | Uniform INT8 | RAMP (JSD) | RAMP speed-up vs. FP32 |
|---|---|---|---|---|
| Apple M1 | 39.0 ms | 27.4 ms | 27.4 ms | 1.42× |
| Raspberry Pi 5 (Cortex-A76) | 167.8 ms | 104.4 ms | 124.8 ms | 1.34× |
| Model | Size |
|---|---|
FP32 (tinyvit_fp32.onnx) |
≈83 MB |
| Uniform INT8 (rebuilt) | ≈33 MB |
| RAMP (JSD) (rebuilt) | ≈33 MB |
Evaluation
Testing Data, Factors & Metrics
Testing Data
Test split of the SuSy Dataset.
Factors
Results are disaggregated by hardware platform (Apple M1, Raspberry Pi 5) and by quantization policy.
Metrics
- Top-1 accuracy (%) on the test split.
- Latency (ms): median end-to-end inference time per image, batch size 1, images decoded one at a time in the main process.
- Collapse: a policy counts as collapsed when its Top-1 accuracy falls more than 15 points below FP32.
Results
| Method | Apple M1 Acc | Apple M1 Lat | Cortex-A76 Acc | Cortex-A76 Lat |
|---|---|---|---|---|
| FP32 | 97.5 | 39.0 | 97.5 | 167.8 |
| Uniform INT8 | 63.2 ❌ | 27.4 | 69.0 ❌ | 104.4 |
| RAMP (JSD) | 97.3 | 27.4 | 97.4 | 124.8 |
| STD | 89.6 | 27.1 | 92.7 | 122.5 |
| HAWQ-V2 | 82.1 ❌ | 27.0 | 80.8 ❌ | 122.5 |
| JSD + Latency (ablation) | 97.3 | 27.4 | 97.4 | 144.1 |
❌ = collapse (> 15 points below FP32). Accuracy in %, latency in ms, as reported in Table 2 of the paper. STD and HAWQ-V2 are the baselines RAMP is compared against; their policies are not part of this repository.
Model Examination
Per-layer INT8 speed-ups differ markedly between CPUs, and between layer types on the same CPU: convolutions gain the most, while normalization and element-wise operations show minimal gains. Section 5 of the paper analyses this; its Figure 3 shows ConvNeXt-Tiny, and the other models behave similarly.
Technical Specifications
Model Architecture and Objective
TinyViT-21M (≈21M parameters) with a 6-way linear classification head, trained with cross-entropy for synthetic image source attribution. The FP32 model and the models rebuilt from the policies share the graph and differ in precision:
- FP32: all nodes in FP32.
- Uniform INT8 (
all_int8): every quantizable node in INT8 through Q/DQ pairs. - RAMP: each of the layers ranked after ONNX Runtime fusion is either INT8 (weights and activations) or FP32, following the policy.
All share the same interface:
- Input:
input,float32[1, 3, 224, 224], RGB in[0, 1] - Output: logits,
float32[1, 6]
Compute Infrastructure
Hardware
- Target / evaluation: Apple MacBook Air M1 and Raspberry Pi 5 (Arm Cortex-A76).
- Not recommended: x86-64 with the default ONNX Runtime CPU provider.
Software
- ONNX Runtime 1.23.2 (CPU execution provider)
- PyTorch, timm and Lightning for FP32 training
- Hydra-based pipeline in ramp-mpq
Citation
@inproceedings{poblacion2026ramp,
title = {{RAMP}: Robust Adaptive Mixed-Precision Quantization for Edge {CPU} Vision Models},
author = {Poblaci{\'o}n-Criado, David and Garcia-Gasulla, Dario and Quinones, Eduardo},
booktitle = {British Machine Vision Conference (BMVC)},
year = {2026},
eprint = {2609.28262},
archivePrefix = {arXiv}
}
Model tree for davidpblcrd/ramp-susy-tinyvit
Base model
timm/tiny_vit_21m_224.in1k