File size: 5,721 Bytes
f02f671 66d104d f02f671 66d104d b70efa7 66d104d b70efa7 e6ae21c b70efa7 e6ae21c b70efa7 e6ae21c b70efa7 66d104d f114d60 66d104d f114d60 66d104d f114d60 66d104d f114d60 66d104d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | ---
license: mit
library_name: onnxruntime
pipeline_tag: image-to-image
tags:
- onnx
- computer-vision
- image-cropping
- saliency
base_model: timm/repvit_m0_9
---
# FocalNet
Appwrite FocalNet is a compact vision model for content-aware image cropping. It
predicts a 64×64 importance map, then ranks crop candidates with a small
composition head trained on human preferences. Training code lives in
[appwrite/focalnet](https://github.com/appwrite/focalnet).
The reference RepViT-M0.9 ranker has 4.76 million parameters. The published
`focalnet-human.onnx` artifact is 19.45 MiB in FP32.
## Examples
These photographs come from [Autogravity](https://github.com/appwrite/autogravity).
A centered 1:1 crop of the park scene keeps grass. FocalNet keeps the dog.

On a two-person portrait, FocalNet holds the nearer subject instead of splitting
the frame.

## Files
| File | Role | SHA-256 |
| --- | --- | --- |
| `focalnet-human.onnx` | Importance map + crop ranking (format v2) | `59164c601c98cea3f62b25166710831dac63e1a872fc64767c65316ad5385439` |
| `focalnet.onnx` | Importance map only (format v1) | `87becceb269a2973c359df789783be49a9f840f47170a015d3776d7c4145a2ce` |
| `focalnet-human.pt` | PyTorch ranking checkpoint | `82c4d695310c220e5f3bcfbf5e64d315a5e826cfd77e945e61e032460014480a` |
| `focalnet.pt` | PyTorch importance checkpoint | `d1942f0652f8ea85f75ffc0cb1bf40d7e70b38e7ad102ab0e6df5c2e07ce52cf` |
JSON sidecars next to the ONNX files record the preprocessing contract, parameter
counts, and export verification.
Use `focalnet-human.onnx` for inference. The PyTorch files are for continued
training, not for the ONNX Runtime path.
## Usage
Download the ranking model and crop a photo:
```sh
mkdir -p artifacts
curl -fsSL \
"https://huggingface.co/appwrite/focalnet/resolve/main/focalnet-human.onnx" \
-o artifacts/focalnet-human.onnx
uv run focalnet predict-human artifacts/focalnet-human.onnx photo.jpg \
--aspect-ratio 16:9
```
From Python:
```python
from pathlib import Path
from urllib.request import urlretrieve
from focalnet.human_runtime import HumanCropPredictor
url = "https://huggingface.co/appwrite/focalnet/resolve/main/focalnet-human.onnx"
path = Path("artifacts/focalnet-human.onnx")
path.parent.mkdir(parents=True, exist_ok=True)
urlretrieve(url, path)
crop = HumanCropPredictor(path).predict("photo.jpg", aspect_ratio=16 / 9)
```
Importance-only inference:
```sh
mkdir -p artifacts
curl -fsSL \
"https://huggingface.co/appwrite/focalnet/resolve/main/focalnet.onnx" \
-o artifacts/focalnet.onnx
uv run focalnet predict artifacts/focalnet.onnx photo.jpg --aspect-ratio 16:9
```
`predict-human` generates up to 125 crops across five positions and five zoom
levels, scores a padded maximum of 128 candidates in one ONNX call, and selects
the highest preference score among crops whose retained importance is within
0.05 of the best candidate. The returned human score is a relative logit.
Compare it only among crops for the same image.
## ONNX contract
| Field | Contract |
| --- | --- |
| Image input | `image`, float32 `[1,3,256,256]`, RGB NCHW |
| Normalization | `(pixel/255 - [0.485,0.456,0.406]) / [0.229,0.224,0.225]` |
| Resize | Apply EXIF orientation, then aspect-fit with bilinear interpolation |
| Alpha and padding | Composite alpha on black; normalized padding values are zero |
| Importance output | `importance`, float32 `[1,1,64,64]`, sigmoid applied |
| Ranking inputs | `boxes` `[1,128,4]` and letterbox `content` `[1,4]` |
| Ranking output | `crop_scores`, float32 `[1,128]` |
## Results
The importance model was trained on 500,000 teacher-labeled Open Images V7
samples (RepViT-M0.9 encoder, 48-channel decoder, 10 epochs, best epoch 7).
Teacher agreement on the group-disjoint 10,000-image validation split:
| Metric | Result |
| --- | ---: |
| Map MAE | 0.09975 |
| Normalized centroid error | 0.06430 |
| 1:1 teacher importance retained | 91.09% |
| 16:9 teacher importance retained | 90.86% |
| 4:5 teacher importance retained | 85.56% |
The ranking head (7,069 parameters) was pretrained on CPC and fine-tuned on
GAICD. The importance network stays frozen. On GAICD's 500-image test split:
| Metric | GAICD only | CPC → GAICD |
| --- | ---: | ---: |
| Spearman correlation | 0.75696 | **0.76450** |
| Pearson correlation | 0.78211 | **0.79192** |
| Top-5 accuracy | 42.48% | **43.70%** |
| Top-10 accuracy | 64.45% | **64.54%** |
| Pairwise accuracy | 86.00% | **86.44%** |
The GAICD test split was inspected during several development iterations. Treat
these figures as engineering benchmarks, not an untouched final test. INT8
post-training quantization did not pass the crop-quality gate; FP32 is the
published artifact.
## Training data
- Open Images V7 (teacher-labeled importance maps from U²-Net + YuNet)
- [Comparative Photo Composition (CPC)](https://www3.cs.stonybrook.edu/~cvl/projects/wei2018goods/VPN_CVPR2018s.html)
- [GAICD](https://github.com/HuiZeng/Grid-Anchor-based-Image-Cropping-Pytorch)
The reviewed CPC and GAICD archives do not state explicit image or annotation
licenses. This checkpoint is published under MIT for the model files; that
license does not grant rights to the third-party datasets or teacher weights
used during training.
## License
Model files are available under the [MIT License](https://github.com/appwrite/focalnet/blob/main/LICENSE),
the same license as the training code. Review dataset and teacher-model terms
before redistribution of derived artifacts.
|