FrancescoMrn/cnn-fkp-baseline

68-point facial keypoint regression, in the iBUG 300-W landmark ordering.

Input and output

  • Input: a single-channel (grayscale) 224x224 face crop, scaled to [0, 1], shape (N, 1, H, W).
  • Output: (N, 136) โ€” 68 (x, y) pairs, normalised to [-0.5, 0.5] relative to the crop. Multiply by the crop size and add half of it to get pixels.

Usage

from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

from fkp.models import build_model

path = hf_hub_download("FrancescoMrn/cnn-fkp-baseline", "model.safetensors")
model = build_model("baseline_cnn")
model.load_state_dict(load_file(path))
model.eval()

Or, with this repository installed, fkp resolves the config automatically:

from fkp.hub import load_from_hub

model, meta = load_from_hub("FrancescoMrn/cnn-fkp-baseline")

Results

Validation NME 0.3172 โ€” mean landmark error as a fraction of the inter-ocular distance. Lower is better; 0.05 means the average landmark sits 5% of the distance between the outer eye corners away from its target.

region NME
jaw 0.3276
right_brow 0.3260
left_brow 0.3251
nose 0.3232
mouth 0.3133
right_eye 0.2979
left_eye 0.2965

The jaw is 17 of the 68 points and is consistently the hardest region, so it dominates the overall average.

Training

architecture baseline_cnn
optimiser adamw @ lr 0.001
loss mse
schedule cosine
initialisation kaiming
batch size 64
epochs 200
seed 0
split 2766 train / 696 validation images
identities 278 / 69, no overlap

The validation split holds out whole identities. The images are YouTube Faces frames, with roughly ten frames per person, so a random split would put near-duplicate frames of the same face on both sides and measure memorisation rather than generalisation.

Data and limitations

Trained on the Udacity extract of the YouTube Faces Database, annotated with 68 landmarks. The dataset is not demographically balanced, and the model inherits that. It expects a reasonably frontal, already-detected face crop; performance on profile views, heavy occlusion or unusual lighting is untested.

Research and educational use.

Downloads last month
10
Safetensors
Model size
15.3M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support