QeyPoint Face - 106-Point Facial Landmark Detection

High-performance real-time facial landmark detection with 106 keypoints. Built with PyTorch and optimized for both CPU and GPU inference.

Features

  • 106 facial landmarks for precise face tracking
  • Real-time performance on both CPU and GPU
  • Simple API with streaming and single-shot modes
  • Automatic face detection using OpenCV Haar Cascades
  • Temporal smoothing for stable video tracking
  • 3.09M parameters - lightweight and efficient

Installation

pip install torch opencv-python numpy

For GPU support, install PyTorch with CUDA:

pip install torch --index-url https://download.pytorch.org/whl/cu118

Quick Start

Basic Usage

from QeyP import QeyPointDetector

# Initialize detector (auto-loads bundled model)
detector = QeyPointDetector()

# Detect from webcam (single shot)
result = detector.detect_single(camera_id=0, mirror=True)

if result:
    landmarks = result["landmarks"]      # (106, 2) numpy array [x, y]
    confidence = result["confidence"]    # (106,) confidence scores
    face_box = result["face_box"]        # [x, y, width, height]
    
    print(f"Detected {len(landmarks)} landmarks")

Real-time Streaming

import cv2
from QeyP import QeyPointDetector

detector = QeyPointDetector()

for result in detector.stream_from_cam(camera_id=0, wait_seconds=0.01):
    if result is None:
        continue
    
    frame = result["frame"]
    landmarks = result["landmarks"]
    face_box = result["face_box"]
    
    # Draw face box
    fx, fy, fw, fh = face_box
    cv2.rectangle(frame, (fx, fy), (fx + fw, fy + fh), (255, 180, 0), 2)
    
    # Draw landmarks
    for (px, py) in landmarks:
        cv2.circle(frame, (int(px), int(py)), 2, (0, 255, 0), -1)
    
    cv2.imshow("QeyPoint Face", frame)
    if cv2.waitKey(1) & 0xFF == ord('q'):
        break

cv2.destroyAllWindows()

Detect from Image

import cv2
from QeyP import QeyPointDetector

detector = QeyPointDetector()
frame = cv2.imread("face.jpg")
result = detector.detect_from_frame(frame)

if result:
    landmarks = result["landmarks"]
    # Process landmarks...

API Reference

QeyPointDetector

Initialization:

QeyPointDetector(
    checkpoint_path=None,      # Auto-loads bundled model
    device=None,               # Auto-detect CUDA/CPU
    min_face_size=80,          # Minimum face size in pixels
    face_padding_x=0.20,       # Horizontal padding around face (20%)
    face_padding_y=0.25        # Vertical padding around face (25%)
)

Methods:

detect_single(camera_id=0, mirror=True)

Capture one frame from camera and detect landmarks.

Parameters:

  • camera_id (int): Camera device ID (default: 0)
  • mirror (bool): Flip frame horizontally (default: True)

Returns: Dict with landmarks, confidence, face_box or None if no face detected.


stream_from_cam(camera_id=0, wait_seconds=0.1, mirror=True, detect_every=5, smoothing=0.65)

Generator that yields landmark detections continuously.

Parameters:

  • camera_id (int): Camera device ID
  • wait_seconds (float): Delay between frames (default: 0.1)
  • mirror (bool): Flip frame horizontally
  • detect_every (int): Re-run face detection every N frames (default: 5)
  • smoothing (float): Temporal smoothing factor 0-1 (default: 0.65)

Yields: Dict with landmarks, confidence, face_box, frame or None


detect_from_frame(frame)

Detect landmarks in a given frame (BGR format).

Parameters:

  • frame (numpy.ndarray): Input image in BGR format (OpenCV standard)

Returns: Dict with landmarks, confidence, face_box or None


Return Format

All detection methods return a dictionary:

{
    "landmarks": np.ndarray,    # Shape: (106, 2) - [x, y] pixel coordinates
    "confidence": np.ndarray,   # Shape: (106,) - confidence per landmark [0-1]
    "face_box": np.ndarray,     # Shape: (4,) - [x, y, width, height]
    "frame": np.ndarray         # Only in stream_from_cam - BGR image
}

Returns None if no face is detected in the frame.

Model Architecture

  • Backbone: U-Net style encoder-decoder with residual blocks
  • Input: 256×256 RGB images
  • Output: 64×64 heatmaps (106 channels)
  • Activation: GELU activations throughout
  • Parameters: 3,090,634 total
  • Inference: Spatial soft-argmax for sub-pixel accuracy

Performance

  • CPU: ~15-30 FPS (depends on processor)
  • GPU: ~60+ FPS on modern GPUs
  • Model Size: ~12 MB
  • Latency: <50ms per frame on GPU

Landmark Indices

The 106 landmarks follow a standard facial landmark convention:

  • 0-32: Face contour
  • 33-37: Right eyebrow
  • 38-42: Left eyebrow
  • 43-47: Nose bridge
  • 48-53: Nose tip
  • 54-58: Right eye
  • 59-63: Left eye
  • 64-81: Outer mouth
  • 82-89: Inner mouth
  • 90-105: Face interior points

License

This model is released under the MIT License.

Requirements

  • Python 3.7+
  • PyTorch 1.9+
  • OpenCV 4.x
  • NumPy

Known Issues

  • Requires OpenCV 4.x (not compatible with OpenCV 5.x due to API changes)
  • Face detection may struggle with extreme poses or occlusions
  • Best performance with frontal or near-frontal faces
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support