QeyPoint Face - 106-Point Facial Landmark Detection
High-performance real-time facial landmark detection with 106 keypoints. Built with PyTorch and optimized for both CPU and GPU inference.
Features
- 106 facial landmarks for precise face tracking
- Real-time performance on both CPU and GPU
- Simple API with streaming and single-shot modes
- Automatic face detection using OpenCV Haar Cascades
- Temporal smoothing for stable video tracking
- 3.09M parameters - lightweight and efficient
Installation
pip install torch opencv-python numpy
For GPU support, install PyTorch with CUDA:
pip install torch --index-url https://download.pytorch.org/whl/cu118
Quick Start
Basic Usage
from QeyP import QeyPointDetector
# Initialize detector (auto-loads bundled model)
detector = QeyPointDetector()
# Detect from webcam (single shot)
result = detector.detect_single(camera_id=0, mirror=True)
if result:
landmarks = result["landmarks"] # (106, 2) numpy array [x, y]
confidence = result["confidence"] # (106,) confidence scores
face_box = result["face_box"] # [x, y, width, height]
print(f"Detected {len(landmarks)} landmarks")
Real-time Streaming
import cv2
from QeyP import QeyPointDetector
detector = QeyPointDetector()
for result in detector.stream_from_cam(camera_id=0, wait_seconds=0.01):
if result is None:
continue
frame = result["frame"]
landmarks = result["landmarks"]
face_box = result["face_box"]
# Draw face box
fx, fy, fw, fh = face_box
cv2.rectangle(frame, (fx, fy), (fx + fw, fy + fh), (255, 180, 0), 2)
# Draw landmarks
for (px, py) in landmarks:
cv2.circle(frame, (int(px), int(py)), 2, (0, 255, 0), -1)
cv2.imshow("QeyPoint Face", frame)
if cv2.waitKey(1) & 0xFF == ord('q'):
break
cv2.destroyAllWindows()
Detect from Image
import cv2
from QeyP import QeyPointDetector
detector = QeyPointDetector()
frame = cv2.imread("face.jpg")
result = detector.detect_from_frame(frame)
if result:
landmarks = result["landmarks"]
# Process landmarks...
API Reference
QeyPointDetector
Initialization:
QeyPointDetector(
checkpoint_path=None, # Auto-loads bundled model
device=None, # Auto-detect CUDA/CPU
min_face_size=80, # Minimum face size in pixels
face_padding_x=0.20, # Horizontal padding around face (20%)
face_padding_y=0.25 # Vertical padding around face (25%)
)
Methods:
detect_single(camera_id=0, mirror=True)
Capture one frame from camera and detect landmarks.
Parameters:
camera_id(int): Camera device ID (default: 0)mirror(bool): Flip frame horizontally (default: True)
Returns: Dict with landmarks, confidence, face_box or None if no face detected.
stream_from_cam(camera_id=0, wait_seconds=0.1, mirror=True, detect_every=5, smoothing=0.65)
Generator that yields landmark detections continuously.
Parameters:
camera_id(int): Camera device IDwait_seconds(float): Delay between frames (default: 0.1)mirror(bool): Flip frame horizontallydetect_every(int): Re-run face detection every N frames (default: 5)smoothing(float): Temporal smoothing factor 0-1 (default: 0.65)
Yields: Dict with landmarks, confidence, face_box, frame or None
detect_from_frame(frame)
Detect landmarks in a given frame (BGR format).
Parameters:
frame(numpy.ndarray): Input image in BGR format (OpenCV standard)
Returns: Dict with landmarks, confidence, face_box or None
Return Format
All detection methods return a dictionary:
{
"landmarks": np.ndarray, # Shape: (106, 2) - [x, y] pixel coordinates
"confidence": np.ndarray, # Shape: (106,) - confidence per landmark [0-1]
"face_box": np.ndarray, # Shape: (4,) - [x, y, width, height]
"frame": np.ndarray # Only in stream_from_cam - BGR image
}
Returns None if no face is detected in the frame.
Model Architecture
- Backbone: U-Net style encoder-decoder with residual blocks
- Input: 256×256 RGB images
- Output: 64×64 heatmaps (106 channels)
- Activation: GELU activations throughout
- Parameters: 3,090,634 total
- Inference: Spatial soft-argmax for sub-pixel accuracy
Performance
- CPU: ~15-30 FPS (depends on processor)
- GPU: ~60+ FPS on modern GPUs
- Model Size: ~12 MB
- Latency: <50ms per frame on GPU
Landmark Indices
The 106 landmarks follow a standard facial landmark convention:
- 0-32: Face contour
- 33-37: Right eyebrow
- 38-42: Left eyebrow
- 43-47: Nose bridge
- 48-53: Nose tip
- 54-58: Right eye
- 59-63: Left eye
- 64-81: Outer mouth
- 82-89: Inner mouth
- 90-105: Face interior points
License
This model is released under the MIT License.
Requirements
- Python 3.7+
- PyTorch 1.9+
- OpenCV 4.x
- NumPy
Known Issues
- Requires OpenCV 4.x (not compatible with OpenCV 5.x due to API changes)
- Face detection may struggle with extreme poses or occlusions
- Best performance with frontal or near-frontal faces