poisedlabs/poseconv3d-pt

PoseConv3D 3D spatio-temporal heatmap CNN embedding model for physical therapy and exercise recognition. Supports both Apple Silicon MLX (native Metal GPU) and PyTorch.

This model maps human motion repetitions into a 64-dimensional unit hypersphere embedding, calibrated against a bank of 1595 physical therapy and fitness exercise prototypes.


Model Architecture

Component Specification
Architecture 3D Spatio-Temporal Heatmap CNN (PoseConv3DEncoder / MLXPoseConv3D)
Input Shape (PyTorch) (Batch, Channels=1, Frames=32, Height=64, Width=64)
Input Shape (MLX) (Batch, Frames=32, Height=64, Width=64, Channels=1)
Heatmap Sigma 1.8
Joints 17 canonical keypoints (COCO/MediaPipe/Apple Vision layout) in [-1, 1]
Output 64-dimensional unit vector ($|z|_2 = 1.0$)
Prototype Bank 1595 calibrated exercise prototypes

1. End-to-End Classification from Video

Classify any exercise repetition video (.mp4 / .mov):

import json
import numpy as np
import torch
from safetensors.torch import load_file
from model import PoseConv3DEncoder
from preprocess import resample_keypoints, rasterize_heatmaps

# 1. Load model & prototypes
model = PoseConv3DEncoder()
model.load_state_dict(load_file("model.safetensors"))
model.eval()

protos = load_file("prototypes.safetensors")["prototypes"].numpy()
with open("prototypes.names.json") as f:
    exercise_names = json.load(f)

# 2. Extract keypoints from video (e.g. via MediaPipe / Apple Vision)
# Raw repetition keypoints: shape (T, 17, 2) in [-1, 1] range
# (Replace with your pose extractor output)
rep_kps = np.random.uniform(-0.8, 0.8, size=(45, 17, 2)).astype(np.float32)

# 3. Resample to 32 frames & rasterize to 3D Gaussian heatmaps
kps_resampled = resample_keypoints(rep_kps, target_t=32)
volume = rasterize_heatmaps(kps_resampled, h=64, w=64, sigma=1.8, format="pytorch")
input_tensor = torch.from_numpy(volume).unsqueeze(0)  # (1, 1, 32, 64, 64)

# 4. Generate 64-d unit embedding
with torch.no_grad():
    emb = model(input_tensor).numpy()[0]  # (64,)

# 5. Cosine similarity against all 1595 prototypes
sims = emb @ protos.T
top5 = np.argsort(-sims)[:5]

print("Top 5 Predicted Exercises:")
for rank, idx in enumerate(top5, 1):
    print(f"{rank}. {exercise_names[idx]} (similarity: {sims[idx]:.4f})")

2. Apple Silicon MLX Quickstart (Native Metal GPU)

Run high-throughput inference on Apple Silicon Macs with zero PyTorch dependency:

import json
import numpy as np
import mlx.core as mx
from safetensors.numpy import load_file
from mlx_model import MLXPoseConv3D

# Load model and weights
model = MLXPoseConv3D()
model.load_weights("weights_mlx.safetensors")

# Load prototypes
protos = load_file("prototypes.safetensors")["prototypes"]
with open("prototypes.names.json") as f:
    exercise_names = json.load(f)

# Input shape: (Batch, Frames=32, Height=64, Width=64, Channels=1)
x = mx.random.normal((1, 32, 64, 64, 1))
emb = model(x)  # (1, 64)

# Cosine similarity against all 1595 exercises
sims = np.array(emb @ mx.array(protos).T)[0]
top5_idx = np.argsort(-sims)[:5]

print("Top 5 MLX Matches:")
for rank, idx in enumerate(top5_idx, 1):
    print(f"{rank}. {exercise_names[idx]}: {sims[idx]:.4f}")

3. Standard PyTorch Quickstart

import json
import torch
import numpy as np
from safetensors.torch import load_file
from model import PoseConv3DEncoder

model = PoseConv3DEncoder()
model.load_state_dict(load_file("model.safetensors"))
model.eval()

protos = load_file("prototypes.safetensors")["prototypes"].numpy()
with open("prototypes.names.json") as f:
    exercise_names = json.load(f)

# Input shape: (Batch, Channels=1, Frames=32, Height=64, Width=64)
x = torch.randn(1, 1, 32, 64, 64)
with torch.no_grad():
    emb = model(x).numpy()[0]

sims = emb @ protos.T
top5_idx = np.argsort(-sims)[:5]

for rank, idx in enumerate(top5_idx, 1):
    print(f"{rank}. {exercise_names[idx]}: {sims[idx]:.4f}")

Output Interpretation

  • Embedding: A 64-dimensional unit vector ($|z|_2 = 1.0$) representing the spatio-temporal kinematic volume of the repetition.
  • Cosine Similarity: The dot product $z \cdot p^\top$ with reference exercise prototypes, bounded between [-1.0, +1.0].
    • $> 0.70$: Strong kinematic and posture alignment with reference demo.
    • $0.50$ – $0.70$: Kinematically related movement (e.g. squat variation vs lunge).
    • $< 0.40$: Unrelated movement pattern.

Citation & License

  • License: Apache 2.0
  • Organization: PoisedLabs
Downloads last month
22
Safetensors
Model size
738k params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support