HRNet: Optimized for AMD ROCm

HRNet (High-Resolution Network, W48 variant) is a human pose-estimation model that predicts COCO body keypoints while maintaining high-resolution feature representations throughout the network. This repository packages inference for human pose estimation using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs, CPUs, and NPUs.

This is based on the implementation of HRNet found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the HRNet AMD scripts to reproduce results or export with custom configurations.


Task Overview

Task: Human pose estimation (keypoint detection)

Dataset: COCO 2017 keypoints (val2017 images + person-keypoint annotations, under dataset/coco/)

Output metrics: Throughput (inferences/sec), latency (mean/P95/P99 ms), per-operator profiling breakdown

NPU note: The NPU (VitisAI) backend accepts FP32 input and auto-quantizes internally; there is no separate FP16/BF16/INT8 NPU path.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:

  • Validated backends: ONNX Runtime across CPU, GPU (MIGraphX), and NPU (VitisAI) execution providers.
  • CPU workflows run on any machine; GPU requires ROCm and a compatible AMD GPU; NPU requires an AMD Ryzen AI device.
Runtime Precision Backend Hardware Notes
ONNX Runtime FP32 / FP16 / BF16 / INT8 CPU Execution Provider AMD CPU
ONNX Runtime FP32 / FP16 / BF16 / INT8 MIGraphX Execution Provider AMD Instinct™ / Radeon™ GPU (ROCm)
ONNX Runtime FP32 VitisAI Execution Provider AMD Ryzen AI NPU Accepts FP32 input; VitisAI quantizes internally

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the HRNet on GitHub.


Model Details

Model Type: Human pose estimation (keypoint detection), HRNet-W48

Base Model: pose_hrnet_w48_384x288.pth (HRNet W48, 384×288 input resolution)

Model Stats:

  • Model variant: W48, 384×288 input resolution
  • Precision tested: FP32, FP16, BF16, INT8 (CPU/GPU); FP32 auto-quantized (NPU)

Accuracy Pipeline

Accuracy evaluation is not yet implemented for this model.


Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Benchmark and profiling scripts for CPU, GPU, and NPU
  • Instructions for downloading pretrained weights and COCO keypoint annotations
  • The upstream HRNet repository clone and native NMS extension build steps
  • Manual PyTorch/ONNX accuracy validation scripts (COCO AP/AR via pycocotools)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support