EMG-GPT
Inference weights for EMG-GPT: Predictive Pretraining on Residual-Quantized EMG Tokens for Hand Pose Estimation.
Paper · Code · Inference guide
Ettore Magni · Rolandos Alexandros Potamias · Stefanos Zafeiriou · Konstantinos Barmpas
The selected GPT checkpoints were pretrained with CUDA on an NVIDIA GH200 120 GB GPU. The inference package accepts CPU or CUDA.
Checkpoints
| Directory | Task | GPT initialization | Selected pose checkpoint |
|---|---|---|---|
regression/ |
Pose estimation without an initial pose | Step 280,000 | Pose warm-start → full fine-tuning, step 2,000 |
tracking/ |
Pose estimation with a boundary pose per window | Step 400,000 | Pose warm-start → full fine-tuning, step 5,000 |
Each directory contains config.json, model.safetensors,
codebooks.safetensors and manifest.json: the adapted GPT backbone, pose head
and codebooks needed for inference. The loader verifies their hashes.
The shared tokenizer is downloaded separately from
ntinosbarmpas/NeuroRVQ at revision
d944b87f44ae0ba2923b2f10d0518f23f6803b76. The helper below downloads and verifies
pretrained_models/tokenizers/NeuroRVQ_EMG_tokenizer_v1.pt.
Regression quick start
Use Python 3.11 or newer (tested on 3.11–3.14). Clone the code repository
and install '.[download]' following its README, then run:
hf download ettoremagni/EMG-GPT \
--revision 812d159b4e7a4fb1c95da865f4f1e2635fa6522f \
--include "regression/*" --include "LICENSE" --local-dir weights
emg-gpt-download-tokenizer --output weights/NeuroRVQ_EMG_tokenizer_v1.pt
emg-gpt-predict \
--model-dir weights/regression \
--tokenizer weights/NeuroRVQ_EMG_tokenizer_v1.pt \
--input recording.npz --output prediction.npz --device cpu
Use --device cuda with a CUDA-enabled PyTorch installation for NVIDIA GPU
inference. MPS is unsupported. Use the current release from the code repository;
its guide includes a runnable synthetic smoke check.
Input NPZ files need raw 2-kHz emg ([samples,16]) in the native emg2pose amplitude
scale, scalar sampling_rate_hz=2000 and Unicode channel_names (c1 through
c16, in column order). A full window needs at least 13,119 samples. Predictions
are joint angles in radians, [windows,250,20] at 50 Hz, with timestamps and a
coverage mask. For Tracking, replace regression/* with tracking/*, keeping
LICENSE included. Supply one boundary pose per window; an all-NaN row explicitly
skips it, retaining timestamps with NaN angles and no coverage. The inference guide
covers HDF5 input, the Python API and alignment.
Scope and verification
This is offline inference: the frontend filters the entire recording and resamples to 1 kHz before tokenization. Outputs exclude warm-up, inter-window gaps and a trailing remainder. Other sensors or amplitude scales are unvalidated.
The code repository includes integration tests using these weights and synthetic reference outputs from the original research implementation. CPU compatibility has also been checked on real recordings. These checks are separate from the paper's benchmark; CUDA numerical parity remains unverified. See the inference guide for the test commands.
Citation
@misc{magni2026emggpt,
title={EMG-GPT: Predictive Pretraining on Residual-Quantized EMG Tokens for Hand Pose Estimation},
author={Ettore Magni and Rolandos Alexandros Potamias and Stefanos Zafeiriou and Konstantinos Barmpas},
year={2026},
eprint={2610.05235},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.05235}
}
License and attribution
EMG-GPT learned weights (regression/model.safetensors and
tracking/model.safetensors) are licensed under CC BY-NC-SA 4.0.
Attribute the EMG-GPT authors and paper cited above.
The NeuroRVQ tokenizer and the bundled codebooks.safetensors are separate
upstream assets and retain their
CC BY-NC 4.0 terms; the EMG-GPT grant does not
relicense them. The inference code
also remains CC BY-NC 4.0. See the
third-party notices
for provenance and component licenses.