Sensor-Language-Action Models
🔥 News
- [2026-10-06] Our paper is available on arXiv.
- [2026-10-06] Code released on GitHub, and model released on HuggingFace.
- [2026-10-06] Project website is live!
📖 Introduction
OpenSLA is a new family of models - Sensor-Langauge-Action (SLA) models. It takes a multi-channel sensor history together with language context and produces a structured action prediction and a sensor-state caption. We release two variants:
- OpenSLA-B: sensor tokens are projected and prepended to the language prompt as a flat token prefix.
- OpenSLA-H: builds on B with a hierarchical sensor encoder that compresses the signal into local, per-channel, and global memory tokens.
📖 Table of Contents
💿 Installation
git clone https://github.com/yang-ai-lab/Sensor-Language-Action-Model.git
cd Sensor-Language-Action-Model
pip install -r requirements.txt
Dependencies
- Python >= 3.10
- PyTorch >= 2.5
- Transformers >= 5.8.1 (for the Qwen3.5 backbone)
- PEFT >= 0.12
🚀 Quick Start
demo.ipynb loads a checkpoint, builds an input batch, and runs prediction.
The same flow in Python:
import torch
from opensla import OpenSLA, SensorBatch
model = OpenSLA.from_checkpoint(
"pretrained_weights/opensla_h_clinical.pt", domain="clinical",
dino_checkpoint="pretrained_weights/waveform_encoder.ckpt",
numeric_config="pretrained_weights/numeric_config.json",
action_group_types="pretrained_weights/action_group_types.json",
)
batch = torch.load("data/preprocessed_batch.pt", weights_only=True)
predictions = model.predict(batch["input_text"], SensorBatch(**batch["sensors"]))
📦 Pretrained Weights
| Model | Domain | Files |
|---|---|---|
| OpenSLA-H | clinical | opensla_h_clinical.pt, waveform_encoder.ckpt, numeric_config.json, action_group_types.json |
Download the files from this repository and put them under pretrained_weights/
in the code repository.
👩💻 Usage
Input Format
The model takes preprocessed signals as a .pt dictionary with domain,
input_text (one text context per sample), and sensors. The sensors are:
- Waveform:
waveformof shape[B, M, C, T], withMone-minute slots,Cphysical channels, andT = 7500samples per slot (60 s at 125 Hz), plus a booleanwaveform_channel_mask([B, M, C]) marking observed channel-minutes and int64waveform_modality_ids([B, C]) indexingopensla.WAVEFORM_MODALITIES[domain]. - Numeric: observed events as parallel
[B, E]tensors (values,rel_time_minin minutes before the decision,measure_ids,source_ids,event_mask), plus per-measure summary features (summary_features,summary_measure_ids).
Command line
Point configs/template.json at your weights and input batch, then:
opensla --config configs/template.json --dry-run # print the resolved config
opensla --config configs/template.json # write outputs/predictions.jsonl
On Slurm:
sbatch --account=YOUR_ACCOUNT --partition=YOUR_PARTITION scripts/run_template.sbatch --config configs/template.json
📊 Datasets
The models are trained and evaluated on six datasets from three healthcare settings, with an additional MIMIC-IV held out as an external clinical cohort. All of them are publicly available and can be accessed through the following links.
| Dataset | Domain | Sensors | Source |
|---|---|---|---|
| MC-MED | Clinical | ECG, plethysmography, respiration, arterial pressure; vital signs, labs, ventilator and other charted measurements | PhysioNet |
| MIMIC-III | Clinical | same as MC-MED | PhysioNet |
| MIMIC-IV | Clinical, external evaluation only | waveform-linked subset | PhysioNet, waveforms |
| MOVER | Operating room | ECG, plethysmography, arterial/central venous pressure, capnography, airway pressure, EEG; vital signs, hemodynamics, ventilation gases, labs | UCI MOVER |
| VitalDB | Operating room | same as MOVER | vitaldb.net, PhysioNet |
| MetaboNet | CGM | glucose trace; basal insulin delivery | metabo-net.org |
| PEDAP | CGM | glucose trace; basal insulin delivery | Jaeb Center |
📝 Citation
If you use this code or models in your research, please cite our paper:
@misc{xu2026opensla,
title = {Sensor-Language-Action Models},
author = {Xu, Yuekai and Shuai, Zitao and Yang, Yuzhe},
journal = {arXiv preprint},
year = {2026}
}
Acknowledgments
The waveform encoder adapts from the sensor encoder from OSF (MIT License).