CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
✨ BMVC 2026 ✨
Summary
We introduce CheXGround, a region-grounded longitudinal chest X-ray language model that represents paired studies through corresponding anatomical regions. CheXGround extracts anatomical regions from current and prior radiographs, encodes them as temporally enhanced Region-of-Interest (ROI) tokens, and combines them with global temporal image context during generation. To connect these region tokens with clinical language representations, we propose Temporal Region–Phrase Alignment, a pretraining objective that aligns temporal anatomical representations with localized report phrases.
Intended use cases
CheXGround is intended for research and educational use, including:
- Studying grounded report generation and visual question answering for chest X-rays.
- Evaluating anatomical localization and reasoning about changes between prior and current studies.
- Developing and comparing multimodal learning methods on appropriately authorized research data.
CheXGround is not meant to be used for clinical practice.
Disclaimer
CheXGround is an experimental research system. Its outputs are not medical advice and must not be relied on for patient care. Although CheXGround shows competitive performance, subtle inaccuracies in anatomical localization and descriptions may still occur.
If you find this work useful, please cite us!
@misc{gebremedhin2026chexgroundanatomicalregiontokens,
title={CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation},
author={Adonay Demewez Gebremedhin and Wessam Shehieb and Sara Alansari and Mohamad Alansari and Muzammal Naseer and Sajid Javed and Naoufel Werghi},
year={2026},
eprint={2608.30758},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.30758},
}
- Downloads last month
- 32