MultimodalAI — TREAT-MMTB 2026 Task 2
TB/Normal classification from chest X-ray PNGs for TREAT-MMTB 2026. Clinical metadata are not used.
Paper: Generalizable Tuberculosis Classification on Chest X-rays through Multi-Source Curation and Model Ensembling, MICCAI 2026 Workshops and Challenges (TREAT-MMTB).
Challenge result
MultimodalAI's official final external leaderboard result:
| Rank | External fullset F1 |
|---|---|
| 2 | 0.8642 |
Installation
Python 3.10, with PyTorch for CUDA 11.8 and dependencies from requirements.txt:
python -m pip install torch==2.4.1 --index-url https://download.pytorch.org/whl/cu118
python -m pip install -r requirements.txt
AutoModel inference
Code and weights load automatically from Hugging Face.
from transformers import AutoModel
model = AutoModel.from_pretrained(
"Deepnoid/TREAT-MMTB-2026-Task2-MultimodalAI",
trust_remote_code=True,
).to("cuda").eval()
result = model.predict("/path/to/image.png", output_dir=None)[0]
label = result["label"]
results = model.predict("/path/to/png_folder", output_dir="output/task2", batch_size=4)
- Input: a PNG file or a folder, read non-recursively with case-insensitive extensions.
- Return: always a list of results containing
filenameandlabel(TBorNormal).output_dir=None(default) writes no output. - Save: setting
output_dirwritesprediction.csvwith columnsfilename,TB/Normal.
Use .to("cpu") for CPU inference. The default batch size is 32; reduce it for smaller GPUs.
CLI inference
Place all four checkpoints in weights/ beside predict.py (or set --weights). Architecture and threshold settings are read from config.json (or set --config). --input accepts a file or folder.
python predict.py --input /path/to/png_folder --output output/task2 --batch-size 4
Models and method
weights/model_0.safetensors through weights/model_3.safetensors are DINOv3 ViT-L/16 classifiers with attention pooling.
Input combines histogram equalization, CLAHE, and grayscale channels, center-padded and resized to 512 × 512. Inference ensembles four models with horizontal-flip augmentation and a decision threshold of 0.35 for TB versus Normal.
Docker
Requires an NVIDIA GPU, driver, and NVIDIA Container Toolkit. Place all four checkpoints in weights/ beside the Dockerfile. Build with network access; inference runs offline.
docker build -t multimodalai-task2:latest .
mkdir -p output
docker run --rm --gpus all --network none \
-v /absolute/path/to/png_folder:/input:ro \
-v "$PWD/output":/output multimodalai-task2:latest
Append --input /input/image.png for a single file or --batch-size 4 for a smaller batch. Results are saved in output/prediction.csv.
Builds on
Vision-language pretraining followed GLINT (Park et al., 2026), using Meta AI / FAIR's DINOv3 encoders (Siméoni et al., 2025), MPNet sentence embeddings (Song et al., 2020), and report labels using Qwen3.6-35B-A3B.
| Dataset | Use |
|---|---|
| TREAT-MMTB 2026 | Classification training |
| MIMIC-CXR | Vision-language pretraining and classification training |
| TB Portals | Classification training |
| VinDr-CXR | Classification training |
| TBX11K | Classification training |
| PadChest | Classification training |
| Shenzhen and Montgomery | Classification training (two of four ensemble members) |
Acknowledgements
This work was supported by the Technology Innovation Program (RS-2025-02221011, Development of Medical-Specialized Multimodal Hyperscale Generative AI Technology for Global Integration) funded by the Ministry of Trade Industry & Energy (MOTIE, South Korea), and by the “Advanced GPU Utilization Support Program” funded by the Government of the Republic of Korea (Ministry of Science and ICT).
Data were obtained from the TB Portals, which is an open-access TB data resource supported by the National Institute of Allergy and Infectious Diseases (NIAID) Office of Cyber Infrastructure and Computational Biology (OCICB) in Bethesda, MD. These data were collected and submitted by members of the TB Portals Consortium. Investigators and other data contributors that originally submitted the data to the TB Portals did not participate in the design or analysis of this study (Rosenthal et al., 2017).
Citation
If you use this code or these weights, please cite:
@InProceedings{LeeSeo_Generalizable_MICCAISAT2026,
author = {Lee, Seongeun AND Yun, Hannah AND Jeong, Taejin AND Jeon, Mingyeong AND Park, Junhyun AND Kim, Hyunwoong AND Park, Jonggwon},
title = {{Generalizable Tuberculosis Classification on Chest X-rays through Multi-Source Curation and Model Ensembling}},
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17265}
}
- Downloads last month
- -