MultimodalAI — TREAT-MMTB 2026 Task 2

Hugging Face GitHub Paper Leaderboard TREAT-MMTB 2026

TB/Normal classification from chest X-ray PNGs for TREAT-MMTB 2026. Clinical metadata are not used.

Paper: Generalizable Tuberculosis Classification on Chest X-rays through Multi-Source Curation and Model Ensembling, MICCAI 2026 Workshops and Challenges (TREAT-MMTB).

Challenge result

MultimodalAI's official final external leaderboard result:

Rank External fullset F1
2 0.8642

Installation

Python 3.10, with PyTorch for CUDA 11.8 and dependencies from requirements.txt:

python -m pip install torch==2.4.1 --index-url https://download.pytorch.org/whl/cu118
python -m pip install -r requirements.txt

AutoModel inference

Code and weights load automatically from Hugging Face.

from transformers import AutoModel

model = AutoModel.from_pretrained(
    "Deepnoid/TREAT-MMTB-2026-Task2-MultimodalAI",
    trust_remote_code=True,
).to("cuda").eval()

result = model.predict("/path/to/image.png", output_dir=None)[0]
label = result["label"]

results = model.predict("/path/to/png_folder", output_dir="output/task2", batch_size=4)
  • Input: a PNG file or a folder, read non-recursively with case-insensitive extensions.
  • Return: always a list of results containing filename and label (TB or Normal). output_dir=None (default) writes no output.
  • Save: setting output_dir writes prediction.csv with columns filename,TB/Normal.

Use .to("cpu") for CPU inference. The default batch size is 32; reduce it for smaller GPUs.

CLI inference

Place all four checkpoints in weights/ beside predict.py (or set --weights). Architecture and threshold settings are read from config.json (or set --config). --input accepts a file or folder.

python predict.py --input /path/to/png_folder --output output/task2 --batch-size 4

Models and method

weights/model_0.safetensors through weights/model_3.safetensors are DINOv3 ViT-L/16 classifiers with attention pooling.

Input combines histogram equalization, CLAHE, and grayscale channels, center-padded and resized to 512 × 512. Inference ensembles four models with horizontal-flip augmentation and a decision threshold of 0.35 for TB versus Normal.

Docker

Requires an NVIDIA GPU, driver, and NVIDIA Container Toolkit. Place all four checkpoints in weights/ beside the Dockerfile. Build with network access; inference runs offline.

docker build -t multimodalai-task2:latest .
mkdir -p output
docker run --rm --gpus all --network none \
  -v /absolute/path/to/png_folder:/input:ro \
  -v "$PWD/output":/output multimodalai-task2:latest

Append --input /input/image.png for a single file or --batch-size 4 for a smaller batch. Results are saved in output/prediction.csv.

Builds on

Vision-language pretraining followed GLINT (Park et al., 2026), using Meta AI / FAIR's DINOv3 encoders (Siméoni et al., 2025), MPNet sentence embeddings (Song et al., 2020), and report labels using Qwen3.6-35B-A3B.

Dataset Use
TREAT-MMTB 2026 Classification training
MIMIC-CXR Vision-language pretraining and classification training
TB Portals Classification training
VinDr-CXR Classification training
TBX11K Classification training
PadChest Classification training
Shenzhen and Montgomery Classification training (two of four ensemble members)

Acknowledgements

This work was supported by the Technology Innovation Program (RS-2025-02221011, Development of Medical-Specialized Multimodal Hyperscale Generative AI Technology for Global Integration) funded by the Ministry of Trade Industry & Energy (MOTIE, South Korea), and by the “Advanced GPU Utilization Support Program” funded by the Government of the Republic of Korea (Ministry of Science and ICT).

Data were obtained from the TB Portals, which is an open-access TB data resource supported by the National Institute of Allergy and Infectious Diseases (NIAID) Office of Cyber Infrastructure and Computational Biology (OCICB) in Bethesda, MD. These data were collected and submitted by members of the TB Portals Consortium. Investigators and other data contributors that originally submitted the data to the TB Portals did not participate in the design or analysis of this study (Rosenthal et al., 2017).

Citation

If you use this code or these weights, please cite:

@InProceedings{LeeSeo_Generalizable_MICCAISAT2026,
    author = {Lee, Seongeun AND Yun, Hannah AND Jeong, Taejin AND Jeon, Mingyeong AND Park, Junhyun AND Kim, Hyunwoong AND Park, Jonggwon},
    title = {{Generalizable Tuberculosis Classification on Chest X-rays through Multi-Source Curation and Model Ensembling}},
    booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
    year = {2026},
    publisher = {Springer Nature Switzerland},
    volume = {LNCS 17265}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support