LibreGTRx-sem
GTR-X Cityscapes semantic segmentation weights (19 classes) converted for LibreYOLO.
GTR support is being prepared for LibreYOLO v1.6.0. Earlier PyPI releases may not include this model family.
Usage
With a LibreYOLO version that includes GTR semantic segmentation:
from libreyolo import LibreYOLO
model = LibreYOLO("LibreGTRx-sem.pt")
result = model.predict("street.jpg")[0]
mask = result.semantic_mask.data # (H, W) Cityscapes train IDs
Images are letterboxed into a 1024x2048 canvas and the network averages overlapping 1024px windows at a 768px stride, as upstream evaluates Cityscapes. A Cityscapes frame runs unresized in three windows.
The source release reports 83.6 mIoU on Cityscapes val with 1024px sliding
windows (32.2M parameters). LibreYOLO has not re-measured that number; it
is quoted from the source release.
Source
Official GTR implementation,
source revision 782e737efe2e6437ac537fbdcee089673d3376c1.
Published checkpoint,
weight repository revision 9fc62c8c2b2c976835d0f1c1ffc544dbc0f9e29f, SHA-256 ee4ade38e2e6e398110cbe909604566b221dbec87743d8f35e09ca2ca6093b55.
Copyright (c) 2026 Intellindust-AI-Lab. The source code is MIT licensed and the
publisher's weight repository explicitly declares MIT.
Modifications
Selected the EMA state dict and added LibreYOLO schema v1.0 metadata.
Learned parameters and state-dict keys are unchanged. Training/optimizer state
was removed. Conversion uses weights/convert_gtr_sem_weights.py in the
LibreYOLO source repository.
Validation
Strict loading was checked, and on CPU the converted weights match the pinned upstream graph exactly (max abs diff 0.0) for a single 1024px window and for upstream's sliding-window inference on a 1024x2048 input, using the portable attention operators on both sides. Cityscapes mIoU, CUDA parity and GPU latency were not re-measured.