ACR-instance-segmentation-RF(FP16)-v1.0.0

ACR instance segmentation용 cmes_RF_960_v1.0.0 RF-DETR Seg2XLarge λͺ¨λΈμ„ RTX 5070κ³Ό RTX 5080μ—μ„œ 각각 λΉŒλ“œν•œ TensorRT FP16 배포 νŒ¨ν‚€μ§€μž…λ‹ˆλ‹€.

Hugging Face μ €μž₯μ†Œ IDλŠ” κ΄„ν˜Έλ₯Ό ν—ˆμš©ν•˜μ§€ μ•ŠμœΌλ―€λ‘œ μ €μž₯μ†Œ 이름은 cmes-deepvision/ACR-instance-segmentation-RF-FP16-v1.0.0을 μ‚¬μš©ν•©λ‹ˆλ‹€.

μ—”μ§„ 선택

λŒ€μƒ GPU μ—”μ§„ 파일 크기 SHA256
RTX 5070 engines/cmes_RF_960_v1.0.0.fp16.RTX5070.trt 81,984,724 bytes 456d4855706691f434eed48ab8950640f6be200ab94d8b74d62183f269911c92
RTX 5080 engines/cmes_RF_960_v1.0.0.fp16.RTX5080.trt 82,088,348 bytes 667d2bc084092501f5e5bd1391cff76aaa640b24e28b2aaa18effd536498c577

TensorRT serialized engine은 GPU와 TensorRT 버전에 μ’…μ†λ©λ‹ˆλ‹€. RTX 5070μ—μ„œλŠ” RTX 5070 엔진을, RTX 5080μ—μ„œλŠ” RTX 5080 엔진을 μ‚¬μš©ν•˜μ‹­μ‹œμ˜€. λ‹€λ₯Έ GPU둜 엔진을 볡사해 μ‚¬μš©ν•˜λŠ” 것은 μ§€μ›ν•˜μ§€ μ•ŠμŠ΅λ‹ˆλ‹€.

검증 ν™˜κ²½

  • NVIDIA driver: 580.173.02
  • CUDA runtime used by PyTorch: 12.8
  • PyTorch: 2.11.0+cu128
  • TensorRT: 10.16.1.11
  • compute capability: 12.0
  • μž…λ ₯: RGB, batch 1, 1 x 3 x 960 x 960
  • κΈ°λ³Έ confidence threshold: 0.50

μ—”μ§„ λ‚΄λΆ€ 계산은 FP16이고 μž…μΆœλ ₯ binding은 FP32μž…λ‹ˆλ‹€.

μ„€μΉ˜

python -m pip install -r requirements.txt

λ‹€μš΄λ‘œλ“œ

from huggingface_hub import hf_hub_download

engine_path = hf_hub_download(
    repo_id="cmes-deepvision/ACR-instance-segmentation-RF-FP16-v1.0.0",
    filename="engines/cmes_RF_960_v1.0.0.fp16.RTX5080.trt",  # GPU에 맞게 λ³€κ²½
)

μΆ”λ‘ 

μ €μž₯μ†Œ 전체λ₯Ό 받은 λ’€ λ‹€μŒκ³Ό 같이 μ‹€ν–‰ν•©λ‹ˆλ‹€. --engine을 μƒλž΅ν•˜λ©΄ ν˜„μž¬ GPU μ΄λ¦„μ—μ„œ RTX 5070/5080 엔진을 μžλ™μœΌλ‘œ μ„ νƒν•©λ‹ˆλ‹€.

python scripts/infer_tensorrt.py image.jpg --threshold 0.50

엔진을 λͺ…μ‹œν•˜λ €λ©΄:

python scripts/infer_tensorrt.py image.jpg \
  --engine engines/cmes_RF_960_v1.0.0.fp16.RTX5070.trt \
  --output-dir outputs

좜λ ₯ ν΄λ”μ—λŠ” mask/bbox overlay JPEG와 detection JSON이 μƒμ„±λ©λ‹ˆλ‹€.

FP16 λ³€ν™˜ μ•ˆμ •μ„±

Test_v2.2.1의 50개 νƒœμŠ€ν¬μ—μ„œ κ³ μ • λŒ€ν‘œ 이미지 1μž₯μ”© 총 50μž₯을 μ‚¬μš©ν–ˆμŠ΅λ‹ˆλ‹€. confidence 0.50, 클래슀 일치 Hungarian matching, bbox/mask IoU >= 0.50 μ‘°κ±΄μ—μ„œ PyTorch FP16 결과와 TensorRT FP16 κ²°κ³Όλ₯Ό λΉ„κ΅ν–ˆμŠ΅λ‹ˆλ‹€.

GPU BBox Precision BBox Recall Mask Precision Mask Recall νŒμ •
RTX 5080 98.33% 98.88% 98.44% 98.99% PASS
RTX 5070 98.33% 98.66% 98.44% 98.77% PASS

톡과 기쀀은 λ„€ μ§€ν‘œκ°€ λͺ¨λ‘ 98% 이상인 κ²½μš°μž…λ‹ˆλ‹€. 이번 λŒ€ν‘œ ν‘œλ³Έμ—μ„œλŠ” FP16 λ³€ν™˜μœΌλ‘œ μΈν•œ μœ μ˜λ―Έν•œ 정확도 μ €ν•˜κ°€ ν™•μΈλ˜μ§€ μ•Šμ•˜μŠ΅λ‹ˆλ‹€.

속도

warm-up 5μž₯ 이후 50μž₯ x 3회, batch 1둜 μΈ‘μ •ν–ˆμŠ΅λ‹ˆλ‹€. E2EλŠ” λ©”λͺ¨λ¦¬ λ‚΄ PIL μ΄λ―Έμ§€μ˜ μ „μ²˜λ¦¬, μΆ”λ‘ , segmentation ν›„μ²˜λ¦¬μ™€ CPU NumPy λ³€ν™˜μ„ ν¬ν•¨ν•˜λ©° λ””μŠ€ν¬ I/O와 λͺ¨λΈ/μ—”μ§„ λ‘œλ“œλŠ” μ œμ™Έν•©λ‹ˆλ‹€.

GPU PyTorch FP16 E2E TensorRT FP16 E2E TensorRT FPS CT κ°μ†Œ μ—”μ§„ 단독
RTX 5080 28.48 ms 16.87 ms 59.28 40.77% 6.63 ms / 150.72 FPS
RTX 5070 40.68 ms 24.72 ms 40.45 39.23% 10.78 ms / 92.74 FPS

원본 μΈ‘μ •μΉ˜λŠ” benchmarks/에 ν¬ν•¨λ˜μ–΄ μžˆμŠ΅λ‹ˆλ‹€.

클래슀 μˆœμ„œ

0  dropping item
1  dumping item
2  item in pb bag
3  item in tote
4  item on buffer
5  item on floor
6  item on plate
7  item out of buffer
8  item out of tote
9  picked item
10 tote

평가 λ²”μœ„

이 μ €μž₯μ†Œμ˜ GPU별 검증은 λ³€ν™˜ μ•ˆμ •μ„±μ„ ν™•μΈν•˜κΈ° μœ„ν•œ 50μž₯ λŒ€ν‘œ ν‘œλ³Έμ˜ λ°±μ—”λ“œ μΌμΉ˜λ„ ν‰κ°€μž…λ‹ˆλ‹€. λͺ¨λΈ 자체의 전체 6,475μž₯ GT 기반 μ„±λŠ₯은 ACR-instance-segmentation-RF-v1.0.0의 평가 κ²°κ³Όλ₯Ό μ°Έκ³ ν•˜μ‹­μ‹œμ˜€. λ³Έ μ €μž₯μ†Œμ˜ μΌμΉ˜λ„ 수치λ₯Ό 전체 ν…ŒμŠ€νŠΈμ…‹ mAP둜 ν•΄μ„ν•˜λ©΄ μ•ˆ λ©λ‹ˆλ‹€.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including cmes-deepvision/ACR-instance-segmentation-RF-FP16-v1.0.0