TelecomGPT-R1-27B-FP8-Dynamic

FP8 Dynamic quantization of KU-DFI/TelecomGPT-R1.

Quantization

  • Scheme: FP8_DYNAMIC
  • Serialization: compressed-tensors
  • Target modules: Linear
  • lm_head: unquantized
  • Calibration dataset: none
  • Source precision: BF16

Intended use

Telecom reasoning, alarm analysis, root-cause analysis, protocol reasoning, and evaluation against operator-specific incident datasets.

A100 note

NVIDIA A100 is an Ampere GPU and does not provide native Hopper-style FP8 Tensor Core execution. The FP8 checkpoint still reduces model-weight memory, and vLLM can use its supported Ampere execution path when loading the model.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ukkathva/TelecomGPT-R1-27B-FP8-Dynamic

Quantized
(1)
this model