YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

GGUF Python Reader Metadata Array DoS PoC

This repository contains benign GGUF metadata-only files that reproduce resource-exhaustion behavior in the official gguf Python package.

Summary

gguf.GGUFReader parses GGUF array metadata by recursing once per array element and appending one numpy view per element. A small GGUF file with a valid UINT8 metadata array therefore causes disproportionate CPU and memory use during model metadata loading. A truncated GGUF that declares an extremely large array length also makes the parser keep iterating without advancing because reads past EOF return empty numpy views.

No file here contains executable code, tensor weights, network callbacks, persistence, credential access, or destructive behavior. The marker string is GGUF_ARRAY_DOS_BENIGN_MARKER.

Tested Versions

  • Python 3.13.13 on Windows
  • gguf==0.19.0
  • numpy==2.4.4
  • llama.cpp source inspected at 78fbbc2c0788efc8857a2c0dc9802ec689fa12c1

Install command:

python -m venv .venv
.venv/Scripts/python -m pip install -r requirements.txt

On Linux/macOS, use .venv/bin/python instead of .venv/Scripts/python.

Reproduce

Generate fresh artifacts:

.venv/Scripts/python make_gguf_array_pocs.py --out artifacts --valid-count 50000

Run the trigger with a timeout-isolated child process:

.venv/Scripts/python trigger_gguf_array_dos.py artifacts/baseline_small_array.gguf artifacts/valid_uint8_array_50000.gguf artifacts/truncated_huge_uint8_array.gguf --timeout 8

Expected result:

  • baseline_small_array.gguf parses quickly.
  • valid_uint8_array_50000.gguf is only about 50 KB but creates 50,005 field parts and tens of MB of Python allocations.
  • truncated_huge_uint8_array.gguf is 140 bytes and times out rather than failing fast.

Included Evidence

array_dos_results_versioned.json:

  • Baseline 141-byte file: elapsed_sec=0.0009319999990111683, tracemalloc_peak_bytes=28813
  • Valid 50,131-byte file: elapsed_sec=1.5484974999999395, tracemalloc_peak_bytes=35335086
  • Truncated 140-byte file: timed out at 8 seconds

array_dos_results_with_memory.json:

  • Valid 250,131-byte file: elapsed_sec=12.744927899999311, tracemalloc_peak_bytes=176169251

Artifact Hashes

  • artifacts/baseline_small_array.gguf: ef66b68d7d9f9690c1c6c27380c5349f435af7ba1d7aebb0dc47cda33f69b12c
  • artifacts/valid_uint8_array_50000.gguf: ead20680b54f62ac1da0cea39fc5f0dc4e0a6ae5e7b0b6b07b22007b7b389e62
  • artifacts/valid_uint8_array_250000.gguf: 530cd494fdaf6485f5d29cc90e0bdd54328c7aa41d50dd16cb4c77ec2697e77e
  • artifacts/truncated_huge_uint8_array.gguf: fd98c113e14b2d0212043da98ee1c3cfb369823d99bc41e143fd14bda1390254

Vulnerable Code Path

In gguf/gguf_reader.py, _get_field_parts() reads alen from the file and then loops over every element:

for idx in range(alen[0]):
    curr_size, curr_parts, curr_idxs, curr_types = self._get_field_parts(offs, raw_itype[0])
    aparts += curr_parts
    data_idxs += (idx + idxs_offs for idx in curr_idxs)
    offs += curr_size

There is no file-size bound check or bulk scalar-array read. For truncated files, _get() returns empty arrays, curr_size stays 0, and the loop continues.

Upload Command

If authenticated:

huggingface-cli upload YOUR_USERNAME/gguf-python-array-dos-poc C:/Users/Pragnyan/dev/huntr-exp1/ggml/hf_ggml_poc . --repo-type model
Downloads last month
23
GGUF
Model size
0 params
Architecture
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support