YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
NumPy NPY header size validation is enforced after reading the declared header
Summary
NumPy's NPY/NPZ parser exposes a denial-of-service issue in numpy/lib/_format_impl.py because the documented max_header_size safety limit is checked after the full declared header is read from the file.
In _read_array_header(), NumPy parses header_length from the file and immediately calls _read_bytes(fp, header_length, "array header"). Only after that read completes does it enforce:
if len(header) > max_header_size:
raise ValueError(...)
This means a malicious .npy file can declare an extremely large header length and force excessive memory allocation / memory consumption before NumPy rejects the file. The same issue also affects .npz files because each member is parsed as an embedded .npy payload through the same code path.
Affected code
File: numpy/lib/_format_impl.py
Function: _read_array_header()
Relevant logic:
header_length = struct.unpack(...)
header = _read_bytes(fp, header_length, "array header")
if len(header) > max_header_size:
raise ValueError(...)
Why this matters
max_header_size is not just a parsing preference; it is a documented security control intended to limit resource use when loading untrusted array files. However, the current validation order enforces that limit too late.
As a result, the parser still performs the expensive read/allocation for the attacker-controlled header_length before applying the protection.
Proof of Concept
Verified with a crafted NPY v2.0 file declaring a header_length of 200_000_000 bytes.
import struct
magic = b'\x93NUMPY'
version = struct.pack('<BB', 2, 0)
header_length = struct.pack('<I', 200_000_000)
When loaded with np.load(), NumPy eventually rejects the file due to max_header_size, but only after consuming substantially more memory during the read. In the verified reproduction, a 200 MB declared header caused process VmPeak to increase by approximately 383 MB before rejection.
A standalone reproducer is included as poc.py in this submission directory.
Impact
This is a denial-of-service issue. An attacker who can supply a malicious .npy or .npz file to software that loads NumPy arrays from untrusted input can trigger excessive memory usage before the documented header-size safeguard aborts parsing.
This can cause process instability, memory pressure, or termination in environments that process attacker-controlled NumPy files.
Affected formats
.npy.npz(because members are parsed as.npyfiles using the same header reader)
Distinction from known CVEs
I checked the following NumPy CVEs and they do not describe this specific issue:
CVE-2019-6446— pickle / object deserialization behaviorCVE-2017-12852CVE-2021-41496
This report is specifically about header size validation order: the resource-limit check exists, but it is applied only after the untrusted length has already been honored.
Suggested fix
Validate header_length against max_header_size before calling _read_bytes().
Conceptually:
header_length = struct.unpack(...)
if header_length > max_header_size:
raise ValueError(...)
header = _read_bytes(fp, header_length, "array header")
That preserves the intended security behavior of max_header_size and prevents attacker-controlled oversized header declarations from causing unnecessary memory allocation before rejection.