You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

NumPy NPY header size validation is enforced after reading the declared header

Summary

NumPy's NPY/NPZ parser exposes a denial-of-service issue in numpy/lib/_format_impl.py because the documented max_header_size safety limit is checked after the full declared header is read from the file.

In _read_array_header(), NumPy parses header_length from the file and immediately calls _read_bytes(fp, header_length, "array header"). Only after that read completes does it enforce:

if len(header) > max_header_size:
    raise ValueError(...)

This means a malicious .npy file can declare an extremely large header length and force excessive memory allocation / memory consumption before NumPy rejects the file. The same issue also affects .npz files because each member is parsed as an embedded .npy payload through the same code path.

Affected code

File: numpy/lib/_format_impl.py

Function: _read_array_header()

Relevant logic:

header_length = struct.unpack(...)
header = _read_bytes(fp, header_length, "array header")
if len(header) > max_header_size:
    raise ValueError(...)

Why this matters

max_header_size is not just a parsing preference; it is a documented security control intended to limit resource use when loading untrusted array files. However, the current validation order enforces that limit too late.

As a result, the parser still performs the expensive read/allocation for the attacker-controlled header_length before applying the protection.

Proof of Concept

Verified with a crafted NPY v2.0 file declaring a header_length of 200_000_000 bytes.

import struct
magic = b'\x93NUMPY'
version = struct.pack('<BB', 2, 0)
header_length = struct.pack('<I', 200_000_000)

When loaded with np.load(), NumPy eventually rejects the file due to max_header_size, but only after consuming substantially more memory during the read. In the verified reproduction, a 200 MB declared header caused process VmPeak to increase by approximately 383 MB before rejection.

A standalone reproducer is included as poc.py in this submission directory.

Impact

This is a denial-of-service issue. An attacker who can supply a malicious .npy or .npz file to software that loads NumPy arrays from untrusted input can trigger excessive memory usage before the documented header-size safeguard aborts parsing.

This can cause process instability, memory pressure, or termination in environments that process attacker-controlled NumPy files.

Affected formats

  • .npy
  • .npz (because members are parsed as .npy files using the same header reader)

Distinction from known CVEs

I checked the following NumPy CVEs and they do not describe this specific issue:

  • CVE-2019-6446 — pickle / object deserialization behavior
  • CVE-2017-12852
  • CVE-2021-41496

This report is specifically about header size validation order: the resource-limit check exists, but it is applied only after the untrusted length has already been honored.

Suggested fix

Validate header_length against max_header_size before calling _read_bytes().

Conceptually:

header_length = struct.unpack(...)
if header_length > max_header_size:
    raise ValueError(...)
header = _read_bytes(fp, header_length, "array header")

That preserves the intended security behavior of max_header_size and prevents attacker-controlled oversized header declarations from causing unnecessary memory allocation before rejection.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support