Dask Warns That Npy Stack Metadata Still Loads Through Pickle


Dask is the Python array library batch jobs use when a NumPy workload does not fit on one machine. On 29 September 2026, main took a single commit that adds a security warning to to_npy_stack and from_npy_stack. The info file beside the .npy blocks is a pickle stream, so an untrusted stack executes code in the process that loads it.

Dhruv Pratap Singh’s commit adds a Sphinx warning to both docstrings in dask/array/core.py. The diff is 10 insertions and 0 deletions. Issue 12578 asked for the note. Pull 12608 merged it. The patch adds the warning text alone.

The writer warning says the format uses pickle to serialize array metadata, and that you should only load stacks you trust. The reader warning is sharper. from_npy_stack uses pickle to load that metadata. The module is not secure against erroneous or maliciously constructed data. The text says never to load a stack that could have come from an untrusted source, or that could have been tampered with.

Both function bodies match the previous revision. pickle.dump writes the metadata. pickle.load reads it. A pipeline exposed before this merge is exposed after it. Both helpers are exported on dask.array. The creation page lists from_npy_stack with the other loaders and to_npy_stack with the store helpers.

to_npy_stack(dirname, x, axis=0) rechunks the array first. The chosen axis keeps its chunk sizes. Every other axis becomes one chunk spanning the full length. It then writes one .npy file per block along that axis, named 0.npy, 1.npy, and so on. The docstring example is a shape (5, 10, 10) array with chunks (2, 4, 4) stacked on axis 0. Those files hold x[0:2], x[2:4], and x[4:5], each written with numpy.save.

A sibling file named info has no extension. The caller writes it before any block is saved.

meta = {"chunks": chunks, "dtype": x.dtype, "axis": axis}
with open(os.path.join(dirname, "info"), "wb") as f:
    pickle.dump(meta, f)

chunks is a tuple of tuples of ints. axis is an int. dtype is a NumPy dtype object. Those values rebuild a Dask array without opening every block. Chunks and axis would fit in a text file. The dtype object is why pickle was the short write path, and why the file behaves like code.

After the dump, to_npy_stack builds a graph of numpy.save tasks and runs it through compute_as_if_collection, which blocks until the saves finish. info reaches disk first. A killed write leaves a readable info and a short set of blocks. The next from_npy_stack unpickles that metadata, then fails when a block path is missing. Write the directory under a temporary name and rename it into place only after every block exists.

open and os.mkdir use the local filesystem. A cloud URL works only after some other layer has mounted it as a path. On the distributed scheduler the workers need that same path. The warning is about who else can write the directory.

from_npy_stack opens dirname/info and unpickles it before it builds the array.

with open(os.path.join(dirname, "info"), "rb") as f:
    info = pickle.load(f)

The call runs in the invoking process and finishes before compute. It stays in the caller. A notebook kernel or a batch driver that maps the helper across incoming directories runs the stream locally, with that process’s privileges. This step finishes even when the workers cannot read the .npy files.

pickle.load runs __reduce__ callables while it rebuilds the object. Lookups of info["dtype"], info["chunks"], and info["axis"] happen only after those callables return. Any pickle payload is enough. Rejecting a suspicious dict afterward misses the code that already ran. Refuse the file before pickle.load, or skip this function.

Block names come from range(len(chunks[axis])). The loader asks for 0.npy, 1.npy, and the rest. Paths are generated integers, so the pickle cannot aim numpy.load at some other file. Code execution is already available from the stream.

The commit subject calls this remote code execution. The code runs in the local interpreter, on whatever bytes sit at dirname/info. On a shared mount, or a directory synced from object storage, the writer and the loader are different identities. Replacing info and leaving the .npy blocks in place is enough.

Each numeric block is a separate task.

(np.load, os.path.join(dirname, f"{i}.npy"), mmap_mode)

Dask calls numpy.load with the path and the memory map mode. allow_pickle stays at the NumPy default. The array extra in pyproject.toml asks for numpy >= 1.24. On that line numpy.load defaults allow_pickle to false, the default since NumPy 1.16.3. An object array .npy, or a pickle renamed to look like one, fails at load time. from_npy_stack leaves allow_pickle unset.

Default mmap_mode is "r". The docstring also allows None. "r" memory maps the block when the task runs. None reads it into memory. Either mode runs later, on the worker or thread that executes the task, after the caller has finished pickle.load on info.

numpy.load on a block follows that NumPy default. from_npy_stack on the parent directory runs pickle.load on info unconditionally. The test test_to_npy_stack in dask/array/tests/test_array_core.py writes a stack, compares one block, and round trips values. It covers the happy path. Deleting the warning leaves CI green. Leaving pickle.load in place does too.

A real fix replaces pickle for this format. Safe contents are the chunks, a dtype string, and the axis, stored as JSON, then loaded with numpy.dtype on the string. Watch from_npy_stack for a change to that pickle.load call. Another docstring edit leaves the same load in place.

If you own the writer and the reader, keep metadata away from pickle.load. A Zarr store does that. So does a directory of .npy blocks plus a JSON index you write yourself. When a job has to keep to_npy_stack, point from_npy_stack only at a directory that same job created and that only that job can write. A replaced info file is a replaced Python module.

Other modules call pickle and cloudpickle. This commit touches two docstrings in one file. The npy stack case stands out because the bytes live in a directory that looks like a dataset. The warning names the bug. The load path is the same as before.