Conversation
…, clamps and metadata Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0, in which case the dimension has zero chunks (ceildiv(0, size) == 0). Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, zarr-developers#972, zarr-developers#1977, zarr-developers#2434, zarr-developers#3711, zarr-developers#4305, zarr-developers#4307, zarr-developers#4328) because the layers disagreed on this invariant and every span-derived chunk spelling clamped on its own: - The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1, but the in-memory FixedDimension allowed size == 0 with four special-case branches left over from zarr-developers#2434, so normalization could build a grid the metadata constructor then rejected. FixedDimension now rejects size < 1 and the four `if self.size == 0` branches are gone. VaryingDimension already required edges > 0 and is unchanged. - `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both the typesize == 0 early return and the np.maximum line) and `shards="auto"` each derived "one chunk covering the axis" independently. They now all go through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is the single definition of that phrase for a possibly zero-length axis. - Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]` document opened fine and read uninitialised memory after a resize. It now raises a clear ValueError at parse time, matching the format 3 grid. - Rectilinear grids had no creation-time spelling for a zero-length axis: normalize_chunks_1d required sum(edges) == span, which no list of positive edges can satisfy for span 0, even though the same state is reachable via resize((0,)) and round-trips through reopen. For span == 0 any non-empty list of positive edges is now accepted verbatim, producing the same VaryingDimension(edges, extent=0) that resize produces; the strict sum check is kept for span > 0. Tests: the per-spelling regression test from zarr-developers#4328 is replaced by one matrix over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0), (0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte budget, explicit shards}, with separate small tests for each error case. Tests that constructed FixedDimension(size=0) now assert it raises, and a zero-extent test covers the behaviour the old special cases were guarding. Assisted-by: ClaudeCode:claude-fable-5-1
zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)` and for `chunks=(0,)`, so stores with that document exist. Rejecting them at open would turn a previously-readable array into an error; leaving the 0 in place read uninitialised memory after a resize. Normalize the edge to 1 with a ZarrUserWarning instead — the same grid every other "one chunk spans the axis" spelling produces — and keep rejecting a zero edge on an axis that has data. Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and `chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write, append, resize and reopen-then-read every raise ZeroDivisionError. There was never a working behaviour to preserve; normalizing the edge to 1 makes such arrays usable for the first time. Say so in the comment and fragment instead of claiming the stores were previously readable. Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: Codex:GPT-6
zarr 3.2.0 and 3.2.1 classified mixed chunk specs like (2, (5, 10, 5)) as regular and stored them as a "regular" grid whose chunk_shape contains an edge list, while laying the chunks out as a rectilinear grid. Newer versions accepted that metadata and failed later with an unrelated TypeError. RegularChunkGridMetadata now rejects edge lists. When reading stored metadata, a "regular" grid with edge lists is read as the rectilinear grid it describes, with a warning explaining how to re-save it; if rectilinear chunks are disabled, the error says what happened and how to enable them. Closes zarr-developers#4374 Assisted-by: ClaudeCode:claude-opus-5
Assisted-by: ClaudeCode:claude-opus-5
…n the 3.2.x shim Review follow-ups for zarr-developers#4375: - RegularChunkGridMetadata now accepts numpy integer scalars and stores Python ints, mirroring parse_shapelike. The strict integer check had reported np.int64(2) as if it were a list of chunk edges. - The mixed "regular" grid reader converts tuple entries to lists before delegating to RectilinearChunkGridMetadata.from_dict, so metadata dicts built in Python are handled the same as parsed JSON. - The error and warning share one message prefix. - Tests cover numpy ints, run-length encoded edges, and tuple input, and the test section header says the reader is broader than the 3.2.x bug. Assisted-by: ClaudeCode:claude-fable-5-1
The rectilinear hypothesis strategy returned one list of edges per dimension, so no property test ever saw a bare-int dimension, a grid mixing bare ints and edge lists, a run-length encoded declaration, or edges overhanging the extent. The first-element-only classifier that zarr 3.2.x shipped (zarr-developers#4374) was invisible to it. The strategies now cover two spaces. `rectilinear_chunks` samples the `chunks=` syntax: bare ints and flat edge lists in any arrangement, at least one list so the grid is rectilinear. `rectilinear_chunk_shape_ declarations` samples the stored metadata: bare-int steps (including larger than the extent), edge lists written in full or run-length encoded in canonical or arbitrary grouping, and overhanging edges. Each draw comes with the chunk_shapes it must parse to. `chunk_grids` and so `array_metadata` draw from the stored space; `arrays` passes grids the list syntax cannot express as the metadata object, and asserts the stored grid equals the declared one. Two property tests pin the properties that would have caught zarr-developers#4374: every stored declaration parses to its expanded edges and re-serializes to an equivalent grid, and every `chunks=` specification is stored in zarr.json as a "rectilinear" grid equal to the specification. test_unified_chunk_grid.py used a private copy of the old strategy; it now draws from the shared one. Assisted-by: ClaudeCode:claude-fable-5-1
Chunk specifications were parsed by `normalize_chunks_nd`, but a separate duck-typed classifier, `_is_rectilinear_chunks`, ran on the raw input first at three sites to decide whether the spec was rectilinear. Two opinions on the same input is the shape of the bug in zarr-developers#4374, and they could disagree: a 0-d numpy array counted as rectilinear because it has `__iter__`. The classifier is gone. Each site normalizes first and asks the resulting `ChunkGrid` (`is_regular`); stored rectilinear metadata passed as `chunks=` counts as rectilinear even when its edges are uniform. The shard resolver sends regular and rectilinear shard specs through the same normalizer. Also fixed on the way: - The legacy v2 branch of `AsyncArray.create` tested `chunks or chunk_shape`, so `zarr.create(chunks=np.array([...]), zarr_format=2)` failed with "truth value of an array is ambiguous". It now uses the `is not None` form the v3 branch already had. - 0-d numpy arrays unwrap to their scalar in both normalizers instead of failing with "len() of unsized object". - A non-integer scalar spec (`2.0`, `np.float64`) raises the normalizer's own TypeError instead of "object has no len()". Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
…k shapes The chunk_shapes parsers checked exact types: `from_dict` accepted a dimension only as `int` or `list`, `expand_rle` accepted an RLE pair only as a `list`, and the reader for 3.2.x mixed grids special-cased `tuple` to convert it to a list before handing it to `from_dict`. A metadata dict built in Python holds tuples where parsed JSON holds lists, and a numpy array is as good as either, so these checks rejected valid input. One structural predicate, `declares_chunk_edges`, now answers "is this a sequence of edges rather than a single integer size" for all of them: any non-integer iterable except `str`/`bytes`. It is a `TypeGuard`, not a `TypeIs`, because it is False for strings, which are iterable. The four sites that make that decision use it: `parse_chunk_grid`'s detection of a mixed grid, the mixed-grid reader, `RectilinearChunkGridMetadata.from_dict`, and `expand_rle`. The shim no longer needs its tuple special case. Integers are `int | np.integer` throughout, matching `_parse_chunk_shape`, and every parser stores Python ints: `_validate_chunk_shapes` now coerces, so constructing a rectilinear grid from numpy values works the way it already did for a regular grid. Assisted-by: ClaudeCode:claude-opus-5
Assisted-by: ClaudeCode:claude-opus-5
Parsing a regular chunk shape repeated its work at two levels: - `_parse_chunk_shape` type-checked and coerced every dimension, then handed the result to `_validate_chunk_shapes`, which re-ran the same isinstance test and coercion and added only the `>= 1` check. Its edge-list branch was unreachable from this caller, which is why a `cast` was needed on the way out. - `RegularChunkGridMetadata.from_dict` parsed the chunk shape and passed it to the constructor, whose `__post_init__` parsed it again. Together that was four passes over the dimensions for `from_dict`. It is now one: `_parse_chunk_shape` checks the range itself and no longer calls the rectilinear validator, and `from_dict` hands the dimensions to the constructor unparsed. Beyond the redundancy, the shared validator was the coupling that let a rectilinear chunk shape be stored as a regular grid (zarr-developers#4374), so the two grid kinds now validate separately. `RectilinearChunkGridMetadata.from_dict` also had its own `>= 1` check for bare-int dimensions, duplicating `_validate_chunk_shapes`, which `__post_init__` runs over the result anyway. It now only puts the JSON into shape (integer vs sequence, RLE expansion), and a bad bare int is reported by the validator, which names the dimension. `expand_rle` keeps its own checks because it is called directly. Assisted-by: ClaudeCode:claude-opus-5
… format 3
The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.
`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.
Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).
Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.
Assisted-by: ClaudeCode:claude-opus-5
… in one routine The legacy zero-chunk policy was written out twice: inline in `ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2 copy also zipped with `strict=False` and re-appended any trailing chunk entries only so that a separate length check, `parse_metadata`, could report a dimensionality mismatch after construction. `parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one place a stored chunk shape is checked against its array's shape, for both formats: one entry per axis, every integer chunk size at least 1, and a size of 0 (or JSON `false`) on a zero-length axis read as 1 with a warning that names the writer and how to re-save. Non-integer entries, such as edge lists, pass through for the caller's own parser. `ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is gone. The Zarr format 3 adapter only locates a regular grid's `chunk_shape` in the stored document and hands it over; it still runs in `ArrayV3Metadata.__init__` because chunk grid metadata has no array shape. The Zarr format 3 import changes that existed only for the old helper are reverted. Tests for the policy now target the routine: one table of valid and legacy inputs, and one test per rejection (dimension mismatch, zero on a non-empty axis, negative). They replace metadata-level tests in test_v2.py and test_v3.py that only re-tested the same rules; the end-to-end tests still cover both formats' wiring against stored arrays. Assisted-by: ClaudeCode:claude-opus-5
…nk grids `parse_stored_chunk_shape` passed non-integer entries through "for the caller's own parser", which made a regular-grid policy look like a general chunk shape routine and let it decide what a 0-length chunk means for grids it does not own. A rectilinear grid, or any other grid, is free to define its own semantics for 0-length chunks. It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with no pass-through, and its docstring says it applies to Zarr format 2 `chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller hands it a chunk shape only when the grid is named `regular` and every entry is an integer (`_is_regular_chunk_shape`); anything else is not a regular chunk shape and goes to the chunk grid parser untouched. Assisted-by: ClaudeCode:claude-opus-5
- Document the `rectilinear_chunks` return type change in the changelog fragment, since `zarr.testing.strategies` is public. - Make the `RectilinearDimDeclaration` alias private. - Encode canonical RLE in the strategy instead of calling `compress_rle`, so the generator does not depend on the code under test. - Assert `rectilinear_chunks` gets a non-empty shape instead of returning a non-rectilinear `[]`. Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts: # tests/test_properties.py
… into test/rectilinear-declaration-space
With zarr-developers#4334 a rectilinear dimension of extent 0 takes any non-empty list of positive edges at creation, the same state 3.2.x wrote and resize(0) produces. The strategies asserted extent > 0 and `chunk_grids` fell back to a regular grid for any empty axis, so no property saw that state. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… into fix/mixed-regular-chunk-grid-4374 # Conflicts: # src/zarr/core/metadata/v3.py
…place zarr-developers#4334 read a zero chunk size on an empty axis in `ArrayV3Metadata.__init__` and zarr-developers#4375 read a 3.2.x mixed grid inside `parse_chunk_grid`, each with its own predicate, warning text and re-save advice. Both are compatibility readings of a stored `regular` grid, so `_read_stored_regular_chunk_grid` now dispatches to both, `parse_chunk_grid` accepts only what the spec allows, and both warnings use `RESAVE_METADATA_HINT`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…panning chunk A stored chunk size of 0 was tolerated only on a zero-length axis and rejected otherwise, because the metadata supposedly could not say how the stored chunks were laid out. But a chunk size of 0 gives a grid of zero chunks, so no release could store a chunk under it, and the writers of that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records a Zarr format 3 resize and the resize half of a failed append. Measured with real installs. Those arrays open in 3.4.0, attributes included; the rejection would have made them unopenable. `parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`) on any axis as one chunk spanning it, `max(extent, 1)`, which is what the `-1`/`False` spec that wrote it meant. On a grown axis the warning also says that data written to it was not saved. Negative sizes are still rejected. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e 2.3 NumPy before 2.3 takes a NumPy boolean as an index with a DeprecationWarning, so `operator.index(np.True_)` returned 1 there, and the normalizer read `chunks=np.True_` as a chunk size of 1 instead of rejecting it as zarr 3.4.0 did. The min-deps CI job failed on this. The integer test now rejects a value whose dtype is NumPy bool before calling `operator.index`. Assisted-by: ClaudeCode:claude-opus-5-5
…an overwrite test Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts: # src/zarr/core/chunk_grids.py
Resolves the import conflict in group.py with zarr-developers#4391: keep both import lines. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves the import conflict in group.py with zarr-developers#4391: keep both import lines. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…was 0 The warning for a stored chunk size of 0 on an axis of positive length now says that the array holds only its fill value, so recreating it with the wanted chunk shape loses nothing, and gives the re-save as the way to keep it instead. A re-save freezes the smallest chunk size into an array that holds no data yet. Each reading now carries its own advice, so mark_upgraded appends no shared hint. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The stored-zero warning now ends with RECREATE_HINT and the mixed-grid warning with RESAVE_HINT; mark_upgraded appends no shared hint. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The stored-zero warning now ends with RECREATE_HINT and the mixed-grid warning with RESAVE_HINT; mark_upgraded appends no shared hint. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ored chunk size of 0 The warning names the call, zarr.from_array(array.store, name=array.path, data=array, chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value, attributes and codecs, and a test pins the recipe on both formats. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with save_new_metadata: keep main's helper and this branch's imports and chunk normalization. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…egular-chunk-grid-4374 Resolves the overlap with zarr-developers#4410: save_new_metadata (main) replaces this branch's encode-then-_prepare_overwrite at the create sites, and save_metadata and save_new_metadata encode through encode_documents so errors name the node. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…chunk-sizes Resolves the overlap with zarr-developers#4410: save_new_metadata (main) replaces this branch's encode-then-_prepare_overwrite at the create sites, and save_metadata and save_new_metadata encode through encode_documents so errors name the node. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…airs, not upgrades The module is zarr.core.metadata.repair: repair_array_document, mark_repaired, ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test module test_repair.py. A reading that turns an invalid stored document into a valid one fixes it; it does not move it to a newer format. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Conflicts resolved by keeping this branch's lines and applying the same rename to them; this branch's own text now says repair as well. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Conflicts resolved by keeping this branch's lines and applying the same rename to them; this branch's own text now says repair as well. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch already contains, so every file it touches keeps this branch's version; the merge brings in only the dependency bump (zarr-developers#4461). Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch already contains, so every file it touches keeps this branch's version; the merge brings in only the dependency bump (zarr-developers#4461). Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…adds over main Rejecting edge lists in a regular chunk grid landed with zarr-developers#4334, and encoding a new node's metadata before deleting the existing one landed with zarr-developers#4410. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b
added a commit
to d-v-b/zarr-python
that referenced
this pull request
Oct 3, 2026
…ut of this patch PR Reverts 47587bc, baddb31 and ad34ae8, and the error-message part of e547401. Rejecting a chunk size of 0 in ArrayV2Metadata and ShardingCodec, and no longer reading a falsy `chunks` as absent in the legacy Zarr format 2 zarr.create, reject inputs zarr 3.4.0 accepted; this PR targets the 3.4.1 patch release, which must not. zarr-developers#4431 makes the same changes, more strictly, for the next minor release. The widened chunk and shard annotations and the zarr.create docstring fix stay. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b
added a commit
that referenced
this pull request
Oct 3, 2026
…izer alone (#4376) * fix: make the chunk normalizer the one judge of a chunk specification Chunk specifications were parsed by `normalize_chunks_nd`, but a separate duck-typed classifier, `_is_rectilinear_chunks`, ran on the raw input first at three sites to decide whether the spec was rectilinear. Two opinions on the same input is the shape of the bug in #4374, and they could disagree: a 0-d numpy array counted as rectilinear because it has `__iter__`. The classifier is gone. Each site normalizes first and asks the resulting `ChunkGrid` (`is_regular`); stored rectilinear metadata passed as `chunks=` counts as rectilinear even when its edges are uniform. The shard resolver sends regular and rectilinear shard specs through the same normalizer. Also fixed on the way: - The legacy v2 branch of `AsyncArray.create` tested `chunks or chunk_shape`, so `zarr.create(chunks=np.array([...]), zarr_format=2)` failed with "truth value of an array is ambiguous". It now uses the `is not None` form the v3 branch already had. - 0-d numpy arrays unwrap to their scalar in both normalizers instead of failing with "len() of unsized object". - A non-integer scalar spec (`2.0`, `np.float64`) raises the normalizer's own TypeError instead of "object has no len()". Assisted-by: ClaudeCode:claude-fable-5-1 * docs: add changelog fragment for #4376 Assisted-by: ClaudeCode:claude-fable-5-1 * refactor(chunk_grids): one integer test for chunk input `_chunk_int` is the normalizer's single integer rule: anything Python's integer protocol accepts (int, numpy integer scalars, 0-d integer arrays), except bool. It replaces the repeated numbers.Integral checks and the two 0-d ndarray unwraps in normalize_chunks_1d / normalize_chunks_nd. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: keep numpy chunk tests that cover real bugs; pin legacy v2 rectilinear rejection The numpy cases that passed before this PR are dropped; the 0-d array and float cases move into the normalizer's table and error tests. The legacy zarr.create Zarr format 2 path gets a truthiness test and its rectilinear rejection is now asserted with a message match. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: note legacy v2 chunks=0/[] and bool chunk sizes in the #4376 changelog Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(chunk_grids): explicit edge lists stay rectilinear in ChunkGrid.from_sizes `ChunkGrid.from_sizes` collapsed uniform edge lists to `FixedDimension` while `normalize_chunks_1d` keeps them as `VaryingDimension`, so `init_array` needed an `isinstance(chunks, RectilinearChunkGridMetadata)` guard to keep uniform stored rectilinear grids under the Zarr format 2 and sharding restrictions. Both now agree that a list declares a rectilinear dimension, and the guard is gone: the normalized grid is the one judge. Also: `normalize_chunks_nd` and the shard spec are typed with `ChunksLike`; a bool chunk size is reported as not a chunk size; an integral float (`10.0`) is pinned as rejected; `None` reaching the normalizer gets the generic non-integer `TypeError`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(chunk_grids): reject every spelling of a boolean chunk size with one error A bool, np.bool_, or boolean array anywhere in a chunk specification now raises the same TypeError from the normalizer's one integer test, instead of three different errors depending on spelling. Document that a RectilinearChunkGridMetadata of bare integers is read as a regular grid. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(chunk_grids): keep zarr 3.4.0's reading of bool chunk sizes and falsy v2 chunks (patch release) Nothing in a patch release may reject a chunk specification that zarr 3.4.0 accepted. The normalizer's integer test reads a Python `bool` as the `int` it is again, so `chunks=(True, 5)` and a `True` edge are a chunk size of 1; `chunks=True` and `chunks=None` raise 3.4.0's `ValueError` again. Numpy booleans stay rejected, as they were. The legacy `zarr.create(..., zarr_format=2)` again reads a falsy `chunks` (0, [], False) as not given and chunks automatically; a numpy array is always taken as given, so its truth value is never tested. The `TypeError` for every boolean spelling returns in the next minor release. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe the 4376 changes for a patch release Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(array): read one-element and empty numpy chunk arrays in legacy v2 create as zarr 3.4.0 did The legacy `zarr.create(..., zarr_format=2)` took every numpy array as a given chunk specification, so `np.array(0)`, `np.array(False)` and `np.array([0])` raised instead of auto-chunking as in zarr 3.4.0. A numpy array with more than one element has no truth value and is always given; a shorter one is read by `.any()`, its truth value, which is False when empty, so `np.array([])` auto-chunks like `[]` (as 3.4.0 did with numpy 2.1). The two legacy v2 tests become one table. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(chunk_grids): let operator.index be the only integer test in _chunk_int `operator.index` raises `TypeError` for everything the `SupportsIndex` check rejected, so the check was redundant (identical results over bool, numpy bools and integers, 0-d and 1-d arrays, float, str, bytes, None, list and a custom `__index__` class). Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(chunk_grids): read bool chunk sizes as rows of the normalizer table Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(array): name the legacy v2 chunks-given test as _v2_chunks_given Replace the inline conditional expression in `AsyncArray._create` with a small named predicate. Results are identical on every probed input. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(array): pin 0-d and all-zero numpy chunks in legacy v2 create Add `np.array(7)` -> `(7, 7)` to `test_legacy_create_v2_chunks` and an error test for `np.array([0, 0])`, killing the mutants that drop either half of `_v2_chunks_given`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(chunk-grids): accept sequences that only define __getitem__ and __len__ The chunk normalizer checked `isinstance(chunks, Iterable)` before calling `list(chunks)`. `Iterable` does not recognize a sequence that implements only `__getitem__` and `__len__`, which `list` accepts and which zarr 3.4.0 took as a chunk specification. The normalizer now calls `list` and turns its `TypeError` into the chunk specification error. Assisted-by: ClaudeCode:claude-opus-5-5 * fix(chunk-grids): reject NumPy booleans as chunk sizes on NumPy before 2.3 NumPy before 2.3 takes a NumPy boolean as an index with a DeprecationWarning, so `operator.index(np.True_)` returned 1 there, and the normalizer read `chunks=np.True_` as a chunk size of 1 instead of rejecting it as zarr 3.4.0 did. The min-deps CI job failed on this. The integer test now rejects a value whose dtype is NumPy bool before calling `operator.index`. Assisted-by: ClaudeCode:claude-opus-5-5 * fix(chunk-grids): declare every accepted chunk form in ShardsLike and zarr.create ShardsLike is now ChunksLike plus the sharding configuration and "auto", so numpy integers and arrays are declared for shards as they are for chunks. zarr.create and AsyncArray._create take ChunksLike for chunks and chunk_shape, matching what the normalizer reads. Tests drop the Any annotations and the type: ignore that worked around the narrower types. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(array): read only None as an absent chunks argument in legacy zarr.create The Zarr format 2 path of AsyncArray._create used the truth value of `chunks` to decide whether it was given, so False, 0, [] and np.array([0]) were silently auto-chunked while Zarr format 3 normalized or rejected them. Both formats now share the None-only `_raw_chunks` and agree on every input. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(sharding): parse the inner chunk shape with the chunk normalizer as a regular grid ShardingCodec read its inner chunk_shape with parse_shapelike, which allows 0, so an inner size of 0 was accepted at construction and failed later with ZeroDivisionError. parse_regular_chunk_shape shares the normalizer's integer test (_chunk_int/_chunk_list) and requires every size to be at least 1; with no axis length to resolve against, -1 and False are rejected, and an explicit edge list (a rectilinear dimension) is rejected because the inner grid of a shard is regular. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(metadata): parse the Zarr format 2 chunk shape as a regular chunk shape ArrayV2Metadata read `chunks` with parse_shapelike, which accepts 0, while RegularChunkGridMetadata rejects it. It now uses parse_regular_chunk_shape, as ShardingCodec does, so a chunk size below 1 raises at construction. Stored documents are unaffected: from_dict repairs a stored 0 before the constructor sees it. The shared chunk-shape tests cover both sites again. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(api): describe what zarr.create does with chunks=None, False and True The docstring said True guesses the chunk shape and is the default; None is the default and True raises. The auto-chunking error now tells create_array callers to pass "auto" rather than the None they just passed. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * revert: keep the strict chunk-size changes in #4431, out of this patch PR Reverts 47587bc, baddb31 and ad34ae8, and the error-message part of e547401. Rejecting a chunk size of 0 in ArrayV2Metadata and ShardingCodec, and no longer reading a falsy `chunks` as absent in the legacy Zarr format 2 zarr.create, reject inputs zarr 3.4.0 accepted; this PR targets the 3.4.1 patch release, which must not. #4431 makes the same changes, more strictly, for the next minor release. The widened chunk and shard annotations and the zarr.create docstring fix stay. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
4 of 7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Metadata built in code is strict about chunk edge lengths:
ArrayV2Metadata(chunks=...),RegularChunkGridMetadata,RectilinearChunkGridMetadataandShardingCodec(chunk_shape=...)take them as Pythonints of at least 1. A chunk size of 0 now raises aValueError, as inArrayV2Metadata(chunks=(0,))for an empty array. Code that passes an array's shape as its chunk shape, such asArrayV2Metadata(chunks=arr.shape)orRegularChunkGridMetadata(chunk_shape=arr.shape), fails for an array with a zero-length axis: usetuple(max(1, s) for s in arr.shape). Aboolor a NumPy integer, including a scalarchunks=np.int64(5), raises aTypeError, as a float already did, and so does a chunk shape that is not a list or tuple. This reverses what 3.4.1 said of these classes:ArrayV2Metadata(chunks=(0,))is no longer accepted, and aboolchunk edge length is no longer read as theintit equals. (This entry assumes 3.4.1 is released before this release.) The array creation functions still accept NumPy integers inchunks=andshards=.This applies to stored metadata too wherever the stored-document upgrades do not read it: a stored chunk shape that is a bare integer instead of a list, as in a Zarr format 2
"chunks": 4or a sharding codec's"chunk_shape": 2, which earlier releases read as a one-dimensional chunk shape, now raises aTypeErrorwhen the array is opened. A group whose consolidated metadata includes such an array does not open from its consolidated metadata at all, and the error names that member: open the group withuse_consolidated=Falseto reach its other members. A storedtrueas a rectilinear chunk grid's bare step ("chunk_shapes": [true]) or as the count of a run-length pair ([[4, true]]), which no known writer stored, raises aTypeErrortoo. The chunk sizes the upgrades read open as before: 0 orfalseas a regular chunk size (in Zarr format 2chunksor aregulargrid'schunk_shape);trueas a regular chunk size, as an edge in a rectilinear edge list, as the size of a run-length pair, or as a sharding codec's inner chunk size; and integral float edges in a rectilinear chunk grid ([[4.0, 2]]).Other than
chunks=Falseorshards=Falseas the whole specification (one chunk spanning every axis), a boolean anywhere in a chunk specification given to the array creation functions (chunks=True,np.True_,np.False_, a boolean array,chunks=(True, 5),chunks=(False, 5), or a boolean edge in a list) raises oneTypeError, instead of being read as a chunk size of 1 or 0 or raising a different error for each spelling.chunks=Trueandcreate_array(chunks=None)raised aValueErrorbefore; both errors say how to ask for automatic chunking. The legacyzarr.create(..., zarr_format=2)no longer reads a falsychunksas not given:chunks=0,np.int64(0)andchunks=[]raise aValueError, andchunks=Falsemeans one chunk spanning the array, as it does for Zarr format 3, instead of automatic chunking;chunks=Nonestill chunks automatically.zarr.testing.strategies.chunks_param_from_rectilinearis removed; nothing in zarr used it afterarrays()began drawing itschunks=argument fromrectilinear_chunks.rectilinear_chunksnow returnslist[int | list[int]]instead oflist[list[int]]: a dimension may be a bare int step, which can exceed the extent, so code that treats every dimension as a list of edges (len(dim),sum(dim)) must handle ints. It also requiresshapeto have at least one dimension.Why a separate release
The 3.4.1 PRs fix what was broken without rejecting, changing or newly warning on anything 3.4.0 accepted and handled correctly. That left the metadata constructors lenient:
ArrayV2Metadata(chunks=(0,)),Trueas a chunk size, NumPy integers and scalar chunk shapes are all still accepted there, and each one is a way to build metadata that is invalid or reads back differently. The patch PRs define "one chunk edge length is anintof at least 1" at the metadata boundary; this PR makes the constructors enforce it with no exceptions. Stored documents are still read leniently, but only inzarr.core.metadata.upgrades.What changes
Strict metadata constructors
ArrayV2Metadata(chunks=...),RegularChunkGridMetadata,RectilinearChunkGridMetadataandShardingCodec(chunk_shape=...)take chunk edge lengths as Pythonints of at least 1, through one rule (parse_chunk_edge) and one chunk-shape type:ArrayV2Metadata(chunks=(0,)),ShardingCodec(chunk_shape=(0,))ValueError: Dimension 0: chunk edge length must be >= 1, got 0True/Falseas a chunk edge lengthTypeError: Dimension 0: chunk edge length must be an int, got Truenp.int64(5))ArrayV2MetadataandShardingCodecTypeErrorchunks=5,chunks=np.int64(5)), or any chunk shape that is not a list or tupleTypeError: A chunk shape must be a list or tuple of chunk edge lengths, got 55.0), including run-length encoded sizes and countsTypeErroralreadyTypeErrorBecause the constructors now reject a chunk size of 0, an array built from a metadata object is taken as built: the 3.4.1 path that read a code-built
ArrayV2Metadatawith a 0 the way a stored 0 is read is removed.Booleans in chunk specifications: one
TypeErrorOther than
chunks=Falseorshards=Falseas the whole specification (one chunk spanning every axis), a boolean anywhere in a specification given to the creation functions raises oneTypeError:chunks=True,np.True_,np.False_, a boolean array,chunks=(True, 5),chunks=(False, 5), or a boolean edge in a list. 3.4.1 read these as a chunk size of 1 or 0, or raised a different error for each spelling.chunks=Trueandcreate_array(chunks=None)share an automatic-chunking hint:Legacy Zarr format 2
zarr.createA falsy
chunksno longer means "not given":chunks=0,np.int64(0),chunks=[]andchunks=()raiseValueError, NumPy booleans raise the booleanTypeError, andchunks=Falsemeans one chunk spanning the array, as it does for Zarr format 3.chunks=Nonestill chunks automatically.zarr.testingchunks_param_from_rectilinearis removed (nothing in zarr used it afterarrays()began drawing itschunks=argument from the rectilinear strategies).rectilinear_chunksnow returns the fullchunks=form,list[int | list[int]]: a dimension may be a bare-int step, which can exceed the extent, so code that treats every dimension as a list (len(dim),sum(dim)) must handle ints. It requiresshapeto have at least one dimension. The rectilinear strategies are experimental.What still reads
Every stored document a known writer produced opens as in 3.4.1, because the upgrades run before the strict constructors:
false(Zarr format 2chunks,regularchunk_shape), including sharded arrays;trueas a regular chunk size, a rectilinear edge, the size of a run-length pair, or a sharding codec's inner chunk size (nested or not);[[4.0, 2]]);regulargrids.Newly rejected on read, because no known writer stored them and the upgrades do not read them:
"chunks": 4, a sharding codec's"chunk_shape": 2), which earlier releases read as one dimension, raises aTypeErrorwhen the array is opened. A group whose consolidated metadata includes such an array does not open from its consolidated metadata; the error names the member, anduse_consolidated=Falsereaches the other members;trueas a rectilinear bare step ("chunk_shapes": [true]) or as a run-length count ([[4, true]]).Downstream
Code that passes an array's shape as its chunk shape fails for an array with a zero-length axis:
Use
tuple(max(1, s) for s in arr.shape)instead. VirtualiZarr's kerchunk writer buildsArrayV2Metadata(chunks=np_arr.shape, ...)for inlined arrays, so an empty inlined array would hit this; the project should get a heads-up before 3.5.0. VirtualiZarr's test suite gave identical results on this branch and on the patch branches. The creation functions still accept NumPy integers inchunks=andshards=.Evidence
TypeError(tests pin each one).Tests
One success table for chunk-edge sites (ints only) and one error test per rejection (0,
bool, NumPy integer, scalar chunk shape, float) across the three metadata classes andShardingCodec; one table for the legacy Zarr format 2chunksargument plus error tests for0and[]; the booleanTypeErroroverTrueandFalsespellings (except the whole-specFalse), including a boolean run-length count; the scalar-chunk consolidated member named in the error.Stack and release notes
Target: 3.5.0 (minor).
Depends on fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334, fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x #4375, fix(chunk-grids): judge every chunk specification by the chunk normalizer alone #4376 and test(chunk-grids): sample the full rectilinear chunk grid declaration space #4377 (the 3.4.1 patch PRs), which are merged into this branch, so until they land this diff includes all of them. After they merge, the diff shrinks to the tightening described above.
The changelog fragment is
changes/4431.removal.md. It assumes 3.4.1 is released first, since it reverses some of what 3.4.1 says.On 2026-09-29 the agent merged the updated fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334 into this branch (5687c55). That brought in current
main(cd1e5b3) and fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334's new reading of a stored chunk size of 0: now 1, or the inner chunk size, however long the axis. The conflicts were resolved as on the patch PRs:ChunkGrid.from_sizes: kept fix(chunk-grids): judge every chunk specification by the chunk normalizer alone #4376's version, which no longer collapses uniform edges to a regular dimension._read_chunk_sizedocstring: combined fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334's and fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x #4375's text.test_upgrades.py:test_array_from_metadata_with_chunk_size_zerostays deleted, since this PR rejects a chunk size of 0 inArrayV2Metadata; the consolidated test now expects chunk size 1; and thelongmixed-grid case expects its stored 0 to read as 1.The full test suite passes (12099 passed). The newer heads of fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x #4375, fix(chunk-grids): judge every chunk specification by the chunk normalizer alone #4376 and test(chunk-grids): sample the full rectilinear chunk grid declaration space #4377 add only merge commits, and merging them would change no files.
On 2026-09-29 the agent merged the review fixes of fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x #4375, fix(chunk-grids): judge every chunk specification by the chunk normalizer alone #4376 and test(chunk-grids): sample the full rectilinear chunk grid declaration space #4377, plus current
main(0ba2ca2), into this branch (2b38f64). The conflicts were resolved as follows:init_array: this PR's normalize-first order is kept, and_prepare_overwritenow comes after the metadata is encoded.test_upgrades.py: fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334's new test runsupgrade_array_documentdirectly, because this PR's strict constructors reject the NumPy integer that the patch version passes tofrom_dict.test_chunk_grids.py: the old-protocol sequence cases are kept, and theboolrows that this PR removes stay removed.ShardingCodec.__init__:main'sBytesCodec(endian="little")defaults are combined with this PR'sChunkShape.The full test suite passes (12198).
🤖 Generated with Claude Code