fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x - #4375
Merged
d-v-b merged 111 commits intoOct 2, 2026
Merged
Conversation
…, clamps and metadata Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0, in which case the dimension has zero chunks (ceildiv(0, size) == 0). Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, zarr-developers#972, zarr-developers#1977, zarr-developers#2434, zarr-developers#3711, zarr-developers#4305, zarr-developers#4307, zarr-developers#4328) because the layers disagreed on this invariant and every span-derived chunk spelling clamped on its own: - The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1, but the in-memory FixedDimension allowed size == 0 with four special-case branches left over from zarr-developers#2434, so normalization could build a grid the metadata constructor then rejected. FixedDimension now rejects size < 1 and the four `if self.size == 0` branches are gone. VaryingDimension already required edges > 0 and is unchanged. - `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both the typesize == 0 early return and the np.maximum line) and `shards="auto"` each derived "one chunk covering the axis" independently. They now all go through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is the single definition of that phrase for a possibly zero-length axis. - Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]` document opened fine and read uninitialised memory after a resize. It now raises a clear ValueError at parse time, matching the format 3 grid. - Rectilinear grids had no creation-time spelling for a zero-length axis: normalize_chunks_1d required sum(edges) == span, which no list of positive edges can satisfy for span 0, even though the same state is reachable via resize((0,)) and round-trips through reopen. For span == 0 any non-empty list of positive edges is now accepted verbatim, producing the same VaryingDimension(edges, extent=0) that resize produces; the strict sum check is kept for span > 0. Tests: the per-spelling regression test from zarr-developers#4328 is replaced by one matrix over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0), (0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte budget, explicit shards}, with separate small tests for each error case. Tests that constructed FixedDimension(size=0) now assert it raises, and a zero-extent test covers the behaviour the old special cases were guarding. Assisted-by: ClaudeCode:claude-fable-5-1
zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)` and for `chunks=(0,)`, so stores with that document exist. Rejecting them at open would turn a previously-readable array into an error; leaving the 0 in place read uninitialised memory after a resize. Normalize the edge to 1 with a ZarrUserWarning instead — the same grid every other "one chunk spans the axis" spelling produces — and keep rejecting a zero edge on an axis that has data. Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and `chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write, append, resize and reopen-then-read every raise ZeroDivisionError. There was never a working behaviour to preserve; normalizing the edge to 1 makes such arrays usable for the first time. Say so in the comment and fragment instead of claiming the stores were previously readable. Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: Codex:GPT-6
zarr 3.2.0 and 3.2.1 classified mixed chunk specs like (2, (5, 10, 5)) as regular and stored them as a "regular" grid whose chunk_shape contains an edge list, while laying the chunks out as a rectilinear grid. Newer versions accepted that metadata and failed later with an unrelated TypeError. RegularChunkGridMetadata now rejects edge lists. When reading stored metadata, a "regular" grid with edge lists is read as the rectilinear grid it describes, with a warning explaining how to re-save it; if rectilinear chunks are disabled, the error says what happened and how to enable them. Closes zarr-developers#4374 Assisted-by: ClaudeCode:claude-opus-5
Assisted-by: ClaudeCode:claude-opus-5
Documentation build overview
|
Documentation build overview
No files changed. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #4375 +/- ##
==========================================
+ Coverage 93.40% 94.69% +1.28%
==========================================
Files 94 94
Lines 13544 13597 +53
==========================================
+ Hits 12651 12875 +224
+ Misses 893 722 -171
🚀 New features to boost your workflow:
|
…n the 3.2.x shim Review follow-ups for zarr-developers#4375: - RegularChunkGridMetadata now accepts numpy integer scalars and stores Python ints, mirroring parse_shapelike. The strict integer check had reported np.int64(2) as if it were a list of chunk edges. - The mixed "regular" grid reader converts tuple entries to lists before delegating to RectilinearChunkGridMetadata.from_dict, so metadata dicts built in Python are handled the same as parsed JSON. - The error and warning share one message prefix. - Tests cover numpy ints, run-length encoded edges, and tuple input, and the test section header says the reader is broader than the 3.2.x bug. Assisted-by: ClaudeCode:claude-fable-5-1
This was referenced Sep 18, 2026
…k shapes The chunk_shapes parsers checked exact types: `from_dict` accepted a dimension only as `int` or `list`, `expand_rle` accepted an RLE pair only as a `list`, and the reader for 3.2.x mixed grids special-cased `tuple` to convert it to a list before handing it to `from_dict`. A metadata dict built in Python holds tuples where parsed JSON holds lists, and a numpy array is as good as either, so these checks rejected valid input. One structural predicate, `declares_chunk_edges`, now answers "is this a sequence of edges rather than a single integer size" for all of them: any non-integer iterable except `str`/`bytes`. It is a `TypeGuard`, not a `TypeIs`, because it is False for strings, which are iterable. The four sites that make that decision use it: `parse_chunk_grid`'s detection of a mixed grid, the mixed-grid reader, `RectilinearChunkGridMetadata.from_dict`, and `expand_rle`. The shim no longer needs its tuple special case. Integers are `int | np.integer` throughout, matching `_parse_chunk_shape`, and every parser stores Python ints: `_validate_chunk_shapes` now coerces, so constructing a rectilinear grid from numpy values works the way it already did for a regular grid. Assisted-by: ClaudeCode:claude-opus-5
Assisted-by: ClaudeCode:claude-opus-5
Parsing a regular chunk shape repeated its work at two levels: - `_parse_chunk_shape` type-checked and coerced every dimension, then handed the result to `_validate_chunk_shapes`, which re-ran the same isinstance test and coercion and added only the `>= 1` check. Its edge-list branch was unreachable from this caller, which is why a `cast` was needed on the way out. - `RegularChunkGridMetadata.from_dict` parsed the chunk shape and passed it to the constructor, whose `__post_init__` parsed it again. Together that was four passes over the dimensions for `from_dict`. It is now one: `_parse_chunk_shape` checks the range itself and no longer calls the rectilinear validator, and `from_dict` hands the dimensions to the constructor unparsed. Beyond the redundancy, the shared validator was the coupling that let a rectilinear chunk shape be stored as a regular grid (zarr-developers#4374), so the two grid kinds now validate separately. `RectilinearChunkGridMetadata.from_dict` also had its own `>= 1` check for bare-int dimensions, duplicating `_validate_chunk_shapes`, which `__post_init__` runs over the result anyway. It now only puts the JSON into shape (integer vs sequence, RLE expansion), and a bad bare int is reported by the validator, which names the dimension. `expand_rle` keeps its own checks because it is called directly. Assisted-by: ClaudeCode:claude-opus-5
… format 3
The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.
`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.
Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).
Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.
Assisted-by: ClaudeCode:claude-opus-5
… in one routine The legacy zero-chunk policy was written out twice: inline in `ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2 copy also zipped with `strict=False` and re-appended any trailing chunk entries only so that a separate length check, `parse_metadata`, could report a dimensionality mismatch after construction. `parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one place a stored chunk shape is checked against its array's shape, for both formats: one entry per axis, every integer chunk size at least 1, and a size of 0 (or JSON `false`) on a zero-length axis read as 1 with a warning that names the writer and how to re-save. Non-integer entries, such as edge lists, pass through for the caller's own parser. `ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is gone. The Zarr format 3 adapter only locates a regular grid's `chunk_shape` in the stored document and hands it over; it still runs in `ArrayV3Metadata.__init__` because chunk grid metadata has no array shape. The Zarr format 3 import changes that existed only for the old helper are reverted. Tests for the policy now target the routine: one table of valid and legacy inputs, and one test per rejection (dimension mismatch, zero on a non-empty axis, negative). They replace metadata-level tests in test_v2.py and test_v3.py that only re-tested the same rules; the end-to-end tests still cover both formats' wiring against stored arrays. Assisted-by: ClaudeCode:claude-opus-5
…nk grids `parse_stored_chunk_shape` passed non-integer entries through "for the caller's own parser", which made a regular-grid policy look like a general chunk shape routine and let it decide what a 0-length chunk means for grids it does not own. A rectilinear grid, or any other grid, is free to define its own semantics for 0-length chunks. It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with no pass-through, and its docstring says it applies to Zarr format 2 `chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller hands it a chunk shape only when the grid is named `regular` and every entry is an integer (`_is_regular_chunk_shape`); anything else is not a regular chunk shape and goes to the chunk grid parser untouched. Assisted-by: ClaudeCode:claude-opus-5
… into fix/mixed-regular-chunk-grid-4374 # Conflicts: # src/zarr/core/metadata/v3.py
…place zarr-developers#4334 read a zero chunk size on an empty axis in `ArrayV3Metadata.__init__` and zarr-developers#4375 read a 3.2.x mixed grid inside `parse_chunk_grid`, each with its own predicate, warning text and re-save advice. Both are compatibility readings of a stored `regular` grid, so `_read_stored_regular_chunk_grid` now dispatches to both, `parse_chunk_grid` accepts only what the spec allows, and both warnings use `RESAVE_METADATA_HINT`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…panning chunk A stored chunk size of 0 was tolerated only on a zero-length axis and rejected otherwise, because the metadata supposedly could not say how the stored chunks were laid out. But a chunk size of 0 gives a grid of zero chunks, so no release could store a chunk under it, and the writers of that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records a Zarr format 3 resize and the resize half of a failed append. Measured with real installs. Those arrays open in 3.4.0, attributes included; the rejection would have made them unopenable. `parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`) on any axis as one chunk spanning it, `max(extent, 1)`, which is what the `-1`/`False` spec that wrote it meant. On a grown axis the warning also says that data written to it was not saved. Negative sizes are still rejected. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The only state machine that touched arrays compared zarr on one store with zarr on a MemoryStore, so a chunk grid bug showed up identically on both sides; it had no append rule, covered Zarr format 3 only, and kept every empty axis at 0 when resizing, which is where the zero-length bugs live. `ArrayLifecycle` checks one array against a NumPy model across both formats, every chunk spelling (-1, False, "auto", ints, sharded, rectilinear) and the stored chunk size of 0 that releases before 3.4 wrote, including on an axis those releases grew. Rules append, resize (growing and shrinking to and from 0), write and re-save the metadata; the invariant reopens the array and compares shape, values and whether the legacy warning is due. Deliberately breaking the grown-axis policy, the legacy warning, or append on an empty axis each fails it. `resize` keeps partly retained chunks whole, so cells cut off by a shrink can come back with their old values when the axis grows (as in 2.x); the model marks such cells unknown until written instead of encoding that. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… into fix/mixed-regular-chunk-grid-4374 # Conflicts: # src/zarr/core/metadata/v3.py
# Conflicts: # src/zarr/core/chunk_grids.py
A stored chunk size of 0 or `false` was read as one chunk spanning the axis at open time. On an axis that grew after the size was written, that made the chunk size depend on when the array was opened, and it could be very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one 800 MB chunk. It is now read as the smallest valid chunk edge length (1, or the inner chunk size for a shard), which is what `chunks=-1` gives on a zero-length axis. A resize by software that kept the stored 0 no longer changes the layout a stale handle expects. Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts: # src/zarr/core/metadata/upgrades.py # tests/test_metadata/test_upgrades.py
…SON input as is A stored `true` chunk size or an integral float edge is read as the value zarr already read it as, so chunks written under the upgraded metadata are where every reader looks. Marking such documents for re-saving turned a plain chunk write into a metadata write: in a ZipStore that adds a second `zarr.json` entry (a `UserWarning`, an error under `-W error`), and it opened a race with concurrent metadata writes that 3.4.0 did not have. Each upgrade now reports whether it moves chunks, and only those mark the metadata. The upgrades also detected changes by comparing JSON encodings, so `ArrayV3Metadata.from_dict` given metadata built in code (a codec instance, a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`. Changes are now tracked where they are made. Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts: # src/zarr/core/metadata/io.py # src/zarr/core/metadata/upgrades.py
With the rectilinear flag checked when metadata is stored rather than when it is built, `zarr.create(..., overwrite=True)` with rectilinear chunks and the flag off deleted the existing array and then raised; zarr 3.4.0 raised before deleting anything. `AsyncArray._create_v2`, `_create_v3` and `init_array` now build and encode the new metadata before `_prepare_overwrite`, so metadata that cannot be stored fails with the store untouched. This also covers `create_array(..., overwrite=True)`, which deleted the existing array before validating its arguments. Assisted-by: ClaudeCode:claude-opus-5-5
…an overwrite test Assisted-by: ClaudeCode:claude-opus-5-5
Resolves the import conflict in group.py with zarr-developers#4391: keep both import lines. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…was 0 The warning for a stored chunk size of 0 on an axis of positive length now says that the array holds only its fill value, so recreating it with the wanted chunk shape loses nothing, and gives the re-save as the way to keep it instead. A re-save freezes the smallest chunk size into an array that holds no data yet. Each reading now carries its own advice, so mark_upgraded appends no shared hint. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The stored-zero warning now ends with RECREATE_HINT and the mixed-grid warning with RESAVE_HINT; mark_upgraded appends no shared hint. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ored chunk size of 0 The warning names the call, zarr.from_array(array.store, name=array.path, data=array, chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value, attributes and codecs, and a test pins the recipe on both formats. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with save_new_metadata: keep main's helper and this branch's imports and chunk normalization. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…egular-chunk-grid-4374 Resolves the overlap with zarr-developers#4410: save_new_metadata (main) replaces this branch's encode-then-_prepare_overwrite at the create sites, and save_metadata and save_new_metadata encode through encode_documents so errors name the node. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…airs, not upgrades The module is zarr.core.metadata.repair: repair_array_document, mark_repaired, ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test module test_repair.py. A reading that turns an invalid stored document into a valid one fixes it; it does not move it to a newer format. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Conflicts resolved by keeping this branch's lines and applying the same rename to them; this branch's own text now says repair as well. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch already contains, so every file it touches keeps this branch's version; the merge brings in only the dependency bump (zarr-developers#4461). Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…adds over main Rejecting edge lists in a regular chunk grid landed with zarr-developers#4334, and encoding a new node's metadata before deleting the existing one landed with zarr-developers#4410. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b
marked this pull request as ready for review
October 2, 2026 08:00
Contributor
Author
|
this is an important bugfix that restores legibility of glitched array metadata older versions of |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Arrays whose stored
regularchunk grid mixes chunk sizes with lists of chunk edge lengths, such as"chunk_shape": [2, [5, 10, 5]](written by zarr 3.2.0 and 3.2.1 forchunks=(2, (5, 10, 5)), and as[2, [5.0, 10.0, 5.0]]for float edges), can be read again, without enablingarray.rectilinear_chunks. The grid is read as the rectilinear chunk grid it describes, with aZarrUserWarning; re-saving the array's metadata (or writing chunks to it) stores that rectilinear grid, which requires the flag. A group's consolidated metadata keeps such an array's metadata as it was stored. A regular chunk grid given edge lists is now rejected, so this metadata is no longer written.The
array.rectilinear_chunksflag now gates reading and storing array metadata that declares a rectilinear chunk grid, including in a group's consolidated metadata, instead of constructingRectilinearChunkGridMetadata.Array.resize, deleting a group member, andcreate_hierarchy(..., overwrite=True)now encode the metadata they will store before deleting anything, so metadata that cannot be stored (for example, a rectilinear chunk grid with the flag off) fails with the store untouched. That error names the array when it is a member of a group's consolidated metadata.Closes #4374.
Problem
With
array.rectilinear_chunksenabled, zarr 3.2.0 and 3.2.1 classified a mixed chunk specification such aschunks=(2, (5, 10, 5))as regular (the classifier looked only at the first element) and stored{"name": "regular", "configuration": {"chunk_shape": [2, [5, 10, 5]]}}while laying the chunks out as a rectilinear grid. 3.2.x also wrote
[2, [5.0, 10.0, 5.0]]for float edges and[true, [5, 10, 5]]for aTruesize. The classifier was fixed in #4218, but 3.4.0 still accepts such a grid inRegularChunkGridMetadata, writes it, and then fails with an unrelatedTypeError: Expected an iterable of integerswhen building the codec pipeline, so these arrays cannot be read.Changes
Read the 3.2.x mixed grid, without the flag
The reading is one more upgrade in
zarr.core.metadata.upgrades, the module #4334 introduces as the only lenient reader of stored documents. A storedregulargrid whosechunk_shapemixes chunk sizes with flat lists of edges is read as the rectilinear grid it describes; its sizes and edges go through the same readings as any other stored chunk size or rectilinear edge (true→ 1,4.0→ 4). Achunk_shapemade only of lists, or with nested lists, was never written by any release and is left for the constructors to reject.This is a compatibility read, so it does not require
array.rectilinear_chunks(3.2.x wrote these with the flag set, but the user reading them now may not have it). It warns, because re-saving needs the flag:RegularChunkGridMetadatanow rejects an edge list with aTypeErrornaming the dimension, so this metadata can no longer be written.The flag gates store boundaries, not a class
In 3.4.0 the flag was checked in
RectilinearChunkGridMetadata.__post_init__. It now gates the two places a rectilinear chunk grid crosses the store boundary:check_storableruns inArrayV3Metadata.to_buffer_dictand, for every array in a group's consolidated metadata, inGroupMetadata.to_buffer_dict. Every serialization of array metadata for a store passes through one of them.ArrayV3Metadata.from_dictchecks a document whose grid is namedrectilinearbefore any upgrade runs (as onmain). The 3.2.x mixed grid is exempt, since its document names aregulargrid.Both raise
RectilinearChunksDisabledError, a new subclass ofValueError(so existingexcept ValueErrorhandlers still catch it), with a note naming the array and saying that nothing was stored or read. A consolidated member that fails the gate is named by its path in the consolidated metadata. Consolidated metadata keeps a mixed-grid member in the form it was stored (#4334's rule for upgraded members), so writing group metadata with the flag off does not fail because of it.Encode before destroy
Any operation that deletes store content now encodes every document it will write first, through one helper (
encode_documents), so metadata that cannot be stored (for example, a rectilinear grid with the flag off) fails with the store untouched:Array.resizeencodes the new metadata before deleting chunks outside the new shape. A resize on a store that cannot delete (ZipStore) still fails before storing new metadata, as in 3.4.0.del group[name]on a group with consolidated metadata encodes the group metadata without the member before deleting it. After the deletion it pops the member in place (so every handle sharing that consolidated metadata sees it, as in 3.4.0) and stores the encoding of the metadata as it then is, so concurrent deletions each store the deletions made before them.create_hierarchy(..., overwrite=True)builds and encodes every node before deleting anything.What users see
TypeError: Expected an iterable of integersupdate_attributesthat array with the flag offRectilinearChunksDisabledErrornaming the array; store untouchedRegularChunkGridMetadata(chunk_shape=(2, (5, 5)))TypeError: Dimension 1: chunk edge length must be an int, got (5, 5)RectilinearChunkGridMetadata(...)with the flag offValueErrorEvidence
zarr==3.2.0andzarr==3.2.1installs (a mixed grid, a mixed grid resized to 0 along its regular axis, a mixed grid with an empty rectilinear axis, a rectilinear grid on an empty axis) open, take an append without losing data, and re-save with the flag to metadata that reopens without a warning. A test fixture is copied verbatim from a 3.2.1-written document.check_patch.py(OK, 0 undocumented changes, 0 warnings-only changes), the byte-identical write matrix, and the xarray and VirtualiZarr runs described in fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334 cover this branch.Tests
tests/test_metadata/test_upgrades.py: one table for the mixed-grid reading (where the edge lists sit,trueand float sizes, a long edge list), one test per rejection (a grid of only edge lists, run-length pairs, a non-integer edge, an edge below 1, edges that do not sum to the extent), a round trip by re-save and by write, the store left untouched without the flag (resize, write,create_hierarchy(overwrite=True)), consolidating a group with the mixed member atmixedandsub/mixed, and group writes (delete, attributes) storing the member as stored with the flag off and on. The fixture is a document copied verbatim from a store zarr 3.2.1 wrote.tests/test_group.py:del group[name]on consolidated groups of both formats (reopened from the store), an aliased subgroup handle seeing the deletion, and concurrent deletions storing the member list the handle holds.tests/test_store/test_zip.py: aZipStoreresize that needs deletes leaves the array unchanged, with no duplicate entries.tests/test_unified_chunk_grid.py,tests/test_metadata/test_v3.py: a regular grid rejects edge lists; the metadata classes are not gated; a stored rectilinear document read without the flag names the array.Known limitations
del group[name]calls on one consolidated group can still store a stale member list, about as often as 3.4.0 does.create_hierarchygiven user-built metadata that already holds a 3.2.x mixed grid can store it as given.Stack and release notes
main(cd1e5b3) into this branch by way of the updated fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334, with no conflicts._read_chunk_sizedocstring and the test module; thelongcase oftest_read_edge_lists_in_regular_gridnow expects the stored 0 on its axis of length 4 to read as 1, not 4. The full test suite passes (12069 passed).zarr.create(..., chunks=[[5, 5]], overwrite=True)deleted the existing array before the store-boundary gate raised; 3.4.0 raised first and kept the array._create_v2,_create_v3andinit_arraynow encode the new metadata before_prepare_overwrite. The same change coverscreate_array(overwrite=True), which already deleted before validating in 3.4.0.test_store_untouched_without_flaggains two cases for this, and both fail without the fix. The mixed-grid reading marks the metadata for re-saving, since only 3.2.x reads the stored grid. The full test suite passes (12075), andcheck_patch.pyis OK.🤖 Generated with Claude Code