Skip to content

fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored - #4334

Merged
d-v-b merged 53 commits into
zarr-developers:mainfrom
d-v-b:fix/zero-length-single-invariant
Oct 1, 2026
Merged

d-v-b merged 53 commits into
zarr-developers:mainfrom
d-v-b:fix/zero-length-single-invariant

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

This PR tightens our representation of chunk grids to disallow the representation of 0-length chunks. This has compatibility implications: previous versions of zarr could access arrays with 0-length chunks and exercise a subset of the Zarr API against them (reading / writing attributes), but not write chunks. I'm still thinking about the right compatibility shim for this, hence the draft status,

Previous PRs that removed 0-length chunks caused issues downstream so I'd like some eyes on this change (cc @TomNicholas)

edit: this PR now includes routines for repairing old broken metadata documents, such as zarr v2 metadata with chunks containing 0 elements, and glitched rectilinear chunks metadata.

🤖 AI text below 🤖

A chunk edge length is now always an integer of at least 1, while an array extent may be 0. Stored metadata that older releases wrote with a chunk size of 0, false or true is read by one repair module, zarr.core.metadata.repair, and nothing that zarr 3.4.0 accepted and handled correctly is rejected, changed, or newly warns.

Situation 3.4.0 This PR
Zarr format 2 array with stored chunks: [0] on an axis that has grown opens; append stores no chunk and reads back fill values (silent data loss) opens with a warning; append stores the data and valid metadata
Zarr format 3 array with stored chunk_shape: [0] / [false] (3.0.x, 3.1.x) does not open opens; silent on an empty axis, warns on a grown one
Sharded array with outer chunk_shape: [0] (3.0.x, 3.1.x) does not open opens; the outer chunk is the inner chunk size
Stored JSON true as a chunk size kept as True, written back as true read as 1, silently
create_array(shape=(0, 20), chunks=(5, 5), shards=-1) ValueError (not divisible) shard shape (5, 20)
create_array(shape=(0,), chunks=[[5, 10, 5]]) (rectilinear, flag on) ValueError; zarr 3.2 allowed it the edges describe the chunks the axis grows into

A stored chunk size of 0 or false is read as 1 (the inner chunk size when sharded), whatever the axis length, so the reading does not depend on how far the axis has grown. On an axis of positive length it warns: no chunk can have been stored, so the array holds only its fill value, and the warning says how to recreate it with the chunk shape you want (zarr.from_array(array.store, name=array.path, data=array, chunks=..., overwrite=True, write_data=False)) or how to keep it by re-saving its metadata. Writing data to such an array first stores the repaired metadata, so every reader finds the chunks. true chunk sizes and the integral float edges zarr 3.2 wrote are read silently as the values they equal.

Design
  • One invariant, one full-span rule. FixedDimension requires size >= 1. "One chunk spanning an axis of length span" is full_span_chunk_size(span, unit) = unit * max(1, ceildiv(span, unit)), where unit is the inner chunk size for a shard and 1 otherwise; every spelling (chunks=-1, False, "auto", shards=-1, shards=False) uses it. A rectilinear dimension of extent 0 keeps any non-empty list of positive edges.
  • One repair module. zarr.core.metadata.repair is the only place invalid stored documents are read leniently. Each repair is a pure function from a stored JSON document to a valid one; ArrayV2Metadata.from_dict and ArrayV3Metadata.from_dict apply them before the constructors, so single arrays, group members and consolidated metadata all read the same way. Stored 0/false → unit (marked, since chunks move); true → 1 and integral float rectilinear edges → int (silent, nothing moves); any other float is rejected. A warning is given once per document, after the repaired document has passed the constructor, naming the array by its path.
  • Strict vs lenient. Constructors keep 3.4.0's acceptance in this patch (ArrayV2Metadata(chunks=(0,)), NumPy integers in ArrayV2Metadata, ShardingCodec(chunk_shape=...)). One rule, parse_chunk_edge, checks every chunk size, edge and run-length size in metadata (int of at least 1, bool read as its int). Strict constructors are fix(metadata)!: require int chunk sizes of at least 1 in metadata constructors and chunk specifications #4431, for 3.5.0.
  • Writing after a repaired read. Metadata read from a document whose repair moves chunks remembers the stored document (_stored_document). Before its first non-empty chunk write, the array re-reads the stored document; if it is gone, chunks are written as in 3.4.0; if its chunk layout differs from the handle's, the write raises asking to reopen and stores nothing; if it still needs repairing, the repair is stored through an internal diff-and-upsert. update_attributes, resize and append store the repair as before.
  • Groups never touch member documents. Group writes store only the group's own documents. The consolidated copy of a repaired member is written back as it was stored, so readers of the consolidated metadata repair it again; the member's own first chunk write is the only place its document changes on disk.
Compatibility evidence
  • Legacy stores: 206 stores written by real installs of zarr 2.18.7, 3.0.10, 3.1.6, 3.2.0, 3.2.1, 3.3.0 and 3.4.0 (both formats, every full-span spelling on empty axes, explicit zeros, grown axes, sharded outer zeros, consolidated groups) all open on this branch, take an append without losing data, and reopen with warnings as errors afterwards; the final documents validate against the spec. zarr 3.4.0 opens 178 of them and leaves 80 documents invalid after the append.
  • Write matrix: 270 chunk/shard spellings per format against a real 3.4.0 install: no combination 3.4.0 created now fails, no stored grid document differs, no document violates the spec; 41 combinations that 3.4.0 rejected now work (zero-length rectilinear axes, non-multiple shards=-1).
  • Downstream: xarray main's zarr backend tests, 1560 passed on both 3.4.0 and this branch; VirtualiZarr main, 476 passed with the same environment failures on both.
  • Other implementations: tensorstore and zarrs read every array this branch writes (19 arrays, both formats, empty and grown axes, sharded); zarrs also reads rectilinear arrays created on an empty axis.
Tests
  • tests/test_metadata/test_repair.py: one table of stored documents, their repairs and warnings; one test per rejection; round trips over both formats; the first-write re-save (stale handle keeping newer metadata, changed chunk layout, empty write, missing document); groups (consolidated copy stored as read, concurrent deletions, member's first write); the recreate recipe from the warning.
  • tests/test_metadata/test_io.py: the internal diff and upsert.
  • tests/test_array_stateful.py (ArrayLifecycle): create, append, resize to and from 0, write and re-save against a NumPy model, over both formats, every chunk spelling, sharding, rectilinear grids and stored zeros; checks that no chunk sits under a still-invalid document. Runs in the slow hypothesis job.
Stack and release notes

🤖 Generated with Claude Code

@d-v-b
d-v-b force-pushed the fix/zero-length-single-invariant branch from 7c5688b to 302b638 Compare September 9, 2026 18:28
…, clamps and metadata

Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0,
in which case the dimension has zero chunks (ceildiv(0, size) == 0).

Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, zarr-developers#972,
zarr-developers#1977, zarr-developers#2434, zarr-developers#3711, zarr-developers#4305, zarr-developers#4307, zarr-developers#4328) because the layers disagreed on
this invariant and every span-derived chunk spelling clamped on its own:

- The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1,
  but the in-memory FixedDimension allowed size == 0 with four special-case
  branches left over from zarr-developers#2434, so normalization could build a grid the
  metadata constructor then rejected. FixedDimension now rejects size < 1
  and the four `if self.size == 0` branches are gone. VaryingDimension
  already required edges > 0 and is unchanged.
- `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both
  the typesize == 0 early return and the np.maximum line) and `shards="auto"`
  each derived "one chunk covering the axis" independently. They now all go
  through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is
  the single definition of that phrase for a possibly zero-length axis.
- Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]`
  document opened fine and read uninitialised memory after a resize. It now
  raises a clear ValueError at parse time, matching the format 3 grid.
- Rectilinear grids had no creation-time spelling for a zero-length axis:
  normalize_chunks_1d required sum(edges) == span, which no list of positive
  edges can satisfy for span 0, even though the same state is reachable via
  resize((0,)) and round-trips through reopen. For span == 0 any non-empty
  list of positive edges is now accepted verbatim, producing the same
  VaryingDimension(edges, extent=0) that resize produces; the strict sum
  check is kept for span > 0.

Tests: the per-spelling regression test from zarr-developers#4328 is replaced by one matrix
over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0),
(0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte
budget, explicit shards}, with separate small tests for each error case.
Tests that constructed FixedDimension(size=0) now assert it raises, and a
zero-extent test covers the behaviour the old special cases were guarding.

Assisted-by: ClaudeCode:claude-fable-5-1
zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)`
and for `chunks=(0,)`, so stores with that document exist. Rejecting them
at open would turn a previously-readable array into an error; leaving the
0 in place read uninitialised memory after a resize. Normalize the edge to
1 with a ZarrUserWarning instead — the same grid every other "one chunk
spans the axis" spelling produces — and keep rejecting a zero edge on an
axis that has data.

Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and
`chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write,
append, resize and reopen-then-read every raise ZeroDivisionError. There
was never a working behaviour to preserve; normalizing the edge to 1 makes
such arrays usable for the first time. Say so in the comment and fragment
instead of claiming the stores were previously readable.

Assisted-by: ClaudeCode:claude-fable-5-1
@d-v-b
d-v-b force-pushed the fix/zero-length-single-invariant branch from 302b638 to 047a92e Compare September 9, 2026 18:28
@codecov

codecov Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.88268% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 94.63%. Comparing base (d18fb50) to head (0992e94).

Files with missing lines Patch % Lines
src/zarr/core/metadata/repair.py 98.12% 3 Missing ⚠️
src/zarr/core/chunk_grids.py 93.33% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4334      +/-   ##
==========================================
+ Coverage   94.47%   94.63%   +0.15%     
==========================================
  Files          93       94       +1     
  Lines       13241    13487     +246     
==========================================
+ Hits        12510    12763     +253     
+ Misses        731      724       -7     
Files with missing lines Coverage Δ
src/zarr/core/_json.py 100.00% <100.00%> (ø)
src/zarr/core/array.py 98.29% <100.00%> (+0.21%) ⬆️
src/zarr/core/common.py 91.80% <100.00%> (+1.21%) ⬆️
src/zarr/core/group.py 95.63% <100.00%> (+0.03%) ⬆️
src/zarr/core/metadata/io.py 100.00% <100.00%> (ø)
src/zarr/core/metadata/v2.py 90.65% <100.00%> (+1.27%) ⬆️
src/zarr/core/metadata/v3.py 96.50% <100.00%> (+1.05%) ⬆️
src/zarr/core/chunk_grids.py 96.74% <93.33%> (-0.04%) ⬇️
src/zarr/core/metadata/repair.py 98.12% <98.12%> (ø)

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@d-v-b d-v-b added this to the 3.5.0 milestone Sep 16, 2026
@TomNicholas

Copy link
Copy Markdown
Member

I was slightly ahead of you here - I think as long as I release my VZ PR before you release this then VZ at least will not break as a result of this change in Zarr-python.

… format 3

The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.

`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.

Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).

Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.

Assisted-by: ClaudeCode:claude-opus-5
… in one routine

The legacy zero-chunk policy was written out twice: inline in
`ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2
copy also zipped with `strict=False` and re-appended any trailing chunk
entries only so that a separate length check, `parse_metadata`, could
report a dimensionality mismatch after construction.

`parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one
place a stored chunk shape is checked against its array's shape, for both
formats: one entry per axis, every integer chunk size at least 1, and a
size of 0 (or JSON `false`) on a zero-length axis read as 1 with a
warning that names the writer and how to re-save. Non-integer entries,
such as edge lists, pass through for the caller's own parser.

`ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is
gone. The Zarr format 3 adapter only locates a regular grid's
`chunk_shape` in the stored document and hands it over; it still runs in
`ArrayV3Metadata.__init__` because chunk grid metadata has no array shape.
The Zarr format 3 import changes that existed only for the old helper are
reverted.

Tests for the policy now target the routine: one table of valid and
legacy inputs, and one test per rejection (dimension mismatch, zero on a
non-empty axis, negative). They replace metadata-level tests in
test_v2.py and test_v3.py that only re-tested the same rules; the
end-to-end tests still cover both formats' wiring against stored arrays.

Assisted-by: ClaudeCode:claude-opus-5
…nk grids

`parse_stored_chunk_shape` passed non-integer entries through "for the
caller's own parser", which made a regular-grid policy look like a
general chunk shape routine and let it decide what a 0-length chunk means
for grids it does not own. A rectilinear grid, or any other grid, is free
to define its own semantics for 0-length chunks.

It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with
no pass-through, and its docstring says it applies to Zarr format 2
`chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller
hands it a chunk shape only when the grid is named `regular` and every
entry is an integer (`_is_regular_chunk_shape`); anything else is not a
regular chunk shape and goes to the chunk grid parser untouched.

Assisted-by: ClaudeCode:claude-opus-5
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
… into fix/mixed-regular-chunk-grid-4374

# Conflicts:
#	src/zarr/core/metadata/v3.py
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
…place

zarr-developers#4334 read a zero chunk size on an empty axis in `ArrayV3Metadata.__init__`
and zarr-developers#4375 read a 3.2.x mixed grid inside `parse_chunk_grid`, each with its
own predicate, warning text and re-save advice. Both are compatibility
readings of a stored `regular` grid, so `_read_stored_regular_chunk_grid`
now dispatches to both, `parse_chunk_grid` accepts only what the spec
allows, and both warnings use `RESAVE_METADATA_HINT`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
With zarr-developers#4334 a rectilinear dimension of extent 0 takes any non-empty list of
positive edges at creation, the same state 3.2.x wrote and resize(0)
produces. The strategies asserted extent > 0 and `chunk_grids` fell back to
a regular grid for any empty axis, so no property saw that state.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b and others added 3 commits September 25, 2026 12:28
…panning chunk

A stored chunk size of 0 was tolerated only on a zero-length axis and
rejected otherwise, because the metadata supposedly could not say how the
stored chunks were laid out. But a chunk size of 0 gives a grid of zero
chunks, so no release could store a chunk under it, and the writers of
that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array
created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records
a Zarr format 3 resize and the resize half of a failed append. Measured
with real installs. Those arrays open in 3.4.0, attributes included; the
rejection would have made them unopenable.

`parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`)
on any axis as one chunk spanning it, `max(extent, 1)`, which is what the
`-1`/`False` spec that wrote it meant. On a grown axis the warning also
says that data written to it was not saved. Negative sizes are still
rejected.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The only state machine that touched arrays compared zarr on one store with
zarr on a MemoryStore, so a chunk grid bug showed up identically on both
sides; it had no append rule, covered Zarr format 3 only, and kept every
empty axis at 0 when resizing, which is where the zero-length bugs live.

`ArrayLifecycle` checks one array against a NumPy model across both
formats, every chunk spelling (-1, False, "auto", ints, sharded,
rectilinear) and the stored chunk size of 0 that releases before 3.4
wrote, including on an axis those releases grew. Rules append, resize
(growing and shrinking to and from 0), write and re-save the metadata;
the invariant reopens the array and compares shape, values and whether
the legacy warning is due. Deliberately breaking the grown-axis policy,
the legacy warning, or append on an empty axis each fails it.

`resize` keeps partly retained chunks whole, so cells cut off by a shrink
can come back with their old values when the axis grows (as in 2.x); the
model marks such cells unknown until written instead of encoding that.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
… into fix/mixed-regular-chunk-grid-4374

# Conflicts:
#	src/zarr/core/metadata/v3.py
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
The warning for a stored chunk size of 0 said which zarr-python releases
wrote it. The check runs on every metadata construction, so metadata built
in code, such as VirtualiZarr's kerchunk writer passing an empty array's
shape as `chunks`, was told it came from zarr-python 2.x. The warning now
says what the chunk size is read as and, on an axis of positive length,
that the axis holds only the fill value. `legacy_writers` is gone, and
the docstrings no longer narrate release history; the changelog keeps it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 25, 2026
… into fix/mixed-regular-chunk-grid-4374

# Conflicts:
#	src/zarr/core/metadata/v2.py
#	src/zarr/core/metadata/v3.py
d-v-b and others added 3 commits September 26, 2026 19:34
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tadata

Setting attributes of an array read from a document that had to be
upgraded stored the upgraded document with the new attributes, but left
the metadata marked with the document it was read from. A consolidated
group handle shares that metadata, so its next write stored the old
document in the consolidated metadata and the new attributes were lost
from it (zarr 3.4.0 kept them).

Every write of an array's own documents (creating, resizing, setting
attributes) now goes through `AsyncArray._save_metadata`, which clears
the mark afterwards, as the first chunk write does after storing the
upgrade. Both clear it through `_stored_document_replaced`, the one
place that declares it: the first chunk write clears it also when it
stores nothing (the store already holds a valid document, or none), so
it cannot be folded into the save.

The deep copy in `mark_upgraded` stays: the metadata's attributes share
objects with the document it was read from, which the caller also
holds. A test pins it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ment mark

Setting attributes on or resizing an array read from a document that needed
an upgrade clears the `_stored_document` mark only after the store accepts the
new metadata. A test now injects a store failure for the array's own document
and checks the mark and the stored bytes survive, for Zarr formats 2 and 3.
The `_save_metadata` docstring now says only what the method does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 26, 2026
zarr-developers#4334 moved stored-document leniency into document upgrades
(`zarr.core.metadata.upgrades`) and made the metadata constructors
strict. Resolution: the zero-chunk-size branch of
`_read_stored_regular_chunk_grid` is dropped (upgrades.py owns it now);
the mixed-grid branch is kept unchanged so this commit changes no
behaviour of zarr-developers#4375, and its warning uses `RESAVE_HINT`. The regular and
rectilinear validators range-check with zarr-developers#4334's `parse_chunk_edge`, and
two tests expect its message.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 zarr-metadata | 🛠️ Build #34778775 | 📁 Comparing 30d7ca6 against latest (cad2561)

  🔍 Preview build  

3 files changed
± api/model/index.html
± api/pydantic/index.html
± api/v2/index.html

# Conflicts:
#	src/zarr/core/chunk_grids.py
@d-v-b d-v-b changed the title fix(chunk-grids): one invariant for zero-length axes across model, clamps, and metadata fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored Sep 29, 2026
A stored chunk size of 0 or `false` was read as one chunk spanning the
axis at open time. On an axis that grew after the size was written, that
made the chunk size depend on when the array was opened, and it could be
very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one
800 MB chunk. It is now read as the smallest valid chunk edge length (1,
or the inner chunk size for a shard), which is what `chunks=-1` gives on
a zero-length axis. A resize by software that kept the stored 0 no longer
changes the layout a stale handle expects.

Assisted-by: ClaudeCode:claude-opus-5-5
…SON input as is

A stored `true` chunk size or an integral float edge is read as the value
zarr already read it as, so chunks written under the upgraded metadata are
where every reader looks. Marking such documents for re-saving turned a
plain chunk write into a metadata write: in a ZipStore that adds a second
`zarr.json` entry (a `UserWarning`, an error under `-W error`), and it
opened a race with concurrent metadata writes that 3.4.0 did not have. Each
upgrade now reports whether it moves chunks, and only those mark the
metadata.

The upgrades also detected changes by comparing JSON encodings, so
`ArrayV3Metadata.from_dict` given metadata built in code (a codec instance,
a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`.
Changes are now tracked where they are made.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b d-v-b modified the milestones: 3.5.0, 3.4.1 Sep 29, 2026
d-v-b and others added 2 commits September 30, 2026 14:41
…was 0

The warning for a stored chunk size of 0 on an axis of positive length now says that
the array holds only its fill value, so recreating it with the wanted chunk shape loses
nothing, and gives the re-save as the way to keep it instead. A re-save freezes the
smallest chunk size into an array that holds no data yet. Each reading now carries its
own advice, so mark_upgraded appends no shared hint.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ored chunk size of 0

The warning names the call, zarr.from_array(array.store, name=array.path, data=array,
chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value,
attributes and codecs, and a test pins the recipe on both formats.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 zarr-indexing | 🛠️ Build #34853483 | 📁 Comparing fcc24bb against latest (9677081)

  🔍 Preview build  

1 file changed
± release-notes/index.html

@d-v-b
d-v-b marked this pull request as ready for review October 1, 2026 08:54
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with
save_new_metadata: keep main's helper and this branch's imports and chunk
normalization.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
…egular-chunk-grid-4374

Resolves the overlap with zarr-developers#4410: save_new_metadata (main) replaces this branch's
encode-then-_prepare_overwrite at the create sites, and save_metadata and
save_new_metadata encode through encode_documents so errors name the node.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
…chunk-sizes

Resolves the overlap with zarr-developers#4410: save_new_metadata (main) replaces this branch's
encode-then-_prepare_overwrite at the create sites, and save_metadata and
save_new_metadata encode through encode_documents so errors name the node.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…airs, not upgrades

The module is zarr.core.metadata.repair: repair_array_document, mark_repaired,
ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test
module test_repair.py. A reading that turns an invalid stored document into a valid one
fixes it; it does not move it to a newer format.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
Conflicts resolved by keeping this branch's lines and applying the same rename to
them; this branch's own text now says repair as well.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
Conflicts resolved by keeping this branch's lines and applying the same rename to
them; this branch's own text now says repair as well.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@d-v-b
d-v-b merged commit a9370dd into zarr-developers:main Oct 1, 2026
40 checks passed
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
Resolves the two hunks with zarr-developers#4334: the single normalizer call takes zarr-developers#4334's unit, and
the test module keeps both imports.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch
already contains, so every file it touches keeps this branch's version; the merge
brings in only the dependency bump (zarr-developers#4461).

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Oct 1, 2026
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch
already contains, so every file it touches keeps this branch's version; the merge
brings in only the dependency bump (zarr-developers#4461).

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
d-v-b added a commit that referenced this pull request Oct 2, 2026
… space (#4377)

* fix(chunk-grids): enforce one zero-length-axis invariant across model, clamps and metadata

Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0,
in which case the dimension has zero chunks (ceildiv(0, size) == 0).

Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, #972,
#1977, #2434, #3711, #4305, #4307, #4328) because the layers disagreed on
this invariant and every span-derived chunk spelling clamped on its own:

- The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1,
  but the in-memory FixedDimension allowed size == 0 with four special-case
  branches left over from #2434, so normalization could build a grid the
  metadata constructor then rejected. FixedDimension now rejects size < 1
  and the four `if self.size == 0` branches are gone. VaryingDimension
  already required edges > 0 and is unchanged.
- `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both
  the typesize == 0 early return and the np.maximum line) and `shards="auto"`
  each derived "one chunk covering the axis" independently. They now all go
  through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is
  the single definition of that phrase for a possibly zero-length axis.
- Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]`
  document opened fine and read uninitialised memory after a resize. It now
  raises a clear ValueError at parse time, matching the format 3 grid.
- Rectilinear grids had no creation-time spelling for a zero-length axis:
  normalize_chunks_1d required sum(edges) == span, which no list of positive
  edges can satisfy for span 0, even though the same state is reachable via
  resize((0,)) and round-trips through reopen. For span == 0 any non-empty
  list of positive edges is now accepted verbatim, producing the same
  VaryingDimension(edges, extent=0) that resize produces; the strict sum
  check is kept for span > 0.

Tests: the per-spelling regression test from #4328 is replaced by one matrix
over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0),
(0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte
budget, explicit shards}, with separate small tests for each error case.
Tests that constructed FixedDimension(size=0) now assert it raises, and a
zero-extent test covers the behaviour the old special cases were guarding.

Assisted-by: ClaudeCode:claude-fable-5-1

* fix(metadata): read a legacy v2 zero chunk edge on an empty axis as 1

zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)`
and for `chunks=(0,)`, so stores with that document exist. Rejecting them
at open would turn a previously-readable array into an error; leaving the
0 in place read uninitialised memory after a resize. Normalize the edge to
1 with a ZarrUserWarning instead — the same grid every other "one chunk
spans the axis" spelling produces — and keep rejecting a zero edge on an
axis that has data.

Assisted-by: ClaudeCode:claude-fable-5-1

* chore: rename changelog fragment to the PR number

Assisted-by: ClaudeCode:claude-fable-5-1

* docs: state what 2.x actually did with a zero chunk edge

Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and
`chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write,
append, resize and reopen-then-read every raise ZeroDivisionError. There
was never a working behaviour to preserve; normalizing the edge to 1 makes
such arrays usable for the first time. Say so in the comment and fragment
instead of claiming the stores were previously readable.

Assisted-by: ClaudeCode:claude-fable-5-1

* docs: qualify zero-chunk compatibility history

Assisted-by: Codex:GPT-6

* test: sample the full rectilinear chunk grid declaration space

The rectilinear hypothesis strategy returned one list of edges per
dimension, so no property test ever saw a bare-int dimension, a grid
mixing bare ints and edge lists, a run-length encoded declaration, or
edges overhanging the extent. The first-element-only classifier that
zarr 3.2.x shipped (#4374) was invisible to it.

The strategies now cover two spaces. `rectilinear_chunks` samples the
`chunks=` syntax: bare ints and flat edge lists in any arrangement, at
least one list so the grid is rectilinear. `rectilinear_chunk_shape_
declarations` samples the stored metadata: bare-int steps (including
larger than the extent), edge lists written in full or run-length
encoded in canonical or arbitrary grouping, and overhanging edges. Each
draw comes with the chunk_shapes it must parse to. `chunk_grids` and so
`array_metadata` draw from the stored space; `arrays` passes grids the
list syntax cannot express as the metadata object, and asserts the
stored grid equals the declared one.

Two property tests pin the properties that would have caught #4374:
every stored declaration parses to its expanded edges and re-serializes
to an equivalent grid, and every `chunks=` specification is stored in
zarr.json as a "rectilinear" grid equal to the specification.

test_unified_chunk_grid.py used a private copy of the old strategy; it
now draws from the shared one.

Assisted-by: ClaudeCode:claude-fable-5-1

* docs: add changelog fragment for #4377

Assisted-by: ClaudeCode:claude-fable-5-1

* fix(metadata): read a legacy zero chunk size on an empty axis in Zarr format 3

The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.

`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.

Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).

Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.

Assisted-by: ClaudeCode:claude-opus-5

* refactor(metadata): check stored chunk shapes against the array shape in one routine

The legacy zero-chunk policy was written out twice: inline in
`ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2
copy also zipped with `strict=False` and re-appended any trailing chunk
entries only so that a separate length check, `parse_metadata`, could
report a dimensionality mismatch after construction.

`parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one
place a stored chunk shape is checked against its array's shape, for both
formats: one entry per axis, every integer chunk size at least 1, and a
size of 0 (or JSON `false`) on a zero-length axis read as 1 with a
warning that names the writer and how to re-save. Non-integer entries,
such as edge lists, pass through for the caller's own parser.

`ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is
gone. The Zarr format 3 adapter only locates a regular grid's
`chunk_shape` in the stored document and hands it over; it still runs in
`ArrayV3Metadata.__init__` because chunk grid metadata has no array shape.
The Zarr format 3 import changes that existed only for the old helper are
reverted.

Tests for the policy now target the routine: one table of valid and
legacy inputs, and one test per rejection (dimension mismatch, zero on a
non-empty axis, negative). They replace metadata-level tests in
test_v2.py and test_v3.py that only re-tested the same rules; the
end-to-end tests still cover both formats' wiring against stored arrays.

Assisted-by: ClaudeCode:claude-opus-5

* refactor(metadata): scope the stored chunk shape check to regular chunk grids

`parse_stored_chunk_shape` passed non-integer entries through "for the
caller's own parser", which made a regular-grid policy look like a
general chunk shape routine and let it decide what a 0-length chunk means
for grids it does not own. A rectilinear grid, or any other grid, is free
to define its own semantics for 0-length chunks.

It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with
no pass-through, and its docstring says it applies to Zarr format 2
`chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller
hands it a chunk shape only when the grid is named `regular` and every
entry is an integer (`_is_regular_chunk_shape`); anything else is not a
regular chunk shape and goes to the chunk grid parser untouched.

Assisted-by: ClaudeCode:claude-opus-5

* test: tighten rectilinear strategy contracts from review

- Document the `rectilinear_chunks` return type change in the changelog
  fragment, since `zarr.testing.strategies` is public.
- Make the `RectilinearDimDeclaration` alias private.
- Encode canonical RLE in the strategy instead of calling `compress_rle`,
  so the generator does not depend on the code under test.
- Assert `rectilinear_chunks` gets a non-empty shape instead of returning
  a non-rectilinear `[]`.

Assisted-by: ClaudeCode:claude-opus-5-5

* test: draw zero-length axes in the rectilinear declaration strategies

With #4334 a rectilinear dimension of extent 0 takes any non-empty list of
positive edges at creation, the same state 3.2.x wrote and resize(0)
produces. The strategies asserted extent > 0 and `chunk_grids` fell back to
a regular grid for any empty axis, so no property saw that state.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a stored zero chunk size on a grown axis as one spanning chunk

A stored chunk size of 0 was tolerated only on a zero-length axis and
rejected otherwise, because the metadata supposedly could not say how the
stored chunks were laid out. But a chunk size of 0 gives a grid of zero
chunks, so no release could store a chunk under it, and the writers of
that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array
created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records
a Zarr format 3 resize and the resize half of a failed append. Measured
with real installs. Those arrays open in 3.4.0, attributes included; the
rejection would have made them unopenable.

`parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`)
on any axis as one chunk spanning it, `max(extent, 1)`, which is what the
`-1`/`False` spec that wrote it meant. On a grown axis the warning also
says that data written to it was not saved. Negative sizes are still
rejected.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: a stateful test of one array's create/append/resize/write life

The only state machine that touched arrays compared zarr on one store with
zarr on a MemoryStore, so a chunk grid bug showed up identically on both
sides; it had no append rule, covered Zarr format 3 only, and kept every
empty axis at 0 when resizing, which is where the zero-length bugs live.

`ArrayLifecycle` checks one array against a NumPy model across both
formats, every chunk spelling (-1, False, "auto", ints, sharded,
rectilinear) and the stored chunk size of 0 that releases before 3.4
wrote, including on an axis those releases grew. Rules append, resize
(growing and shrinking to and from 0), write and re-save the metadata;
the invariant reopens the array and compares shape, values and whether
the legacy warning is due. Deliberately breaking the grown-axis policy,
the legacy warning, or append on an empty axis each fails it.

`resize` keeps partly retained chunks whole, so cells cut off by a shrink
can come back with their old values when the axis grows (as in 2.x); the
model marks such cells unknown until written instead of encoding that.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): stop the zero chunk size warning from naming writers

The warning for a stored chunk size of 0 said which zarr-python releases
wrote it. The check runs on every metadata construction, so metadata built
in code, such as VirtualiZarr's kerchunk writer passing an empty array's
shape as `chunks`, was told it came from zarr-python 2.x. The warning now
says what the chunk size is read as and, on an axis of positive length,
that the axis holds only the fill value. `legacy_writers` is gone, and
the docstrings no longer narrate release history; the changelog keeps it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(metadata): drop an orphaned comment in ArrayV2Metadata.__init__

The comment introduced a consistency check (`parse_metadata`) that this
branch removed; the check now lives in `parse_stored_regular_chunk_shape`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(chunk-grids): size a full-span shard as a multiple of the inner chunk

One definition of "one chunk spanning the axis", full_span_chunk_size(span, unit),
is the smallest positive multiple of unit covering span. shards=-1 and shards=False
now use the inner chunk size as unit, so a zero-length axis, or one whose length is
not a multiple of the inner chunk, gets a valid shard instead of a divisibility error.

The zero-length array test drops the explicit-size spellings and keeps the spellings
that derive a chunk from the span.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): read invalid stored chunk sizes in one module of document upgrades

The metadata constructors are strict again: a chunk edge length is an int of at
least 1, so metadata built in code with 0 or True raises, without a warning.

Stored documents are read leniently only in zarr.core.metadata.upgrades, whose
upgrades map a stored array document to a valid one before the constructors run.
ArrayV2Metadata.from_dict and ArrayV3Metadata.from_dict apply them, so opening an
array and parsing consolidated metadata take the same path. The first upgrade reads
a regular chunk size of 0 or JSON false as one chunk spanning the axis, a multiple of
the inner chunk size when the array is sharded (this opens sharded stores with an
outer chunk size of 0), and JSON true, which real releases wrote for chunks=(True,),
as 1. Its one warning says what is invalid, how it is read, and how to re-save,
including zarr.consolidate_metadata.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: rewrite the 4334 changelog fragment and trim restating docstrings

The fragment now covers the full-span shard rule, strict constructors, and the
stored-document upgrade (including JSON true) in two short paragraphs.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: model exactly what the store holds in the array lifecycle state machine

The model now tracks cells beyond the array's shape: a shrinking resize deletes
exactly the chunks outside the new grid, kept chunks keep their out-of-bounds cells,
and a write covering every in-bounds cell of an unsharded chunk resets its
out-of-bounds cells to the fill value. A resize that never deletes chunks now fails
the test.

Sharding is its own chunk spelling, a stored chunk size of 0 applies to sharded
arrays too (in the outer grid), re-saving metadata only runs when the store warns,
and the test takes its settings from the repository's hypothesis profiles under the
slow_hypothesis marker, like the other stateful tests.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(strategies): drop the overhang parameter and trim rectilinear events

No caller passed overhang=False, so edge lists that sum past the extent are
always drawn. rectilinear_chunk_grids now builds the metadata from the drawn
chunk shapes instead of parsing a drawn declaration with the code under test,
and only the events whose reach is reported remain.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: check rectilinear properties against the declaration, not zarr's parsing

test_create_array_stores_declared_rectilinear_chunks compared the stored grid
with a RectilinearChunkGridMetadata built from the same chunks=, so a
constructor that collapsed uniform edges to a bare int passed on both sides.
It now expands the stored JSON chunk_shapes with a test-local RLE expansion
and compares them with the raw chunks= list. The block indexing properties
work their block extents out from the declared chunk sizes instead of from
the chunk grid under test.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: round-trip data through a rectilinear zero-length axis as it grows

Draws rectilinear chunks= over a shape with a zero-length axis, grows that
axis by append or resize, also past the declared edges, and checks the values
against NumPy before and after reopening.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: draw the state machine's rectilinear chunks from the shared strategy

Replaces the local _rectilinear_dim copy with zarr.testing's
rectilinear_chunks, now that this branch builds on the state machine. Reach
for zero-length rectilinear axes is unchanged or higher.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): warn about an upgraded document only once it validates

An upgrade now returns the upgraded document and how it read it, and
`from_dict` warns with those readings after the metadata constructor
accepts the upgraded document. A stored document that is still invalid
after an upgrade raises its own error instead of first warning how it
was read.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): one per-axis rule for stored chunk sizes, one rule for edges

Stored regular chunk shapes, in both formats and in a sharding codec's
inner chunk shape, are read entry by entry by one rule in upgrades.py:
an int >= 1 is kept, `true` is read as 1 (zarr 3.0.10 also wrote it as
the inner chunk size of a shard, which the sharding codec used to read
silently), and 0 or `false` as one chunk spanning the axis. Integral
floats are not read: no release wrote a regular grid with them. A
document warns once, naming the array where the caller knows its path.

Metadata constructors check every chunk edge length with one function,
`parse_chunk_edge`, which tells a wrong type (TypeError) from a wrong
value (ValueError); this covers bare sizes, rectilinear edges, RLE sizes
and the sharding codec's inner chunk shape.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(chunk-grids): start chunk guessing from the one full-span rule

`_guess_regular_chunks` clamped zero-length axes with `np.maximum`,
restating `full_span_chunk_size` with unit 1. Fold the zero-span
normalizer tests into the `normalize_chunks_1d` tables.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: run the array lifecycle state machine in the slow Hypothesis job

Add it to `just hypothesis`. Re-saving metadata is always enabled and
must leave valid metadata unchanged; appends favour an axis stored with
chunk size 0, which raises how often a legacy axis grows before it is
re-saved. A fixed example pins that a write to a shard kept by a shrink
leaves its cells beyond the shape alone.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: draw distinct rectilinear dividers without rejection

`st.lists(..., unique=True)` rejected about 12% of draws per dimension
(Hypothesis invalid cases); drawing each divider by index into the unused
positions never rejects.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(testing): type rectilinear declarations with RectilinearDimSpecJSON

Reuse the JSON type from the metadata module instead of a private copy,
and type the test's stored grid document with RectilinearChunkGridMetadataJSON
so its type: ignore goes.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: draw the list form of rectilinear chunks= from rectilinear_chunks

`arrays()` decided whether a drawn rectilinear grid could be passed as
`chunks=` lists by restating the normalizer's sum rule. Draw the list
form from `rectilinear_chunks` instead, which also reaches it about
three times as often. Drop issue references from shipped testing code.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): store upgraded metadata before writing the first chunk

An array read from a stored document that the upgrades had to correct
(a chunk size of 0, `false` or `true`) wrote chunks under the corrected
layout while the store kept the old document, so zarr 3.2.1-3.4.0 then
read only the fill value, and tensorstore and zarrs could not open it.

`from_dict` now marks the metadata it read from an upgraded document
(`_stored_document_upgraded`, a field outside the document and
equality), and `AsyncArray._set_selection`, which every chunk write
goes through (`setitem` now included), stores that metadata first and
keeps the stored copy. Concurrent first writes store the same document.

The array lifecycle state machine now ends the legacy state on a write
and checks that no chunk is stored under a document that is still
upgraded on read; `resave_metadata` runs only while the stored document
is invalid, and re-saving valid metadata is checked once, at creation.
The changelog also says which releases stored a chunk size of 0 on an
axis of positive length.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): name an array one way in upgrade warnings

Warnings about an upgraded document name the array by its store path,
`str(StorePath(store, path))`, on every path that reads one: opening an
array, `AsyncArray.from_dict`, a group read with or without
consolidated metadata, and consolidated members, which were named
relative to the group. `GroupMetadata.from_dict` and
`ConsolidatedMetadata.from_dict` take the group's path for that.

The consolidated test now also opens the group with
`use_consolidated=False` and checks each array's name, so dropping the
path on either read path fails a test.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): one rule for bare sizes and edge lists; one noun in messages

- `_validate_chunk_shapes` decides "edge list or bare size" with one
  `match`: a list or tuple is an edge list, anything else is a bare size
  checked by `parse_chunk_edge`, so a string dimension such as `"10"` is
  rejected as not an int instead of being iterated per character.
  Stored rectilinear `from_dict` makes the same JSON split and passes the
  dimension to `expand_rle`, so an invalid edge or RLE count inside a
  list names its dimension.
- Messages say "dimension" throughout, lowercase after the colon
  ("Dimension 0: chunk edge length must be >= 1, got 0"); the upgrade
  warnings read "0 in dimension 0 as one chunk spanning the dimension".
- The upgrades test for JSON arrays with `list` only (a stored document
  holds no tuples), and `_read_chunk_size` says what `span=None` means.
- `ArrayV2Metadata.from_dict` copies the document once.
- The rejected stored chunk shapes get one test per error case.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(testing)!: remove chunks_param_from_rectilinear

`arrays()` now draws its rectilinear `chunks=` argument from
`rectilinear_chunks`, so nothing uses the metadata-to-`chunks=` converter.
The rectilinear strategies are experimental; the removal is noted in the
changelog.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(testing): drop the extent-1 special case in rectilinear_dim_edges

Both modes already draw [1] for an extent of 1.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(testing): bound the chunks a drawn rectilinear grid declares in total

Every dimension could draw up to 20 chunks, so a 3-d grid reached 8000
chunks; rejection-free divider draws made the large counts common, and
indexing property tests visit every chunk. Mean chunks per drawn grid went
from ~23 (before the rewrite) to ~75, with a tail past 5000, and
test_mask_indexing went from 3.2s to 11.8s CPU under the ci profile.

Cap the product at 400 chunks: 20 per dimension up to 2-d, 7 for 3-d,
4 for 4-d. Every declaration kind is still drawn.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): store the upgrade of the current stored document before writing

A handle read from an upgraded document upserts the upgrade of what the store
holds on its first non-empty chunk write: a stale handle no longer rolls back
newer metadata, and an empty write stores nothing. consolidate_metadata does
the same for each upgraded member before writing the consolidated document.

Both go through two internal primitives in zarr.core.metadata.io:
diff_documents compares a node's stored documents with those its metadata
would store, value by value, and upsert_metadata encodes first, then stores
only the documents that differ and returns the changes.

The upgraded-document flag is a ClassVar set on the instance by one helper,
mark_upgraded, which also warns; the upgrades are keyed by Zarr format; the
Metadata.to_dict and V2 from_dict changes are reverted. _append goes through
AsyncArray.setitem and _setitem is removed. A chunk shape that is not a list
or tuple is rejected as a whole, and every expand_rle error names its
dimension.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: prefer zero-length axes for stored chunk size 0 in the lifecycle machine

An empty write stores no metadata, so it no longer ends the legacy state.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(testing): declare the rectilinear chunk limits and per-axis draw once

The per-dimension limit of 20 chunks is now one constant, used by both
`_max_chunks_per_dim` and the `rectilinear_dim_edges` default, and
`rectilinear_chunk_grids` and `rectilinear_chunk_shape_declarations` draw
their `chunk_shapes` through one private helper. The strategies draw
exactly what they drew before.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): keep accepting the chunk sizes zarr 3.4.0 accepted in metadata constructors (patch release)

Nothing in a patch release may reject a chunk size that zarr 3.4.0 accepted.
The one chunk edge rule (`parse_chunk_edge`, which also reads RLE repeat
counts) now reads any integral number as the `int` it equals: an `int`, a
`bool`, a NumPy integer or an integral float, so rectilinear grids written
with float edges (`[[4.0, 2]]`) read again and the grid constructors accept
NumPy integers; a fractional float is still rejected. A regular chunk shape
may again be any iterable. `ArrayV2Metadata(chunks=...)` and
`ShardingCodec(chunk_shape=...)` read their chunk shape with
`parse_shapelike`, as in 3.4.0: a scalar, NumPy integers, bools and 0 are
accepted, and a 0 is written back as given (the kerchunk writer in
VirtualiZarr relies on `chunks=(0,)` for empty inlined arrays). Stored
documents with a chunk size of 0 are still read by the upgrades.

The strict rule returns in the next minor release.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): leave a valid stored document as written on a stale handle's first write

A handle read from an upgraded document re-reads the stored document before its
first chunk write. If that document no longer needs an upgrade, store nothing:
re-encoding it could differ from how another implementation wrote it, and
nothing about it needs fixing.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe the 4334 changes to metadata constructors for a patch release

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(testing): keep chunks_param_from_rectilinear in zarr.testing (patch release)

Revert 0d640f8: a patch release removes nothing from `zarr.testing`. The
removal returns in the next minor release.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(testing): keep rectilinear_chunks drawing an edge list per dimension (patch release)

A patch release keeps what `zarr.testing` strategies return. `rectilinear_chunks`
again returns `list[list[int]]`, an explicit edge list for every dimension, and
`[]` for a 0-d shape; it now also draws zero-length dimensions. The form mixing
bare-int steps and edge lists moves to the private `_rectilinear_chunks`, which
`arrays()`, `rectilinear_arrays()`, `block_test_arrays()` and the tests draw
from. The public strategy takes the mixed form in the next minor release.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe the 4377 strategy changes as additive for a patch release

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read float edges only in stored rectilinear documents

The patch release widened the one chunk edge rule to accept integral
floats so that stored rectilinear grids written with float edges
(`[[4.0, 2]]`, as zarr 3.2 wrote for float edges) keep opening. That also
made a stored regular chunk shape `[10.0]` open, which zarr 3.4.0 rejected
and the next minor release would reject again.

The constructor rule now accepts integers of any integer type (`int`,
`bool`, NumPy integers) and rejects floats. A new document upgrade reads
an integral JSON float of at least 1 as an `int` where zarr 3.2 stored
one: the explicit edges and run-length encoded sizes of a rectilinear
chunk grid, with the standard warning. Bare sizes, repeat counts, regular
chunk shapes and inner chunk shapes stay for the constructors to reject.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(codecs): pickle a ShardingCodec with an inner chunk size of 0

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): build every node before create_hierarchy deletes or stores anything

A node whose metadata no array or group can be built from now fails with
the store untouched, even with overwrite=True.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): warn only where the user must act; guard and refresh upgraded copies

- Readings that give what zarr 3.4.0 read (0/false on an empty axis, JSON
  true, float rectilinear edges) are silent; the metadata is still marked
  upgraded, so the first write re-saves it.
- JSON true is read as 1 in stored rectilinear edges and RLE sizes and in
  nested sharding codecs' inner chunk shapes.
- A write through a handle whose stored document now lays out chunks
  differently raises and stores nothing; with no stored document it writes
  as before.
- Storing group metadata refreshes each upgraded consolidated member from its
  own stored document (the one place: save_metadata); consolidate_metadata
  no longer needs its own pass.
- Metadata built in code with chunk size 0 is read through the upgrades, so
  an array can be built from it, as in 3.4.0.
- Chunk edges in metadata follow the 3.4.0 integer rule (int; bool read as
  its value); a string or mapping is rejected as a chunk shape as a whole.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: expect the lifecycle machine's upgrade warning only on non-empty axes

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: pin chunks_param_from_rectilinear returning lists, as zarr 3.4.0 did

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): leave a sharded chunk size of 0 unread when the inner shape is invalid

A stored outer chunk size of 0 of a sharded array is read in multiples of the
inner chunk size. When that inner size is itself 0 or `false`, the unit is
unknown: the upgrade now leaves the 0 for the constructor, which rejects it
with the same `ValueError` zarr 3.4.0 raised, instead of a ZeroDivisionError.
`_read_codec` now matches the sharding codec like `_inner_chunk_shape` does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(metadata): read each upgraded member once, concurrently, and adopt it

A group handle kept the consolidated copies of upgraded members flagged, so
every later group write re-read every such member, one after another, and
read each twice on the first write (once to parse, once to diff).

- `save_metadata` returns the metadata it stored; `AsyncGroup._save_metadata`
  (attrs, `update_attributes`, `delitem`, `consolidate_metadata`) and
  `Group.update_attributes_async` adopt it, so later writes read no member.
- The refresh visits only nested groups and flagged members, and reads them
  with one `asyncio.gather`. The documents it reads are the ones the upsert
  diffs against (`read_documents` + `upsert_metadata(..., stored)`), and the
  array write path does the same.
- `parse_stored_array` (documents -> silently marked metadata) replaces
  `read_stored_array`'s `(metadata, bool)` tuple, and is the one "read, mark,
  don't warn" path, also for code-built metadata.
- A flagged member whose own document is gone or cannot be read (invalid,
  replaced by a group) keeps its consolidated copy instead of failing every
  group write.
- `parse_array_metadata` of a metadata object reads it as a stored document
  only when `ArrayV2Metadata.chunks` holds a 0, the one chunk size that
  constructors still accept and the upgrades change. This removes the
  per-construction `to_dict` + upgrade (AsyncArray() back to 3.4.0 speed)
  and building arrays from codec configurations holding NumPy scalars
  works again, as in 3.4.0.
- `create_hierarchy` documents that it stores a group's consolidated
  metadata as given.
- Tests: second group write reads no member; nested member refresh; member
  without a readable document; NumPy-scalar codec configuration; stale
  handle whose inner chunk shape alone changed (kills `_chunk_layout`
  returning only the outer grid).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): adopt refreshed consolidated members in place

A group write replaced the handle's metadata with a copy holding the
consolidated members it read again, so the handle no longer shared its
consolidated dicts with subgroup handles taken earlier: an array deleted
through such a subgroup was stored again by the parent's next write, and
reappeared on reopen.

`_refresh_consolidated` now returns, along with the copy it encodes, the
members it replaced (with the consolidated dict that holds each), and
`save_metadata` writes them back into those dicts once the store succeeds,
unless the member was deleted or replaced meanwhile. `save_metadata` returns
nothing again, and the group callers keep their metadata objects. A failed
store leaves the members flagged, so a retry reads them again.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): raise the rectilinear flag error when refreshing a consolidated member

A group write reads each upgraded consolidated member again from its own
document, and kept the consolidated copy when that document could not be
read. That also swallowed the rectilinear chunks flag error, so a member
whose document now declares a rectilinear grid had its stale copy stored
as valid.

The flag error is now `RectilinearChunksDisabledError`, a `ValueError`,
and the refresh lets it propagate: the group write raises and stores
nothing. Unreadable documents keep their consolidated copy as before.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): store a group's documents only; keep upgraded members as stored

Group writes (attributes, `update_attributes`, deleting a member,
`consolidate_metadata`, `create_hierarchy`) read each upgraded member of
the consolidated metadata again from its own document and stored its
upgrade. Those member writes raced concurrent deletions: two concurrent
`delitem`s on a consolidated group recreated a deleted array's metadata
document in most runs, which zarr 3.4.0 never did.

A group write now stores only the group's own documents and reads none.
Array metadata read from a document that had to be upgraded keeps that
document (`_stored_document`, set by `mark_upgraded`, replacing the
`_stored_document_upgraded` flag), and consolidated metadata stores such
a member as it was stored. Every reader of the consolidated metadata then
reads the member as upgraded again, and the array's own first chunk write
stores the upgrade of its current document (or refuses a changed chunk
grid) as before: the only write that upgrades a member document.

Removed: the refresh and adoption of consolidated members in
`save_metadata` (`_refresh_consolidated`, `_refresh_array`,
`_Refreshed`), and `Group.update_attributes_async`'s detour through
`save_metadata`. Tests pinning the refresh are replaced by tests that a
group write reads nothing and writes only the group's documents, stores
the member's consolidated copy byte for byte as stored, that concurrent
deletions leave no member, and that the first write through consolidated
metadata stores the member's upgrade or refuses a changed grid.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(group): say what consolidated metadata stores for an upgraded array

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): clear the stored document whenever an array stores its metadata

Setting attributes of an array read from a document that had to be
upgraded stored the upgraded document with the new attributes, but left
the metadata marked with the document it was read from. A consolidated
group handle shares that metadata, so its next write stored the old
document in the consolidated metadata and the new attributes were lost
from it (zarr 3.4.0 kept them).

Every write of an array's own documents (creating, resizing, setting
attributes) now goes through `AsyncArray._save_metadata`, which clears
the mark afterwards, as the first chunk write does after storing the
upgrade. Both clear it through `_stored_document_replaced`, the one
place that declares it: the first chunk write clears it also when it
stores nothing (the store already holds a valid document, or none), so
it cannot be folded into the save.

The deep copy in `mark_upgraded` stays: the metadata's attributes share
objects with the document it was read from, which the caller also
holds. A test pins it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(metadata): pin that a failed metadata save keeps the stored document mark

Setting attributes on or resizing an array read from a document that needed
an upgrade clears the `_stored_document` mark only after the store accepts the
new metadata. A test now injects a store failure for the array's own document
and checks the mark and the stored bytes survive, for Zarr formats 2 and 3.
The `_save_metadata` docstring now says only what the method does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a stored chunk size of 0 as 1, however long the axis

A stored chunk size of 0 or `false` was read as one chunk spanning the
axis at open time. On an axis that grew after the size was written, that
made the chunk size depend on when the array was opened, and it could be
very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one
800 MB chunk. It is now read as the smallest valid chunk edge length (1,
or the inner chunk size for a shard), which is what `chunks=-1` gives on
a zero-length axis. A resize by software that kept the stored 0 no longer
changes the layout a stale handle expects.

Assisted-by: ClaudeCode:claude-opus-5-5

* fix(metadata): re-save only upgrades that move chunks, and read non-JSON input as is

A stored `true` chunk size or an integral float edge is read as the value
zarr already read it as, so chunks written under the upgraded metadata are
where every reader looks. Marking such documents for re-saving turned a
plain chunk write into a metadata write: in a ZipStore that adds a second
`zarr.json` entry (a `UserWarning`, an error under `-W error`), and it
opened a race with concurrent metadata writes that 3.4.0 did not have. Each
upgrade now reports whether it moves chunks, and only those mark the
metadata.

The upgrades also detected changes by comparing JSON encodings, so
`ArrayV3Metadata.from_dict` given metadata built in code (a codec instance,
a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`.
Changes are now tracked where they are made.

Assisted-by: ClaudeCode:claude-opus-5-5

* fix(metadata): recommend recreating an array whose stored chunk size was 0

The warning for a stored chunk size of 0 on an axis of positive length now says that
the array holds only its fill value, so recreating it with the wanted chunk shape loses
nothing, and gives the re-save as the way to keep it instead. A re-save freezes the
smallest chunk size into an array that holds no data yet. Each reading now carries its
own advice, so mark_upgraded appends no shared hint.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(metadata): give the recipe for recreating an array read from a stored chunk size of 0

The warning names the call, zarr.from_array(array.store, name=array.path, data=array,
chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value,
attributes and codecs, and a test pins the recipe on both formats.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(metadata): call the lenient readings of stored documents repairs, not upgrades

The module is zarr.core.metadata.repair: repair_array_document, mark_repaired,
ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test
module test_repair.py. A reading that turns an invalid stored document into a valid one
fixes it; it does not move it to a newer format.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b added a commit that referenced this pull request Oct 2, 2026
…by zarr 3.2.x (#4375)

* fix(chunk-grids): enforce one zero-length-axis invariant across model, clamps and metadata

Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0,
in which case the dimension has zero chunks (ceildiv(0, size) == 0).

Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, #972,
#1977, #2434, #3711, #4305, #4307, #4328) because the layers disagreed on
this invariant and every span-derived chunk spelling clamped on its own:

- The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1,
  but the in-memory FixedDimension allowed size == 0 with four special-case
  branches left over from #2434, so normalization could build a grid the
  metadata constructor then rejected. FixedDimension now rejects size < 1
  and the four `if self.size == 0` branches are gone. VaryingDimension
  already required edges > 0 and is unchanged.
- `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both
  the typesize == 0 early return and the np.maximum line) and `shards="auto"`
  each derived "one chunk covering the axis" independently. They now all go
  through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is
  the single definition of that phrase for a possibly zero-length axis.
- Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]`
  document opened fine and read uninitialised memory after a resize. It now
  raises a clear ValueError at parse time, matching the format 3 grid.
- Rectilinear grids had no creation-time spelling for a zero-length axis:
  normalize_chunks_1d required sum(edges) == span, which no list of positive
  edges can satisfy for span 0, even though the same state is reachable via
  resize((0,)) and round-trips through reopen. For span == 0 any non-empty
  list of positive edges is now accepted verbatim, producing the same
  VaryingDimension(edges, extent=0) that resize produces; the strict sum
  check is kept for span > 0.

Tests: the per-spelling regression test from #4328 is replaced by one matrix
over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0),
(0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte
budget, explicit shards}, with separate small tests for each error case.
Tests that constructed FixedDimension(size=0) now assert it raises, and a
zero-extent test covers the behaviour the old special cases were guarding.

Assisted-by: ClaudeCode:claude-fable-5-1

* fix(metadata): read a legacy v2 zero chunk edge on an empty axis as 1

zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)`
and for `chunks=(0,)`, so stores with that document exist. Rejecting them
at open would turn a previously-readable array into an error; leaving the
0 in place read uninitialised memory after a resize. Normalize the edge to
1 with a ZarrUserWarning instead — the same grid every other "one chunk
spans the axis" spelling produces — and keep rejecting a zero edge on an
axis that has data.

Assisted-by: ClaudeCode:claude-fable-5-1

* chore: rename changelog fragment to the PR number

Assisted-by: ClaudeCode:claude-fable-5-1

* docs: state what 2.x actually did with a zero chunk edge

Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and
`chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write,
append, resize and reopen-then-read every raise ZeroDivisionError. There
was never a working behaviour to preserve; normalizing the edge to 1 makes
such arrays usable for the first time. Say so in the comment and fragment
instead of claiming the stores were previously readable.

Assisted-by: ClaudeCode:claude-fable-5-1

* docs: qualify zero-chunk compatibility history

Assisted-by: Codex:GPT-6

* fix: read mixed regular/rectilinear chunk grids written by zarr 3.2.x

zarr 3.2.0 and 3.2.1 classified mixed chunk specs like (2, (5, 10, 5)) as
regular and stored them as a "regular" grid whose chunk_shape contains an
edge list, while laying the chunks out as a rectilinear grid. Newer versions
accepted that metadata and failed later with an unrelated TypeError.

RegularChunkGridMetadata now rejects edge lists. When reading stored
metadata, a "regular" grid with edge lists is read as the rectilinear grid
it describes, with a warning explaining how to re-save it; if rectilinear
chunks are disabled, the error says what happened and how to enable them.

Closes #4374

Assisted-by: ClaudeCode:claude-opus-5

* docs: add changelog fragment for #4375

Assisted-by: ClaudeCode:claude-opus-5

* fix: accept numpy integers in regular chunk grids; normalize tuples in the 3.2.x shim

Review follow-ups for #4375:

- RegularChunkGridMetadata now accepts numpy integer scalars and stores
  Python ints, mirroring parse_shapelike. The strict integer check had
  reported np.int64(2) as if it were a list of chunk edges.
- The mixed "regular" grid reader converts tuple entries to lists before
  delegating to RectilinearChunkGridMetadata.from_dict, so metadata dicts
  built in Python are handled the same as parsed JSON.
- The error and warning share one message prefix.
- Tests cover numpy ints, run-length encoded edges, and tuple input, and
  the test section header says the reader is broader than the 3.2.x bug.

Assisted-by: ClaudeCode:claude-fable-5-1

* refactor: branch on integer vs sequence when parsing rectilinear chunk shapes

The chunk_shapes parsers checked exact types: `from_dict` accepted a
dimension only as `int` or `list`, `expand_rle` accepted an RLE pair only
as a `list`, and the reader for 3.2.x mixed grids special-cased `tuple` to
convert it to a list before handing it to `from_dict`. A metadata dict
built in Python holds tuples where parsed JSON holds lists, and a numpy
array is as good as either, so these checks rejected valid input.

One structural predicate, `declares_chunk_edges`, now answers "is this a
sequence of edges rather than a single integer size" for all of them:
any non-integer iterable except `str`/`bytes`. It is a `TypeGuard`, not a
`TypeIs`, because it is False for strings, which are iterable. The four
sites that make that decision use it: `parse_chunk_grid`'s detection of a
mixed grid, the mixed-grid reader, `RectilinearChunkGridMetadata.from_dict`,
and `expand_rle`. The shim no longer needs its tuple special case.

Integers are `int | np.integer` throughout, matching `_parse_chunk_shape`,
and every parser stores Python ints: `_validate_chunk_shapes` now
coerces, so constructing a rectilinear grid from numpy values works the
way it already did for a regular grid.

Assisted-by: ClaudeCode:claude-opus-5

* docs: note the rectilinear parsing change in the changelog fragment

Assisted-by: ClaudeCode:claude-opus-5

* refactor: validate each chunk grid kind once, in one place

Parsing a regular chunk shape repeated its work at two levels:

- `_parse_chunk_shape` type-checked and coerced every dimension, then
  handed the result to `_validate_chunk_shapes`, which re-ran the same
  isinstance test and coercion and added only the `>= 1` check. Its
  edge-list branch was unreachable from this caller, which is why a
  `cast` was needed on the way out.
- `RegularChunkGridMetadata.from_dict` parsed the chunk shape and passed
  it to the constructor, whose `__post_init__` parsed it again.

Together that was four passes over the dimensions for `from_dict`. It
is now one: `_parse_chunk_shape` checks the range itself and no longer
calls the rectilinear validator, and `from_dict` hands the dimensions to
the constructor unparsed. Beyond the redundancy, the shared validator
was the coupling that let a rectilinear chunk shape be stored as a
regular grid (#4374), so the two grid kinds now validate separately.

`RectilinearChunkGridMetadata.from_dict` also had its own `>= 1` check
for bare-int dimensions, duplicating `_validate_chunk_shapes`, which
`__post_init__` runs over the result anyway. It now only puts the JSON
into shape (integer vs sequence, RLE expansion), and a bad bare int is
reported by the validator, which names the dimension. `expand_rle` keeps
its own checks because it is called directly.

Assisted-by: ClaudeCode:claude-opus-5

* fix(metadata): read a legacy zero chunk size on an empty axis in Zarr format 3

The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.

`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.

Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).

Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.

Assisted-by: ClaudeCode:claude-opus-5

* refactor(metadata): check stored chunk shapes against the array shape in one routine

The legacy zero-chunk policy was written out twice: inline in
`ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2
copy also zipped with `strict=False` and re-appended any trailing chunk
entries only so that a separate length check, `parse_metadata`, could
report a dimensionality mismatch after construction.

`parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one
place a stored chunk shape is checked against its array's shape, for both
formats: one entry per axis, every integer chunk size at least 1, and a
size of 0 (or JSON `false`) on a zero-length axis read as 1 with a
warning that names the writer and how to re-save. Non-integer entries,
such as edge lists, pass through for the caller's own parser.

`ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is
gone. The Zarr format 3 adapter only locates a regular grid's
`chunk_shape` in the stored document and hands it over; it still runs in
`ArrayV3Metadata.__init__` because chunk grid metadata has no array shape.
The Zarr format 3 import changes that existed only for the old helper are
reverted.

Tests for the policy now target the routine: one table of valid and
legacy inputs, and one test per rejection (dimension mismatch, zero on a
non-empty axis, negative). They replace metadata-level tests in
test_v2.py and test_v3.py that only re-tested the same rules; the
end-to-end tests still cover both formats' wiring against stored arrays.

Assisted-by: ClaudeCode:claude-opus-5

* refactor(metadata): scope the stored chunk shape check to regular chunk grids

`parse_stored_chunk_shape` passed non-integer entries through "for the
caller's own parser", which made a regular-grid policy look like a
general chunk shape routine and let it decide what a 0-length chunk means
for grids it does not own. A rectilinear grid, or any other grid, is free
to define its own semantics for 0-length chunks.

It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with
no pass-through, and its docstring says it applies to Zarr format 2
`chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller
hands it a chunk shape only when the grid is named `regular` and every
entry is an integer (`_is_regular_chunk_shape`); anything else is not a
regular chunk shape and goes to the chunk grid parser untouched.

Assisted-by: ClaudeCode:claude-opus-5

* refactor(metadata): read both legacy regular chunk grid forms in one place

#4334 read a zero chunk size on an empty axis in `ArrayV3Metadata.__init__`
and #4375 read a 3.2.x mixed grid inside `parse_chunk_grid`, each with its
own predicate, warning text and re-save advice. Both are compatibility
readings of a stored `regular` grid, so `_read_stored_regular_chunk_grid`
now dispatches to both, `parse_chunk_grid` accepts only what the spec
allows, and both warnings use `RESAVE_METADATA_HINT`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a stored zero chunk size on a grown axis as one spanning chunk

A stored chunk size of 0 was tolerated only on a zero-length axis and
rejected otherwise, because the metadata supposedly could not say how the
stored chunks were laid out. But a chunk size of 0 gives a grid of zero
chunks, so no release could store a chunk under it, and the writers of
that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array
created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records
a Zarr format 3 resize and the resize half of a failed append. Measured
with real installs. Those arrays open in 3.4.0, attributes included; the
rejection would have made them unopenable.

`parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`)
on any axis as one chunk spanning it, `max(extent, 1)`, which is what the
`-1`/`False` spec that wrote it meant. On a grown axis the warning also
says that data written to it was not saved. Negative sizes are still
rejected.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: a stateful test of one array's create/append/resize/write life

The only state machine that touched arrays compared zarr on one store with
zarr on a MemoryStore, so a chunk grid bug showed up identically on both
sides; it had no append rule, covered Zarr format 3 only, and kept every
empty axis at 0 when resizing, which is where the zero-length bugs live.

`ArrayLifecycle` checks one array against a NumPy model across both
formats, every chunk spelling (-1, False, "auto", ints, sharded,
rectilinear) and the stored chunk size of 0 that releases before 3.4
wrote, including on an axis those releases grew. Rules append, resize
(growing and shrinking to and from 0), write and re-save the metadata;
the invariant reopens the array and compares shape, values and whether
the legacy warning is due. Deliberately breaking the grown-axis policy,
the legacy warning, or append on an empty axis each fails it.

`resize` keeps partly retained chunks whole, so cells cut off by a shrink
can come back with their old values when the axis grows (as in 2.x); the
model marks such cells unknown until written instead of encoding that.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): stop the zero chunk size warning from naming writers

The warning for a stored chunk size of 0 said which zarr-python releases
wrote it. The check runs on every metadata construction, so metadata built
in code, such as VirtualiZarr's kerchunk writer passing an empty array's
shape as `chunks`, was told it came from zarr-python 2.x. The warning now
says what the chunk size is read as and, on an axis of positive length,
that the axis holds only the fill value. `legacy_writers` is gone, and
the docstrings no longer narrate release history; the changelog keeps it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): stop the mixed chunk grid messages from naming releases

The warning and error for a stored `regular` grid that lists chunk edges
said which zarr releases wrote it. They now say what is wrong with the
metadata, that it lists chunk edges only a rectilinear grid can declare,
and how it is read. The reader's docstrings no longer narrate release
history either; the changelog fragment keeps it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(metadata): drop an orphaned comment in ArrayV2Metadata.__init__

The comment introduced a consistency check (`parse_metadata`) that this
branch removed; the check now lives in `parse_stored_regular_chunk_shape`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(chunk-grids): size a full-span shard as a multiple of the inner chunk

One definition of "one chunk spanning the axis", full_span_chunk_size(span, unit),
is the smallest positive multiple of unit covering span. shards=-1 and shards=False
now use the inner chunk size as unit, so a zero-length axis, or one whose length is
not a multiple of the inner chunk, gets a valid shard instead of a divisibility error.

The zero-length array test drops the explicit-size spellings and keeps the spellings
that derive a chunk from the span.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): read invalid stored chunk sizes in one module of document upgrades

The metadata constructors are strict again: a chunk edge length is an int of at
least 1, so metadata built in code with 0 or True raises, without a warning.

Stored documents are read leniently only in zarr.core.metadata.upgrades, whose
upgrades map a stored array document to a valid one before the constructors run.
ArrayV2Metadata.from_dict and ArrayV3Metadata.from_dict apply them, so opening an
array and parsing consolidated metadata take the same path. The first upgrade reads
a regular chunk size of 0 or JSON false as one chunk spanning the axis, a multiple of
the inner chunk size when the array is sharded (this opens sharded stores with an
outer chunk size of 0), and JSON true, which real releases wrote for chunks=(True,),
as 1. Its one warning says what is invalid, how it is read, and how to re-save,
including zarr.consolidate_metadata.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: rewrite the 4334 changelog fragment and trim restating docstrings

The fragment now covers the full-span shard rule, strict constructors, and the
stored-document upgrade (including JSON true) in two short paragraphs.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: model exactly what the store holds in the array lifecycle state machine

The model now tracks cells beyond the array's shape: a shrinking resize deletes
exactly the chunks outside the new grid, kept chunks keep their out-of-bounds cells,
and a write covering every in-bounds cell of an unsharded chunk resets its
out-of-bounds cells to the fill value. A resize that never deletes chunks now fails
the test.

Sharding is its own chunk spelling, a stored chunk size of 0 applies to sharded
arrays too (in the outer grid), re-saving metadata only runs when the store warns,
and the test takes its settings from the repository's hypothesis profiles under the
slow_hypothesis marker, like the other stateful tests.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): warn about an upgraded document only once it validates

An upgrade now returns the upgraded document and how it read it, and
`from_dict` warns with those readings after the metadata constructor
accepts the upgraded document. A stored document that is still invalid
after an upgrade raises its own error instead of first warning how it
was read.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a regular grid listing chunk edges as a document upgrade

A `regular` chunk grid whose `chunk_shape` holds flat lists of integer
edge lengths, e.g. `[2, [5, 10, 5]]`, is now read by one more entry in
`V3_ARRAY_UPGRADES` instead of by a reader in `ArrayV3Metadata.__init__`.
It applies on every path that parses a stored document, including
consolidated metadata, and does not require `array.rectilinear_chunks`.
Only what such documents hold is upgraded (integer sizes and flat integer
edge lists); run-length encoded or non-integer entries are left for the
regular grid to reject. The warning quotes a bounded prefix of the chunk
shape and is emitted only after the upgraded document validates.

The rectilinear flag moves off the chunk grid metadata constructor to
the stored-document boundary: `ArrayV3Metadata.from_dict` checks the
document as stored (before upgrades) and `ArrayV3Metadata.to_buffer_dict`
checks the document about to be stored. Reading or creating a genuine
rectilinear array still requires the flag, as before; re-saving an
upgraded array stores a rectilinear grid and so requires it too, which
the warning says.

Tests: the upgrade's table test and one test per rejected form, using a
document copied verbatim from a store zarr 3.2.1 wrote.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): drop numpy and tuple widening from chunk grid metadata

The chunk grid metadata classes take typed Python values again: a
regular chunk shape of `int`s, rectilinear dimensions of `int` or
`list` in stored documents. `declares_chunk_edges` and the numpy and
tuple handling in `expand_rle`, `_validate_chunk_shapes` and
`RectilinearChunkGridMetadata.from_dict` are removed with their tests;
numpy input belongs to the `chunks=` normalizer. The regular grid still
rejects edge lists (gh-4374), now naming the offending type instead of
echoing the value.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe reading mixed regular grids without the rectilinear flag

Rewrite the 4375 changelog fragment for the upgrade and the moved flag
gate, and drop the numpy/tuple parsing claims that no longer hold. Trim
the history from the regular chunk shape docstring.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): warn about an upgraded document only once it validates

An upgrade now returns the upgraded document and how it read it, and
`from_dict` warns with those readings after the metadata constructor
accepts the upgraded document. A stored document that is still invalid
after an upgrade raises its own error instead of first warning how it
was read.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): one per-axis rule for stored chunk sizes, one rule for edges

Stored regular chunk shapes, in both formats and in a sharding codec's
inner chunk shape, are read entry by entry by one rule in upgrades.py:
an int >= 1 is kept, `true` is read as 1 (zarr 3.0.10 also wrote it as
the inner chunk size of a shard, which the sharding codec used to read
silently), and 0 or `false` as one chunk spanning the axis. Integral
floats are not read: no release wrote a regular grid with them. A
document warns once, naming the array where the caller knows its path.

Metadata constructors check every chunk edge length with one function,
`parse_chunk_edge`, which tells a wrong type (TypeError) from a wrong
value (ValueError); this covers bare sizes, rectilinear edges, RLE sizes
and the sharding codec's inner chunk shape.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(chunk-grids): start chunk guessing from the one full-span rule

`_guess_regular_chunks` clamped zero-length axes with `np.maximum`,
restating `full_span_chunk_size` with unit 1. Fold the zero-span
normalizer tests into the `normalize_chunks_1d` tables.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: run the array lifecycle state machine in the slow Hypothesis job

Add it to `just hypothesis`. Re-saving metadata is always enabled and
must leave valid metadata unchanged; appends favour an axis stored with
chunk size 0, which raises how often a legacy axis grows before it is
re-saved. A fixed example pins that a write to a shard kept by a shrink
leaves its cells beyond the shape alone.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a regular grid mixing sizes and edge lists per axis

A stored regular chunk shape is read by the one per-axis rule of
upgrades.py; a flat list is one more case of it, kept as that axis's
chunk edges. A shape mixing chunk sizes and edge lists is read as the
rectilinear grid it describes, so `[true, [5, 10, 5]]`, which zarr 3.2.x
wrote and read, reads as `[1, [5, 10, 5]]`. A shape of only lists was
never stored in a regular grid and is rejected.

The rectilinear grid's constructor checks each edge with the one edge
rule, so a float or bool in an edge list is reported as itself. The
warning says that re-saving stores the rectilinear grid and so needs
the flag, before the re-save steps it lists.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): check that metadata can be stored before touching the store

`check_storable` is the one store-write gate for the rectilinear chunks
flag. `ArrayV3Metadata.to_buffer_dict` calls it, and so does
`GroupMetadata.to_buffer_dict` for each array in the group's
consolidated metadata, so `consolidate_metadata` can no longer store a
rectilinear grid with the flag off and leave the group unreadable
without it.

Every metadata write serializes before it writes. `resize` used to
delete the chunks outside the new shape before storing the new
metadata, so metadata that could not be stored lost data; it now stores
the metadata first, for all metadata, and deletes afterwards.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): store upgraded metadata before writing the first chunk

An array read from a stored document that the upgrades had to correct
(a chunk size of 0, `false` or `true`) wrote chunks under the corrected
layout while the store kept the old document, so zarr 3.2.1-3.4.0 then
read only the fill value, and tensorstore and zarrs could not open it.

`from_dict` now marks the metadata it read from an upgraded document
(`_stored_document_upgraded`, a field outside the document and
equality), and `AsyncArray._set_selection`, which every chunk write
goes through (`setitem` now included), stores that metadata first and
keeps the stored copy. Concurrent first writes store the same document.

The array lifecycle state machine now ends the legacy state on a write
and checks that no chunk is stored under a document that is still
upgraded on read; `resave_metadata` runs only while the stored document
is invalid, and re-saving valid metadata is checked once, at creation.
The changelog also says which releases stored a chunk size of 0 on an
axis of positive length.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): name an array one way in upgrade warnings

Warnings about an upgraded document name the array by its store path,
`str(StorePath(store, path))`, on every path that reads one: opening an
array, `AsyncArray.from_dict`, a group read with or without
consolidated metadata, and consolidated members, which were named
relative to the group. `GroupMetadata.from_dict` and
`ConsolidatedMetadata.from_dict` take the group's path for that.

The consolidated test now also opens the group with
`use_consolidated=False` and checks each array's name, so dropping the
path on either read path fails a test.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(metadata): one rule for bare sizes and edge lists; one noun in messages

- `_validate_chunk_shapes` decides "edge list or bare size" with one
  `match`: a list or tuple is an edge list, anything else is a bare size
  checked by `parse_chunk_edge`, so a string dimension such as `"10"` is
  rejected as not an int instead of being iterated per character.
  Stored rectilinear `from_dict` makes the same JSON split and passes the
  dimension to `expand_rle`, so an invalid edge or RLE count inside a
  list names its dimension.
- Messages say "dimension" throughout, lowercase after the colon
  ("Dimension 0: chunk edge length must be >= 1, got 0"); the upgrade
  warnings read "0 in dimension 0 as one chunk spanning the dimension".
- The upgrades test for JSON arrays with `list` only (a stored document
  holds no tuples), and `_read_chunk_size` says what `span=None` means.
- `ArrayV2Metadata.from_dict` copies the document once.
- The rejected stored chunk shapes get one test per error case.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: encode metadata before deleting store content

`Array.resize`, `AsyncGroup.delitem` and `create_hierarchy(overwrite=True)`
now encode every document they will store (so the rectilinear gate runs)
before deleting anything, then store the encoded documents. A group
member deletion or hierarchy overwrite that would store a rectilinear grid
with the flag off no longer deletes data first. `_resize` goes back to
deleting before storing, which restores the behaviour of stores that
cannot delete (a ZipStore resize fails before writing duplicate metadata).

The gate error raised from `GroupMetadata.to_buffer_dict` names the
consolidated member. Tests: one table of operations that must leave the
store untouched without the flag (resize, chunk write, member deletion,
hierarchy overwrite), the consolidation test for a member in a subgroup,
a write with the flag that stores the rectilinear document, an
abbreviated long chunk shape, and a ZipStore resize.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: reject a run-length entry that is not a pair

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): store the upgrade of the current stored document before writing

A handle read from an upgraded document upserts the upgrade of what the store
holds on its first non-empty chunk write: a stale handle no longer rolls back
newer metadata, and an empty write stores nothing. consolidate_metadata does
the same for each upgraded member before writing the consolidated document.

Both go through two internal primitives in zarr.core.metadata.io:
diff_documents compares a node's stored documents with those its metadata
would store, value by value, and upsert_metadata encodes first, then stores
only the documents that differ and returns the changes.

The upgraded-document flag is a ClassVar set on the instance by one helper,
mark_upgraded, which also warns; the upgrades are keyed by Zarr format; the
Metadata.to_dict and V2 from_dict changes are reverted. _append goes through
AsyncArray.setitem and _setitem is removed. A chunk shape that is not a list
or tuple is rejected as a whole, and every expand_rle error names its
dimension.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: prefer zero-length axes for stored chunk size 0 in the lifecycle machine

An empty write stores no metadata, so it no longer ends the legacy state.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: the caller that knows the node names it in gate errors

`check_storable` and `_check_rectilinear_chunks_enabled` take no subject.
`encode_documents(store_path, metadata)` in `metadata/io.py` encodes the
documents an operation stores and, on failure, adds a note naming the node
at `store_path` in the full-path form the upgrade warnings use; every
encode-before-write site goes through it (`save_metadata`,
`upsert_metadata`, `_resize`, `Group.delitem`, group attribute updates,
`_encode_nodes` for `create_hierarchy`/`create_nodes`).
`ArrayV3Metadata.from_dict` names the array at `path` when the flag refuses
a rectilinear document, which is where `consolidate_metadata` and a first
write to an upgraded 3.2.x mixed grid are refused without the flag.

The upsert gate test now uses the real gate instead of a monkeypatched
`to_buffer_dict`, and consolidating a 3.2.x mixed grid is checked both
ways: refused with the store byte-identical without the flag, stored as
rectilinear (member and consolidated document) with it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): delete a consolidated member in place, after encoding without it

`Group.delitem` encodes the group documents from a copy without the member
(so the gate runs with the store and the group untouched on failure),
deletes the member, then pops it from the shared consolidated metadata in
place, as main does, so a parent or subgroup handle sharing that
consolidated metadata sees the deletion, then stores the documents.

Tests: the consolidated delitem test runs for both Zarr formats and
reopens from the store (V2 `.zmetadata` is rewritten); a subgroup handle
read from its parent's consolidated metadata sees a deletion made through
it; a deletion refused without the flag leaves the group listing the
member.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): keep accepting the chunk sizes zarr 3.4.0 accepted in metadata constructors (patch release)

Nothing in a patch release may reject a chunk size that zarr 3.4.0 accepted.
The one chunk edge rule (`parse_chunk_edge`, which also reads RLE repeat
counts) now reads any integral number as the `int` it equals: an `int`, a
`bool`, a NumPy integer or an integral float, so rectilinear grids written
with float edges (`[[4.0, 2]]`) read again and the grid constructors accept
NumPy integers; a fractional float is still rejected. A regular chunk shape
may again be any iterable. `ArrayV2Metadata(chunks=...)` and
`ShardingCodec(chunk_shape=...)` read their chunk shape with
`parse_shapelike`, as in 3.4.0: a scalar, NumPy integers, bools and 0 are
accepted, and a 0 is written back as given (the kerchunk writer in
VirtualiZarr relies on `chunks=(0,)` for empty inlined arrays). Stored
documents with a chunk size of 0 are still read by the upgrades.

The strict rule returns in the next minor release.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): leave a valid stored document as written on a stale handle's first write

A handle read from an upgraded document re-reads the stored document before its
first chunk write. If that document no longer needs an upgrade, store nothing:
re-encoding it could differ from how another implementation wrote it, and
nothing about it needs fixing.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe the 4334 changes to metadata constructors for a patch release

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: expect the patch release's wording for non-integral chunk edges (patch release)

The chunk edge rule reads any integral number (loop/4334), so a float or bool
edge in a stored mixed regular grid is read as the int it equals; a fractional
or string edge is still rejected, with "must be an integer".

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read float edges only in stored rectilinear documents

The patch release widened the one chunk edge rule to accept integral
floats so that stored rectilinear grids written with float edges
(`[[4.0, 2]]`, as zarr 3.2 wrote for float edges) keep opening. That also
made a stored regular chunk shape `[10.0]` open, which zarr 3.4.0 rejected
and the next minor release would reject again.

The constructor rule now accepts integers of any integer type (`int`,
`bool`, NumPy integers) and rejects floats. A new document upgrade reads
an integral JSON float of at least 1 as an `int` where zarr 3.2 stored
one: the explicit edges and run-length encoded sizes of a rectilinear
chunk grid, with the standard warning. Bare sizes, repeat counts, regular
chunk shapes and inner chunk shapes stay for the constructors to reject.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: read float edges in a regular grid that lists chunk edge lengths

zarr 3.2 stored `[2, [5.0, 10.0, 5.0]]` for float edges given to a
regular grid. The mixed-grid upgrade reads it as a rectilinear grid, and
the float edge upgrade that follows reads its edges as integers, both in
one warning.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(codecs): pickle a ShardingCodec with an inner chunk size of 0

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): build every node before create_hierarchy deletes or stores anything

A node whose metadata no array or group can be built from now fails with
the store untouched, even with overwrite=True.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): warn only where the user must act; guard and refresh upgraded copies

- Readings that give what zarr 3.4.0 read (0/false on an empty axis, JSON
  true, float rectilinear edges) are silent; the metadata is still marked
  upgraded, so the first write re-saves it.
- JSON true is read as 1 in stored rectilinear edges and RLE sizes and in
  nested sharding codecs' inner chunk shapes.
- A write through a handle whose stored document now lays out chunks
  differently raises and stores nothing; with no stored document it writes
  as before.
- Storing group metadata refreshes each upgraded consolidated member from its
  own stored document (the one place: save_metadata); consolidate_metadata
  no longer needs its own pass.
- Metadata built in code with chunk size 0 is read through the upgrades, so
  an array can be built from it, as in 3.4.0.
- Chunk edges in metadata follow the 3.4.0 integer rule (int; bool read as
  its value); a string or mapping is rejected as a chunk shape as a whole.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: expect the lifecycle machine's upgrade warning only on non-empty axes

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): encode a group's consolidated metadata before storing member upgrades

After the merge, storing group metadata refreshed each upgraded
consolidated member and stored its upgrade before encoding the group, and
`Group.delitem` (which encodes before deleting) skipped the refresh.

`encode_node` refreshes the consolidated members (reading only) and
encodes the group, so the rectilinear gate runs for every member before
anything is stored; `store_node` then stores the group's documents and the
members' upgrades. `save_metadata` and `Group.delitem` both use the pair,
so a group write refused without the flag stores no member upgrade either.

Tests: group attribute writes join the store-untouched table; a refused
group write stores no other member's upgrade; with the flag, deleting a
member or setting a group attribute stores the rectilinear grid in the
member and consolidated documents. Error-message expectations follow the
patch release's integer rule wording, and the mixed-grid rows expect the
JSON `true` and float edge readings to be silent.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): leave a sharded chunk size of 0 unread when the inner shape is invalid

A stored outer chunk size of 0 of a sharded array is read in multiples of the
inner chunk size. When that inner size is itself 0 or `false`, the unit is
unknown: the upgrade now leaves the 0 for the constructor, which rejects it
with the same `ValueError` zarr 3.4.0 raised, instead of a ZeroDivisionError.
`_read_codec` now matches the sharding codec like `_inner_chunk_shape` does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(metadata): read each upgraded member once, concurrently, and adopt it

A group handle kept the consolidated copies of upgraded members flagged, so
every later group write re-read every such member, one after another, and
read each twice on the first write (once to parse, once to diff).

- `save_metadata` returns the metadata it stored; `AsyncGroup._save_metadata`
  (attrs, `update_attributes`, `delitem`, `consolidate_metadata`) and
  `Group.update_attributes_async` adopt it, so later writes read no member.
- The refresh visits only nested groups and flagged members, and reads them
  with one `asyncio.gather`. The documents it reads are the ones the upsert
  diffs against (`read_documents` + `upsert_metadata(..., stored)`), and the
  array write path does the same.
- `parse_stored_array` (documents -> silently marked metadata) replaces
  `read_stored_array`'s `(metadata, bool)` tuple, and is the one "read, mark,
  don't warn" path, also for code-built metadata.
- A flagged member whose own document is gone or cannot be read (invalid,
  replaced by a group) keeps its consolidated copy instead of failing every
  group write.
- `parse_array_metadata` of a metadata object reads it as a stored document
  only when `ArrayV2Metadata.chunks` holds a 0, the one chunk size that
  constructors still accept and the upgrades change. This removes the
  per-construction `to_dict` + upgrade (AsyncArray() back to 3.4.0 speed)
  and building arrays from codec configurations holding NumPy scalars
  works again, as in 3.4.0.
- `create_hierarchy` documents that it stores a group's consolidated
  metadata as given.
- Tests: second group write reads no member; nested member refresh; member
  without a readable document; NumPy-scalar codec configuration; stale
  handle whose inner chunk shape alone changed (kills `_chunk_layout`
  returning only the outer grid).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): store the consolidated metadata as it is after a deletion

`delitem` on a consolidated group stored the group metadata it had
encoded before deleting, so concurrent deletions through one handle
each stored a member list that still held the others' members. The
up-front encode now only validates (metadata that cannot be stored still
fails with the store untouched); after the deletion the handle's current
metadata is stored and adopted, as zarr 3.4.0 stored it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: name a node in errors as "Array 'path': what happened"

The rectilinear flag's read note, the stale-handle write error and the
encode note now share the shape the stored-document warnings use: the
node, its path, then what was (not) done. The consolidated-member note
keeps its location qualifier ("in the consolidated metadata"), followed
by the group's note.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): adopt refreshed consolidated members in place

A group write replaced the handle's metadata with a copy holding the
consolidated members it read again, so the handle no longer shared its
consolidated dicts with subgroup handles taken earlier: an array deleted
through such a subgroup was stored again by the parent's next write, and
reappeared on reopen.

`_refresh_consolidated` now returns, along with the copy it encodes, the
members it replaced (with the consolidated dict that holds each), and
`save_metadata` writes them back into those dicts once the store succeeds,
unless the member was deleted or replaced meanwhile. `save_metadata` returns
nothing again, and the group callers keep their metadata objects. A failed
store leaves the members flagged, so a retry reads them again.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): raise the rectilinear flag error when refreshing a consolidated member

A group write reads each upgraded consolidated member again from its own
document, and kept the consolidated copy when that document could not be
read. That also swallowed the rectilinear chunks flag error, so a member
whose document now declares a rectilinear grid had its stale copy stored
as valid.

The flag error is now `RectilinearChunksDisabledError`, a `ValueError`,
and the refresh lets it propagate: the group write raises and stores
nothing. Unreadable documents keep their consolidated copy as before.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): say nothing was stored when a consolidated member needs the flag

Storing a group's consolidated metadata reads each upgraded member again
from its own document; one read as a rectilinear chunk grid now raises the
flag error there (naming the array) instead of being kept stale. That
error now also says the group stored nothing, as the encoding error does.
`zarr.consolidate_metadata` without the flag reports the member this way.
The refresh test gains a mixed regular grid row.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(group): read each upgraded member once when deleting a member

`delitem` encodes the group metadata without the member before deleting
it, which reads each upgraded consolidated member again, then stored the
metadata through `save_metadata`, which read them all a second time.

It now adopts the members the first encoding read, after the deletion and
in place, and stores the metadata as it then is with those members'
upgrades (`store_node`), so the first deletion reads each member once.
The handle's consolidated metadata is no longer replaced by writes, so the
deletion pops from it directly.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(group): store a group's documents only; keep upgraded members as stored

Group writes (attributes, `update_attributes`, deleting a member,
`consolidate_metadata`, `create_hierarchy`) read each upgraded member of
the consolidated metadata again from its own document and stored its
upgrade. Those member writes raced concurrent deletions: two concurrent
`delitem`s on a consolidated group recreated a deleted array's metadata
document in most runs, which zarr 3.4.0 never did.

A group write now stores only the group's own documents and reads none.
Array metadata read from a document that had to be upgraded keeps that
document (`_stored_document`, set by `mark_upgraded`, replacing the
`_stored_document_upgraded` flag), and consolidated metadata stores such
a member as it was stored. Every reader of the consolidated metadata then
reads the member as upgraded again, and the array's own first chunk write
stores the upgrade of its current document (or refuses a changed chunk
grid) as before: the only write that upgrades a member document.

Removed: the refresh and adoption of consolidated members in
`save_metadata` (`_refresh_consolidated`, `_refresh_array`,
`_Refreshed`), and `Group.update_attributes_async`'s detour through
`save_metadata`. Tests pinning the refresh are replaced by tests that a
group write reads nothing and writes only the group's documents, stores
the member's consolidated copy byte for byte as stored, that concurrent
deletions leave no member, and that the first write through consolidated
metadata stores the member's upgrade or refuses a changed grid.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(group): say what consolidated metadata stores for an upgraded array

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(array): clear the stored document whenever an array stores its metadata

Setting attributes of an array read from a document that had to be
upgraded stored the upgraded document with the new attributes, but left
the metadata marked with the document it was read from. A consolidated
group handle shares that metadata, so its next write stored the old
document in the consolidated metadata and the new attributes were lost
from it (zarr 3.4.0 kept them).

Every write of an array's own documents (creating, resizing, setting
attributes) now goes through `AsyncArray._save_metadata`, which clears
the mark afterwards, as the first chunk write does after storing the
upgrade. Both clear it through `_stored_document_replaced`, the one
place that declares it: the first chunk write clears it also when it
stores nothing (the store already holds a valid document, or none), so
it cannot be folded into the save.

The deep copy in `mark_upgraded` stays: the metadata's attributes share
objects with the document it was read from, which the caller also
holds. A test pins it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(metadata): pin that a failed metadata save keeps the stored document mark

Setting attributes on or resizing an array read from a document that needed
an upgrade clears the `_stored_document` mark only after the store accepts the
new metadata. A test now injects a store failure for the array's own document
and checks the mark and the stored bytes survive, for Zarr formats 2 and 3.
The `_save_metadata` docstring now says only what the method does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(metadata): read a stored chunk size of 0 as 1, however long the axis

A stored chunk size of 0 or `false` was read as one chunk spanning the
axis at open time. On an axis that grew after the size was written, that
made the chunk size depend on when the array was opened, and it could be
very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one
800 MB chunk. It is now read as the smallest valid chunk edge length (1,
or the inner chunk size for a shard), which is what `chunks=-1` gives on
a zero-length axis. A resize by software that kept the stored 0 no longer
changes the layout a stale handle expects.

Assisted-by: ClaudeCode:claude-opus-5-5

* fix(metadata): re-save only upgrades that move chunks, and read non-JSON input as is

A stored `true` chunk size or an integral float edge is read as the value
zarr already read it as, so chunks written under the upgraded metadata are
where every reader looks. Marking such documents for re-saving turned a
plain chunk write into a metadata write: in a ZipStore that adds a second
`zarr.json` entry (a `UserWarning`, an error under `-W error`), and it
opened a race with concurrent metadata writes that 3.4.0 did not have. Each
upgrade now reports whether it moves chunks, and only those mark the
metadata.

The upgrades also detected changes by comparing JSON encodings, so
`ArrayV3Metadata.from_dict` given metadata built in code (a codec instance,
a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`.
Changes are now tracked where they are made.

Assisted-by: ClaudeCode:claude-opus-5-5

* fix(array): encode new array metadata before deleting an existing node

With the rectilinear flag checked when metadata is stored rather than when
it is built, `zarr.create(..., overwrite=True)` with rectilinear chunks
and the flag off deleted the existing array and then raised; zarr 3.4.0
raised before deleting anything. `AsyncArray._create_v2`, `_create_v3`
and `init_array` now build and encode the new metadata before
`_prepare_overwrite`, so metadata that cannot be stored fails with the
store untouched. This also covers `create_array(..., overwrite=True)`,
which deleted the existing array before validating its arguments.

Assisted-by: ClaudeCode:claude-opus-5-5

* test(metadata): type the rectilinear chunks passed to zarr.create in an overwrite test

Assisted-by: ClaudeCode:claude-opus-5-5

* fix(metadata): recommend recreating an array whose stored chunk size was 0

The warning for a stored chunk size of 0 on an axis of positive length now says that
the array holds only its fill value, so recreating it with the wanted chunk shape loses
nothing, and gives the re-save as the way to keep it instead. A re-save freezes the
smallest chunk size into an array that holds no data yet. Each reading now carries its
own advice, so mark_upgraded appends no shared hint.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(metadata): widen the abbreviated-message bound for the two hints

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(metadata): give the recipe for recreating an array read from a stored chunk size of 0

The warning names the call, zarr.from_array(array.store, name=array.path, data=array,
chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value,
attributes and codecs, and a test pins the recipe on both formats.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(metadata): widen the abbreviated-message bound for the recipe

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(metadata): call the lenient readings of stored documents repairs, not upgrades

The module is zarr.core.metadata.repair: repair_array_document, mark_repaired,
ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test
module test_repair.py. A reading that turns an invalid stored document into a valid one
fixes it; it does not move it to a newer format.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(changes): trim the #4375 fragment to what the PR adds over main

Rejecting edge lists in a regular chunk grid landed with #4334, and encoding a new
node's metadata before deleting the existing one landed with #4410.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants