Skip to content

fix(chunk-grids): judge every chunk specification by the chunk normalizer alone - #4376

Merged
d-v-b merged 30 commits into
zarr-developers:mainfrom
d-v-b:fix/single-chunk-normalizer
Oct 3, 2026
Merged

d-v-b merged 30 commits into
zarr-developers:mainfrom
d-v-b:fix/single-chunk-normalizer

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

This AI-written PR consolidates the logic for parsing a chunk request into one function, where previously we had 2 paths that separately handled chunk- and shard-like requests, and that led to bugs as the two paths disagreed. one function should at least keep the bugs in one place :)

🤖 AI text below 🤖

Chunk specifications are judged only by the chunk normalizer: the separate duck-typed check for rectilinear input is gone, so the Zarr format 2 and sharding restrictions apply to what the normalized grid says was declared. A chunk size is an integer by Python's integer protocol: numpy integer scalars and 0-d integer arrays are accepted wherever a chunk or shard shape is. The legacy zarr.create(..., zarr_format=2) no longer tests the truth value of a numpy array given as chunks, so numpy arrays work there. A non-integer scalar chunk specification raises a TypeError naming the problem. An explicit list of chunk edges declares a rectilinear dimension everywhere, including a stored rectilinear chunk grid passed as chunks= whose edges happen to be uniform, so the Zarr format 2 and sharding restrictions treat it the same whether it is given as lists or as metadata, and zarr.create stores it as a rectilinear chunk grid, as create_array does. A RectilinearChunkGridMetadata made only of bare integers declares no explicit edge list, so it is read as the regular grid it describes: it is accepted for Zarr format 2 and as the chunk shape of a sharded array.

Follow-up to #4374.

Problem

The chunk normalizer (normalize_chunks_nd → normalize_chunks_1d) was not the only judge of a chunks=/shards= specification. A separate duck-typed classifier, _is_rectilinear_chunks, ran on the raw input at three sites (AsyncArray._create, init_array, resolve_outer_and_inner_chunks) to decide whether the specification was rectilinear, so that the Zarr format 2 and sharding restrictions could fire. Two opinions on the same input is the shape of the bug in #4374, and they could disagree: a 0-d NumPy array counted as rectilinear because it has __iter__, and ChunkGrid.from_sizes collapsed a uniform edge list to a regular dimension while normalize_chunks_1d kept it rectilinear.

Changes

  • One judge. _is_rectilinear_chunks is deleted. Each site normalizes first and asks the resulting ChunkGrid (is_regular). Regular and rectilinear shard specifications go through the same normalizer.
  • One integer rule in the normalizer (_chunk_int): a chunk size is anything Python's integer protocol (operator.index) accepts: int, NumPy integer scalars and 0-d integer arrays. Floats, arrays with dimensions and NumPy booleans are not integers. A Python bool is still an int (0 or 1), as in 3.4.0.
  • An explicit edge list declares a rectilinear dimension everywhere. ChunkGrid.from_sizes no longer collapses uniform edges, matching normalize_chunks_1d. So a stored RectilinearChunkGridMetadata passed as chunks= is treated the same whether its edges are uniform or not: zarr.create stores it as a rectilinear grid, as create_array does (3.4.0 stored RectilinearChunkGridMetadata(((5, 5),)) as a regular [5], A rectilinear chunk spec with a short trailing chunk is silently normalized to a regular grid, changing resize semantics #4272). A RectilinearChunkGridMetadata made only of bare integers declares no edge list, so it is the regular grid it describes (accepted for Zarr format 2 and as the chunk shape of a sharded array).
  • Legacy zarr.create(..., zarr_format=2) used chunks or chunk_shape, so a NumPy array such as chunks=np.array([5, 3]) failed with "truth value of an array is ambiguous". A small predicate (_v2_chunks_given) now decides whether chunks was given: falsy values (None, 0, [], (), False, np.int64(0), np.array(0), np.array([0])) still mean automatic chunking, exactly as in 3.4.0; a NumPy array with more than one element is always given.
  • Errors: a non-integer scalar specification (2.0, np.float64(2.0), a 0-d float array) raises the normalizer's own TypeError naming it (was "object has no len()" or "len() of unsized object").
  • Typing: normalize_chunks_nd(chunks: ChunksLike | None, ...); shard_spec uses the existing aliases instead of Any.

Not changed in this patch release: chunks=True still raises ValueError, and a bool inside a specification is still read as 1 or 0. Rejecting booleans with one TypeError is part of the 3.5.0 follow-up.

Compatibility with 3.4.0

In the 244-row behaviour table run against a real 3.4.0 install, this branch changes no result that 3.4.0 handled correctly. The rows it changes: zarr.create(chunks=np.array([5, 5]), zarr_format=2) and np.array([]) now work (3.4.0 raised a truth-value error with current NumPy); a uniform RectilinearChunkGridMetadata passed to zarr.create is stored as rectilinear (gh-4272). check_patch.py: OK, 0 undocumented changes, 0 warnings-only changes.

Tests

  • tests/test_unified_chunk_grid.py: one table for the legacy Zarr format 2 chunks argument (every falsy spelling auto-chunks, np.array(7) → (7, 7), np.array([5, 3]) → (5, 3)) and one error test for np.array([0, 0]); the classifier's tests are replaced by one test that the Zarr format 2 and sharding restrictions recognize every rectilinear specification (whichever dimension carries the list, uniform edges included) through create_array and zarr.create; the uniform-edge resize test covers edges given as lists and as metadata.
  • tests/test_chunk_grids.py: NumPy booleans are rejected by the normalizer; NumPy inputs join the existing normalizer tables.
Stack and release notes

🤖 Generated with Claude Code

Chunk specifications were parsed by `normalize_chunks_nd`, but a separate
duck-typed classifier, `_is_rectilinear_chunks`, ran on the raw input first
at three sites to decide whether the spec was rectilinear. Two opinions on
the same input is the shape of the bug in zarr-developers#4374, and they could disagree:
a 0-d numpy array counted as rectilinear because it has `__iter__`.

The classifier is gone. Each site normalizes first and asks the resulting
`ChunkGrid` (`is_regular`); stored rectilinear metadata passed as
`chunks=` counts as rectilinear even when its edges are uniform. The shard
resolver sends regular and rectilinear shard specs through the same
normalizer.

Also fixed on the way:

- The legacy v2 branch of `AsyncArray.create` tested `chunks or
  chunk_shape`, so `zarr.create(chunks=np.array([...]), zarr_format=2)`
  failed with "truth value of an array is ambiguous". It now uses the
  `is not None` form the v3 branch already had.
- 0-d numpy arrays unwrap to their scalar in both normalizers instead of
  failing with "len() of unsized object".
- A non-integer scalar spec (`2.0`, `np.float64`) raises the normalizer's
  own TypeError instead of "object has no len()".

Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
@github-actions github-actions Bot added needs release notes Automatically applied to PRs which haven't added release notes and removed needs release notes Automatically applied to PRs which haven't added release notes labels Sep 18, 2026
@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 zarr-indexing | 🛠️ Build #34831131 | 📁 Comparing 37fc614 against latest (1187a43)

  🔍 Preview build  

No files changed.

@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.68%. Comparing base (f078bd3) to head (81d97e0).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4376      +/-   ##
==========================================
- Coverage   94.69%   94.68%   -0.01%     
==========================================
  Files          94       94              
  Lines       13604    13596       -8     
==========================================
- Hits        12882    12874       -8     
  Misses        722      722              
Files with missing lines Coverage Δ
src/zarr/api/asynchronous.py 96.46% <ø> (ø)
src/zarr/api/synchronous.py 95.77% <ø> (ø)
src/zarr/core/array.py 98.39% <100.00%> (+0.09%) ⬆️
src/zarr/core/chunk_grids.py 96.35% <100.00%> (-0.39%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

d-v-b and others added 13 commits September 25, 2026 21:39
`_chunk_int` is the normalizer's single integer rule: anything Python's
integer protocol accepts (int, numpy integer scalars, 0-d integer arrays),
except bool. It replaces the repeated numbers.Integral checks and the two
0-d ndarray unwraps in normalize_chunks_1d / normalize_chunks_nd.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ilinear rejection

The numpy cases that passed before this PR are dropped; the 0-d array and
float cases move into the normalizer's table and error tests. The legacy
zarr.create Zarr format 2 path gets a truthiness test and its rectilinear
rejection is now asserted with a message match.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…elopers#4376 changelog

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rid.from_sizes

`ChunkGrid.from_sizes` collapsed uniform edge lists to `FixedDimension`
while `normalize_chunks_1d` keeps them as `VaryingDimension`, so
`init_array` needed an `isinstance(chunks, RectilinearChunkGridMetadata)`
guard to keep uniform stored rectilinear grids under the Zarr format 2 and
sharding restrictions. Both now agree that a list declares a rectilinear
dimension, and the guard is gone: the normalized grid is the one judge.

Also: `normalize_chunks_nd` and the shard spec are typed with `ChunksLike`;
a bool chunk size is reported as not a chunk size; an integral float
(`10.0`) is pinned as rejected; `None` reaching the normalizer gets the
generic non-integer `TypeError`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…one error

A bool, np.bool_, or boolean array anywhere in a chunk specification now
raises the same TypeError from the normalizer's one integer test, instead
of three different errors depending on spelling. Document that a
RectilinearChunkGridMetadata of bare integers is read as a regular grid.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alsy v2 chunks (patch release)

Nothing in a patch release may reject a chunk specification that zarr 3.4.0
accepted. The normalizer's integer test reads a Python `bool` as the `int` it
is again, so `chunks=(True, 5)` and a `True` edge are a chunk size of 1;
`chunks=True` and `chunks=None` raise 3.4.0's `ValueError` again. Numpy
booleans stay rejected, as they were. The legacy `zarr.create(...,
zarr_format=2)` again reads a falsy `chunks` (0, [], False) as not given and
chunks automatically; a numpy array is always taken as given, so its truth
value is never tested.

The `TypeError` for every boolean spelling returns in the next minor release.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2 create as zarr 3.4.0 did

The legacy `zarr.create(..., zarr_format=2)` took every numpy array as a
given chunk specification, so `np.array(0)`, `np.array(False)` and
`np.array([0])` raised instead of auto-chunking as in zarr 3.4.0. A numpy
array with more than one element has no truth value and is always given; a
shorter one is read by `.any()`, its truth value, which is False when empty,
so `np.array([])` auto-chunks like `[]` (as 3.4.0 did with numpy 2.1).

The two legacy v2 tests become one table.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… _chunk_int

`operator.index` raises `TypeError` for everything the `SupportsIndex`
check rejected, so the check was redundant (identical results over bool,
numpy bools and integers, 0-d and 1-d arrays, float, str, bytes, None,
list and a custom `__index__` class).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iven

Replace the inline conditional expression in `AsyncArray._create` with a
small named predicate. Results are identical on every probed input.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add `np.array(7)` -> `(7, 7)` to `test_legacy_create_v2_chunks` and an
error test for `np.array([0, 0])`, killing the mutants that drop either
half of `_v2_chunks_given`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	src/zarr/core/chunk_grids.py
@d-v-b d-v-b changed the title fix: make the chunk normalizer the one judge of a chunk specification fix(chunk-grids): judge every chunk specification by the chunk normalizer alone Sep 29, 2026
…_len__

The chunk normalizer checked `isinstance(chunks, Iterable)` before calling
`list(chunks)`. `Iterable` does not recognize a sequence that implements
only `__getitem__` and `__len__`, which `list` accepts and which zarr 3.4.0
took as a chunk specification. The normalizer now calls `list` and turns
its `TypeError` into the chunk specification error.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b d-v-b added this to the 3.4.1 milestone Sep 29, 2026
…e 2.3

NumPy before 2.3 takes a NumPy boolean as an index with a
DeprecationWarning, so `operator.index(np.True_)` returned 1 there, and
the normalizer read `chunks=np.True_` as a chunk size of 1 instead of
rejecting it as zarr 3.4.0 did. The min-deps CI job failed on this. The
integer test now rejects a value whose dtype is NumPy bool before calling
`operator.index`.

Assisted-by: ClaudeCode:claude-opus-5-5
d-v-b and others added 11 commits October 1, 2026 11:09
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with
save_new_metadata: keep main's helper and this branch's imports and chunk
normalization.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves the two hunks with zarr-developers#4334: the single normalizer call takes zarr-developers#4334's unit, and
the test module keeps both imports.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… zarr.create

ShardsLike is now ChunksLike plus the sharding configuration and "auto", so
numpy integers and arrays are declared for shards as they are for chunks.
zarr.create and AsyncArray._create take ChunksLike for chunks and
chunk_shape, matching what the normalizer reads. Tests drop the Any
annotations and the type: ignore that worked around the narrower types.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r.create

The Zarr format 2 path of AsyncArray._create used the truth value of
`chunks` to decide whether it was given, so False, 0, [] and np.array([0])
were silently auto-chunked while Zarr format 3 normalized or rejected them.
Both formats now share the None-only `_raw_chunks` and agree on every input.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…as a regular grid

ShardingCodec read its inner chunk_shape with parse_shapelike, which
allows 0, so an inner size of 0 was accepted at construction and failed
later with ZeroDivisionError. parse_regular_chunk_shape shares the
normalizer's integer test (_chunk_int/_chunk_list) and requires every
size to be at least 1; with no axis length to resolve against, -1 and
False are rejected, and an explicit edge list (a rectilinear dimension)
is rejected because the inner grid of a shard is regular.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… shape

ArrayV2Metadata read `chunks` with parse_shapelike, which accepts 0, while
RegularChunkGridMetadata rejects it. It now uses parse_regular_chunk_shape,
as ShardingCodec does, so a chunk size below 1 raises at construction.
Stored documents are unaffected: from_dict repairs a stored 0 before the
constructor sees it. The shared chunk-shape tests cover both sites again.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… True

The docstring said True guesses the chunk shape and is the default; None is
the default and True raises. The auto-chunking error now tells
create_array callers to pass "auto" rather than the None they just passed.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ut of this patch PR

Reverts 47587bc, baddb31 and ad34ae8, and the error-message part of
e547401. Rejecting a chunk size of 0 in ArrayV2Metadata and
ShardingCodec, and no longer reading a falsy `chunks` as absent in the
legacy Zarr format 2 zarr.create, reject inputs zarr 3.4.0 accepted; this
PR targets the 3.4.1 patch release, which must not. zarr-developers#4431 makes the same
changes, more strictly, for the next minor release. The widened chunk and
shard annotations and the zarr.create docstring fix stay.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@d-v-b
d-v-b marked this pull request as ready for review October 3, 2026 17:42
@d-v-b
d-v-b merged commit 25d390b into zarr-developers:main Oct 3, 2026
40 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant