fix(chunk-grids): judge every chunk specification by the chunk normalizer alone - #4376
Merged
Merged
Conversation
Chunk specifications were parsed by `normalize_chunks_nd`, but a separate duck-typed classifier, `_is_rectilinear_chunks`, ran on the raw input first at three sites to decide whether the spec was rectilinear. Two opinions on the same input is the shape of the bug in zarr-developers#4374, and they could disagree: a 0-d numpy array counted as rectilinear because it has `__iter__`. The classifier is gone. Each site normalizes first and asks the resulting `ChunkGrid` (`is_regular`); stored rectilinear metadata passed as `chunks=` counts as rectilinear even when its edges are uniform. The shard resolver sends regular and rectilinear shard specs through the same normalizer. Also fixed on the way: - The legacy v2 branch of `AsyncArray.create` tested `chunks or chunk_shape`, so `zarr.create(chunks=np.array([...]), zarr_format=2)` failed with "truth value of an array is ambiguous". It now uses the `is not None` form the v3 branch already had. - 0-d numpy arrays unwrap to their scalar in both normalizers instead of failing with "len() of unsized object". - A non-integer scalar spec (`2.0`, `np.float64`) raises the normalizer's own TypeError instead of "object has no len()". Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Documentation build overview
|
Documentation build overview
No files changed. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #4376 +/- ##
==========================================
- Coverage 94.69% 94.68% -0.01%
==========================================
Files 94 94
Lines 13604 13596 -8
==========================================
- Hits 12882 12874 -8
Misses 722 722
🚀 New features to boost your workflow:
|
d-v-b
force-pushed
the
fix/single-chunk-normalizer
branch
from
September 18, 2026 17:00
22d828c to
9d0f413
Compare
`_chunk_int` is the normalizer's single integer rule: anything Python's integer protocol accepts (int, numpy integer scalars, 0-d integer arrays), except bool. It replaces the repeated numbers.Integral checks and the two 0-d ndarray unwraps in normalize_chunks_1d / normalize_chunks_nd. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ilinear rejection The numpy cases that passed before this PR are dropped; the 0-d array and float cases move into the normalizer's table and error tests. The legacy zarr.create Zarr format 2 path gets a truthiness test and its rectilinear rejection is now asserted with a message match. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…elopers#4376 changelog Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rid.from_sizes `ChunkGrid.from_sizes` collapsed uniform edge lists to `FixedDimension` while `normalize_chunks_1d` keeps them as `VaryingDimension`, so `init_array` needed an `isinstance(chunks, RectilinearChunkGridMetadata)` guard to keep uniform stored rectilinear grids under the Zarr format 2 and sharding restrictions. Both now agree that a list declares a rectilinear dimension, and the guard is gone: the normalized grid is the one judge. Also: `normalize_chunks_nd` and the shard spec are typed with `ChunksLike`; a bool chunk size is reported as not a chunk size; an integral float (`10.0`) is pinned as rejected; `None` reaching the normalizer gets the generic non-integer `TypeError`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…one error A bool, np.bool_, or boolean array anywhere in a chunk specification now raises the same TypeError from the normalizer's one integer test, instead of three different errors depending on spelling. Document that a RectilinearChunkGridMetadata of bare integers is read as a regular grid. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alsy v2 chunks (patch release) Nothing in a patch release may reject a chunk specification that zarr 3.4.0 accepted. The normalizer's integer test reads a Python `bool` as the `int` it is again, so `chunks=(True, 5)` and a `True` edge are a chunk size of 1; `chunks=True` and `chunks=None` raise 3.4.0's `ValueError` again. Numpy booleans stay rejected, as they were. The legacy `zarr.create(..., zarr_format=2)` again reads a falsy `chunks` (0, [], False) as not given and chunks automatically; a numpy array is always taken as given, so its truth value is never tested. The `TypeError` for every boolean spelling returns in the next minor release. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2 create as zarr 3.4.0 did The legacy `zarr.create(..., zarr_format=2)` took every numpy array as a given chunk specification, so `np.array(0)`, `np.array(False)` and `np.array([0])` raised instead of auto-chunking as in zarr 3.4.0. A numpy array with more than one element has no truth value and is always given; a shorter one is read by `.any()`, its truth value, which is False when empty, so `np.array([])` auto-chunks like `[]` (as 3.4.0 did with numpy 2.1). The two legacy v2 tests become one table. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… _chunk_int `operator.index` raises `TypeError` for everything the `SupportsIndex` check rejected, so the check was redundant (identical results over bool, numpy bools and integers, 0-d and 1-d arrays, float, str, bytes, None, list and a custom `__index__` class). Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iven Replace the inline conditional expression in `AsyncArray._create` with a small named predicate. Results are identical on every probed input. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add `np.array(7)` -> `(7, 7)` to `test_legacy_create_v2_chunks` and an error test for `np.array([0, 0])`, killing the mutants that drop either half of `_v2_chunks_given`. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # src/zarr/core/chunk_grids.py
…_len__ The chunk normalizer checked `isinstance(chunks, Iterable)` before calling `list(chunks)`. `Iterable` does not recognize a sequence that implements only `__getitem__` and `__len__`, which `list` accepts and which zarr 3.4.0 took as a chunk specification. The normalizer now calls `list` and turns its `TypeError` into the chunk specification error. Assisted-by: ClaudeCode:claude-opus-5-5
…e 2.3 NumPy before 2.3 takes a NumPy boolean as an index with a DeprecationWarning, so `operator.index(np.True_)` returned 1 there, and the normalizer read `chunks=np.True_` as a chunk size of 1 instead of rejecting it as zarr 3.4.0 did. The min-deps CI job failed on this. The integer test now rejects a value whose dtype is NumPy bool before calling `operator.index`. Assisted-by: ClaudeCode:claude-opus-5-5
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with save_new_metadata: keep main's helper and this branch's imports and chunk normalization. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves the two hunks with zarr-developers#4334: the single normalizer call takes zarr-developers#4334's unit, and the test module keeps both imports. Assisted-by: ClaudeCode:claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… zarr.create ShardsLike is now ChunksLike plus the sharding configuration and "auto", so numpy integers and arrays are declared for shards as they are for chunks. zarr.create and AsyncArray._create take ChunksLike for chunks and chunk_shape, matching what the normalizer reads. Tests drop the Any annotations and the type: ignore that worked around the narrower types. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r.create The Zarr format 2 path of AsyncArray._create used the truth value of `chunks` to decide whether it was given, so False, 0, [] and np.array([0]) were silently auto-chunked while Zarr format 3 normalized or rejected them. Both formats now share the None-only `_raw_chunks` and agree on every input. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…as a regular grid ShardingCodec read its inner chunk_shape with parse_shapelike, which allows 0, so an inner size of 0 was accepted at construction and failed later with ZeroDivisionError. parse_regular_chunk_shape shares the normalizer's integer test (_chunk_int/_chunk_list) and requires every size to be at least 1; with no axis length to resolve against, -1 and False are rejected, and an explicit edge list (a rectilinear dimension) is rejected because the inner grid of a shard is regular. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… shape ArrayV2Metadata read `chunks` with parse_shapelike, which accepts 0, while RegularChunkGridMetadata rejects it. It now uses parse_regular_chunk_shape, as ShardingCodec does, so a chunk size below 1 raises at construction. Stored documents are unaffected: from_dict repairs a stored 0 before the constructor sees it. The shared chunk-shape tests cover both sites again. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… True The docstring said True guesses the chunk shape and is the default; None is the default and True raises. The auto-chunking error now tells create_array callers to pass "auto" rather than the None they just passed. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ut of this patch PR Reverts 47587bc, baddb31 and ad34ae8, and the error-message part of e547401. Rejecting a chunk size of 0 in ArrayV2Metadata and ShardingCodec, and no longer reading a falsy `chunks` as absent in the legacy Zarr format 2 zarr.create, reject inputs zarr 3.4.0 accepted; this PR targets the 3.4.1 patch release, which must not. zarr-developers#4431 makes the same changes, more strictly, for the next minor release. The widened chunk and shard annotations and the zarr.create docstring fix stay. Assisted-by: ClaudeCode:claude-opus-5-5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b
marked this pull request as ready for review
October 3, 2026 17:42
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This AI-written PR consolidates the logic for parsing a chunk request into one function, where previously we had 2 paths that separately handled chunk- and shard-like requests, and that led to bugs as the two paths disagreed. one function should at least keep the bugs in one place :)
🤖 AI text below 🤖
Chunk specifications are judged only by the chunk normalizer: the separate duck-typed check for rectilinear input is gone, so the Zarr format 2 and sharding restrictions apply to what the normalized grid says was declared. A chunk size is an integer by Python's integer protocol: numpy integer scalars and 0-d integer arrays are accepted wherever a chunk or shard shape is. The legacy
zarr.create(..., zarr_format=2)no longer tests the truth value of a numpy array given aschunks, so numpy arrays work there. A non-integer scalar chunk specification raises aTypeErrornaming the problem. An explicit list of chunk edges declares a rectilinear dimension everywhere, including a stored rectilinear chunk grid passed aschunks=whose edges happen to be uniform, so the Zarr format 2 and sharding restrictions treat it the same whether it is given as lists or as metadata, andzarr.createstores it as a rectilinear chunk grid, ascreate_arraydoes. ARectilinearChunkGridMetadatamade only of bare integers declares no explicit edge list, so it is read as the regular grid it describes: it is accepted for Zarr format 2 and as the chunk shape of a sharded array.Follow-up to #4374.
Problem
The chunk normalizer (
normalize_chunks_nd→normalize_chunks_1d) was not the only judge of achunks=/shards=specification. A separate duck-typed classifier,_is_rectilinear_chunks, ran on the raw input at three sites (AsyncArray._create,init_array,resolve_outer_and_inner_chunks) to decide whether the specification was rectilinear, so that the Zarr format 2 and sharding restrictions could fire. Two opinions on the same input is the shape of the bug in #4374, and they could disagree: a 0-d NumPy array counted as rectilinear because it has__iter__, andChunkGrid.from_sizescollapsed a uniform edge list to a regular dimension whilenormalize_chunks_1dkept it rectilinear.Changes
_is_rectilinear_chunksis deleted. Each site normalizes first and asks the resultingChunkGrid(is_regular). Regular and rectilinear shard specifications go through the same normalizer._chunk_int): a chunk size is anything Python's integer protocol (operator.index) accepts:int, NumPy integer scalars and 0-d integer arrays. Floats, arrays with dimensions and NumPy booleans are not integers. A Pythonboolis still anint(0 or 1), as in 3.4.0.ChunkGrid.from_sizesno longer collapses uniform edges, matchingnormalize_chunks_1d. So a storedRectilinearChunkGridMetadatapassed aschunks=is treated the same whether its edges are uniform or not:zarr.createstores it as a rectilinear grid, ascreate_arraydoes (3.4.0 storedRectilinearChunkGridMetadata(((5, 5),))as a regular[5], A rectilinear chunk spec with a short trailing chunk is silently normalized to a regular grid, changing resize semantics #4272). ARectilinearChunkGridMetadatamade only of bare integers declares no edge list, so it is the regular grid it describes (accepted for Zarr format 2 and as the chunk shape of a sharded array).zarr.create(..., zarr_format=2)usedchunks or chunk_shape, so a NumPy array such aschunks=np.array([5, 3])failed with "truth value of an array is ambiguous". A small predicate (_v2_chunks_given) now decides whetherchunkswas given: falsy values (None,0,[],(),False,np.int64(0),np.array(0),np.array([0])) still mean automatic chunking, exactly as in 3.4.0; a NumPy array with more than one element is always given.2.0,np.float64(2.0), a 0-d float array) raises the normalizer's ownTypeErrornaming it (was "object has no len()" or "len() of unsized object").normalize_chunks_nd(chunks: ChunksLike | None, ...);shard_specuses the existing aliases instead ofAny.Not changed in this patch release:
chunks=Truestill raisesValueError, and aboolinside a specification is still read as 1 or 0. Rejecting booleans with oneTypeErroris part of the 3.5.0 follow-up.Compatibility with 3.4.0
In the 244-row behaviour table run against a real 3.4.0 install, this branch changes no result that 3.4.0 handled correctly. The rows it changes:
zarr.create(chunks=np.array([5, 5]), zarr_format=2)andnp.array([])now work (3.4.0 raised a truth-value error with current NumPy); a uniformRectilinearChunkGridMetadatapassed tozarr.createis stored as rectilinear (gh-4272).check_patch.py: OK, 0 undocumented changes, 0 warnings-only changes.Tests
tests/test_unified_chunk_grid.py: one table for the legacy Zarr format 2chunksargument (every falsy spelling auto-chunks,np.array(7)→(7, 7),np.array([5, 3])→(5, 3)) and one error test fornp.array([0, 0]); the classifier's tests are replaced by one test that the Zarr format 2 and sharding restrictions recognize every rectilinear specification (whichever dimension carries the list, uniform edges included) throughcreate_arrayandzarr.create; the uniform-edge resize test covers edges given as lists and as metadata.tests/test_chunk_grids.py: NumPy booleans are rejected by the normalizer; NumPy inputs join the existing normalizer tables.Stack and release notes
src/zarr/core/chunk_grids.py,resolve_outer_and_inner_chunks: keep this PR's single normalizer call and add fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334's unit,outer = normalize_chunks_nd(shard_spec, array_shape, unit=chunks.chunk_shape).tests/test_unified_chunk_grid.py: keep both imports (import refrom fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334,from functools import partialfrom this PR).main(cd1e5b3) into this branch. The one conflict was inChunkGrid.from_sizes, wheremainhad switched the uniform-edge collapse toceildiv_int; the resolution keeps this PR's removal of that collapse._chunk_listno longer checksisinstance(chunks, Iterable)beforelist(chunks). That check rejected sequences that define only__getitem__and__len__, which 3.4.0 accepted aschunks=. Two cases intest_normalize_chunkscover this. The full test suite passes (11765).🤖 Generated with Claude Code