Skip to content

feat: read-only builtin zip: URL pipeline adapter - #4477

Draft
jhamman wants to merge 16 commits into
zarr-developers:mainfrom
jhamman:feature/url-pipeline-zip
Draft

jhamman wants to merge 16 commits into
zarr-developers:mainfrom
jhamman:feature/url-pipeline-zip

Conversation

@jhamman

@jhamman jhamman commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

(not ready for review yet)

Summary

Adds a builtin, read-only zip: URL pipeline adapter
(schemes/zip.md). This makes
the headline example from #2943 work:

zarr.open_group("/data/my-archive.zip|zip:inner-dir|zarr3:subgroup", mode="r")
zarr.open_group("s3://bucket/data.zip|zip:path/in/archive", storage_options={"anon": True})
zarr.open_array("file:/data/outer.zip|zip:inner.zip|zip:x")          # nested archives

How it works (src/zarr/storage/_url_adapters/_zip.py):

  • The adapter resolves the preceding pipeline with context.resolve_preceding(mode="r").
    The archive is the value at preceding.path in preceding.store, or the store's root
    object when the path is empty (a local file or an fsspec object).
  • ZipReaderStore is a thin, read-only store over that resource. It uses async byte-range
    requests only
    :
    • At open, it reads the archive's size and the last 64 KiB + 22 bytes. It parses the
      central directory with the stdlib zipfile in asyncio.to_thread, over a sparse
      in-memory view; any range the view lacks is fetched and the parse retried. This handles
      central directories larger than the initial read.
    • On get, it reads the member's local header once (the offset is cached) and then the
      member's data.
    • Stored (uncompressed) members get true range reads. Deflate and bzip2 members are fetched
      whole and decompressed, in a thread above 1 MiB. CRC-32 is verified. Other methods (for
      example LZMA) and encrypted members raise NotImplementedError.
    • Directory entries (dir/) are not keys.
    • close() closes the preceding store, which the reader owns.
  • Body: one leading / is ignored, and //... is invalid. A query, . or .. segments
    raise URLPipelineError. The body becomes the residual path, and every other preceding
    field (for example zarr_format) is carried forward with dataclasses.replace.
  • Modes: "w", "w-" and "r+" raise URLPipelineError("zip: is read-only until writable ZIP support lands; ..."). "r", None and "a" open read-only; "a" (the zarr.open
    default) serves the "open" half, per the contract from feat: url-pipeline core — parser, adapter ABC, registry, store hooks #4192.
  • Errors: a resource that is not a ZIP archive, a directory, or an unreadable root raises
    URLPipelineError chained from the original exception, and the preceding store is closed.
  • zip is an adapter-only scheme, like zarr* in feat: builtin zarr:, zarr2: and zarr3: URL pipeline adapters #4476. A string without |, such as
    fsspec's zip://... or zip::file://..., behaves exactly as before.

Why not ZipStore?

The adapter must work for local and remote (fsspec) roots and must not block zarr's I/O
event loop. ZipStore needs a local path or a sync file object, and it does blocking
zipfile I/O inside its async methods. Over an fsspec root that means either downloading the
archive or making sync, loop-blocking reads through an fsspec file object. The reader here
uses only the preceding store's async get(..., byte_range=...) and getsize. So one code
path serves local files, fsspec objects, memory: entries, and archives nested in other
archives. For an uncompressed archive (ZipStore's default), a chunk read is a single range
request.

The get("") quirk (F2)

Reading an archive that is the pipeline root relies on LocalStore / FsspecStore treating
the empty key as the root object. That reliance is isolated in two helpers,
_read_resource_range and _resource_size. The file-resource primitive in the next PR
replaces both.

Depends on

Closes

Decision needed

  • Thin range-reading reader vs. ZipStore. Implemented: the thin reader described above.
    Alternative: open ZipStore on a local path (in a thread) for local roots, and on an fsspec
    file object for remote roots. That is less code, but it blocks the event loop and does not
    compose with nested archives.
  • Mode "a" on zip:. Implemented: open read-only, following the open-or-create
    contract. Alternative: reject every mode except "r" until writes are supported. That is
    stricter, but zarr.open(url) (default "a") would then fail on a perfectly readable
    archive.

For reviewers

  • _SparseFile / _parse_central_directory: is lazily fetching ranges into the stdlib
    zipfile parser robust enough? It raises a private exception on a missing range and
    retries after fetching it.
  • Compressed members are not cached. A compressed nested archive is therefore decompressed
    once per read of the inner archive. Stored members, the common case for Zarr ZIP stores,
    are read with ranges.

Tests

uv sync --frozen
uv run --frozen ruff check src tests
uv run --frozen mypy src
uv run --frozen --with pytest-examples pytest -n auto tests
pre-commit run --all-files

Full suite: 13053 passed, 1295 skipped, 9 xfailed. ruff, mypy and pre-commit are clean.

New tests are in tests/test_url_pipeline/test_zip_adapter.py:

  • archives written by ZipStore, opened as file:/…/a.zip|zip:, file://…,
    file://localhost… and schemeless paths;
  • /…/a.zip|zip:inner|zarr3:sub and the headline example;
  • fsspec roots: local:// and an in-memory fsspec filesystem standing in for a remote
    object store;
  • a memory: entry, and nested zip: (stored and deflated outer archives);
  • stored, deflate and bzip2 members; directory entries; a central directory larger than the
    first read; byte ranges; a CRC mismatch;
  • mode rules, malformed bodies, non-archives, and closing the preceding store on error and on
    close().

Author attestation

  • I am a human, these are my changes, and I have reviewed and understood every change and can explain why each is correct.

TODO

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.md
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

🤖 Generated with Claude Code

jhamman and others added 16 commits August 19, 2026 09:16
Implements URL pipeline support (https://github.com/jbms/url-pipeline):
'|'-chained URLs resolve through pluggable adapters registered under the
'zarr.url_adapters' entry-point group (entry-point name = URL scheme).

- zarr.abc.url_pipeline: PipelineSegment, AdapterResolution,
  PipelineContext, URLPipelineAdapter (single-classmethod contract)
- zarr.storage._url_pipeline: parse_pipeline / resolve_pipeline; the root
  sub-URL delegates to make_store so existing file/memory/fsspec routing
  is unchanged
- registry: register_url_adapter / get_url_adapter /
  list_url_adapter_schemes (name check only; no adapter imports)
- make_store/make_store_path route strings containing '|' (or a
  registered root scheme) through the resolver; residual store paths
  combine with the user-supplied path
- StorePath gains a zarr_format attribute (populated by format segments
  in a follow-up)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e-core

# Conflicts:
#	src/zarr/errors.py
#	src/zarr/registry.py
…thoring

- make_store validates the access mode before routing a URL pipeline
  string to adapters, matching every other StoreLike branch
- make_store docstring lists URL pipeline strings; StorePath.open
  documents the invalid-mode ValueError
- user guide: StoreLike bullet for pipeline strings and an adapter
  authoring section in extending.md (runnable example)
- changelog: drop PR-relative wording, name the public parser entry points

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
get_url_adapter took a non-reentrant lock across entry_point.load(), so an
adapter module that resolved another scheme at import time deadlocked the
process. Matching entry points are now taken off the pending list under the
lock and imported outside it; import failures are wrapped in URLPipelineError
and leave the entry point discoverable for a retry.

A same-named entry point is discarded with a ZarrUserWarning when the scheme
is already registered (e.g. by a builtin), instead of staying pending forever;
duplicate entry-point names warn and the first wins.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ent contracts

- The resolver closes the store an adapter returned when it cannot be made
  read-only for mode 'r', instead of leaking it.
- resolve_pipeline documents every error type it lets through; the
  resolve_preceding docstring states the local-file-root limitation.
- AdapterResolution and PipelineContext compare and hash by identity, since
  a Store is not hashable and frozen=True implied otherwise.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… caller

The zarr2:/zarr3: format carried on StorePath.zarr_format was merged only in
zarr.api.asynchronous; create_array, from_array, Group.from_store,
Group.open, Array.open and the deprecated AsyncArray._create dropped it and
swallowed explicit conflicts. The merge now lives on
StorePath.resolve_zarr_format and is applied at each site.

To let the pipeline's format apply when the caller does not specify one,
create_array (sync and async) and Group.from_store default zarr_format to
None, resolving to the configured default (3) as before; Array.open gains a
zarr_format parameter mirroring Group.open.

Tests: registry deadlock/race/import-failure/shadowing, resolver leak,
zarr_format merge at each core site, and pins for the documented memory:/
file: divergences and the local-file-root limitation (xfail, strict).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- The extending guide's wrapper example dropped the preceding residual path
  while its comment claimed the opposite; it now joins the two and a tested
  snippet shows root|a|b resolving to a/b.
- Changelog and storage guide state the memory:// routing split with fsspec,
  the undecoded percent-escapes in file: roots, the local-file-root
  limitation, the entry-point collision warnings, and the zarr_format
  default changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ature/url-pipeline-core

# Conflicts:
#	src/zarr/api/asynchronous.py
#	src/zarr/core/group.py
…leSystemWrapper

fsspec < 2024.12.0 (the min_deps job) cannot open a plain memory:// URL at
all, so the routing comparison only applies with fsspec>=2024.12.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`Group.open`, `Array.open` and their async variants defaulted to
`zarr_format=3`, so a pipeline with a `zarr2:` segment raised a conflict
unless the caller passed `zarr_format=None`.

The default is now a private "unspecified" sentinel. When the caller does
not pass a format, a URL pipeline decides: its format segment if any,
otherwise auto-detection. Every other store keeps the default of 3, so the
metadata probing of non-pipeline callers is unchanged. An explicit format
(including None) behaves as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The format adapters from the URL pipeline spec (schemes/zarr.md). The
segment body is a path within the preceding resource and is joined onto
the preceding residual path; `zarr2:` / `zarr3:` record the Zarr format and
`zarr:` leaves it to auto-detection (keeping a format pinned by an earlier
segment). Every other field of the preceding resolution is carried forward
with `dataclasses.replace`, so a format segment also works as an
intermediate segment.

The builtins are registered as lazily loaded entry points ahead of the
third-party ones in `_collect_entrypoints`, so importing zarr does not
import them, and a third-party `zarr.url_adapters` entry point with the
same name is ignored with a warning. They are adapter-only schemes: a
string without `|` such as `zarr3:foo` is never routed to them as a
pipeline root, so its meaning is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uide

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`zip:` (schemes/zip.md) addresses an entry or directory inside the ZIP
archive that the preceding pipeline resolves to. The preceding pipeline is
resolved with mode "r", and the archive is read through that store with
async byte-range requests only: the central directory is fetched from the
end of the archive (parsed in a worker thread), and each member is fetched
and decompressed on demand, with true range reads for stored members.

A thin read-only reader over the preceding store's bytes is used instead
of `ZipStore`, because `ZipStore` does blocking file I/O and needs a local
path or a sync file object. The reader works the same way for local files,
fsspec objects, `memory:` entries and nested `zip:` archives, and never
blocks the IO event loop.

Reading a root that *is* the archive relies on `get("")` addressing the
root object (follow-up F2); that is isolated in `_read_resource_range` /
`_resource_size` so the file-resource primitive can replace it.

Writable modes ("w", "w-", "r+") raise `URLPipelineError`; "a" and None
open read-only. One leading "/" in the body is ignored and more than one is
invalid. The preceding store is closed when the archive cannot be opened,
and the reader closes it on `close()`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant