Skip to content

CI: fail when the ephy_testing_data HEAD hash cannot be resolved - #1912

Merged
alejoe91 merged 1 commit into
NeuralEnsemble:masterfrom
AxelNoun:ci/fail-fast-ephy-hash
Oct 8, 2026
Merged

alejoe91 merged 1 commit into
NeuralEnsemble:masterfrom
AxelNoun:ci/fail-fast-ephy-hash

Conversation

@AxelNoun

@AxelNoun AxelNoun commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Closes #1896.

The step that computes the ephy_testing_data hash writes dataset_hash=$(git ls-remote ... | cut -f1) with echo, so its exit status is that of echo. When git ls-remote fails against gin, the step stays green, dataset_hash is empty and the cache key becomes Linux-datasets-.

This PR replaces the body of that step in io-test.yml, caches_cron_job.yml and plexon2-testing.yml:

set -o pipefail
if ! hash="$(git ls-remote https://gin.g-node.org/NeuralEnsemble/ephy_testing_data.git HEAD | cut -f1)" || [ -z "$hash" ]; then
  echo "::error::could not resolve ephy_testing_data HEAD on gin; refusing to build a degenerate cache key"
  exit 1
fi
echo "dataset_hash=$hash" >> "$GITHUB_OUTPUT"

Step ids (ephy_testing_data) and step names are unchanged, so steps.ephy_testing_data.outputs.dataset_hash still resolves. The assignment is inside the if on purpose: caches_cron_job.yml runs steps under bash -e, where, with pipefail on, a bare assignment exits with git's code before the ::error:: line. The snippet I first posted in #1896 had that problem and the wrong step id; the issue body is now corrected.

What the logs show

Counts use the latest attempt of each run, as of 2026-10-07.

  • io-test (NeoIoTest-automatic-trigger), runs created 2026-07-13 to 2026-10-06: in 9 runs (8 failed, 1 cancelled), one job got HTTP 403 on git ls-remote, computed Linux-datasets-, got an exact cache hit on it, and pytest then reported 94 to 100 errors, all datalad IncompleteResultsError (595 to 631 error: 403 lines per job). In all 9, the other job of the same run computed a valid key and logged no 403. None of the 112 jobs of the 56 green runs in that period had an empty key. Recent examples: master, run 37334011073 (2026-10-05); Release 0.14.6 #1910, run 37214229194 (2026-10-04). The first attempt of run 37517882872 (Read the lower word of PLX timestamps as unsigned #1904, 2026-10-06) showed the same pattern; its rerun on 2026-10-07 passed, so that run is counted as green.
  • caches_cron_job, the 80 runs created 2026-07-26 to 2026-10-06: 12 computed an empty key and concluded success (HTTP 403: 6; connection reset: 3; connection timeout after about 134 to 136 s: 3). The run of 2026-09-27 (36334511093) logged Cache saved with key: Linux-datasets-; that entry (882,953,306 bytes) is still listed.
  • That entry was also restored on a primary-key miss. In the rerun of 36097083501 on 2026-10-06 (attempt 3), both jobs computed Linux-datasets-43e19611…, and both logs show Cache hit for restore-key: Linux-datasets- and a cache size of 882,953,306 bytes, i.e. the degenerate entry rather than the most recent regular entry at that time (Linux-datasets-d7796f59…, 745,114,516 bytes). Tests passed there; download_dataset runs dataset.update(merge=True) on an existing clone before fetching files.

Cost

  • This does not make tests pass when gin refuses a runner (Gin/G-node is being flaky in actions #1897). The affected job now stops at the hash step with an explicit error instead of failing later in pytest. With fail-fast: true, the other job of the run is then cancelled within seconds rather than after 130 to 199 s (the duration of the failing jobs above); it was already cancelled in all 9 io-test runs above.
  • The cache cron will turn red on those days: the 12 runs above would have failed instead of succeeding. If you prefer that workflow to stay green, an alternative for it alone is to skip the cache and download steps when the hash is empty, at the cost of making the failure silent again.

Why no retry loop

In the two failing jobs I read in full (111844044441 and 111471365590), every access to gin got a 403, from the hash step to the last datalad fetch about 2 min 18 s and 2 min 8 s later (no successful fetch), while the other job of the same run reached gin normally. In the timeout mode, one git ls-remote attempt took about 134 to 136 s. A retry inside the same job would therefore mostly add time.

After merge

Could a maintainer delete the Linux-datasets- entry (id 8181158804, created 2026-09-27)? I do not have the rights to do so. After this change no job computes that key any more, but the entry can still be restored through restore-keys, as shown above. The command would be gh cache delete 8181158804 -R NeuralEnsemble/python-neo.

Testing

  • Locally (bash 5.2.37 under MSYS2 on Windows, not a GitHub runner), the step body extracted from the YAML, run with bash -e, bash -eo pipefail, bash -l and bash --noprofile --norc -eo pipefail, with a temporary GITHUB_OUTPUT:
    • connection refused, and a local HTTP server answering 403: exit 1, git's fatal: line followed by the ::error:: line, nothing written to GITHUB_OUTPUT;
    • gin itself (reachable at the time): exit 0, dataset_hash=43e19611e863ec40fa40243c4ce42bdff1ba1320 written;
    • a command that succeeds with empty output (ls-remote of a nonexistent ref): exit 1 with the ::error:: line;
    • for comparison, the current form exits 0 and writes dataset_hash= in every failure case, under all four invocations.
  • The three files load with yaml.safe_load, and each step keeps id: ephy_testing_data.
  • In CI, the io-test run of this PR should use the modified io-test.yml (called from io-test_trigger.yml through a local uses:), which covers the success path. caches_cron_job.yml runs on push to master, so the merge will exercise it. plexon2-testing.yml only runs weekly or on dispatch, and currently fails for an unrelated reason (CI: drop Python 3.9 from the NeoPlexon2Test matrix #1911).

The hash step wrote `dataset_hash=$(git ls-remote ... | cut -f1)` with
echo, so its exit status was the one of echo. When git ls-remote failed
against gin (HTTP 403, connection failure), the step stayed green, wrote
an empty dataset_hash and the cache key degraded to `Linux-datasets-`.

Enable pipefail, test both the assignment and the emptiness of the result,
emit an ::error:: annotation and exit 1. The step id and names are
unchanged, so steps.ephy_testing_data.outputs.dataset_hash still resolves.

See NeuralEnsemble#1896.

@h-mayorquin h-mayorquin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, sorry for taking so long to review this. Thanks for your contribution.

@h-mayorquin

Copy link
Copy Markdown
Contributor

Once this merges, I will erase the other cache that was created accidentally.

@alejoe91
alejoe91 merged commit 9dd5d29 into NeuralEnsemble:master Oct 8, 2026
73 of 75 checks passed
@AxelNoun

AxelNoun commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

LGTM, sorry for taking so long to review this. Thanks for your contribution.

No problem, it's with pleasure! Do not hesitate to contact me if you need me!

@AxelNoun
AxelNoun deleted the ci/fail-fast-ephy-hash branch October 9, 2026 17:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Dataset cache-key step silently succeeds when gin is unreachable

3 participants