Skip to content

fix: retry uploads that receive an HTTP 408 (request timeout) - #618

Merged
ffumero2003 merged 2 commits into
masterfrom
retry-408-request-timeout
Oct 1, 2026
Merged

ffumero2003 merged 2 commits into
masterfrom
retry-408-request-timeout

Conversation

@ffumero2003

Copy link
Copy Markdown
Contributor

Context

An upload that receives an HTTP 408 request_timeout response fails hard: the client raises
UnknownError to the caller after exactly one attempt, with no retry and no fresh upload URL.
Every other transient upload status — 401, 500, 503 — is retried. Backblaze's Integration
Checklist lists 408 as a recognized, recoverable upload failure after which the client should
fetch a new upload URL and retry, so customers on flaky networks currently see spurious upload
failures the SDK is meant to absorb.

Root cause

In b2sdk/_internal/exception.py, interpret_b2_error has an elif branch for each status it
treats specially (the 4xx codes, 429, and the whole 5xx range) but no branch for 408, so a
408 falls through to the unconditional return UnknownError(...). UnknownError is a plain
B2SimpleError — it does not mix in TransientErrorMixin, so both should_retry_upload() and
should_retry_http() inherit False from the base. The upload manager's retry gate checks
should_retry_upload() and raises immediately, so the upload is never retried.

Fix

  • Add class RequestTimeout(TransientErrorMixin, B2Error) next to ServiceError (the 5xx
    class), so both retry predicates are True.
  • Add elif status == 408: return RequestTimeout('%d %s %s' % (status, code, message))
    immediately before the 429 branch, grouping 408 with the other transient/retryable statuses.

No change to the upload manager or try_count wiring — the existing retry gate is correct; it
simply never received a retryable exception for 408.

Verification

  • New unit test in test/unit/test_exception.py: interpret_b2_error(408, 'request_timeout', 'request timeout', {}) returns RequestTimeout with should_retry_upload() and
    should_retry_http() both True. Red before the change, green after.
  • Full unit suite green across all apivers (nox -s unit).
  • Independently confirmed against a B2 simulator: the upload.retry_408 recovery scenario goes
    FAIL → PASS (the client now fetches a new upload URL and retries).

Fixes #605

@sophiecarreras sophiecarreras left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, this fixes the right thing: uploads use try_count=1, so the retry happens in UploadManager with a fresh upload URL, which is what we want after a 408. Three non-blocking requests: (1) please export RequestTimeout from b2sdk.v3.exception next to ServiceError so callers can catch it; (2) the name sits very close to the existing B2RequestTimeout/B2RequestTimeoutDuringUpload (client-side timeouts), so a docstring cross-reference or a clearer name would help; (3) consider a test that drives the upload manager through a 408 followed by success and asserts a new upload URL was fetched. CI: all unit sessions are green; the red jobs are the known test_encryption integration failures (#614), unrelated to this change.

@ffumero2003
ffumero2003 merged commit 5b89c9d into master Oct 1, 2026
10 of 28 checks passed
@ffumero2003
ffumero2003 deleted the retry-408-request-timeout branch October 1, 2026 17:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A 408 (request_timeout) response is never retried on upload, unlike 401/500/503

2 participants