Skip to content

fix(delayed_job): stop losing jobs to SIGTERM, report permanent failures - #2947

Merged
mroderick merged 3 commits into
masterfrom
fix/delayed-job-sigterm-and-failure-visibility
Sep 28, 2026
Merged

mroderick merged 3 commits into
masterfrom
fix/delayed-job-sigterm-and-failure-visibility

Conversation

@mroderick

Copy link
Copy Markdown
Collaborator

Closes #2942

Removing Delayed::Worker.raise_signal_exceptions = :term changes what a SIGTERM does to the job pipeline. Reviewers should focus on that behaviour change, summarised here.

Behaviour change

Before: a TERM signal raised a SignalException inside the worker, interrupting the job in flight. The interrupted job consumed one of three attempts; after three such events the job was permanently marked failed and its side effect (typically an invitation email) was silently lost.

After: on TERM the worker finishes the current job, then exits. Any jobs still queued are picked up by the next Heroku Scheduler run (rake jobs:workoff, every 10 minutes), so a kill during an invitation send delays delivery by minutes instead of losing emails. Heroku's SIGKILL (30s after an unhandled TERM) is also recoverable: Delayed::Plugins::ClearLocks clears stale locks at worker boot.

Why the setting can be removed safely here

  • The scheduler one-off dynos exit on their own when out of jobs, so nothing relies on TERM forcing a fast exit.
  • Per-job work is short (render + SendGrid call, well inside the 30s TERM-to-SIGKILL window).
  • The setting predates the scheduler architecture: it was added in April 2015 (commit dddc1a3e, PR Revamp #230) for a persistent worker-dyno setup, and PR Procfile: Drop the Worker process type, and use Heroku Scheduler instead #2294's move to scheduler draining changed its blast radius without anyone revisiting it.

Also in this PR

  • A Delayed::Worker.lifecycle.after(:failure) hook reports permanently failed jobs (attempts exhausted) to Rollbar with the job id, attempt count, first line of the error, and first line of the handler, so silent accumulation like the current 2,184-row pile cannot happen unobserved.
  • A delayed_jobs:prune_failed[before] rake task deletes failed jobs older than an optional cutoff (one year by default), for housekeeping of the existing pile.

Deliberately not done

Verification detail

The :failure callback signature (worker, job) and its post-fail! timing were verified against delayed_job 4.2.0 (lib/delayed/worker.rb#failed, lib/delayed/lifecycle.rb). In production the only TERM sources for job-draining dynos are the scheduler's 1-hour one-off cap and manual dyno stops; deploys do not restart scheduler dynos.

New specs cover: the removed setting, the Rollbar hook invocation, and the rake task (default cutoff, explicit cutoff, cutoff before all failures). Full suite: 1,557 examples, 0 failures; RuboCop clean on changed files.

Removing `raise_signal_exceptions = :term` makes TERM finish the running
job before exit, so restarts and scheduler one-off runtime kills no longer
interrupt in-flight jobs and count them against max_attempts. Queued jobs
simply wait for the next scheduler run (issue #2942).

The `:failure` lifecycle hook sends a Rollbar warning when a job exhausts
max_attempts, so permanently lost jobs are visible. Also fixes a
Rails/FilePath offense on the logger line.
delayed_jobs accumulates failed jobs (destroy_failed_jobs = false);
production currently holds 2,184, nearly all from a single 2020 incident.
The task deletes failed jobs older than an optional cutoff date, one year
by default.
@mroderick
mroderick marked this pull request as ready for review September 27, 2026 09:20
@mroderick

Copy link
Copy Markdown
Collaborator Author

Follow-up: pruning the existing failed-jobs pile (for after this merges and deploys)

Preview the count before deleting anything:

heroku pg:psql --app codebar-production -c "SELECT count(*) FROM delayed_jobs WHERE failed_at IS NOT NULL AND failed_at < '2025-01-01';"

Then delete with the new rake task:

heroku run "bundle exec rake delayed_jobs:prune_failed[2025-01-01]" -a codebar-production

The 2025-01-01 cutoff keeps the recent SIGTERM failures referenced in #2942 (evidence for this PR) while dropping the bulk of the pile, most of which is from a single 2020 incident.

@olleolleolle olleolleolle left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good!

@mroderick
mroderick enabled auto-merge September 28, 2026 11:17
@mroderick
mroderick merged commit ea7445b into master Sep 28, 2026
10 checks passed
@mroderick
mroderick deleted the fix/delayed-job-sigterm-and-failure-visibility branch September 28, 2026 11:19
@mroderick

Copy link
Copy Markdown
Collaborator Author

I ran the commands after deployment

heroku pg:psql --app codebar-production -c "SELECT count(*) FROM delayed_jobs WHERE failed_at IS NOT NULL AND failed_at < '2025-01-01';"
--> Connecting to ⛁ postgresql-infinite-55100
 count
-------
  2113
(1 row)

heroku run "bundle exec rake delayed_jobs:prune_failed[2025-01-01]" -a codebar-production
Running bundle exec rake delayed_jobs:prune_failed[2025-01-01] on ⬢ codebar-production... up, run.4857
deleted=2113 cutoff=2025-01-01

heroku pg:psql --app codebar-production -c "SELECT count(*) FROM delayed_jobs WHERE failed_at IS NOT NULL AND failed_at < '2025-01-01';"
--> Connecting to ⛁ postgresql-infinite-55100
 count
-------
     0
(1 row)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Delayed jobs marked failed on SIGTERM — invitation emails silently lost during restarts

2 participants