fix(delayed_job): stop losing jobs to SIGTERM, report permanent failures - #2947
Merged
mroderick merged 3 commits intoSep 28, 2026
Merged
Conversation
Removing `raise_signal_exceptions = :term` makes TERM finish the running job before exit, so restarts and scheduler one-off runtime kills no longer interrupt in-flight jobs and count them against max_attempts. Queued jobs simply wait for the next scheduler run (issue #2942). The `:failure` lifecycle hook sends a Rollbar warning when a job exhausts max_attempts, so permanently lost jobs are visible. Also fixes a Rails/FilePath offense on the logger line.
delayed_jobs accumulates failed jobs (destroy_failed_jobs = false); production currently holds 2,184, nearly all from a single 2020 incident. The task deletes failed jobs older than an optional cutoff date, one year by default.
mroderick
marked this pull request as ready for review
September 27, 2026 09:20
Collaborator
Author
|
Follow-up: pruning the existing failed-jobs pile (for after this merges and deploys) Preview the count before deleting anything: heroku pg:psql --app codebar-production -c "SELECT count(*) FROM delayed_jobs WHERE failed_at IS NOT NULL AND failed_at < '2025-01-01';"Then delete with the new rake task: heroku run "bundle exec rake delayed_jobs:prune_failed[2025-01-01]" -a codebar-productionThe 2025-01-01 cutoff keeps the recent SIGTERM failures referenced in #2942 (evidence for this PR) while dropping the bulk of the pile, most of which is from a single 2020 incident. |
mroderick
enabled auto-merge
September 28, 2026 11:17
mroderick
deleted the
fix/delayed-job-sigterm-and-failure-visibility
branch
September 28, 2026 11:19
Collaborator
Author
|
I ran the commands after deployment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2942
Removing
Delayed::Worker.raise_signal_exceptions = :termchanges what a SIGTERM does to the job pipeline. Reviewers should focus on that behaviour change, summarised here.Behaviour change
Before: a TERM signal raised a
SignalExceptioninside the worker, interrupting the job in flight. The interrupted job consumed one of three attempts; after three such events the job was permanently marked failed and its side effect (typically an invitation email) was silently lost.After: on TERM the worker finishes the current job, then exits. Any jobs still queued are picked up by the next Heroku Scheduler run (
rake jobs:workoff, every 10 minutes), so a kill during an invitation send delays delivery by minutes instead of losing emails. Heroku's SIGKILL (30s after an unhandled TERM) is also recoverable:Delayed::Plugins::ClearLocksclears stale locks at worker boot.Why the setting can be removed safely here
dddc1a3e, PR Revamp #230) for a persistent worker-dyno setup, and PR Procfile: Drop the Worker process type, and use Heroku Scheduler instead #2294's move to scheduler draining changed its blast radius without anyone revisiting it.Also in this PR
Delayed::Worker.lifecycle.after(:failure)hook reports permanently failed jobs (attempts exhausted) to Rollbar with the job id, attempt count, first line of the error, and first line of the handler, so silent accumulation like the current 2,184-row pile cannot happen unobserved.delayed_jobs:prune_failed[before]rake task deletes failed jobs older than an optional cutoff (one year by default), for housekeeping of the existing pile.Deliberately not done
send_*_emailsjobs resumable — the September 2026 InvitationManager changes (commit26d94993) already make re-runs safe.Verification detail
The
:failurecallback signature(worker, job)and its post-fail!timing were verified against delayed_job 4.2.0 (lib/delayed/worker.rb#failed,lib/delayed/lifecycle.rb). In production the only TERM sources for job-draining dynos are the scheduler's 1-hour one-off cap and manual dyno stops; deploys do not restart scheduler dynos.New specs cover: the removed setting, the Rollbar hook invocation, and the rake task (default cutoff, explicit cutoff, cutoff before all failures). Full suite: 1,557 examples, 0 failures; RuboCop clean on changed files.