Skip to content

Close job lifecycle gaps left after stuck-job detection #888

Description

@bencap

Problem

Job lifecycle gaps remain after the stuck-job work in #788.

  • A score set's mapping_state stays processing forever after the mapping job is cancelled on timeout. The reset in worker/jobs/variant_processing/mapping.py runs only on ordinary exceptions, and Redesign stuck-RUNNING-job detection: progress-heartbeat + abort-first recovery #788 fixes only the job-run row.
  • A job outside a pipeline goes back to PENDING on retry, and nothing re-enqueues it (worker/lib/decorators/job_management.py). Cleanup (worker/jobs/system/cleanup.py) then counts it as stalled and uses up another retry, so max_retries=3 gives one real retry.
  • scripts/run_job.py enqueues a job before committing its row, so the worker can pick up a row that doesn't exist yet.
  • Cleanup holds a row lock across a Redis await, and checks QUEUED jobs with no grace period, which can enqueue a job twice.

Scope

  • Reset mapping_state on cancellation as well as on exceptions.
  • Re-enqueue a retried job that isn't in a pipeline.
  • Commit the job row before enqueueing in run_job.py.
  • Release the row lock before the Redis await in cleanup, and give QUEUED jobs a grace period before they count as stalled.

Acceptance criteria

  • A test cancels a mapping job on timeout and asserts mapping_state isn't left processing.
  • A test asserts a standalone job with max_retries=3 runs up to four times.
  • A test asserts run_job.py commits before it enqueues.
  • worker/pipeline_management.md and the comments in cleanup.py match the new behavior.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    app: workerTask implementation touches the worker

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions