Skip to content

No supported way to cancel a job that is already executing! #806

Description

@epugh

Summary

SolidQueue supports cancelling a queued (ready/scheduled) job cleanly via job.discard (as clarified in #395). But there is no supported way to cancel a job that is already executing (claimed) — ClaimedExecution#discard explicitly raises:

def discard
  raise UndiscardableError, "Can't discard a job in progress"
end

In Quepid we have very long "LLM as a judge" type jobs taht could run for many many minutes or hours... And you might say "oh, crap, it's not what I want" and then it's awkward. We wrote a bunch of janky code to support this.

Image

Current behavior

Applications that want a user-facing "Cancel" button for a long-running job (ours: an AI judging run scoped to one book+judge, potentially processing hundreds of records) have no sanctioned way to request cancellation of an already-claimed job. Our workaround bypasses the guard directly:

def self.cancel book, judge
  active_for(book, judge).each do |job|
    if job.claimed_execution.present?
      # Job is actively running — force destroy it. The job's own #perform
      # loop checks for its own SolidQueue row on every iteration and stops
      # as soon as it notices this row is gone.
      job.claimed_execution.destroy
      job.destroy
    else
      job.discard
    end
  end
end

This only works because ClaimedExecution#finalize's unless_already_finalized check (self.class.unscoped.lock.find_by(id: id)) happens to tolerate the claimed_execution row already being gone by the time the job actually finishes — but that's an internal implementation detail we're relying on, not a documented contract, and it could change between versions without notice.

It also means the running job's #perform never gets any signal that cancellation was requested other than a self-written polling loop:

cancellable = SolidQueue::Job.exists?(active_job_id: job_id)
loop do
  break if cancellable && !SolidQueue::Job.exists?(active_job_id: job_id)
  # ... do one unit of work ...
end

There's no cooperative "cancellation requested" flag to check cheaply, and no built-in helper for this pattern either — every app doing cooperative cancellation re-derives the same polling idiom.

Why this belongs in SolidQueue, not application code

ClaimedExecution, #finalize, and the UndiscardableError guard are all internal to SolidQueue; there's no supported extension point for "cancel this specific already-running job" without reaching past that guard into internals that could change between versions.

Related

#395 covers cancelling a scheduled (not-yet-executing) job via discard — this issue is specifically about the claimed/executing case, which that thread doesn't touch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions