Skip to content

Worker outbound calls have no timeout #883

Description

@bencap

Summary

Several external calls made by worker jobs set no timeout. A stalled upstream holds a thread, a process-pool slot or the event loop indefinitely. ARQ's job timeout cancels the awaiting coroutine but not the thread or process making the call. Stuck calls pile up until a pool is exhausted, and then every job that uses that pool queues forever.

Problem

Call Timeout today Runs on
Allele Registry GET, lib/clingen/allele_registry.py (run_in_executor(None, requests.get, url)) none default thread pool
dcd-mapping POST, lib/mapping/client.py none process pool (max_workers=4), shared with LDH submission
ClinVar TSV streamed download, lib/clinvar/utils.py none, and it holds a FileLock that has no timeout thread pool
UniProt ID mapping, lib/uniprot/id_mapping.py none, plus a blocking time.sleep(15) event loop
gnomAD Athena query, lib/gnomad.py no client-side limit event loop

The ClinGen CAR and LDH calls in lib/clingen/services.py already use (5, 300), and the Ensembl/VEP calls are bounded.

Failure scenarios:

  • ClinGen stalls. warm_clingen_cache workers plus the ClinVar allele ID lookups fill the default thread pool, and every run_in_executor(None, …) in the worker queues behind them.
  • dcd-mapping hangs. The job's wait_for ends the job, but the pool process stays stuck. Four stuck processes starve LDH submission.
  • NCBI stalls mid-download. The thread keeps the FileLock, which blocks every ClinVar job for that release.
  • UniProt or Athena is slow. The event loop blocks, so every job on the worker stops, including the heartbeat and cleanup cron.

Proposed behavior

Dependency Timeout
Allele Registry GET (5, 30)
dcd-mapping (10, _map_budget_seconds). The response isn't streamed, so the read timeout is effectively the total.
UniProt (5, 60). Replace time.sleep with await asyncio.sleep.
ClinVar TSV (10, 120). The read timeout applies between chunks, so large files still finish. Also FileLock(timeout=1800).
Athena botocore Config(connect_timeout=10, read_timeout=60). Run the query off the loop inside asyncio.wait_for.
  • Add lib/http.py with a requests.Session that applies a default timeout when a call doesn't pass one. Give each dependency a timeout constant that an environment variable can override, following the existing CLINGEN_* settings. Route the calls above through it.
  • Enable ruff S113 to catch bare requests.* calls without a timeout.
  • Delete request_with_backoff in lib/utils.py, which has no callers.

Tests

Assert the timeout each module passes, using requests_mock (last_request.timeout) or a patched session:

  • tests/lib/clingen/test_allele_registry.py
  • tests/lib/clingen/test_services.py
  • tests/lib/clinvar/test_utils.py
  • tests/lib/uniprot/test_id_mapping.py
  • a new tests/lib/mapping/test_client.py

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    app: workerTask implementation touches the worker

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions