Summary
Several external calls made by worker jobs set no timeout. A stalled upstream holds a thread, a process-pool slot or the event loop indefinitely. ARQ's job timeout cancels the awaiting coroutine but not the thread or process making the call. Stuck calls pile up until a pool is exhausted, and then every job that uses that pool queues forever.
Problem
| Call |
Timeout today |
Runs on |
Allele Registry GET, lib/clingen/allele_registry.py (run_in_executor(None, requests.get, url)) |
none |
default thread pool |
dcd-mapping POST, lib/mapping/client.py |
none |
process pool (max_workers=4), shared with LDH submission |
ClinVar TSV streamed download, lib/clinvar/utils.py |
none, and it holds a FileLock that has no timeout |
thread pool |
UniProt ID mapping, lib/uniprot/id_mapping.py |
none, plus a blocking time.sleep(15) |
event loop |
gnomAD Athena query, lib/gnomad.py |
no client-side limit |
event loop |
The ClinGen CAR and LDH calls in lib/clingen/services.py already use (5, 300), and the Ensembl/VEP calls are bounded.
Failure scenarios:
- ClinGen stalls.
warm_clingen_cache workers plus the ClinVar allele ID lookups fill the default thread pool, and every run_in_executor(None, …) in the worker queues behind them.
- dcd-mapping hangs. The job's
wait_for ends the job, but the pool process stays stuck. Four stuck processes starve LDH submission.
- NCBI stalls mid-download. The thread keeps the
FileLock, which blocks every ClinVar job for that release.
- UniProt or Athena is slow. The event loop blocks, so every job on the worker stops, including the heartbeat and cleanup cron.
Proposed behavior
| Dependency |
Timeout |
| Allele Registry GET |
(5, 30) |
| dcd-mapping |
(10, _map_budget_seconds). The response isn't streamed, so the read timeout is effectively the total. |
| UniProt |
(5, 60). Replace time.sleep with await asyncio.sleep. |
| ClinVar TSV |
(10, 120). The read timeout applies between chunks, so large files still finish. Also FileLock(timeout=1800). |
| Athena |
botocore Config(connect_timeout=10, read_timeout=60). Run the query off the loop inside asyncio.wait_for. |
- Add
lib/http.py with a requests.Session that applies a default timeout when a call doesn't pass one. Give each dependency a timeout constant that an environment variable can override, following the existing CLINGEN_* settings. Route the calls above through it.
- Enable ruff
S113 to catch bare requests.* calls without a timeout.
- Delete
request_with_backoff in lib/utils.py, which has no callers.
Tests
Assert the timeout each module passes, using requests_mock (last_request.timeout) or a patched session:
tests/lib/clingen/test_allele_registry.py
tests/lib/clingen/test_services.py
tests/lib/clinvar/test_utils.py
tests/lib/uniprot/test_id_mapping.py
- a new
tests/lib/mapping/test_client.py
Summary
Several external calls made by worker jobs set no timeout. A stalled upstream holds a thread, a process-pool slot or the event loop indefinitely. ARQ's job timeout cancels the awaiting coroutine but not the thread or process making the call. Stuck calls pile up until a pool is exhausted, and then every job that uses that pool queues forever.
Problem
lib/clingen/allele_registry.py(run_in_executor(None, requests.get, url))lib/mapping/client.pymax_workers=4), shared with LDH submissionlib/clinvar/utils.pyFileLockthat has no timeoutlib/uniprot/id_mapping.pytime.sleep(15)lib/gnomad.pyThe ClinGen CAR and LDH calls in
lib/clingen/services.pyalready use(5, 300), and the Ensembl/VEP calls are bounded.Failure scenarios:
warm_clingen_cacheworkers plus the ClinVar allele ID lookups fill the default thread pool, and everyrun_in_executor(None, …)in the worker queues behind them.wait_forends the job, but the pool process stays stuck. Four stuck processes starve LDH submission.FileLock, which blocks every ClinVar job for that release.Proposed behavior
(5, 30)(10, _map_budget_seconds). The response isn't streamed, so the read timeout is effectively the total.(5, 60). Replacetime.sleepwithawait asyncio.sleep.(10, 120). The read timeout applies between chunks, so large files still finish. AlsoFileLock(timeout=1800).Config(connect_timeout=10, read_timeout=60). Run the query off the loop insideasyncio.wait_for.lib/http.pywith arequests.Sessionthat applies a default timeout when a call doesn't pass one. Give each dependency a timeout constant that an environment variable can override, following the existingCLINGEN_*settings. Route the calls above through it.S113to catch barerequests.*calls without a timeout.request_with_backoffinlib/utils.py, which has no callers.Tests
Assert the timeout each module passes, using
requests_mock(last_request.timeout) or a patched session:tests/lib/clingen/test_allele_registry.pytests/lib/clingen/test_services.pytests/lib/clinvar/test_utils.pytests/lib/uniprot/test_id_mapping.pytests/lib/mapping/test_client.py