Skip to content

Reduce ClinVar control refresh memory and runtime #890

Description

@bencap

The ClinVar refresh job uses up to 5 GB of memory and may run past its time limit on large score sets. This issue makes it read only the ClinVar rows it needs, and look up each allele once instead of twelve times.

Problem

refresh_clinvar_controls (worker/jobs/external_services/clinvar.py) is heavy on memory and time. A 2026 ClinVar release takes about 3 GB in memory once parsed (lib/clinvar/constants.py), and one job peaks around 5 GB. The CAID-to-ClinVar allele ID lookup repeats for each of the 12 releases, so large score sets likely run past the 150-minute job limit.

Scope

  • Stream each release's TSV and keep only the rows for the job's alleles, instead of loading the whole release.
  • Resolve allele IDs once per job, before the release loop. ClinVar control refresh can freeze the worker on a row lock #882 does this as part of its lock fix, so land after it.
  • Record peak memory and runtime for the largest score set before and after, in this issue.

Acceptance criteria

  • Peak memory for one job stays under 1 GB on the largest score set.
  • The largest score set's refresh finishes inside the job limit.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    app: workerTask implementation touches the worker

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions