Skip to content

CI Intermittent Windows fatal exception: access violation in tests/empirical_scope_observation.py #261

Description

@lesteve

Partly generated by an agent, but I did double-check with my brain on 😉.

cc @itamarst in case you have some insights or you have seen something similar when putting this empirical tests together.

I may try to investigate further on my Windows VM to see if I can at least reproduce.

The problem

The Windows pylatest_conda_forge_openblas CI job intermittently crashes while running:

PYTHONPATH=. python -X faulthandler tests/empirical_scope_observation.py blas

+ python -X faulthandler tests/empirical_scope_observation.py blas
Windows fatal exception: access violation

Thread 0x00000f48 [ThreadPoolExecutor-0_0Windows fatal exception: ]access violation

 (most recent call first):

Current thread's C stack trace (most recent call first):
  <cannot get C stack on this system>
  File "D:\a\threadpoolctl\threadpoolctl\tests\empirical_scope_observation.py", line 58 in blas_math
  File "C:\Users\runneradmin\miniconda3\envs\testenv\Lib\concurrent\futures\thread.py", line 73 in run
  File "C:\Users\runneradmin\miniconda3\envs\testenv\Lib\concurrent\futures\thread.py", line 86 in run
  File "C:\Users\runneradmin\miniconda3\envs\testenv\Lib\concurrent\futures\thread.py", line 119 in _worker
  File "C:\Users\runneradmin\miniconda3\envs\testenv\Lib\threading.py", line 1024 in run
Windows fatal exception: access violation


Current thread's C stack trace (most recent call first):
  <cannot get C stack on this system>
./continuous_integration/run_tests.sh: line 34:  1491 Segmentation fault         PYTHONPATH=. python -X faulthandler tests/empirical_scope_observation.py blas
Error: Process completed with exit code 139.

Unfortunately, faulthandler doesn't show anything useful on Windows hence the cannot get C stack on this system

Environment

The September 24 failure and September 26 successful run both use:

  • Python 3.14.7, with the GIL enabled
  • NumPy 2.5.3
  • conda-forge OpenBLAS 0.3.34, pthreads build
  • The same threadpoolctl commit, f6bef5dca966327941a87f4a92673cd9c0096add

CI logs

It seems to happen quite often

To investigate further

It could be useful to capture a native crash dump with ProcDump and inspect it with WinDbg.

LLM suggestion

For this CI crash, use ProcDump to capture a dump, then inspect it with WinDbg. After downloading Microsoft ProcDump, run this in PowerShell with the conda environment activated and the repository as the working directory:

$env:PYTHONPATH = "."
New-Item -ItemType Directory -Force dumps | Out-Null

.\procdump64.exe -accepteula -ma -e 1 -f C0000005 -n 1 `
    -x dumps "$env:CONDA_PREFIX\python.exe" `
    -X faulthandler tests/empirical_scope_observation.py blas

This captures a full dump at the first access violation (C0000005), before application exception handlers run. Upload dumps/*.dmp as a CI artifact, using if: always() on the upload step. ProcDump options

Open the dump in WinDbg and run:

.symfix
.reload
!analyze -v
.ecxr
kv
~* kb

.ecxr selects the exception context; kv shows its native stack; ~* kb shows every thread’s stack.

Microsoft debugging guidance

Symbol availability determines the detail: without matching symbols, you may get DLL names and offsets rather than internal function names or source lines. Preserve the exact Python, NumPy and OpenBLAS binaries from the failing environment, plus any matching debug symbols. Even an unsymbolized trace can help establish which native library is crashing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions