Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 11 additions & 4 deletions .github/workflows/_checks.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ jobs:

pytest:
runs-on: ubuntu-latest
continue-on-error: ${{ matrix.prerelease }}
strategy:
fail-fast: false
matrix:
Expand All @@ -28,6 +29,12 @@ jobs:
- "3.13"
- "3.14"
- "3.14t"
prerelease: [false]
include:
- python-version: "3.15"
prerelease: true
- python-version: "3.15t"
prerelease: true
steps:
- uses: actions/checkout@v6
- uses: extractions/setup-just@v4
Expand All @@ -39,12 +46,12 @@ jobs:
- run: uv python pin ${{ matrix.python-version }}
- run: just install
- name: Confirm the interpreter is free-threaded
if: matrix.python-version == '3.14t'
if: endsWith(matrix.python-version, 't')
run: uv run --no-sync python -c "import sys; assert not sys._is_gil_enabled(), 'GIL is enabled on a t-build'"
- run: just test-ci
- name: Stress the concurrency test (free-threaded only)
if: matrix.python-version == '3.14t'
run: just test tests/test_free_threading.py --count=50 -W error::RuntimeWarning
- name: Stress the thread-race tests (free-threaded only)
if: endsWith(matrix.python-version, 't')
run: just test-race --count=50

docs:
runs-on: ubuntu-latest
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ jobs:
steps:
- uses: actions/checkout@v6
- uses: extractions/setup-just@v4
- uses: astral-sh/setup-uv@v7
- uses: astral-sh/setup-uv@v8.2.0

# PyPI is irreversible, so it runs FIRST: if it fails the job stops and no
# GitHub Release is created advertising a version that never reached PyPI.
Expand Down
5 changes: 3 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,9 @@ repository** and ships as a separate PyPI package, `modern-di-pytest` included.

`just` (task runner) and `uv` (package manager). The [`justfile`](justfile) is the source of truth —
`just --list`, or read it. Every recipe carries its intent as a comment. The one thing it does not
say: nothing validates Markdown links outside `docs/`. `just docs-build` runs `mkdocs --strict` over
the site only, and root Markdown, `.github/`, and `docs/agents/` are unchecked.
say: `just docs-build` runs `mkdocs --strict` over the site only. Local links in every Markdown file,
root and `docs/agents/` included, are checked by the `links` job in CI (lychee, offline), which has
no `just` recipe.

Run `just install` before `just lint`, in every checkout and worktree. `uv.lock` is gitignored and
CI's `just install` runs `uv lock --upgrade` before it syncs, so CI always lints with the newest
Expand Down
13 changes: 10 additions & 3 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,8 @@ cost. Runs in CI (informational, non-gating) and locally via `just bench`.
| G13b | Batch of K=100 request cycles, 10 finalizer-less cached REQUEST providers, `await close_async()` | the async close loop when there is nothing to finalize |
| G14 | Concurrent cached-hit throughput, N threads (lock-free read) | free-threaded read scaling |
| G15 | Concurrent first-resolve, N threads (per-item double-checked creation lock) | free-threaded creation-lock contention |
| G15b | Concurrent first-resolve, N threads, each in its own REQUEST child | that sibling children's creations do not contend |
| G15c | Control: an empty job on the G14/G15 worker pool, N threads | barrier floor inside every G14/G15/G15b batch |
| G16 | Warm by-type `resolve(SomeType)`, small graph | `find_provider` lookup on the integration/`@inject` path |
| G17 | Warm by-type `resolve(SomeType)`, 200-provider registry | lookup cost at realistic registry scale |
| G18 | Warm resolve through an `Alias` to a cached source | the alias hop, read against G2 |
Expand Down Expand Up @@ -91,9 +93,14 @@ threshold this low workable at all.

### Concurrency (G14/G15)

G14/G15 use a custom N-thread harness (`test_guard_concurrency.py`) — pytest-benchmark
times a parallel batch of worker threads released together behind a barrier,
parametrized over thread count `{1, 2, 4}` so the scaling trend shows within one run.
G14, G15 and G15b use a custom N-thread harness (`test_guard_concurrency.py`). pytest-benchmark
times a parallel batch: one job run by each of N persistent worker threads, released together
behind a barrier and parametrized over thread count `{1, 2, 4}` so the scaling trend shows within
one run. The workers start once per benchmark, outside the timed call, so thread start-up and
join are not in the number; G15c times an empty job on the same pool, which is the floor left in
every batch. G15 compiles its resolvers once and empties the cache in an untimed per-round
setup, so each round times creation and nothing else. Each scenario asserts on what the timed
batch resolved.
The GIL vs free-threaded (PEP 703) comparison comes from running the file under each
build (same version/arch):

Expand Down
166 changes: 109 additions & 57 deletions benchmarks/test_guard_concurrency.py
Original file line number Diff line number Diff line change
@@ -1,10 +1,11 @@
# ruff: noqa: ANN001, ANN201
"""Guard tier — concurrent-resolution throughput (custom N-thread harness).

pytest-benchmark measures single-thread wall time, so these time a *parallel batch* (N worker
threads released together behind a barrier) as one unit, parametrized over thread count so the
scaling curve is visible **within a single interpreter run** (no cross-interpreter comparison
needed). Two sub-cases:
pytest-benchmark measures single-thread wall time, so these time a *parallel batch* (one job run
by each of N persistent worker threads, released together behind a barrier) as one unit,
parametrized over thread count so the scaling curve is visible **within a single interpreter run**
(no cross-interpreter comparison needed). The workers are started once per benchmark, outside the
timed call, so a batch costs two barrier crossings and no thread start-up. Sub-cases:

- G14 concurrent cached-hit: a fixed total number of reads of a warm cached singleton, split
across N threads. The cached-hit path is lock-free, so on a free-threaded build (PEP 703) the
Expand All @@ -13,6 +14,9 @@
they contend on each item's double-checked creation lock (`CacheItem.get_or_create`). Creating
one singleton is serialized by design, so this is expected *not* to scale even free-threaded: the
measured cost is the contention itself (the known trade-off vs lock-free-slot rivals).
- G15b concurrent first-resolve in sibling children: the same K cold misses, each thread in its
own REQUEST child, so no two threads share a cache item.
- G15c control: an empty job on the same pool, the harness floor inside every batch above.

Read the batch-time-vs-thread-count trend, not the absolutes. The GIL vs free-threaded comparison
comes from running the whole file under each build (same version/arch), e.g.:
Expand All @@ -25,28 +29,70 @@

import dataclasses
import threading
import typing

import pytest

from modern_di import Container, Group, Scope, providers


_THREAD_COUNTS = [1, 2, 4]


def _run_parallel(worker, n_threads: int) -> None:
# Release all threads together (barrier) so the work overlaps maximally.
barrier = threading.Barrier(n_threads)

def _target() -> None:
barrier.wait()
worker()

threads = [threading.Thread(target=_target) for _ in range(n_threads)]
for thread in threads:
thread.start()
for thread in threads:
thread.join()
_TIMEOUT = 30


class _WorkerPool:
"""N persistent threads that each run the current job once per ``run``."""

def __init__(self, n_threads: int) -> None:
self._start = threading.Barrier(n_threads + 1, timeout=_TIMEOUT)
self._done = threading.Barrier(n_threads + 1, timeout=_TIMEOUT)
self._job: typing.Callable[[int], object] = lambda _: None
self._stopping = False
self._errors: list[BaseException] = []
self.results: list[object] = [None] * n_threads
self._threads = [threading.Thread(target=self._loop, args=(i,), daemon=True) for i in range(n_threads)]
for thread in self._threads:
thread.start()

def _loop(self, index: int) -> None:
while True:
try:
self._start.wait()
except threading.BrokenBarrierError:
return
if self._stopping:
return
try:
self.results[index] = self._job(index)
except BaseException as exc: # noqa: BLE001
self._errors.append(exc)
try:
self._done.wait()
except threading.BrokenBarrierError:
return

def run(self, job: typing.Callable[[int], object]) -> None:
self._job = job
self._start.wait()
self._done.wait()
if self._errors:
raise self._errors[0]

def stop(self) -> None:
self._stopping = True
if not (self._start.broken or self._done.broken):
self._start.wait()
self._start.abort()
self._done.abort()
for thread in self._threads:
thread.join(timeout=_TIMEOUT)


@pytest.fixture
def pool(n_threads: int) -> typing.Iterator[_WorkerPool]:
worker_pool = _WorkerPool(n_threads)
yield worker_pool
worker_pool.stop()


# --- G14: concurrent cached-hit (lock-free read path, fixed total work) -----
Expand All @@ -63,19 +109,20 @@ class CachedGroup(Group):


@pytest.mark.parametrize("n_threads", _THREAD_COUNTS)
def test_g14_concurrent_cached_hit(benchmark, n_threads):
def test_g14_concurrent_cached_hit(benchmark, n_threads, pool):
# Fixed total reads split across N threads: batch time drops with N iff the read path scales.
container = Container(scope=Scope.APP, groups=[CachedGroup])
container.open()
warm = container.resolve_provider(CachedGroup.obj)
reads_per_thread = _TOTAL_READS // n_threads

def _worker() -> None:
def _job(_: int) -> object:
result = None
for _ in range(reads_per_thread):
container.resolve_provider(CachedGroup.obj)
result = container.resolve_provider(CachedGroup.obj)
return result

benchmark(_run_parallel, _worker, n_threads)
assert container.resolve_provider(CachedGroup.obj) is warm # same cached instance
benchmark(pool.run, _job)
assert all(result is warm for result in pool.results)


# --- G15: concurrent first-resolve (creation under the double-checked lock) --
Expand All @@ -90,26 +137,28 @@ def _worker() -> None:


@pytest.mark.parametrize("n_threads", _THREAD_COUNTS)
def test_g15_concurrent_first_resolve(benchmark, n_threads):
def test_g15_concurrent_first_resolve(benchmark, pool):
# All N threads race to first-resolve the SAME K cold singletons -> contention on each
# creation lock. Fresh container per round (untimed setup) so every round actually creates.
check = Container(scope=Scope.APP, groups=[_COLD_GROUP])
check.open()
assert all(check.resolve_provider(p) is not None for p in _COLD_PROVIDERS)

def _setup() -> "tuple[tuple[Container], dict[str, object]]":
container = Container(scope=Scope.APP, groups=[_COLD_GROUP])
# creation lock. Resolvers are compiled once up front; the untimed per-round setup only empties
# the cache, so every round creates and none compiles.
container = Container(scope=Scope.APP, groups=[_COLD_GROUP])
for provider in _COLD_PROVIDERS:
container.resolve_provider(provider)

def _setup() -> None:
container.close_sync()
container.open()
return (container,), {}

def _batch(container) -> None:
def _worker() -> None:
for provider in _COLD_PROVIDERS:
container.resolve_provider(provider)

_run_parallel(_worker, n_threads)
def _job(_: int) -> list[object]:
return [container.resolve_provider(provider) for provider in _COLD_PROVIDERS]

benchmark.pedantic(_batch, setup=_setup, rounds=120, iterations=1)
benchmark.pedantic(pool.run, args=(_job,), setup=_setup, rounds=120, iterations=1)
first = typing.cast("list[object]", pool.results[0])
assert [type(obj) for obj in first] == _COLD_TYPES
assert all(
all(mine is theirs for mine, theirs in zip(typing.cast("list[object]", result), first, strict=True))
for result in pool.results
)


# --- G15b: concurrent first-resolve in sibling children ---------------------
Expand All @@ -123,29 +172,32 @@ def _worker() -> None:


@pytest.mark.parametrize("n_threads", _THREAD_COUNTS)
def test_g15b_concurrent_first_resolve_sibling_children(benchmark, n_threads):
def test_g15b_concurrent_first_resolve_sibling_children(benchmark, n_threads, pool):
"""Each thread builds its own REQUEST child and first-resolves K cached providers in it.

Every creation is a cold miss in a container no other thread touches. Each cache item has its
own lock, so these creations never contend; under a lock shared by the tree they would
serialize. G15 does not cover this: it races on one root's items, whose locks are shared.
"""
check = Container(scope=Scope.APP, groups=[_REQUEST_GROUP])
check.open()
probe = check.build_child_container(scope=Scope.REQUEST)
assert all(probe.resolve_provider(p) is not None for p in _REQUEST_PROVIDERS)
container = Container(scope=Scope.APP, groups=[_REQUEST_GROUP])
with container.build_child_container(scope=Scope.REQUEST) as probe:
for provider in _REQUEST_PROVIDERS:
probe.resolve_provider(provider)

def _setup() -> "tuple[tuple[Container], dict[str, object]]":
container = Container(scope=Scope.APP, groups=[_REQUEST_GROUP])
container.open()
return (container,), {}
def _job(_: int) -> list[object]:
child = container.build_child_container(scope=Scope.REQUEST)
return [child.resolve_provider(provider) for provider in _REQUEST_PROVIDERS]

def _batch(container) -> None:
def _worker() -> None:
child = container.build_child_container(scope=Scope.REQUEST)
for provider in _REQUEST_PROVIDERS:
child.resolve_provider(provider)
benchmark.pedantic(pool.run, args=(_job,), rounds=120, iterations=1)
results = [typing.cast("list[object]", result) for result in pool.results]
assert all([type(obj) for obj in result] == _REQUEST_TYPES for result in results)
assert len({id(obj) for result in results for obj in result}) == n_threads * _K_COLD

_run_parallel(_worker, n_threads)

benchmark.pedantic(_batch, setup=_setup, rounds=120, iterations=1)
# --- G15c: control, the harness floor ----------------------------------------
@pytest.mark.parametrize("n_threads", _THREAD_COUNTS)
def test_g15c_worker_pool_floor_control(benchmark, n_threads, pool):
# Harness floor: the same batch with an empty job, so the barrier cost inside every
# G14/G15/G15b number is visible in the same run.
benchmark.pedantic(pool.run, args=(lambda index: index,), rounds=120, iterations=1)
assert pool.results == list(range(n_threads))
2 changes: 1 addition & 1 deletion docs/dev/contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ CI runs the coverage-enforcing recipe `just test-ci` along with `just lint-ci`.

## Submitting changes
1. Fork the repo and branch off `main`.
2. Make your change with tests; keep **100% line coverage** (CI runs `just test-ci`, which fails below the `report.fail_under = 100` gate in `pyproject.toml`).
2. Make your change with tests; keep **100% line and branch coverage** of `modern_di` (CI runs `just test-ci`, which fails below the `report.fail_under = 100` gate in `pyproject.toml`).
3. Run `just lint` and `just test` locally before pushing (CI runs the non-fixing variants `just lint-ci` / `just test-ci`).
4. For non-trivial changes, the PR body is the spec; the pull-request template walks you through it (why, design, non-goals, verification).
5. Open a pull request upstream.
12 changes: 6 additions & 6 deletions justfile
Original file line number Diff line number Diff line change
Expand Up @@ -34,13 +34,13 @@ adr-check:
test *args:
uv run --no-sync pytest {{ args }}

# The gated full run: 100% line coverage required. CI runs this.
test-ci:
uv run --no-sync pytest --cov=. --cov-report term-missing --cov-report xml
# Run the thread_race tests only, collecting just the files that hold them. Passes args through.
test-race *args:
uv run --no-sync pytest -m thread_race tests/test_free_threading.py tests/providers/test_cached_factory.py tests/registries/test_providers_registry.py {{ args }}

# Branch-coverage run (diagnostic; line coverage is the enforced gate, not branch).
test-branch:
uv run --no-sync pytest --cov=. --cov-branch
# The gated full run: 100% line and branch coverage of modern_di required. CI runs this.
test-ci:
uv run --no-sync pytest --cov --cov-report term-missing --cov-report xml

# Run the guard-tier benchmark suite (zero-dep; pytest-benchmark). Excludes the
# comparative tier, whose deps live in benchmarks/comparative and are not in this env.
Expand Down
2 changes: 2 additions & 0 deletions modern_di/dependency_graph.py
Original file line number Diff line number Diff line change
Expand Up @@ -207,4 +207,6 @@ def collect_errors(container: "Container", registry: "ProvidersRegistry") -> lis
)
case Cycle(providers):
errors.append(build_cycle_error(providers, container))
case _:
typing.assert_never(event)
return errors
Loading
Loading