Skip to content

perf: registry hooks, container-owned cache and context, cheaper child build (#605) - #610

Merged
lesnik512 merged 17 commits into
mainfrom
issue-605
Oct 6, 2026
Merged

lesnik512 merged 17 commits into
mainfrom
issue-605

Conversation

@lesnik512

@lesnik512 lesnik512 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Closes #605.

Summary

One refactor commit, twelve perf commits and a docs commit, then three review-fix commits (a typing.Generic fix and two refactors). Each perf commit body carries its own A/B delta against its parent. No public API changes.

Refactor (a2846ce, no behaviour change):

  • The provider hooks and the dependency_graph functions take a ProvidersRegistry, not a Container. _defaultless_context_terminal reuses terminal_chain, so the test that kept two alias walks in sync is gone. Factory._wiring_plan is gone.
  • ContextRegistry is folded into a Container._context dict.
  • STEP_ERRORS is gone; the container and the template catch ResolutionError.
  • AbstractProvider.__init__ resolves an unset bound_type from an inferred_bound_type argument, once.
  • SignatureItem.raw_annotation is now unresolvable_generic; param_hints is now params.
  • resolve uses find_provider in its RecursionError branch. register and add_providers share one locked _add.

The close-path cleanup (refactor item 2) lands with the perf commits it needs. CacheRegistry is gone, and both close functions take the container's creation-order list and update it in place. The async loop calls finalizers itself, which removes CacheItem.close_async and the finalizer-less short-circuit.

The two loops still differ, on purpose. The async loop inlines the owed-finalizer check and can run every finalizer, so it never keeps an item and empties the list. The sync loop calls CacheItem.close_sync per item, because that method has to refuse an async finalizer (by the _is_async_finalizer flag, or by an awaitable result, whose coroutine it closes) and leave that item owed. The loop collects those items and puts them back in the list for a later close_async. Inlining the per-item method into the sync loop was measured as part of #605's research and did not pay for itself.

The cache code now lives in modern_di/cache.py, out of registries/. It holds CacheItem and the functions over a container's own items and creation order (fetch_cache_item, cached_count, close_async, close_sync), and none of that is a registry. Container.__repr__ calls cached_count.

Measurements

Interleaved A/B, 12 to 16 alternating subprocess rounds, min of 30 repeats each, paired median, against main ffe365c:

Scenario main this PR change
G1 transient 220 ns 216 ns −1.8%
G2 cached 134 ns 133 ns −1.2%
G3 chain 6 703 ns 706 ns +0.1%
G4 wide 10 958 ns 973 ns +1.5%
G6 child 428 ns 295 ns −31.1%
G6b child, auto scope 500 ns 314 ns −37.2%
G7 batch of 100 212.5 µs 177.1 µs −16.6%
G7b request cycle 1848 ns 1403 ns −24.1%
G8 cold 20.3 µs 17.8 µs −11.8%
G8b cold, cached 25.5 µs 22.6 µs −11.2%
G10 validate chain 18.6 µs 16.9 µs −9.1%
G11 validate wide 32.1 µs 28.7 µs −10.6%
G13 ten sync finalizers 12.3 µs 9.0 µs −26.6%
G13b async, no finalizers 1018 µs 871 µs −14.6%
C6-shaped request cycle 903 ns 709 ns −21.3%
Factory(dataclass) 16.3 µs 12.4 µs −24.1%

A child container takes 504 bytes instead of 584. The comparative tables are regenerated: C4 2.28 to 1.85 µs (now 0.88 against dishka, was 1.12), C6 1.09 µs to 806 ns.

Design decisions

  • typing.Generic stays unresolvable. Review found that the plain-class fast path in SignatureItem.from_type took a bare typing.Generic annotation for a plain class (its type is type), so Factory(...) stopped rejecting it and the failure moved to resolve time. get_origin(Generic) is Generic, which is why main treats it as an unresolvable generic. The fast path now excludes it, and the NoneType branch runs first. Covered by two tests that pass on main and failed on the branch before the fix. Factory(...) measured +0.5 to +0.8%, within noise.

  • Perf item 1 is partial. _set_state stays the one place that sets container slots, as asked. What moved is the validation: the root constructor checks the scope type, build_child_container checks an explicit scope, and an auto-scoped child skips both. Setting the slots inline in build_child_container would be worth another ~9% on G6 and G6b, but it would duplicate _set_state. object.__new__ instead of cls.__new__ measured as noise and is not done.

  • Perf item 8 drops the parent-scope memo. It helped G11 (−1.6%) and cost G10 (+2.9%), where every parent has one edge. The type dispatch and the visited-root skip are kept. The four event classes are typing.final, so ty narrows each type(event) is branch and the last branch is typed Cycle.

  • Perf item 5 is kept. _signature reads __init__ directly only for a metaclass-type class with no __signature__ or __wrapped__, whose first MRO entry defining __new__ or __init__ defines a plain-function __init__ alone. Everything else goes to inspect.signature. test_class_signature_matches_inspect_signature compares the two, ValueError included, over 23 creators, and the full suite passes on 3.11, 3.12, 3.13, 3.14 and 3.15.

  • Perf item 9 inlines the pending-finalizer check in the async loop. Calling _pending_finalizer() there cost G13b +2.9% against the old short-circuit; the inline condition is flat on G13b.

  • The 14 assertions on _cache_registry.cached_count() now check observable behaviour: the cached=N in the container repr where nothing should have been built, and finalizer effects or instance identity after reopening where something was.

Open questions

The review-fix commits were re-measured against the last perf commit (12 rounds): every guard scenario is within ±1.2%, so the table above stands. A fresh run against main gives the same picture (G6 −31.8%, G7 −17.6%, G7b −24.5%, G13 −27.3%, G4 +1.4%).

Test plan

Rebased onto main e44d4df (#609). collect_errors keeps this branch's type dispatch, so the typing.assert_never\( coverage exclusion is gone from pyproject.toml. The per-file SLF001 list drops test_cached_factory.py, test_cache_registry.py and test_free_threading.py, which no longer read internals, and tests/test_dependency_graph.py builds its ProvidersRegistry directly instead of reaching into a container. #609's cancelled-close test passes unchanged.

  • just lint-ci after just install
  • just test-ci: 100% line and branch on modern_di
  • Every commit passes lint and test-ci on its own
  • Full suite on Python 3.11, 3.12, 3.13, 3.15 (3.14 is the dev env)
  • just test-race --count=50 on 3.14t (250 passed), full suite on 3.14t
  • mkdocs build --strict
  • just bench-report, 5 runs, tables regenerated

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark

Details
Benchmark suite Current: 4668331 Previous: e44d4df Ratio
benchmarks/test_guard_by_type.py::test_g16_resolve_by_type 3095945.902195591 iter/sec (stddev: 9.625943455494879e-9) 2905410.5252089053 iter/sec (stddev: 1.2061833223870489e-8) 0.94
benchmarks/test_guard_by_type.py::test_g17_resolve_by_type_large_registry 3101303.425110006 iter/sec (stddev: 9.884409974094577e-9) 2805572.8215196584 iter/sec (stddev: 3.7956419944988586e-8) 0.90
benchmarks/test_guard_cold.py::test_g8_cold_first_resolve 22304.774687233643 iter/sec (stddev: 0.000040745352589451265) 19996.984827408312 iter/sec (stddev: 0.000041624937723947144) 0.90
benchmarks/test_guard_cold.py::test_g8b_cold_first_resolve_cached 16603.85414368776 iter/sec (stddev: 0.00012676579068549628) 15151.142982238001 iter/sec (stddev: 0.00010996858875858591) 0.91
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[1] 443.2114849607255 iter/sec (stddev: 0.00009530051572167489) 430.7231777039997 iter/sec (stddev: 0.00005758187504697698) 0.97
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[2] 437.8507034063758 iter/sec (stddev: 0.00008000737198276956) 410.18591982941007 iter/sec (stddev: 0.00010169868183643517) 0.94
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[4] 430.0729234779525 iter/sec (stddev: 0.00007188810072877915) 391.62060352392854 iter/sec (stddev: 0.00012892801828008377) 0.91
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[1] 13592.076635092597 iter/sec (stddev: 0.000007388072127698506) 10077.473093786855 iter/sec (stddev: 0.0000133295061452886) 0.74
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[2] 6934.498575629695 iter/sec (stddev: 0.000022537259960107885) 5506.060405965939 iter/sec (stddev: 0.00002236044175930745) 0.79
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[4] 3535.3024103772664 iter/sec (stddev: 0.00004213327824624536) 3071.9690148874884 iter/sec (stddev: 0.000035883206450868244) 0.87
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[1] 8066.840224667448 iter/sec (stddev: 0.000010492619448736243) 7478.378139177093 iter/sec (stddev: 0.000011707848799076186) 0.93
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[2] 3664.879172748289 iter/sec (stddev: 0.00002321106751811644) 3255.140782802764 iter/sec (stddev: 0.00003141389402621282) 0.89
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[4] 1786.311610233649 iter/sec (stddev: 0.000038138079481191376) 1497.830766591896 iter/sec (stddev: 0.00005440344389832282) 0.84
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[1] 41388.85832453229 iter/sec (stddev: 0.00000441312498451148) 32143.643513929273 iter/sec (stddev: 0.000005522914124538603) 0.78
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[2] 18591.953867494063 iter/sec (stddev: 0.000014414375156563837) 13663.874605714103 iter/sec (stddev: 0.000014745265224910578) 0.73
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[4] 6512.114649746871 iter/sec (stddev: 0.0000330939014221707) 5925.609308185773 iter/sec (stddev: 0.0000370934675298238) 0.91
benchmarks/test_guard_lifecycle.py::test_g6_build_child_container 1624178.7593803701 iter/sec (stddev: 3.273385729318145e-8) 977867.3417337334 iter/sec (stddev: 2.445760611113942e-7) 0.60
benchmarks/test_guard_lifecycle.py::test_g6b_build_child_container_auto_scope 1693344.1752469374 iter/sec (stddev: 2.38857791720706e-8) 1019374.5642661753 iter/sec (stddev: 5.926862677491684e-8) 0.60
benchmarks/test_guard_lifecycle.py::test_g7_request_lifecycle_batch 2810.0221840699924 iter/sec (stddev: 0.000012541723245411969) 2304.7474277478386 iter/sec (stddev: 0.000029503621138277164) 0.82
benchmarks/test_guard_lifecycle.py::test_g7c_event_loop_floor_control 58796.696246732055 iter/sec (stddev: 0.0000018486299143935733) 56022.748390396264 iter/sec (stddev: 0.000003884835826610305) 0.95
benchmarks/test_guard_lifecycle.py::test_g7b_request_cycle_sync 333220.89905065735 iter/sec (stddev: 7.87443883045732e-8) 302921.62301393447 iter/sec (stddev: 1.0988644704546988e-7) 0.91
benchmarks/test_guard_lifecycle.py::test_g13_teardown_at_scale 49103.09585317861 iter/sec (stddev: 0.000001709611261622255) 39039.23917109726 iter/sec (stddev: 0.000002906134789713714) 0.80
benchmarks/test_guard_lifecycle.py::test_g13b_teardown_at_scale_async_no_finalizers 478.34449798256196 iter/sec (stddev: 0.0009712107356158805) 429.0679874725799 iter/sec (stddev: 0.0009883401688103434) 0.90
benchmarks/test_guard_resolve.py::test_g1_transient_resolve 2032919.4716033482 iter/sec (stddev: 2.4219281837580794e-8) 2125106.9592946237 iter/sec (stddev: 5.442645728993484e-8) 1.05
benchmarks/test_guard_resolve.py::test_g2_cached_resolve 3069948.154721133 iter/sec (stddev: 1.095443729769591e-8) 3151184.174047065 iter/sec (stddev: 1.871886747624135e-8) 1.03
benchmarks/test_guard_resolve.py::test_g3_deep_chain 709024.2479198356 iter/sec (stddev: 3.890110703206501e-8) 720215.2334414327 iter/sec (stddev: 5.354678685470155e-8) 1.02
benchmarks/test_guard_resolve.py::test_g4_wide_resolve 447375.7697886735 iter/sec (stddev: 3.680810055288055e-7) 440512.2267365485 iter/sec (stddev: 3.9881197297577387e-7) 0.98
benchmarks/test_guard_resolve.py::test_g5_cross_scope 1878271.1265782553 iter/sec (stddev: 3.024333537610421e-8) 1903799.03572997 iter/sec (stddev: 3.749786098021958e-8) 1.01
benchmarks/test_guard_resolve.py::test_g9_context_resolve 1334386.4311178997 iter/sec (stddev: 1.3835434373768218e-7) 1299429.1055680322 iter/sec (stddev: 1.3625289680845453e-7) 0.97
benchmarks/test_guard_resolve.py::test_g12_override_active_resolve 712545.6523507192 iter/sec (stddev: 4.883001690568414e-8) 704915.0270757646 iter/sec (stddev: 5.875407583018373e-8) 0.99
benchmarks/test_guard_resolve.py::test_g18_alias_hop 3063977.2545130015 iter/sec (stddev: 9.952864464910664e-9) 2905826.6619156906 iter/sec (stddev: 7.916311368966162e-9) 0.95
benchmarks/test_guard_validate.py::test_g10_validate_deep_chain 28321.121910850696 iter/sec (stddev: 0.000020214553603235905) 25964.11285330336 iter/sec (stddev: 0.000021302687495626346) 0.92
benchmarks/test_guard_validate.py::test_g11_validate_wide 16814.893220095506 iter/sec (stddev: 0.000025107240524120107) 15245.120344405394 iter/sec (stddev: 0.000028051744327008157) 0.91

This comment was automatically generated by workflow using github-action-benchmark.

…als (#605)

- The provider hooks (`_get_dependencies`, `_redirect_target`,
  `_iter_validation_issues`) and the `dependency_graph` functions take a
  `ProvidersRegistry` instead of a `Container`; they only ever used its
  registry. `collect_errors` takes the registry alone.
- `_defaultless_context_terminal` reuses `terminal_chain` instead of its own
  alias walk, so the test that kept the two walks in sync is gone.
  `Factory._wiring_plan` is gone; callers use `registry.plan_for`.
- `ContextRegistry` is folded into a `Container._context` dict.
- `STEP_ERRORS` is gone; the container and the resolver template catch
  `ResolutionError` directly.
- `AbstractProvider.__init__` resolves an unset `bound_type` from the
  subclass's `inferred_bound_type`, once.
- `SignatureItem.raw_annotation` is renamed `unresolvable_generic`.
- `Container.resolve` looks the provider up with `find_provider` in its
  `RecursionError` branch too, and an unregistered type there re-raises the
  `RecursionError`. `register` and `add_providers` share one locked `_add`.

No behaviour change. Folding the context dict into the container takes one
object off child construction: G6 -11.4%, G6b -9.0%, C6 -5.8% (A/B against
ffe365c, 10 interleaved rounds).
`CacheRegistry` is gone. A container holds its cache items and creation order
in two slots, and `fetch_cache_item`, `close_async` and `close_sync` are module
functions in `cache_registry` over that dict and list. Both close functions
now update the list in place, so the container no longer reaches into another
object's private list to decide whether to close. Child construction allocates
one object fewer and the cached template drops an attribute hop.

Tests that counted cached items through the registry now check observable
behaviour: the container repr, finalizer effects, or whether a resolve after
reopening returns the same instance.

A/B against the parent commit, 12 interleaved rounds, paired median:
G6 -11.6%, G6b -9.9%, G7 -4.8%, G7b -3.3%, G13 -2.9%, C6 -6.2%;
G1-G4, G8, G10, G11 within noise.
`{**parent_map, parent_scope: parent}` builds the dict through the generic
unpacking path. A `copy()` of the parent's map plus one item assignment does
the same work faster.

A/B against the parent commit, 12 interleaved rounds, paired median:
G6 -8.5%, G6b -7.5%, C6 -4.2%, G7 -1.1%.
A child built with `context={...}` copied it through `copy.copy`, which looks
up the type's copier before it lands on `dict.copy`. A plain dict now calls
`dict.copy` directly; any other mapping still goes through `copy.copy`, so a
dict subclass keeps its type.

A/B against the parent commit, 12 interleaved rounds, paired median:
C6 (request cycle with a context value) -4.7%; G6 unchanged, as it passes no
context.
`_set_state` checked the scope type and the parent ordering for every
container, including an auto-scoped child whose scope came from
`next_deeper` and so is always a deeper member of the same enum. The checks
now sit at the two entry points: the root constructor checks the type, and
`build_child_container` checks an explicit scope. `_set_state` stays the one
place that sets a container's slots and does nothing else.

A/B against the parent commit, 12 interleaved rounds, paired median:
G6b -17.2%, G6 -4.4%, C6 -2.0%; G7 and G7b within noise.

Tried and dropped: `object.__new__(cls)` instead of `cls.__new__(cls)`
(G6 -0.2%, noise). Setting the child's slots inline in
`build_child_container` is worth another ~9% on G6 and G6b, but it would
duplicate `_set_state`, so it is not done.
`inspect.isawaitable` is a Python-level function that runs on every finalizer
result, and a sync finalizer almost always returns `None`. A `None` check in
front of it skips the call for that case.

A/B against the parent commit, 12 interleaved rounds, paired median:
G13 -14.7%, G7b -9.5%; G7 within noise (its finalizer is async).
`close_async` awaited `CacheItem.close_async` for every item that had a
finalizer, a coroutine per item, and short-circuited the items without one.
The loop now calls the owed finalizer itself and awaits only an awaitable
result, so a sync finalizer costs no coroutine, and one condition replaces
both the short-circuit and `CacheItem.close_async`. The sync and async loops
now have the same shape: finalize what is owed, collect failures, clear.

A/B against the parent commit, 12 interleaved rounds, paired median:
G7 -4.2%; G13b (no finalizers), G13 and G7b within noise.
A cache miss built `functools.partial(build, target)` only to call it once
under the item lock. `get_or_create(build, target, create)` takes the target
as an argument, which allocates nothing and still keeps `target` out of a
closure cell on the warm path.

A/B against the parent commit, 12 interleaved rounds, paired median:
G13b -10.6%, G13 -9.8%, G7b -6.3%, G7 -5.8%, G8b -4.1%.
…605)

`_can_call_positionally` rebuilt the parameter-name tuple and rescanned for
keyword-only parameters each time a resolver compiled, and a resolver
compiles once per container tree, so a suite that builds a container per test
paid it per test. The part that depends only on the signature is now
computed in `Factory.__init__`; compiling compares the plan's names against
it.

A/B against the parent commit, 12 interleaved rounds, paired median:
G8 -8.0%, G8b -6.3%. The work moves to construction: `Factory(...)` +1.2%,
paid once per provider.
…605)

Every compiled factory rebuilt the same eight module-level entries of its
`exec` namespace from a dict literal, plus a merged comprehension for the
argument resolvers. The constant entries now live in one module-level dict
that each compile copies, then adds the factory's own names.

A/B against the parent commit, 12 interleaved rounds, paired median:
G8 -4.5%, G8b -3.7%.
`collect_errors` matched each event with class patterns, which run an
`isinstance` and a `__match_args__` unpack per case. It now compares
`type(event)` against the four event classes, which are marked
`typing.final` so `ty` narrows each branch and the last one is known to be a
`Cycle`. `walk` skips a root it already visited before starting a generator
for it, and `_walk_from` drops its own, now unreachable, visited check.

A/B against the parent commit, 12 interleaved rounds, paired median:
G10 -8.6%, G11 -8.6%.

Tried and dropped: memoizing each parent's effective scope across its
edges. It helps the wide graph (G11 -1.6%) and costs the chain, where every
parent has one edge (G10 +2.9%).
Most creator parameters are annotated with a plain class. `from_type` ran
every such annotation through `typing.get_origin` and the union and generic
checks before landing on the plain-class branch. An annotation whose type is
exactly `type` (and is not `NoneType`) now returns there directly; a class
with a custom metaclass, a generic alias, a union or a `NewType` takes the
full path as before.

A/B against the parent commit, 12 interleaved rounds, paired median:
`Factory(...)` on a two-parameter dataclass -5.3%, on a three-parameter
plain class -5.9%.
`inspect.signature(SomeClass)` checks the metaclass for `__call__`, looks
`__new__` and `__init__` up statically, unwraps both and walks the MRO before
it reads `__init__` as a bound method. `_signature` takes that last step
directly when the earlier ones cannot change the answer: the metaclass is
`type`, the class has no `__signature__` or `__wrapped__`, and the first MRO
entry that defines `__new__` or `__init__` defines a plain-function
`__init__` with no `__wrapped__` and no `__new__` beside it. Anything else
goes to `inspect.signature` as before.

`test_class_signature_matches_inspect_signature` checks the result, or the
`ValueError`, against `inspect.signature` for plain, inherited, dataclass,
generated, `*args`-first, self-less, wrapped, `__signature__`, NamedTuple,
`__new__`-only, `__init__`-over-`__new__`, metaclass, Generic, exception,
staticmethod and no-`__init__` classes, plus non-class creators. It passes on
3.11, 3.12, 3.13, 3.14 and 3.15.

A/B against the parent commit, 12 interleaved rounds, paired median:
`Factory(...)` on a dataclass -21.8%, on a plain class -19.1%; G8 and G8b
within noise, since their providers are built outside the timed call.
Re-ran `just bench-report` (5 runs, Apple M2, CPython 3.14.7, the same rival
versions) at the last perf commit in this series. C4 fell from 2.28 to
1.85 µs and now beats dishka (0.88, was 1.12); C6 fell from 1.09 µs to
806 ns. C1-C3 did not move. "What moved in this publication" gains a column
for the previous 4.0 run, and the performance history gains a paragraph for
#605)

The plain-class fast path in `SignatureItem.from_type` took any annotation
whose type is exactly `type`. `typing.Generic` is one, but
`typing.get_origin(Generic)` is `Generic`, so on main it is an unresolvable
generic and `Factory(...)` rejects a parameter annotated with it. The fast
path took it for a plain class and moved the failure to resolve time. It now
excludes `Generic`, and the `NoneType` branch runs first so the fast path no
longer needs its own `NoneType` check.

A/B against the parent commit, 12 interleaved rounds: `Factory(...)` +0.5%
and +0.8%, within noise.
…605)

`registries/cache_registry.py` stopped holding a registry when the cache
items moved onto the container. It is now `modern_di/cache.py`: `CacheItem`
plus the functions over a container's own items and creation order. A new
`cached_count(items)` there replaces the inline count in `Container.__repr__`.
Its tests move to `tests/test_cache.py`.
…e spots (#605)

- `_navigate` and `_FACTORY_GLOBALS` sit just above `_compile_factory`, the
  one function that reads them, instead of at the end of the module.
- `Factory._positional_names` is set with a plain `if` instead of a
  multi-line conditional expression.
- `collect_errors` calls `walk(roots=registry, registry=registry)`, so it is
  clear the registry is both the root set and the lookup.
- The `resolver_compiler` docstring lists `_closed` among the container slots
  the generated code reads.

No behaviour change.
@lesnik512
lesnik512 merged commit 9131e81 into main Oct 6, 2026
11 checks passed
@lesnik512
lesnik512 deleted the issue-605 branch October 6, 2026 17:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Refactor and perf: registry hooks, container-owned cache and context, cheaper child build

1 participant