Skip to content

test: gate branch coverage, time concurrency benchmarks on persistent workers, pay down test debt (#606) - #609

Merged
lesnik512 merged 2 commits into
mainfrom
issue-606
Oct 6, 2026
Merged

lesnik512 merged 2 commits into
mainfrom
issue-606

Conversation

@lesnik512

Copy link
Copy Markdown
Member

Closes #606.

Summary

Baseline on ffe365c: 963 tests, 100% line coverage, 99.90% branch coverage of modern_di (two partial branches). After: 965 tests, 100% line and 100% branch coverage, and CI now enforces both.

  1. Branch coverage gate. just test-ci measures modern_di only (run.source) with run.branch = true. The two partial branches are closed: collect_errors gets case _: typing.assert_never(event) (excluded from the report via exclude_also; ty fails if an Event member loses its case), and a new test cancels close_async after an earlier item finalized, then closes again, which reaches CacheItem.close_async with nothing owed. The test-branch recipe is gone because test-ci now does what it did.
  2. G14/G15/G15b. The harness keeps N worker threads alive for the whole benchmark and releases them with two barriers per batch, so thread start-up and join are out of the timed call. G15 compiles its resolvers once and its per-round setup only closes and reopens the container, so each round times creation alone. G15b reuses one root and builds the children inside the job. New G15c row: an empty job on the same pool. Every scenario now asserts on what the timed batch resolved (G15: every thread got the same K instances; G15b: N x K distinct instances). README gets G15b and G15c rows and a rewritten harness paragraph. tests/test_bench_report.py only parses the comparative tier and still passes.
  3. Free-threading tests. Every Barrier and join in test_free_threading.py and the two race tests in registries/test_providers_registry.py has a timeout, joins are followed by an is_alive assertion, and the threads are daemons so a hang fails the test instead of the run. A new thread_race marker covers that module plus the three race tests named in the issue; the 3.14t step runs just test -m thread_race --count=50 without the redundant -W error::RuntimeWarning.
  4. Private API in tests. All 27 _providers_registry.add_providers/register(...) calls in tests now use container.add_providers(...); every registered type matched the provider's bound_type, so they are equivalent. The blanket tests/** SLF001 ignore is narrowed to the 13 files that still read internals with no public equivalent.
  5. Coverage config. With only modern_di measured, all 54 # pragma: no cover comments and the 12 "exercise the body for coverage" asserts in tests are removed. run.concurrency, run.omit and addopts = "" are dropped.
  6. Weak assertions. Removed the duplicate simplefilter("error") blocks; the partial-creator test now asserts the warning on 3.11-3.13 and its absence on 3.14+; the inspect.signature patch test asserts the patched call happened.
  7. Lint. Dropped the ISC001 ignore.
  8. CI. release.yml uses astral-sh/setup-uv@v8.2.0 like the other workflows. New allowed-to-fail matrix entries for 3.15 and 3.15t (continue-on-error), and the free-threaded steps now key on a trailing t. uv fetches cpython-3.15.0rc1; both builds pass the full gated suite locally after adding a repr(UNSET) test, which was only covered incidentally before 3.15.
  9. AGENTS.md now says the lychee links job checks local links in every Markdown file.

Design decisions

  • Measuring modern_di instead of . (the issue's exclude_also alternative) because branch coverage over tests reports 139 partial branches that are coverage.py artefacts of class X: ... bodies, and the pragmas existed only to satisfy line coverage of test code. The cost: a test helper that never runs is no longer reported.
  • The 16 _cache_registry.cached_count() asserts are untouched; the registry refactor in Refactor and perf: registry hooks, container-owned cache and context, cheaper child build #605 replaces them.
  • Free-threading stress step runtime on 3.14t, measured locally: 0.70 s of pytest time (1.2 s wall) before, 3.2 s (4.0 s wall) after. Most of the increase is collecting 48,000 repeated items to select 250.
  • Local GIL numbers (3.14, median, 1/2/4 threads): G15 went from 232/280/382 us to 39/61/106 us, G15b from 225/319/499 us to 62/128/264 us; G15c is 11/24/53 us. The benchmark IDs are unchanged, so the first CI comparison will show G14/G15/G15b as large improvements.

Open questions

Test plan

  • just lint, just lint-ci
  • just test-ci (965 passed, 100% line and branch)
  • Gated suite on 3.11, 3.12, 3.13, 3.14t, 3.15rc1 and 3.15rc1t locally
  • just test -m thread_race --count=50 on 3.14t and 3.15t
  • Guard benchmarks on 3.14 and 3.14t; tests/test_bench_report.py
  • mkdocs build --strict, just adr-check
  • CI green

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark

Details
Benchmark suite Current: cf0caa8 Previous: ffe365c Ratio
benchmarks/test_guard_by_type.py::test_g16_resolve_by_type 3950185.9431616464 iter/sec (stddev: 7.745686381423656e-9) 4650912.939324458 iter/sec (stddev: 1.706440348160392e-8) 1.18
benchmarks/test_guard_by_type.py::test_g17_resolve_by_type_large_registry 3959333.683742874 iter/sec (stddev: 6.0933195179902515e-9) 4564403.042262547 iter/sec (stddev: 2.130887895421523e-8) 1.15
benchmarks/test_guard_cold.py::test_g8_cold_first_resolve 22511.995070418634 iter/sec (stddev: 0.00004111685903610907) 18672.690009299124 iter/sec (stddev: 0.000040957900608557336) 0.83
benchmarks/test_guard_cold.py::test_g8b_cold_first_resolve_cached 17528.177338987985 iter/sec (stddev: 0.00014840046722346283) 14435.165757881416 iter/sec (stddev: 0.00015649027293176276) 0.82
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[1] 512.1402568351706 iter/sec (stddev: 0.000057403048352041066) 606.7458247729055 iter/sec (stddev: 0.00006403397442405826) 1.18
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[2] 499.36872452706854 iter/sec (stddev: 0.00008212224430016972) 566.5625364040568 iter/sec (stddev: 0.00006131135700297531) 1.13
benchmarks/test_guard_concurrency.py::test_g14_concurrent_cached_hit[4] 477.667401533961 iter/sec (stddev: 0.00003973808132966421) 512.1861640209803 iter/sec (stddev: 0.00005466146361602645) 1.07
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[1] 12361.0947728639 iter/sec (stddev: 0.000010075445713161572) 1756.8118292460897 iter/sec (stddev: 0.0001580997113539773) 0.14
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[2] 7152.323327958439 iter/sec (stddev: 0.000018633341471109837) 1424.340746556056 iter/sec (stddev: 0.00017645930230460275) 0.20
benchmarks/test_guard_concurrency.py::test_g15_concurrent_first_resolve[4] 3712.139770477795 iter/sec (stddev: 0.00003179413836037132) 999.3575047320206 iter/sec (stddev: 0.00019075327136160947) 0.27
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[1] 8715.81792393906 iter/sec (stddev: 0.000009341772940330055) 1776.3768085999434 iter/sec (stddev: 0.00019073682152782808) 0.20
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[2] 3609.1900445262654 iter/sec (stddev: 0.000020340743363126515) 1244.4859807766236 iter/sec (stddev: 0.00018810472467390178) 0.34
benchmarks/test_guard_concurrency.py::test_g15b_concurrent_first_resolve_sibling_children[4] 1828.0617997614563 iter/sec (stddev: 0.000035023646802588695) 723.4024060255196 iter/sec (stddev: 0.0014598867005660315) 0.40
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[1] 40243.39203514915 iter/sec (stddev: 0.000003460489734697166)
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[2] 15869.07379347525 iter/sec (stddev: 0.000013541609614824049)
benchmarks/test_guard_concurrency.py::test_g15c_worker_pool_floor_control[4] 6608.72321723006 iter/sec (stddev: 0.00002684098432508923)
benchmarks/test_guard_lifecycle.py::test_g6_build_child_container 1204208.6490166876 iter/sec (stddev: 4.1114870543639516e-8) 1219360.5929127778 iter/sec (stddev: 4.021730185866024e-8) 1.01
benchmarks/test_guard_lifecycle.py::test_g6b_build_child_container_auto_scope 1088356.686696722 iter/sec (stddev: 2.1866267986318034e-8) 1165899.4125525632 iter/sec (stddev: 2.7625866083506307e-8) 1.07
benchmarks/test_guard_lifecycle.py::test_g7_request_lifecycle_batch 2533.6686100802854 iter/sec (stddev: 0.00001139311912002462) 2236.288706117797 iter/sec (stddev: 0.00001689205206591054) 0.88
benchmarks/test_guard_lifecycle.py::test_g7c_event_loop_floor_control 57115.703312321224 iter/sec (stddev: 0.0000016443527337652109) 67638.75445199345 iter/sec (stddev: 0.000001407229405623694) 1.18
benchmarks/test_guard_lifecycle.py::test_g7b_request_cycle_sync 310800.7165076886 iter/sec (stddev: 1.631130904295121e-7) 300382.96424878866 iter/sec (stddev: 9.599708707416182e-8) 0.97
benchmarks/test_guard_lifecycle.py::test_g13_teardown_at_scale 40906.247841908604 iter/sec (stddev: 0.0000017852769299060014) 41864.52786889268 iter/sec (stddev: 0.0000019744511004598563) 1.02
benchmarks/test_guard_lifecycle.py::test_g13b_teardown_at_scale_async_no_finalizers 459.329246138679 iter/sec (stddev: 0.001540877709414442) 469.4101679223813 iter/sec (stddev: 0.0010765366169991888) 1.02
benchmarks/test_guard_resolve.py::test_g1_transient_resolve 2370395.373655168 iter/sec (stddev: 2.719325034407428e-8) 2643693.8483676505 iter/sec (stddev: 3.622929811410649e-8) 1.12
benchmarks/test_guard_resolve.py::test_g2_cached_resolve 3826361.6959387376 iter/sec (stddev: 6.881111680281884e-9) 4533963.42125545 iter/sec (stddev: 1.7932439514885506e-8) 1.18
benchmarks/test_guard_resolve.py::test_g3_deep_chain 804851.499967102 iter/sec (stddev: 6.221465859700483e-8) 874097.5270822884 iter/sec (stddev: 1.3204596479966074e-7) 1.09
benchmarks/test_guard_resolve.py::test_g4_wide_resolve 474597.19690068497 iter/sec (stddev: 4.500930080375432e-7) 473981.3057505384 iter/sec (stddev: 6.531655111675316e-7) 1.00
benchmarks/test_guard_resolve.py::test_g5_cross_scope 2117031.0650465 iter/sec (stddev: 2.421736471148955e-8) 2097795.777795691 iter/sec (stddev: 1.4347586989334604e-7) 0.99
benchmarks/test_guard_resolve.py::test_g9_context_resolve 1599748.4555455411 iter/sec (stddev: 1.7486836072420347e-7) 1717618.5322124795 iter/sec (stddev: 1.6797887345620352e-7) 1.07
benchmarks/test_guard_resolve.py::test_g12_override_active_resolve 814359.9013097215 iter/sec (stddev: 3.83673928040184e-8) 912090.6924636723 iter/sec (stddev: 1.0201256745872507e-7) 1.12
benchmarks/test_guard_resolve.py::test_g18_alias_hop 3555794.3785150237 iter/sec (stddev: 9.776863197907824e-9) 4347042.296334007 iter/sec (stddev: 2.597541863228709e-8) 1.22
benchmarks/test_guard_validate.py::test_g10_validate_deep_chain 29421.917286472395 iter/sec (stddev: 0.00002187082673368506) 20347.922617500906 iter/sec (stddev: 0.00024790180939011935) 0.69
benchmarks/test_guard_validate.py::test_g11_validate_wide 17064.86456847256 iter/sec (stddev: 0.000027105569824449245) 13773.313381664582 iter/sec (stddev: 0.000027510882864299994) 0.81

This comment was automatically generated by workflow using github-action-benchmark.

@lesnik512
lesnik512 merged commit e44d4df into main Oct 6, 2026
11 checks passed
@lesnik512
lesnik512 deleted the issue-606 branch October 6, 2026 17:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tests and tooling: branch coverage gate, concurrency benchmarks, test debt

1 participant