GROOVY-12191: Scope indy SwitchPoint invalidation - #2736
Conversation
5c5a20a to
793abf8
Compare
There was a problem hiding this comment.
⚠️ Performance Alert ⚠️
Possible performance regression was detected for benchmark.
Benchmark result of this commit is worse than the previous benchmark result exceeding threshold 1.50.
| Benchmark suite | Current: 637bdaf | Previous: 00f7489 | Ratio |
|---|---|---|---|
org.apache.groovy.bench.AckermannBench.java ( {"n":"5"} ) |
0.06997222463418132 ms/op |
0.04000714974250017 ms/op |
1.75 |
org.apache.groovy.bench.AckermannBench.java ( {"n":"6"} ) |
0.28848271731001324 ms/op |
0.14918095442974466 ms/op |
1.93 |
org.apache.groovy.bench.AckermannBench.java ( {"n":"7"} ) |
1.1498743466781267 ms/op |
0.649040920815141 ms/op |
1.77 |
org.apache.groovy.bench.AryBench.groovyCS ( {"n":"1000"} ) |
0.06445658594417991 ms/op |
0.037069165224466585 ms/op |
1.74 |
org.apache.groovy.bench.AryBench.java ( {"n":"1000"} ) |
0.09092693016658131 ms/op |
0.03428444328401786 ms/op |
2.65 |
org.apache.groovy.bench.DynamicDispatchColdBench.dynamicMono_java ( {"n":"20000"} ) |
1193.21304 us/op |
726.0382800000001 us/op |
1.64 |
org.apache.groovy.bench.StaticMethodCallIndyColdBench.staticSum_groovyCS ( {"n":"20000"} ) |
2872.8105250000003 us/op |
1890.7597374999993 us/op |
1.52 |
org.apache.groovy.bench.StaticMethodCallIndyColdBench.staticSum_java ( {"n":"2000"} ) |
50.88577500000001 us/op |
30.64482499999999 us/op |
1.66 |
org.apache.groovy.bench.StaticMethodCallIndyColdBench.staticSum_java ( {"n":"20000"} ) |
452.51471249999975 us/op |
238.3599875 us/op |
1.90 |
This comment was automatically generated by workflow using github-action-benchmark.
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #2736 +/- ##
==================================================
+ Coverage 69.3657% 69.9707% +0.6050%
- Complexity 35082 35427 +345
==================================================
Files 1553 1558 +5
Lines 131490 131465 -25
Branches 24045 24120 +75
==================================================
+ Hits 91209 91987 +778
+ Misses 32045 31161 -884
- Partials 8236 8317 +81
🚀 New features to boost your workflow:
|
JMH summary — classic (commit
|
| Group | Speedup | Calibrated | n |
|---|---|---|---|
| bench | 0.987 × | 0.969 × | 92 |
| core | 1.147 × | 1.175 × | 83 |
| grails | 0.903 × | 0.869 × | 80 |
Runner calibration (this run vs baseline hardware): bench 1.01× (26 rulers) · core-ag 0.99× (3 rulers) · core-hz 0.96× (3 rulers) · grails-ad 0.94× (3 rulers) · grails-ez 1.13× (3 rulers)
Baseline: dev/bench/jmh/<part>/classic/data.js on gh-pages, trailing 90 days. Daily dashboard · Per-suite raw data
JMH summary — indy (commit
|
| Group | Speedup | Calibrated | n |
|---|---|---|---|
| bench | 0.931 × | 0.962 × | 92 |
| core | 3.008 × | 2.673 × | 83 |
| grails | 4.349 × | 4.233 × | 80 |
Runner calibration (this run vs baseline hardware): bench 0.96× (26 rulers) · core-ag 1.13× (3 rulers) · core-hz 1.12× (3 rulers) · grails-ad 1.13× (3 rulers) · grails-ez 0.95× (3 rulers)
Baseline: dev/bench/jmh/<part>/indy/data.js on gh-pages, trailing 90 days. Daily dashboard · Per-suite raw data
8e3d5cf to
9f118c0
Compare
9f118c0 to
ae8e497
Compare
GROOVY-12191 Performance Verification ReportScoped indy SwitchPoint invalidation
1. Executive summaryThis commit changes Groovy’s invokedynamic MOP On the same host, JDK, and JMH annotation settings, A/B measurement shows:
2. Change mechanism and performance hypotheses2.1 Previous model (parent)
2.2 New model (
|
| Component | Role |
|---|---|
SwitchPointInvalidator |
Single-domain lifecycle: CAS get / detachLive / invalidate |
IndyInvalidation |
Policy: invalidateClass (+ hierarchy), invalidateCategory (bulk), guardWithMopSwitchPoints |
ClassInfo |
Holds one indySwitchPoint per class; coordinates version bumps with domain retirement |
Selector / IndyCompoundAssign / ColdReflectiveMethodHandleWrapper |
Install a per-receiver-class SwitchPoint guard at link time |
Hot-path shape (intentionally monomorphic):
handle = classSp.guardWithTest(fast, fallback)
Still a single guard—matching the pre-6.0 monomorphic shape. There is no second category SwitchPoint on the hot path.
Invalidation policy:
| Event | Behavior | Expected hot-path cost |
|---|---|---|
| Unrelated-type MetaClass change | Retire that type’s domain (and loaded subtypes) only | Hot sites do not re-link |
| Same-type / parent MetaClass change | Retire class + hierarchy fan-out | Must re-link (correctness) |
Category enter/leave, invalidateCallSites() |
Bulk-retire all loaded class domains | Same order of cost as old global invalidation |
| Unattributed MetaClass event | invalidateUnscoped() bulk |
Conservative, correctness-first |
2.3 Hypotheses under test
| ID | Hypothesis | How verified |
|---|---|---|
| H1 | Baseline hot loop has no material regression | A/B ratio ≈ 1 |
| H2 | Cross-type churn throughput/latency improves substantially | Unrelated ≫ parent; on HEAD, unrelated ≫ same-type |
| H3 | Pure hot loop after unrelated burst approaches baseline | afterUnrelatedBurst ≈ baseline |
| H4 | Same-type / hierarchy / category still pay re-link tax | Clearly slower than baseline |
| H5 | Correctness suite stays green | Unit / integration tests |
3. Methodology
3.1 Environment
| Item | Value |
|---|---|
| OS | Linux 6.15.5 x86_64 (hostname: hera) |
| CPU | AMD EPYC 7763 (6 vCPUs visible in container) |
| Memory | 23 GiB |
| JDK | Amazon Corretto 25.0.2 (25.0.2+10-LTS) |
| Build | Gradle wrapper, ./gradlew :perf:jmh, indy enabled by default |
| Trees | HEAD: project worktree; parent: worktree at 0ca665dad0 |
| Bench source parity | ScopedInvalidationBench workload identical on parent (Javadoc-only diff) |
3.2 Benchmark suites
-
org.apache.groovy.bench.ScopedInvalidationBench(added by this commit; Throughput, ops/ms)- 100 000 monomorphic calls per op; MetaClass write every 1000 iterations in churn cases.
- JMH:
@Warmup(3×2s) @Measurement(5×2s) @Fork(2)→ 10 samples per benchmark.
-
org.apache.groovy.perf.grails.CallSiteInvalidationBench(existing Grails-oriented suite; AverageTime, ms/op)- Same 100 000-iteration loops; cross-type / same-type / burst patterns.
- Same JMH annotation settings.
3.3 Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest --rerun-tasks
Result: 56 passed / 0 failed / 0 error.
| Test class | Cases |
|---|---|
bugs.Groovy12191 |
8 |
IndyInvalidationTest |
27 |
SwitchPointInvalidatorTest |
12 |
IndyScopedSwitchPointTest |
9 |
3.4 Statistics and artifacts
- Scores are JMH means;
±is JMH’s default 99.9% CI. - Throughput ratio = HEAD / parent (>1 is better).
- AverageTime speedup = parent / HEAD (>1 is better).
- Builds ran sequentially to avoid dual-JVM core contention; no CPU pinning. Conclusions rest on order-of-magnitude and structural separation, not ±1% micro-deltas.
Raw JSON (local run artifacts, not committed):
/tmp/groovy-12191-perf-results/head/scoped-results.json/tmp/groovy-12191-perf-results/parent/scoped-results.json/tmp/groovy-12191-perf-results/head/callsite-results.json/tmp/groovy-12191-perf-results/parent/callsite-results.json
4. Results
4.1 ScopedInvalidationBench (Throughput, ops/ms — higher is better)
| Benchmark | Parent | HEAD | HEAD/Parent | Interpretation |
|---|---|---|---|---|
hotLoop_baseline |
3.891 ± 0.224 | 3.869 ± 0.219 | 0.99× | H1: no hot-path regression |
hotLoop_afterUnrelatedBurst |
3.774 ± 0.201 | 4.522 ± 0.168 | 1.20× | H3: full speed after unrelated burst |
hotLoop_unrelatedMetaClassChurn |
0.00247 ± 0.00017 | 0.386 ± 0.078 | ≈156× | H2 primary win |
hotLoop_sameTypeMetaClassChurn |
0.00249 ± 0.00017 | 0.0188 ± 0.0012 | ≈7.6× | Still forced re-link; narrower blast radius than parent |
hotLoop_categoryEnterLeave |
0.00225 ± 0.00045 | 0.00231 ± 0.00028 | ≈1.03× | H4: bulk semantics retained |
hierarchy_baseline |
3.325 ± 0.118 | 2.896 ± 0.113 | 0.87× | Noise / load variance; both full-speed order |
hierarchy_parentMetaClassChurn |
0.00233 ± 0.00021 | 0.00957 ± 0.00080 | ≈4.1× | H4: parent change still invalidates child sites |
Structural fingerprint on HEAD only:
baseline ≈ 3.87 ops/ms
unrelated churn ≈ 0.39 ops/ms (~20× slower than baseline: mostly MetaClass write cost)
same-type churn ≈ 0.019 ops/ms (~200× slower than baseline: write + same-class re-link)
category ≈ 0.002 ops/ms (bulk re-link tax)
On parent, unrelated and same-type both collapse to ~0.002, showing that the old global SwitchPoint made “unrelated churn” equivalent to “global deopt.” HEAD separates them by ~20× (0.386 / 0.019)—a direct fingerprint that scoping is live.
4.2 CallSiteInvalidationBench (AverageTime, ms/op — lower is better)
| Benchmark | Parent | HEAD | Speedup (P/H) | Interpretation |
|---|---|---|---|---|
baselineHotLoop |
0.257 ± 0.017 | 0.248 ± 0.014 | 1.04× | H1 |
baselineListSize |
0.242 ± 0.010 | 0.236 ± 0.006 | 1.03× | H1 |
baselineSteadyStateNoBurst |
0.612 ± 0.029 | 0.583 ± 0.033 | 1.05× | H1 |
baselineMultipleCallSites |
20.42 ± 1.03 | 18.86 ± 0.54 | 1.08× | H1 |
crossTypeInvalidationEvery10000 |
288.5 ± 24.3 | 2.73 ± 1.20 | ≈106× | H2 |
crossTypeInvalidationEvery1000 |
405.6 ± 27.3 | 2.77 ± 0.88 | ≈146× | H2 primary win |
crossTypeInvalidationEvery100 |
137.4 ± 10.2 | 12.5 ± 1.0 | ≈11× | Still large under denser churn |
listSizeWithCrossTypeInvalidation |
330.7 ± 25.0 | 2.21 ± 0.53 | ≈150× | JDK-type hot sites benefit too |
multipleCallSitesWithInvalidation |
1098 ± 89 | 25.9 ± 1.7 | ≈42× | Multi-site still much improved |
burstThenSteadyState |
145.8 ± 30.6 | 1.93 ± 0.16 | ≈76× | “Framework MC extend + request steady state” |
sameTypeInvalidationEvery1000 |
407.5 ± 29.4 | 53.3 ± 3.2 | ≈7.6× | Same-type still slow; local invalidation cheaper than global rotate |
Penalty vs baseline on HEAD:
| Scenario | Relative cost | Meaning |
|---|---|---|
| crossType @1000 | 2.77 / 0.25 ≈ 11× baselineHotLoop | Dominated by ~100 ColdType MetaClass writes, not hot-site deopt |
| sameType @1000 | 53.3 / 0.25 ≈ 215× baselineHotLoop | Clear re-link tax |
| burstThenSteady | 1.93 / 0.58 ≈ 3.3× steady baseline | Burst write cost dominates; steady calls recovered |
On parent, crossType@1000 and sameType@1000 are both ~406 ms/op, again confirming global invalidation collapsed cross-type into full deopt.
5. Why it is faster
Old path New path
ColdType.metaClass.foo = ... ColdType.metaClass.foo = ...
│ │
▼ ▼
invalidate global SwitchPoint invalidate ClassInfo(ColdType) SP
│ │
▼ ▼
ALL sites: guard fails ColdType sites only
HotTarget.compute → full re-select HotTarget.compute → stays linked
(MH rebuild / re-guard / cold path) (single SwitchPoint check passes)
Cost breakdown:
- Avoided re-links — After each global invalidation, hot sites re-run
Selector, reinstall guards, and rewriteCallSites. At 100 churn events × many sites this dominates (parent’s 400+ ms/op). - Avoided JIT churn — Frequent
SwitchPoint.invalidateAllbreaks monomorphic inlining assumptions and amplifies deopt/recompile cost (especially under dense cross-type churn). - Unchanged hot-path guard count — Still one
guardWithTest, so baselines stay flat (H1). - Same-type / hierarchy still correctly charged — Prevents “fast but wrong” semantics (H4 + 56 tests).
- Category remains bulk — No second hot-path guard; category visibility stays simple and correct; cost matches the old global path (measured ~1.0×).
Same-type still improves ~7–8× on HEAD vs parent, not because re-link is skipped, but because the invalidation surface shrinks from “all process sites” to “this class hierarchy” and batch invalidateAll is cheaper—a secondary benefit, not the primary goal.
6. Risks and boundaries
| Topic | Assessment |
|---|---|
| Hot-path regression | Not observed (baselines 0.99–1.08×) |
| Semantic regression | 56/56 tests for this commit; Category / hierarchy / per-instance MC covered |
| Dense Category enter/leave | Still expensive (by design); Category-heavy workloads do not get faster |
| Non-final parent MetaClass change | Hierarchy scan over ClassInfo.getAllClassInfo() is O(loaded classes); cold path, correctness-first |
Legacy IndyInterface.switchPoint |
Rotated for binary compatibility only; no longer the real call-site MOP guard; external readers should migrate to IndyInvalidation |
| Environment noise | Shared/container CPUs; conclusions use order-of-magnitude gaps (10¹–10²×), robust to ±10% noise |
| Not covered here | Concurrent multi-thread churn throughput, full Grails end-to-end, JDK 17/21 matrix (recommended for CI) |
7. Hypothesis scorecard
| Hypothesis | Result | Evidence |
|---|---|---|
| H1 baseline parity | Holds | Scoped baseline 0.99×; CallSite baselines 1.03–1.08× |
| H2 cross-type speedup | Holds | ≈156× thrpt; ≈146× avgt (@1000) |
| H3 post-unrelated-burst full speed | Holds | afterUnrelatedBurst ≥ baseline on HEAD |
| H4 necessary invalidation still paid | Holds | same-type / hierarchy / category ≪ baseline |
| H5 correctness | Holds | 56/56 tests |
8. Conclusions and recommendations
Verdict
ae8e497 (GROOVY-12191) keeps a single-guard monomorphic hot path while scoping indy MOP invalidation to per-class domains (+ hierarchy). For unrelated-type MetaClass churn with hot monomorphic call sites—a Grails/framework-shaped load—measured improvement is on the order of 10¹–10²×. Steady-state undisturbed paths show no material regression. Correctness costs for Category and same-type changes remain visible and intentional.
Performance verification: PASS.
Recommendations
- Keep
ScopedInvalidationBenchand the cross-type / burst rows ofCallSiteInvalidationBenchin regular JMH regression (README already documents the suite) so a return to a global SwitchPoint cannot land silently. - In release notes, state clearly that Category and unscoped MetaClass events may still bulk-invalidate; the win is type-local MetaClass change.
- Optionally re-run CallSite
crossType@1000and baselines on JDK 17/21 to confirm cross-LTS consistency. - Document migration away from reading
IndyInterface.switchPoint(field is@Deprecated).
Appendix A — Reproduction commands
# Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest
# Performance (HEAD)
./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSON
# Performance (parent @ 0ca665dad0; mount the same ScopedInvalidationBench sources for a fair A/B)
cd /path/to/parent && ./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
cd /path/to/parent && ./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSONAppendix B — Files touched by the commit
- New:
org.apache.groovy.runtime.indy.IndyInvalidation,SwitchPointInvalidator - Updated:
ClassInfo,IndyInterface,Selector,ColdReflectiveMethodHandleWrapper,IndyCompoundAssign - Tests:
Groovy12191,IndyInvalidationTest,SwitchPointInvalidatorTest,IndyScopedSwitchPointTest - Benchmarks:
ScopedInvalidationBench; docs:subprojects/performance/README.adoc
Report generated from local JMH A/B measurements on 2026-07-26. Raw JSON and run logs: /tmp/groovy-12191-perf-results/.
|
AI read:
|
|
Thanks for your review.
Happy to adjust if anything still looks off. |
GROOVY-12191 Performance Verification Report V2Scoped indy SwitchPoint invalidation (re-verification after review follow-ups)
1. Executive summaryThis verification covers the full GROOVY-12191 stack on the current tip of PR #2736:
On the same host, JDK, and JMH annotation settings, A/B measurement of HEAD (
2. Change mechanism and performance hypotheses2.1 Previous model (baseline
|
| Component | Role |
|---|---|
SwitchPointInvalidator |
Single-domain lifecycle: CAS get / detachLive / invalidate |
IndyInvalidation |
Policy: invalidateClass (+ hierarchy), invalidateCategory (bulk), guardWithMopSwitchPoints |
ClassHierarchyIndex |
Weak ancestor→descendant index; O(|subtypes|) fan-out incl. JLS array lattice |
ClassInfo |
Holds one indySwitchPoint per class; registers into the index; coordinates version bumps |
Selector / IndyCompoundAssign / ColdReflectiveMethodHandleWrapper |
Install a per-receiver-class SwitchPoint guard at link time |
Hot-path shape (intentionally monomorphic):
handle = classSp.guardWithTest(fast, fallback)
Still a single guard—matching the pre-6.0 monomorphic shape. There is no second category SwitchPoint on the hot path.
Invalidation policy:
| Event | Behavior | Expected hot-path cost |
|---|---|---|
| Unrelated-type MetaClass change | Retire that type’s domain (and indexed subtypes) only | Hot sites do not re-link |
| Same-type / parent MetaClass change | Retire class + hierarchy fan-out | Must re-link (correctness) |
Category enter/leave, invalidateCallSites() |
Bulk-retire all loaded class domains | Same order of cost as old global invalidation |
| Unattributed MetaClass event | invalidateUnscoped() bulk |
Conservative, correctness-first |
Array root (Object[], interface arrays) |
Fan-out via full JLS array covariance lattice in the index | Correct subtype retirement (review fix) |
What f48ab19 specifically adds beyond ae8e497:
| Review item | Resolution in f48ab19 |
|---|---|
Array fan-out hole (Object[] treated as final leaf) |
ClassHierarchyIndex indexes full array lattice |
Legacy switchPoint lost lock |
Restored synchronized (IndyInterface.class) on bulk rotation |
switchPoint semantics over-claim |
Docs: bulk-only rotation; @Deprecated(since="6.0.0", forRemoval=false) |
incVersion() no longer global flush |
Documented; bulk callers → invalidateCallSites() / invalidateCategory() |
| Scalability O(N×M) per typed MC change | Weak subtype index → O(|subtypes of T|); category still full walk |
2.3 Hypotheses under test
| ID | Hypothesis | How verified |
|---|---|---|
| H1 | Baseline hot loop has no material regression | A/B ratio ≈ 1 |
| H2 | Cross-type churn throughput/latency improves substantially | Unrelated ≫ parent; on HEAD, unrelated ≫ same-type |
| H3 | Pure hot loop after unrelated burst approaches baseline | afterUnrelatedBurst ≈ baseline |
| H4 | Same-type / hierarchy / category still pay re-link tax | Clearly slower than baseline |
| H5 | Correctness suite stays green (incl. index / array cases) | Unit / integration tests |
3. Methodology
3.1 Environment
| Item | Value |
|---|---|
| OS | Linux 6.15.5 x86_64 (hostname: hera) |
| CPU | AMD EPYC 7763 (6 vCPUs visible) |
| Memory | 23 GiB |
| JDK | Amazon Corretto 25.0.2 (25.0.2+10-LTS) |
| Build | Gradle wrapper, ./gradlew :perf:jmh, indy enabled by default |
| Trees | HEAD: project worktree at f48ab19; baseline: worktree at 0ca665da (/tmp/groovy-12191-parent) |
| Bench source parity | ScopedInvalidationBench + CallSiteInvalidationBench workloads identical on parent (sources copied into parent worktree for fair A/B) |
| Run order | Sequential (HEAD then parent per suite) to avoid dual-JVM core contention; no CPU pinning |
3.2 Benchmark suites
-
org.apache.groovy.bench.ScopedInvalidationBench(added byae8e497; Throughput, ops/ms)- 100 000 monomorphic calls per op; MetaClass write every 1000 iterations in churn cases.
- JMH:
@Warmup(3×2s) @Measurement(5×2s) @Fork(2)→ 10 samples per benchmark.
-
org.apache.groovy.perf.grails.CallSiteInvalidationBench(existing Grails-oriented suite; AverageTime, ms/op)- Same 100 000-iteration loops; cross-type / same-type / burst patterns.
- Same JMH annotation settings.
3.3 Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.apache.groovy.runtime.indy.ClassHierarchyIndexTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest --rerun-tasksResult: 76 passed / 0 failed / 0 skipped / 0 error.
| Test class | Cases (approx.) | Notes |
|---|---|---|
bugs.Groovy12191 |
10 | End-to-end MOP / category / incVersion scoping |
IndyInvalidationTest |
34 | Policy, bulk, unscoped, hierarchy, concurrency probes |
SwitchPointInvalidatorTest |
12 | CAS lifecycle, concurrent detach/get races |
ClassHierarchyIndexTest |
9 | Class/interface/array lattice indexing (f48ab19) |
IndyScopedSwitchPointTest |
11 | Guard install, cold tier, logging child process |
3.4 Statistics and artifacts
- Scores are JMH means;
±is JMH’s default 99.9% CI. - Throughput ratio = HEAD / parent (>1 is better).
- AverageTime speedup = parent / HEAD (>1 is better).
- Conclusions rest on order-of-magnitude and structural separation, not ±1% micro-deltas.
Raw JSON (local run artifacts, not committed):
/tmp/groovy-12191-perf-v2/head/scoped-results.json/tmp/groovy-12191-perf-v2/parent/scoped-results.json/tmp/groovy-12191-perf-v2/head/callsite-results.json/tmp/groovy-12191-perf-v2/parent/callsite-results.json/tmp/groovy-12191-perf-v2/summary.json
4. Results
4.1 ScopedInvalidationBench (Throughput, ops/ms — higher is better)
| Benchmark | Parent | HEAD | HEAD/Parent | Interpretation |
|---|---|---|---|---|
hotLoop_baseline |
3.931 ± 0.207 | 4.068 ± 0.113 | 1.03× | H1: no hot-path regression |
hotLoop_afterUnrelatedBurst |
3.889 ± 0.242 | 4.360 ± 0.289 | 1.12× | H3: full speed after unrelated burst |
hotLoop_unrelatedMetaClassChurn |
0.002278 ± 0.000316 | 0.4486 ± 0.124 | ≈197× | H2 primary win |
hotLoop_sameTypeMetaClassChurn |
0.002327 ± 0.000251 | 0.01892 ± 0.00132 | ≈8.1× | Still forced re-link; narrower blast radius than parent |
hotLoop_categoryEnterLeave |
0.002350 ± 0.000136 | 0.002194 ± 0.000421 | ≈0.93× | H4: bulk semantics retained (parity within noise) |
hierarchy_baseline |
3.295 ± 0.142 | 2.791 ± 0.117 | 0.85× | Both full-speed order; container noise |
hierarchy_parentMetaClassChurn |
0.002352 ± 0.000169 | 0.01021 ± 0.000709 | ≈4.3× | H4: parent change still invalidates child sites |
Structural fingerprint on HEAD only:
baseline ≈ 4.07 ops/ms
after burst ≈ 4.36 ops/ms (≥ baseline: no residual deopt)
unrelated churn ≈ 0.45 ops/ms (~9× slower than baseline: mostly MetaClass write cost)
same-type churn ≈ 0.019 ops/ms (~215× slower than baseline: write + same-class re-link)
category ≈ 0.002 ops/ms (bulk re-link tax)
On parent, unrelated and same-type both collapse to ~0.002, showing that the old global SwitchPoint made “unrelated churn” equivalent to “global deopt.” HEAD separates them by ~24× (0.449 / 0.019)—a direct fingerprint that scoping is live.
4.2 CallSiteInvalidationBench (AverageTime, ms/op — lower is better)
| Benchmark | Parent | HEAD | Speedup (P/H) | Interpretation |
|---|---|---|---|---|
baselineHotLoop |
0.264 ± 0.021 | 0.252 ± 0.015 | 1.05× | H1 |
baselineListSize |
0.241 ± 0.004 | 0.250 ± 0.018 | 0.96× | H1 (noise) |
baselineSteadyStateNoBurst |
0.608 ± 0.021 | 0.586 ± 0.020 | 1.04× | H1 |
baselineMultipleCallSites |
20.82 ± 1.56 | 19.01 ± 1.09 | 1.10× | H1 |
crossTypeInvalidationEvery10000 |
306.1 ± 33.1 | 2.503 ± 0.653 | ≈122× | H2 |
crossTypeInvalidationEvery1000 |
428.4 ± 49.0 | 2.250 ± 0.894 | ≈190× | H2 primary win |
crossTypeInvalidationEvery100 |
139.2 ± 12.8 | 7.916 ± 0.701 | ≈18× | Still large under denser churn |
listSizeWithCrossTypeInvalidation |
323.4 ± 56.4 | 1.750 ± 0.394 | ≈185× | JDK-type hot sites benefit too |
multipleCallSitesWithInvalidation |
1103 ± 56.2 | 25.68 ± 5.97 | ≈43× | Multi-site still much improved |
burstThenSteadyState |
154.7 ± 23.0 | 1.361 ± 0.112 | ≈114× | “Framework MC extend + request steady state” |
sameTypeInvalidationEvery1000 |
410.6 ± 23.5 | 51.45 ± 2.06 | ≈8.0× | Same-type still slow; local invalidation cheaper than global rotate |
Penalty vs baseline on HEAD:
| Scenario | Relative cost | Meaning |
|---|---|---|
| crossType @1000 | 2.25 / 0.25 ≈ 9× baselineHotLoop |
Dominated by ~100 ColdType MetaClass writes, not hot-site deopt |
| sameType @1000 | 51.5 / 0.25 ≈ 205× baselineHotLoop |
Clear re-link tax |
| burstThenSteady | 1.36 / 0.59 ≈ 2.3× steady baseline | Burst write cost dominates; steady calls recovered |
On parent, crossType@1000 and sameType@1000 are both ~410–428 ms/op, again confirming global invalidation collapsed cross-type into full deopt.
4.3 Comparison to the earlier ae8e497-only report
An earlier local report (posted on PR #2736 for ae8e497 alone) measured ~146× on crossType@1000 and ~156× on scoped unrelated thrpt. This re-run on f48ab19 measures ~190× / ~197× on the same primary rows. Differences of tens of percent on multi-hundred-× ratios are expected on a 6-vCPU shared host; the structural story is unchanged:
- baselines stay ~1×,
- cross-type improves by 10¹–10²×,
- same-type / category remain correctly expensive,
- HEAD separates unrelated from same-type by ~20× where parent could not.
f48ab19 is primarily a correctness / scalability / review-hygiene commit (index, array lattice, lock restore, docs). It is not expected to move monomorphic hot-path baselines; the A/B above confirms that.
5. Why it is faster
Old path New path
ColdType.metaClass.foo = ... ColdType.metaClass.foo = ...
│ │
▼ ▼
invalidate global SwitchPoint invalidate ClassInfo(ColdType) SP
│ (+ ClassHierarchyIndex descendants)
▼ ▼
ALL sites: guard fails ColdType (+ subtypes) sites only
HotTarget.compute → full re-select HotTarget.compute → stays linked
(MH rebuild / re-guard / cold path) (single SwitchPoint check passes)
Cost breakdown:
- Avoided re-links — After each global invalidation, hot sites re-run
Selector, reinstall guards, and rewriteCallSites. At 100 churn events × many sites this dominates (parent’s 400+ ms/op). - Avoided JIT churn — Frequent
SwitchPoint.invalidateAllbreaks monomorphic inlining assumptions and amplifies deopt/recompile cost (especially under dense cross-type churn). - Unchanged hot-path guard count — Still one
guardWithTest, so baselines stay flat (H1). - Same-type / hierarchy still correctly charged — Prevents “fast but wrong” semantics (H4 + 76 tests).
- Category remains bulk — No second hot-path guard; category visibility stays simple and correct; cost matches the old global path (measured ~1×).
- Indexed fan-out (
f48ab19) — Typed MetaClass invalidation walks only registered descendants, not every loaded class. Correctness forObject[]/ interface arrays / multi-dim arrays is restored without reintroducing a global SP.
Same-type still improves ~8× on HEAD vs parent, not because re-link is skipped, but because the invalidation surface shrinks from “all process sites” to “this class hierarchy” and batch invalidateAll is cheaper—a secondary benefit, not the primary goal.
6. Risks and boundaries
| Topic | Assessment |
|---|---|
| Hot-path regression | Not observed (baselines 0.96–1.10×) |
| Semantic regression | 76/76 tests for this tip; Category / hierarchy / per-instance MC / array lattice covered |
| Dense Category enter/leave | Still expensive (by design); Category-heavy workloads do not get faster |
| Hierarchy baseline 0.85× | Same full-speed band as parent; treat as noise on shared 6-vCPU host, not a design regression |
| Non-final parent MetaClass change | Fan-out is O(|indexed subtypes|) after f48ab19 (was full ClassInfo walk in ae8e497) |
| Category bulk | Still walks all loaded ClassInfos; detach is a cheap no-op when no SP was allocated |
Legacy IndyInterface.switchPoint |
Rotated for binary compatibility (bulk paths only, under lock again); no longer the real call-site MOP guard |
| Environment noise | Shared/container CPUs; conclusions use order-of-magnitude gaps (10¹–10²×), robust to ±10% noise |
| Not covered here | Concurrent multi-thread churn throughput, full Grails end-to-end, JDK 17/21 matrix (recommended for CI) |
7. Hypothesis scorecard
| Hypothesis | Result | Evidence |
|---|---|---|
| H1 baseline parity | Holds | Scoped baseline 1.03×; CallSite baselines 0.96–1.10× |
| H2 cross-type speedup | Holds | ≈197× thrpt; ≈190× avgt (@1000) |
| H3 post-unrelated-burst full speed | Holds | afterUnrelatedBurst ≥ baseline on HEAD |
| H4 necessary invalidation still paid | Holds | same-type / hierarchy / category ≪ baseline |
| H5 correctness | Holds | 76/76 tests (incl. ClassHierarchyIndexTest) |
8. Conclusions and recommendations
Verdict
ae8e497 + f48ab19 (GROOVY-12191, PR #2736) keep a single-guard monomorphic hot path while scoping indy MOP invalidation to per-class domains (+ indexed hierarchy / array lattice). For unrelated-type MetaClass churn with hot monomorphic call sites—a Grails/framework-shaped load—measured improvement is on the order of 10¹–10²× (primary rows ≈ 190–197× on this host). Steady-state undisturbed paths show no material regression. Correctness costs for Category and same-type changes remain visible and intentional. Review follow-ups (lock restore, docs, array fan-out, subtype index) do not undo the win.
Performance verification: PASS.
Recommendations
- Keep
ScopedInvalidationBenchand the cross-type / burst rows ofCallSiteInvalidationBenchin regular JMH regression (README already documents the suite) so a return to a global SwitchPoint cannot land silently. - In release notes, state clearly that Category and unscoped MetaClass events may still bulk-invalidate; the win is type-local MetaClass change.
- Document migration away from reading
IndyInterface.switchPoint(field is@Deprecated; bulk-only rotation). - Optionally re-run CallSite
crossType@1000and baselines on JDK 17/21 to confirm cross-LTS consistency. - CI “Performance Alert” on throughput benches that report higher ops/ms as “worse” remains a bot false-positive class; prefer calibrated geomean summaries for PR triage.
Appendix A — Reproduction commands
# Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.apache.groovy.runtime.indy.ClassHierarchyIndexTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest
# Performance (HEAD @ f48ab19)
./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSON
# Performance (parent @ 0ca665da; same ScopedInvalidationBench + CallSiteInvalidationBench sources)
cd /path/to/parent && ./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
cd /path/to/parent && ./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSONAppendix B — Files touched (combined)
ae8e497 — scoping core
- New:
org.apache.groovy.runtime.indy.IndyInvalidation,SwitchPointInvalidator, package-info - Updated:
ClassInfo,IndyInterface,Selector,ColdReflectiveMethodHandleWrapper,IndyCompoundAssign - Tests:
Groovy12191,IndyInvalidationTest,SwitchPointInvalidatorTest,IndyScopedSwitchPointTest - Benchmarks:
ScopedInvalidationBench; docs:subprojects/performance/README.adoc
f48ab19 — review follow-ups + index
- New:
ClassHierarchyIndex,ClassHierarchyIndexTest - Updated:
IndyInvalidation(index-backed fan-out),ClassInfo(register),IndyInterface(sync + deprecation metadata), docs/tests for array /incVersion/ bulk paths
*Report generated from local JMH A/B measurements on 2026-07-26 (re-run on tip f48ab19 vs baseline 0ca665da).
blackdrag
left a comment
There was a problem hiding this comment.
There is one point that escapes me right now. Why do we need the hierarchy at all? I don´t think this is for MetaClassImpl cases.
| */ | ||
| protected static SwitchPoint switchPoint = new SwitchPoint(); | ||
| @Deprecated(since = "6.0.0", forRemoval = false) | ||
| protected static volatile SwitchPoint switchPoint = new SwitchPoint(); |
There was a problem hiding this comment.
all under vmplugin is supposed to be Internal, thus removing instead of deprecating is imho very much allowed. Also why do you make it volatile if you deprecate it?
There was a problem hiding this comment.
Agreed — thank you.
org.codehaus.groovy.vmplugin.* is internal-by-intent. Keeping a dead process-wide SwitchPoint only for binary compatibility was weaker than a clean removal: after scoping, the field was no longer the MOP guard and was rotated only on bulk paths, which was easy to misread.
Action taken
- Removed
protected static SwitchPoint switchPointentirely. invalidateSwitchPoints()now only bulk-retires per-class domains viaIndyInvalidation.invalidateCategory()(no legacy field rotation, novolatile, no extra lock).- Docs/
package-infoupdated; tests no longer assert on the removed field.
The earlier volatile was only there to make concurrent bulk rotation of that legacy field observable; with the field gone, that concern disappears.
| */ | ||
| static MethodHandle applyMopSwitchPoints(final MethodHandle handle, final MethodHandle fallback, final Object receiver) { | ||
| return IndyInvalidation.guardWithMopSwitchPoints(handle, fallback, receiver); | ||
| } |
There was a problem hiding this comment.
what is the rationale for having the guard creation in IndyInvalidation instead of here.
There was a problem hiding this comment.
Good question — the split is intentional, and the previous one-line delegate made it look accidental.
Rationale
| Concern | Owner |
|---|---|
Domain resolution (receiver → class → ClassInfo SwitchPoint) |
IndyInvalidation (also used by ClassInfo, cold tier, tests) |
| Invalidation policy (class / hierarchy / category / unscoped) | IndyInvalidation |
Link-time guardWithTest install on production call sites |
IndyInterface.applyMopSwitchPoints (Selector, IndyCompoundAssign) |
Putting invalidation policy next to bootstrap wiring would grow IndyInterface further and pull ClassInfo toward vmplugin. Putting only the bind step in IndyInterface keeps the bytecode surface as the place that wires handles, while the runtime.indy package owns domains and retire rules.
Action taken
-
applyMopSwitchPointsnow performs the bind itself:IndyInvalidation.classSwitchPointFor(receiver).guardWithTest(handle, fallback) -
IndyInvalidation.guardWithMopSwitchPointsremains the public / test entry (same semantics), documented as such. -
Javadoc on both sides spells out the division of responsibility.
…tion Remove the legacy process-wide IndyInterface.switchPoint (vmplugin is internal-by-intent), bind MOP guards at link time in applyMopSwitchPoints, document hierarchy fan-out for cross-class MOP visibility, and clarify the cold-path PIC sentinel policy. Extend unit coverage for these paths.
You are right that this is not because Hierarchy fan-out exists for cross-class MOP visibility that already-linked subtype sites must re-observe:
So: fan-out is about inherited / covariant MOP state, not about rewriting Action taken
Happy to discuss a narrower fan-out (e.g. EMC-only) as a follow-up optimization; the current rule prefers correctness and matches pre-6.0 “something in the MOP above me changed → re-link” behaviour, but scoped to the subtype cone instead of the whole process. |
This all sounds all like there is some problem deeper down. Technically none of these operations would have to invalidate the callsite, since fooC is unchanged, but of course that is not how it works in reality. But the question is when do we invalidate and why. For me this is not yet the stage to propose a follow-up, first I would like to better understand the old and new behavior and think about if the difference matters. Also of interest is when and if the those foo methods becomes visible on the instance c as well as the metaClass of C. |
…name Narrow hierarchy invalidation to an allow-list of cross-class MOP cases (EMC, modified custom MetaClass, interface/array, global EMC) while pure MetaClassImpl replaces and per-instance MetaClass stay exact-class. Rename NULL_METHOD_HANDLE_WRAPPER to UNCACHEABLE_PIC_SENTINEL and document the invalidation matrix. Extend unit coverage for blackdrag's scenario and the refined policy.
|
@blackdrag Thank you for pressing on the “when / why” of hierarchy fan-out. That is the right question, and the previous reply was still too coarse. Below is the refined model after re-checking MOP selection against your scenario. 1. When we invalidate (and when we do not)Monomorphic indy sites guard a class-domain SwitchPoint on the receiver’s runtime class (via For your ladder
So for Regression coverage for this matrix is in 2. Why hierarchy exists at all (and when it does not)Hierarchy fan-out is not “ It exists because selection can observe ancestor MetaClass state:
So hierarchy is required for cross-class MOP visibility, not for ordinary Refined policy (exact-class is an allow-list)
This matches your intuition: without EMC (and without a modified mutable MetaClass), hierarchy is unnecessary. Pure 3. “Should the MetaClass own the SwitchPoint?”Agreed in principle. The current class-domain SP is a stand-in for class-level MetaClass generation, not a claim that every MetaClass kind shares the same invalidation logic. A MetaClass-owned guard would be the natural place for:
That is a larger redesign than this PR. For 6.0 we kept a single monomorphic hot-path guard and made invalidation MetaClass-aware at the registry boundary. MetaClass-owned SwitchPoints remain a follow-up once the “when / why” matrix above is agreed. 4. Update vs replace, and
|
Now I feel kind of bad, because I made a major mistake in my example. not [...]
but the important point here is that this happens only for a missing method. And it is actually why I think we are having a major semantic gap here, that we have to decide on: The case of C and D show how broken the current system is. The point of time a MetaClassImpl instance comes into existence should not decide about what it sees. In fact my POV is that MetaClassImpl should be immune to changes of other meta classes. This means especially that no hierarchy check is needed for it and that find method should not exist on MetaClassImpl. Now to produce a version with callsite caching: Most of this should not be caught by the switch point, but by the fact, that the metaclass changed. But I also think I captured cases where a hierarchy version would break the code in case of MetaClassImpl. Also I think it should behave the same if classes used are Java based and not from Groovy. [...]
+1
ok, agreed... if a follow-up issue is created |
…nario tests Refine scoped SwitchPoint docs to match MetaClassImpl selection: hierarchy fan-out is for the missing-method walk under modified MutableMetaClass, not a shared MetaClassImpl table; construction-time C/D visibility remains pre-existing MOP behaviour. Add regression coverage for parent→child miss/hit paths, present- method stability after re-link, callsite MetaClass-change matrix, and Java subtype EMC fan-out.
|
@blackdrag Thank you for the correction and for the detailed parent→child matrix — that is exactly the right level of precision. 1. Hierarchy directionUnderstood: the scenario of interest is superclass MetaClass changes observed from a child receiver, not the reverse. Sorry the earlier matrix inverted that ladder. The tests now follow your corrected direction ( 2. Missing-method-only walk (and the C/D gap)Agreed on the important refinement:
6.0 decision for this PR: preserve current MOP visibility rules; scope invalidation to match those rules. Making Documented on 3. Callsite caching vs hierarchy versionAgreed that most of your second script is driven by MetaClass change on the receiver, not by a hierarchy version:
Hierarchy SwitchPoint is only there so already-linked miss sites on subtypes re-select when an ancestor becomes / stops being a modified mutable MetaClass (e.g. linked miss → parent adds method; linked hit → parent EMC removed). Measured note on the last line of your draft ( Pure Java inheritance matches the miss path ( 4. Policy (your +1)Unchanged and reaffirmed:
5. MetaClass-owned SwitchPoint — follow-upAgreed, and we will file follow-up JIRAs (not landing in this PR):
I will link the issue keys in 6. Regression coverage added
Happy to adjust further if any row of the matrix still disagrees with the behaviour you want for 6.0. |
|
[...]
which means we add a lot of machinery for a small feature. Sure, considering how this is implemented it makes sense, but from a user side I think it does not.
[...]
So if class A extends B and A or B have a method foo, and we have a callsite new B().foo(x), then we do not need to invalidate it if the meta class of A changes. Only if the metaClass for A changes (new or mutates) then we have to invalidate the callsite. If this is the case, then why the hierarchy logic? If I take the above together I get for example something like this on a pure meta class level: Only replacing the meta class of E made the more specific foo in D visible and it stays visible, even though we removed the meta class of D. Is it because of the declared method? No it is not. To sum it up: In what case must the mutation/replacement of a parent meta class invalidate a callsite if the meta class of the current receiver stays unchanged? |
|
@blackdrag Thank you — this is exactly the right question, and your A/B/C, D/E, F/G scripts make the boundary sharp. Short answerHierarchy SwitchPoint fan-out is needed in one live-selection situation only:
Concrete dual:
That is all hierarchy SwitchPoint is for. It is not a general “parent MetaClass version” for child tables, and it does not rebuild a child’s method index. Mapping your examplesThey are consistent with that rule; most of them do not require hierarchy SP:
class Parent {}
class Child extends Parent {}
def call(x) {
try { x.onlyOnParent(); 'hit' }
catch (MissingMethodException e) { 'miss' }
}
def c = new Child()
assert call(c) == 'miss' // monomorphic miss linked on Child
Parent.metaClass.onlyOnParent = { -> 'now-visible' }
assert call(c) == 'hit' // must re-select; Child MetaClass unchangedand the inverse (linked hierarchy hit → parent EMC removed → miss again). Without fan-out, those already-linked subtype sites stay wrong until something else retires the child’s domain. “A lot of machinery for a small feature”Fair point. The feature is small in description (missing-method hierarchy + EMC), but it is the same feature that makes What we did simplify on the domain side (current tip of this PR):
So hierarchy logic is not “version every parent change for every child site”; it is “when parent can change what a linked miss/hit on the missing-method path means, re-select subtype sites.” Bottom line
Happy to tighten docs or tests further if any row of this matrix still disagrees with the behaviour you want for 6.0. |
|
🚨 TestLens detected 1 failed test 🚨Here is what you can do:
Test SummaryBuild and test / lts (17, windows-latest, 1) > :test
🏷️ Commit: 637bdaf Test FailuresClassTagExtensionModuleTest (:test in Build and test / lts (17, windows-latest, 1))Muted TestsSelect tests to mute in this pull request:
Reuse successful test results:
Click the checkbox to trigger a rerun:
Learn more about TestLens at testlens.app. |
GROOVY-12191 Performance Verification Report V3Scoped indy SwitchPoint invalidation (MetaClass identity-map domains)
1. Executive summaryThis verification compares the full GROOVY-12191 stack at At HEAD, indy MOP invalidation uses:
On the same host, JDK, and JMH annotation settings, A/B measurement of HEAD (
2. Change mechanism and performance hypotheses2.1 Previous model (baseline
|
| Component | Role |
|---|---|
SwitchPointInvalidator |
Single-domain lifecycle: CAS get / detachLive / invalidate |
IndyInvalidation |
Policy + weak identity map MetaClass → SwitchPointInvalidator; hierarchy / category / exact-class |
ClassInfo |
Pending domain for pre-MetaClass link only; version bumps; hierarchy index registration |
ClassHierarchyIndex |
Weak ancestor→descendant index; O(|subtypes|) fan-out incl. JLS array lattice |
Selector / IndyCompoundAssign / ColdReflectiveMethodHandleWrapper |
Install MetaClass-domain (or pending) SwitchPoint guard at link time |
Hot-path shape (intentionally monomorphic):
handle = mopSwitchPoint.guardWithTest(fast, fallback)
Still a single guard—matching the pre-6.0 monomorphic shape. There is no second category SwitchPoint on the hot path.
Invalidation policy:
| Event | Behavior | Expected hot-path cost |
|---|---|---|
| Unrelated-type MetaClass change | Retire that MetaClass domain (and indexed subtypes when fan-out applies) only | Hot sites do not re-link |
| Same-type / parent EMC MetaClass change | Retire class + hierarchy fan-out | Must re-link (correctness) |
Pure unmodified MetaClassImpl class replace |
Exact class only | Subtypes not force-relinked |
Category enter/leave, invalidateCallSites() |
Bulk-retire class-level / pending domains of loaded types | Same order of cost as old global invalidation |
| Unattributed MetaClass event | invalidateUnscoped() bulk |
Conservative, correctness-first |
2.3 Hypotheses under test
| ID | Hypothesis | How verified |
|---|---|---|
| H1 | Baseline hot loop has no material regression | A/B ratio ≈ 1 |
| H2 | Cross-type churn throughput/latency improves substantially | Unrelated ≫ parent; on HEAD, unrelated ≫ same-type |
| H3 | Pure hot loop after unrelated burst approaches baseline | afterUnrelatedBurst ≈ baseline |
| H4 | Same-type / hierarchy / category still pay re-link tax | Clearly slower than baseline |
| H5 | Correctness suite stays green | Unit / integration tests |
3. Methodology
3.1 Environment
| Item | Value |
|---|---|
| OS | Linux 6.15.5 x86_64 (hostname: hera) |
| CPU | AMD EPYC 7763 (6 vCPUs visible) |
| Memory | 23 GiB |
| JDK | Amazon Corretto 25.0.2 (25.0.2+10-LTS) |
| Build | Gradle wrapper, ./gradlew :perf:jmh, indy enabled by default |
| Trees | HEAD worktree @ 637bdaf68a; baseline worktree @ 0ca665dad0 |
| Bench source parity | ScopedInvalidationBench copied into parent worktree for fair A/B (not present on parent history); CallSiteInvalidationBench workloads identical |
| Run order | Sequential: HEAD scoped → HEAD callsite → parent scoped → parent callsite (avoid dual-JVM core contention); no CPU pinning |
3.2 Benchmark suites
-
org.apache.groovy.bench.ScopedInvalidationBench(Throughput, ops/ms)- 100 000 monomorphic calls per op; MetaClass write every 1000 iterations in churn cases.
- JMH:
@Warmup(3×2s) @Measurement(5×2s) @Fork(2)→ 10 samples per benchmark.
-
org.apache.groovy.perf.grails.CallSiteInvalidationBench(AverageTime, ms/op)- Same 100 000-iteration loops; cross-type / same-type / burst patterns.
- Same JMH annotation settings.
3.3 Correctness
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.apache.groovy.runtime.indy.ClassHierarchyIndexTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest --rerun-tasksResult (HEAD 637bdaf68a): 107 passed / 0 failed / 0 skipped / 0 error.
| Test class | Cases |
|---|---|
bugs.Groovy12191 |
24 |
IndyInvalidationTest |
47 |
SwitchPointInvalidatorTest |
12 |
ClassHierarchyIndexTest |
12 |
IndyScopedSwitchPointTest |
12 |
| Total | 107 |
3.4 Statistics and artifacts
- Scores are JMH means;
±is JMH’s default 99.9% CI. - Throughput ratio = HEAD / parent (>1 is better).
- AverageTime speedup = parent / HEAD (>1 is better).
- Conclusions rest on order-of-magnitude and structural separation, not ±1% micro-deltas.
Raw JSON (local run artifacts, not committed):
/tmp/groovy-12191-perf-637bdaf/head/scoped-results.json/tmp/groovy-12191-perf-637bdaf/parent/scoped-results.json/tmp/groovy-12191-perf-637bdaf/head/callsite-results.json/tmp/groovy-12191-perf-637bdaf/parent/callsite-results.json/tmp/groovy-12191-perf-637bdaf/summary.json
4. Results
4.1 ScopedInvalidationBench (Throughput, ops/ms — higher is better)
| Benchmark | Parent | HEAD | HEAD/Parent | Interpretation |
|---|---|---|---|---|
hotLoop_baseline |
3.756 ± 0.151 | 4.015 ± 0.274 | 1.07× | H1: no hot-path regression |
hotLoop_afterUnrelatedBurst |
3.693 ± 0.439 | 4.352 ± 0.423 | 1.18× | H3: full speed after unrelated burst |
hotLoop_unrelatedMetaClassChurn |
0.002336 ± 0.000187 | 0.352 ± 0.064 | ≈151× | H2 primary win |
hotLoop_sameTypeMetaClassChurn |
0.002334 ± 0.000234 | 0.0195 ± 0.0013 | ≈8.4× | Still forced re-link; narrower blast radius than parent |
hotLoop_categoryEnterLeave |
0.00221 ± 0.00026 | 0.00209 ± 0.00045 | ≈0.95× | H4: bulk semantics retained (parity within noise) |
hierarchy_baseline |
3.133 ± 0.122 | 2.817 ± 0.103 | 0.90× | Both full-speed order; container noise |
hierarchy_parentMetaClassChurn |
0.00217 ± 0.00025 | 0.00929 ± 0.00066 | ≈4.3× | H4: parent change still invalidates child sites |
Structural fingerprint on HEAD only:
baseline ≈ 4.01 ops/ms
after burst ≈ 4.35 ops/ms (≥ baseline: no residual deopt)
unrelated churn ≈ 0.35 ops/ms (~11× slower than baseline: mostly MetaClass write cost)
same-type churn ≈ 0.019 ops/ms (~210× slower than baseline: write + same-class re-link)
category ≈ 0.002 ops/ms (bulk re-link tax)
On parent, unrelated and same-type both collapse to ~0.002, showing that the old global SwitchPoint made “unrelated churn” equivalent to “global deopt.” HEAD separates them by ~18× (0.352 / 0.019)—a direct fingerprint that scoping is live.
4.2 CallSiteInvalidationBench (AverageTime, ms/op — lower is better)
| Benchmark | Parent | HEAD | Speedup (P/H) | Interpretation |
|---|---|---|---|---|
baselineHotLoop |
0.262 ± 0.016 | 0.253 ± 0.027 | 1.03× | H1 |
baselineListSize |
0.242 ± 0.010 | 0.246 ± 0.024 | 0.98× | H1 (noise) |
baselineSteadyStateNoBurst |
0.630 ± 0.052 | 0.593 ± 0.053 | 1.06× | H1 |
baselineMultipleCallSites |
21.56 ± 1.39 | 19.03 ± 1.57 | 1.13× | H1 |
crossTypeInvalidationEvery10000 |
321.5 ± 36.3 | 2.689 ± 0.678 | ≈120× | H2 |
crossTypeInvalidationEvery1000 |
443.1 ± 29.3 | 2.594 ± 1.353 | ≈171× | H2 primary win |
crossTypeInvalidationEvery100 |
148.7 ± 18.8 | 8.123 ± 1.167 | ≈18× | Still large under denser churn |
listSizeWithCrossTypeInvalidation |
354.1 ± 50.6 | 1.738 ± 0.315 | ≈204× | JDK-type hot sites benefit too |
multipleCallSitesWithInvalidation |
1214 ± 181 | 25.34 ± 2.19 | ≈48× | Multi-site still much improved |
burstThenSteadyState |
159.7 ± 19.3 | 1.432 ± 0.104 | ≈112× | “Framework MC extend + request steady state” |
sameTypeInvalidationEvery1000 |
422.8 ± 20.0 | 54.46 ± 4.38 | ≈7.8× | Same-type still slow; local invalidation cheaper than global rotate |
Penalty vs baseline on HEAD:
| Scenario | Relative cost | Meaning |
|---|---|---|
| crossType @1000 | 2.59 / 0.25 ≈ 10× baselineHotLoop |
Dominated by ~100 ColdType MetaClass writes, not hot-site deopt |
| sameType @1000 | 54.5 / 0.25 ≈ 215× baselineHotLoop |
Clear re-link tax |
| burstThenSteady | 1.43 / 0.59 ≈ 2.4× steady baseline | Burst write cost dominates; steady calls recovered |
On parent, crossType@1000 and sameType@1000 are both ~430 ms/op, again confirming global invalidation collapsed cross-type into full deopt.
5. Why it is faster
Old path New path
ColdType.metaClass.foo = ... ColdType.metaClass.foo = ...
│ │
▼ ▼
invalidate global SwitchPoint retire MetaClass(ColdType) domain
│ (+ hierarchy only if policy requires)
▼ │
ALL sites: guard fails ▼
HotTarget.compute → full re-select ColdType sites only
(MH rebuild / re-guard / cold path) HotTarget.compute → stays linked
(single SwitchPoint check passes)
Cost breakdown:
- Avoided re-links — After each global invalidation, hot sites re-run
Selector, reinstall guards, and rewriteCallSites. At 100 churn events × many sites this dominates (parent’s 400+ ms/op). - Avoided JIT churn — Frequent process-wide
SwitchPoint.invalidateAllbreaks monomorphic inlining assumptions and amplifies deopt/recompile cost. - Unchanged hot-path guard count — Still one
guardWithTest, so baselines stay flat (H1). - Same-type / hierarchy still correctly charged — Prevents “fast but wrong” semantics (H4 + 107 tests).
- Category remains bulk — No second hot-path guard; category visibility stays simple and correct; cost matches the old global path (measured ~1.0×).
Same-type still improves ~8× on HEAD vs parent, not because re-link is skipped, but because the invalidation surface shrinks from “all process sites” to “this class hierarchy” and batch invalidateAll is cheaper—a secondary benefit, not the primary goal.
6. Risks and boundaries
| Topic | Assessment |
|---|---|
| Hot-path regression | Not observed (baselines 0.98–1.13×) |
| Semantic regression | 107/107 scoped suite green; Category / hierarchy / per-instance / MetaClass identity domain covered |
| Dense Category enter/leave | Still expensive (by design); Category-heavy workloads do not get faster |
| Parent EMC / hierarchy | Fan-out still paid for missing-method correctness; pure MetaClassImpl replace is exact-class |
Legacy IndyInterface.switchPoint |
Removed on this branch (vmplugin internal-by-intent); external code must use IndyInvalidation |
| Environment noise | Shared/container CPUs; conclusions use order-of-magnitude gaps (10¹–10²×), robust to ±10% noise |
| Not covered here | Concurrent multi-thread churn throughput, full Grails end-to-end, JDK 17/21 matrix (recommended for CI) |
7. Hypothesis scorecard
| Hypothesis | Result | Evidence |
|---|---|---|
| H1 baseline parity | Holds | Scoped baseline 1.07×; CallSite baselines 0.98–1.13× |
| H2 cross-type speedup | Holds | ≈151× thrpt; ≈171× avgt (@1000) |
| H3 post-unrelated-burst full speed | Holds | afterUnrelatedBurst ≥ baseline on HEAD |
| H4 necessary invalidation still paid | Holds | same-type / hierarchy / category ≪ baseline |
| H5 correctness | Holds | 107/107 tests |
8. Conclusions and recommendations
Verdict
637bdaf68a (GROOVY-12191) keeps a single-guard monomorphic hot path while scoping indy MOP invalidation to per-MetaClass domains (+ hierarchy fan-out where MOP visibility requires it). For unrelated-type MetaClass churn with hot monomorphic call sites—a Grails/framework-shaped load—measured improvement is on the order of 10¹–10²×. Steady-state undisturbed paths show no material regression. Correctness costs for Category and same-type changes remain visible and intentional.
Performance verification: PASS.
Recommendations
- Keep
ScopedInvalidationBenchand the cross-type / burst rows ofCallSiteInvalidationBenchin regular JMH regression so a return to a global SwitchPoint cannot land silently. - In release notes, state clearly that Category and unscoped MetaClass events may still bulk-invalidate; the win is type-/MetaClass-local MetaClass change.
- Optionally re-run CallSite
crossType@1000and baselines on JDK 17/21 to confirm cross-LTS consistency. - Document that MOP guards are resolved via
IndyInvalidation(MetaClass identity map + ClassInfo pending), not a process-wideIndyInterface.switchPoint.
Appendix A — Reproduction commands
# Correctness (HEAD @ 637bdaf68a)
./gradlew :test --tests Groovy12191 \
--tests org.apache.groovy.runtime.indy.IndyInvalidationTest \
--tests org.apache.groovy.runtime.indy.SwitchPointInvalidatorTest \
--tests org.apache.groovy.runtime.indy.ClassHierarchyIndexTest \
--tests org.codehaus.groovy.vmplugin.v8.IndyScopedSwitchPointTest --rerun-tasks
# Performance (HEAD)
./gradlew :perf:jmh -PbenchInclude=ScopedInvalidation -PjmhResultFormat=JSON
./gradlew :perf:jmh -PbenchInclude=CallSiteInvalidation -PjmhResultFormat=JSON
# Performance (parent @ 0ca665dad0; copy ScopedInvalidationBench sources for fair A/B)
# Run the same two :perf:jmh invocations in a clean worktree at the parent commit.Appendix B — Commits between baseline and HEAD
637bdaf68a GROOVY-12191: Attach indy MOP SwitchPoints to MetaClass instances
3558689a06 GROOVY-12191: Document missing-method hierarchy and add blackdrag scenario tests
0303da79dc GROOVY-12191: MetaClass-aware SwitchPoint fan-out and PIC sentinel rename
1a31f276ca GROOVY-12191: Address blackdrag review on scoped SwitchPoint invalidation
f48ab19256 GROOVY-12191: Address PR review and add index hierarchy for SwitchPoint fan-out
ae8e49741e GROOVY-12191: Scope indy SwitchPoint invalidation
Appendix C — Key files (HEAD)
org.apache.groovy.runtime.indy.IndyInvalidation— MetaClass identity domains + policyorg.apache.groovy.runtime.indy.SwitchPointInvalidator,ClassHierarchyIndexClassInfo— pending domain + version coordinationIndyInterface/Selector/ColdReflectiveMethodHandleWrapper/IndyCompoundAssign— link-time guards- Tests:
Groovy12191,IndyInvalidationTest,SwitchPointInvalidatorTest,ClassHierarchyIndexTest,IndyScopedSwitchPointTest - Benchmarks:
ScopedInvalidationBench,CallSiteInvalidationBench
Report generated from local JMH A/B measurements on 2026-07-31 (hostname hera, Corretto 25.0.2). Raw JSON and run logs: /tmp/groovy-12191-perf-637bdaf/.
Does this single MissingMethodException case really justify the SwitchPoint usage? Are we even caching in that case? If we are then caching that would be new in this pull request I assume. And in that case I would question that decision. [...]
not quite, it is: missing-method hierarchy + EMC + MissingMethodException + Method exists later. That is a lot smaller. How often does this appear in a typical code base in for example Grails? @paulk-asert maybe you can help here in answering that? I simply think that the hierarchy is overkill for this and a simple uncached path through the metaclass would be sufficient. If I oversee something here, I would love for you to correct me. Switchpoints add overhead too. |




https://issues.apache.org/jira/browse/GROOVY-12191