Performance gate: per-PR baseline overlays, so baselines stop conflicting - #5930
Conversation
…ting Every baseline sat in one vm/selfhost/perf-baseline.json. Any branch that met a new runner CPU or moved a benchmark edited it, unrelated rows sat within git's context of each other, and two branches calibrating the same CPU conflicted by construction -- eleven open branches carried edits to it. vm/selfhost/perf-baseline/ now holds policy.json, base/<platform@cpu>.json and pr/<PR number>.json. A pull request writes only its own overlay, so the change stays in its diff and no other branch can touch the file. Calibrations of the same CPU combine; a rebaseline records the value it replaces, so two changes moving one benchmark are reported by name instead of as a JSON conflict. A nightly fold moves merged overlays into base/ and refuses to commit unless the resolved baseline is unchanged. The migration is lossless (asserted). An improvement past tolerance now fails the gate until it is rebaselined, and calibration tolerances learn the spread in both directions. The Port Status page gains a ParparVM vs JDK 25 table rendered from these baselines at site-build time; the absolute-time table stays for ports with no JDK arm. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#5927 added a Windows x64 Intel Family 6 Model 173 calibration and re-measured the Windows ARM64 d49 objectAllocation row in the old single file. base/ is regenerated from master's file, asserted lossless; the old file stays deleted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cloudflare Preview
|
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
|
✅ Continuous Quality ReportTest & Coverage
Static Analysis
Generated automatically by the PR CI workflow. |
|
Compared 193 screenshots: 193 matched. |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows x64Runner CPU: AMD64 Family 25 Model 1 Stepping 1, AuthenticAMD (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux x64Runner CPU: AMD EPYC 7763 64-Core Processor (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux arm64Runner CPU: Neoverse-N2 (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows arm64Runner CPU: ARMv8 (64-bit) Family 8 Model D49 Revision 0, MICROSOFT CORPORATION (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in
Result: no regression |
✅ ByteCodeTranslator Quality ReportTest & Coverage
Benchmark Results
Static Analysis
Generated automatically by the PR CI workflow. |
|
Compared 223 screenshots: 223 matched. |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
Compared 154 screenshots: 154 matched. Benchmark Results
Detailed Performance Metrics
|
|
Compared 166 screenshots: 166 matched. Benchmark Results
Detailed Performance Metrics
ParparVM vs HotSpot (JDK 25): macOS arm64Runner CPU: Apple M1 (Virtual) (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in
Result: no regression |
|
Compared 150 screenshots: 150 matched. |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 709d4e600c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… stricter check - A pull request may edit or delete an overlay already on its base branch: two merged rebaselines of one row otherwise leave master unresolvable with no check-compliant repair. - Combined calibrations get a tolerance covering every contributing row's own band around the combined median, so none of the runs they came from fails. - import-legacy merges a row field by field against the merge base, keeping master's changes and refusing fields both sides changed differently. - policy.json values are type- and range-checked; a change to the retired perf-baseline.json fails the check with migration instructions. - README shows the import command with the --ref flag it requires. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0f3252a1b4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…re stale check CI on this branch failed macOS and Windows x64 (AMD Family 25 Model 1) as IMPROVED on stringBuilding RAM. Not an improvement: across 79 CI runs that row reads 0.267x or 0.322x on that Windows CPU, and 0.389x-0.54x on macOS, with unchanged VM code, and the baselines were calibrated from runs on the upper mode. pr/5930.json recalibrates the RAM of those two rows from all their runs; replaying the 30 runs whose VM and benchmark sources equal master's now gives 0 non-ok verdicts of 780. - calibrate-perf-baseline.py feeds a run to the row that JUDGED it (macOS reports 'Apple M1 (Virtual)' but is judged by the plain macos-arm64 row), and gains --metric to recalibrate one metric without touching the other. - A rebaseline's 'from' records the replaced row's tolerance too, so a later rebaseline cannot silently restore a tolerance an earlier one changed. - import-legacy refuses rows the branch deleted (the old --fresh), which an overlay cannot express, instead of reporting no change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b2d883fa6a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…egacy policy edits CI failed Windows x64 (AMD Family 25 Model 1) as IMPROVED on translator time, 0.48x against 0.61x, on this branch's unchanged VM code. The row is bimodal: rounds alternate 0.42x-0.49x and 0.59x-0.64x inside single runs, so a run's median lands on either mode (this branch's own runs read 0.594x and 0.479x). hello time on the same CPU is as wide (rounds 0.87x-1.57x) and had already fallen outside on a master run. pr/5930.json recalibrates the time of both rows from all 11 of that CPU's runs, leaving RAM alone; the 32 runs whose VM and benchmark sources equal master's now give 0 non-ok verdicts of 832. import-legacy now refuses a branch that changed the old file's tolerance or floor, which an overlay cannot carry, instead of reporting no change and losing the edit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f6138aaeb7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…tricter values A pull request whose own rebaseline goes stale (another merged change moved the row) was told to re-measure on top of that change, but the gate then refused to measure and the calibrator refused to load, so no run existed to re-measure from. The gate now judges such a run against the baseline without the pull request's own overlay and fails afterwards, naming it; the calibrator does the same and rewrites every stale own row from the fresh runs, refusing if the runs do not cover one. Also: ratios and tolerances must be finite (JSON's 1e999 is infinity), a rebaseline's 'from' values are validated like the row, and calibrations of one CPU that combine into an out-of-range tolerance are refused, naming the overlays, instead of resolving to a row that gates nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7fe248f5c7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…folded row When another pull request's calibration of the same CPU was folded into base/ first, this pull request's own calibration of that row is superseded; a fresh run past the folded row was still written as a calibration, which resolve() ignored the same way, so the gate could never pass. Such a row is now a rebaseline from the folded row, and writing it drops the stale calibration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 071a909e86
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…he fold A pull request measured on top of a merged but not yet folded rebaseline names that one's result as its 'from', yet resolve() refused any row two overlays touched, so it was blocked until the nightly fold. Rebaselines of a row are now applied as a chain, each from the value the previous one left; two from the same value (changes that never saw each other) and one whose 'from' the row never reaches are still refused. The order comes from the 'from' values, not from which merged first. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4ea343750c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…l rows baseline_key() selects no row for a runner whose CPU model cannot be decoded on a platform with per-model rows, so the plain-platform row the calibrator used to write there could never be read and the gate stayed uncalibrated. It now refuses and points at perf-gate.cpu_model(). The fold's one-shot site rebuild is kept, with a comment on why: website-docs.yml rebuilds daily on its own. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1e4a8971da
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
build.sh renders the Port Status JDK 25 table with perf_baseline.py summary, but website-docs.yml's path filters did not include it, so a change to its output shape skipped the Hugo build and validate_port_status until a scheduled production build. The renderer is now a trigger; the baseline data is not (a row changes numbers, not shape, and the nightly fold dispatches a rebuild). Comments record why two recovery paths are deliberately absent: recalibrating a rebaseline another overlay is chained on would make that one stale, and a CPU calibrated an order of magnitude apart is a broken measurement to delete. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 79c7a685e8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…n the table The overlay carried ONE reason for the whole file, so once a pull request had written a rebaseline, any row it rebaselined later passed under that earlier row's explanation -- a regression accepted for one cause, explained by another. The reason now belongs to each rebaseline row, like its 'from': required and validated per row, reused only when that same row is re-measured, never inherited by a row new to the overlay; a file-level reason is refused. pr/5930.json gives each of its four rows its own. The JDK 25 summary recorded one CPU list per platform while each cell's median could cover fewer models (one calibrated for some benchmarks only), so published numbers could shift as models were filled in with nothing to show it. Each cell now records the models it covers, and its tooltip says how many of the platform's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ba36b34d61
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… moved a row The calibrator processes the whole results file, so a run that both fills a missing row and moves another past its tolerance writes a calibration AND a rebaseline -- and refuses without a reason for the rebaseline. The verdict step and the PR comment printed the reason-less command in that case; both now print the one that works. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@codex review |
|
Codex Review: Didn't find any major issues. You're on a roll. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…ilation Master retired vm/selfhost/perf-baseline.json for per-PR overlays (#5930). This branch's own edits to it -- the linux-x64 AMD EPYC 9V45 rows measured with this branch's VM changes, and the Windows AMD Family 26 objectAllocation RAM row -- are imported into perf-baseline/pr/5883.json with perf_baseline.py import-legacy (explicit --legacy/--original: run mid-merge, --ref resolves the merge base to the branch itself and imports nothing). Conflicts: - .gitignore: both generated data files. - port-status.html: both sections, the Flutter comparison and master's ParparVM vs JDK 25 table. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… into A pull_request run checks out GitHub's merge of the pull request into the current base branch, but the check diffed from pull_request.base.sha, which can be an older base commit. Every base/ row the nightly fold wrote after it was then charged to the pull request: this one failed on master's own fold of #5930 (a8efcce). The merge commit's first parent is the branch it was merged into, so HEAD^1...HEAD is exactly the pull request's change. Reproduced on a local merge of this branch into master: the old base reports both base/ files, HEAD^1 reports OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Why
Every performance-gate baseline lived in one
vm/selfhost/perf-baseline.json. The gate fails any PR that lands on an uncalibrated runner CPU, so unrelated feature branches carried calibration rows, and branches conflicted on that file at every merge. Eleven open branches carried edits to it, and two branches calibrating the same CPU conflicted by construction.What changes
vm/selfhost/perf-baseline/holdspolicy.json,base/<platform>@<cpu>.jsonandpr/<PR number>.json. The migration from the single file is lossless; the gate reads identical values.calibrate-perf-baseline.py --pr Nwritespr/N.json, so a baseline change is still in the PR's diff, but no other branch can touch that file.calibrateis a row for a new CPU model. Overlays calibrating the same CPU are combined, and one already folded intobase/wins.rebaselineis a deliberate move, better or worse, and requires areason. It records the value it replaces. If another merged change moved that row first, the gate names both PRs instead of producing a JSON conflict..github/workflows/perf-baseline.yml). The fold job moves merged overlays intobase/and refuses to commit unless the resolved baseline is unchanged. The check job runs on every PR, with no paths filter. It rejects edits tobase/or to another PR's overlay. This PR is exempt as the migration, sincebase/does not exist at its base.scripts/website/build.shand never committed, so it cannot go stale or conflict. The absolute-time table stays for ports with no JDK arm.perf_baseline.py import-legacy --pr N --ref origin/<branch>converts that branch's own edits to the old file. It reads the merge base, so master's later changes are not mistaken for the branch's.Verification
vm/selfhost/test_perf_gate.py. Probes that removed the stale-fromcheck, the improvement failure and the two-sided spread each failed a test.validate_port_status.mjsagainst the real baselines: 13 workloads by 5 platforms.perf-results.jsonfiles against the new rules. Four would fail as improvements, all on Windows x64amd64-family-25-model-1:hellotime at -15.3% against a 15% tolerance learned from upward spread only. Expect occasional failures like this until such rows are recalibrated with--all.Not exercised before merge: the gate's full CI run, the fold job and the website workflow.
🤖 Generated with Claude Code