Skip to content

Performance gate: per-PR baseline overlays, so baselines stop conflicting - #5930

Merged
shai-almog merged 12 commits into
masterfrom
perf-gate-overlays
Oct 2, 2026
Merged

shai-almog merged 12 commits into
masterfrom
perf-gate-overlays

Conversation

@shai-almog

Copy link
Copy Markdown
Collaborator

Why

Every performance-gate baseline lived in one vm/selfhost/perf-baseline.json. The gate fails any PR that lands on an uncalibrated runner CPU, so unrelated feature branches carried calibration rows, and branches conflicted on that file at every merge. Eleven open branches carried edits to it, and two branches calibrating the same CPU conflicted by construction.

What changes

  • Layout. vm/selfhost/perf-baseline/ holds policy.json, base/<platform>@<cpu>.json and pr/<PR number>.json. The migration from the single file is lossless; the gate reads identical values.
  • PRs write only their own overlay. calibrate-perf-baseline.py --pr N writes pr/N.json, so a baseline change is still in the PR's diff, but no other branch can touch that file.
    • calibrate is a row for a new CPU model. Overlays calibrating the same CPU are combined, and one already folded into base/ wins.
    • rebaseline is a deliberate move, better or worse, and requires a reason. It records the value it replaces. If another merged change moved that row first, the gate names both PRs instead of producing a JSON conflict.
  • Nightly fold (.github/workflows/perf-baseline.yml). The fold job moves merged overlays into base/ and refuses to commit unless the resolved baseline is unchanged. The check job runs on every PR, with no paths filter. It rejects edits to base/ or to another PR's overlay. This PR is exempt as the migration, since base/ does not exist at its base.
  • Improvements fail the gate until they are rebaselined, so a later change cannot silently give an improvement back. Calibrated tolerances now learn the spread in both directions.
  • Port Status gets a ParparVM vs JDK 25 table with time and RAM ratios per OS/architecture. Where a platform has several CPU models it shows the median and the range. The table is rendered from these baselines by scripts/website/build.sh and never committed, so it cannot go stale or conflict. The absolute-time table stays for ports with no JDK arm.
  • Migrating an open branch: perf_baseline.py import-legacy --pr N --ref origin/<branch> converts that branch's own edits to the old file. It reads the merge base, so master's later changes are not mistaken for the branch's.

Verification

  • 51 unit tests in vm/selfhost/test_perf_gate.py. Probes that removed the stale-from check, the improvement failure and the two-sided spread each failed a test.
  • Local Hugo build plus validate_port_status.mjs against the real baselines: 13 workloads by 5 platforms.
  • actionlint on the touched workflows, and the control-character gate.
  • Replayed 22 recent CI perf-results.json files against the new rules. Four would fail as improvements, all on Windows x64 amd64-family-25-model-1:
    • One is noise on master code: hello time at -15.3% against a 15% tolerance learned from upward spread only. Expect occasional failures like this until such rows are recalibrated with --all.
    • Three are branch-specific changes.

Not exercised before merge: the gate's full CI run, the fold job and the website workflow.

🤖 Generated with Claude Code

shai-almog and others added 2 commits October 1, 2026 19:55
…ting

Every baseline sat in one vm/selfhost/perf-baseline.json. Any branch that met a
new runner CPU or moved a benchmark edited it, unrelated rows sat within git's
context of each other, and two branches calibrating the same CPU conflicted by
construction -- eleven open branches carried edits to it.

vm/selfhost/perf-baseline/ now holds policy.json, base/<platform@cpu>.json and
pr/<PR number>.json. A pull request writes only its own overlay, so the change
stays in its diff and no other branch can touch the file. Calibrations of the
same CPU combine; a rebaseline records the value it replaces, so two changes
moving one benchmark are reported by name instead of as a JSON conflict. A
nightly fold moves merged overlays into base/ and refuses to commit unless the
resolved baseline is unchanged. The migration is lossless (asserted).

An improvement past tolerance now fails the gate until it is rebaselined, and
calibration tolerances learn the spread in both directions. The Port Status page
gains a ParparVM vs JDK 25 table rendered from these baselines at site-build
time; the absolute-time table stays for ports with no JDK arm.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#5927 added a Windows x64 Intel Family 6 Model 173 calibration and re-measured the
Windows ARM64 d49 objectAllocation row in the old single file. base/ is regenerated
from master's file, asserted lossless; the old file stays deleted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Cloudflare Preview

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 172 screenshots: 172 matched.
Native Windows port, REAL shipping pipeline: the hellocodenameone screenshot suite rendered by a binary CROSS-COMPILED on Linux (clang-cl + xwin, WebView2 linked) and RUN on a Windows x64 runner. Compared against the in-repo baseline in scripts/windows/screenshots.

Benchmark Results

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 57ms / native 6ms = 9.5x speedup
SIMD float-mul (64K x300) java 56ms / native 4ms = 14.0x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 native bridge unavailable (CN1 + SIMD + image benchmarks only)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path gated to scalar (CPU autovectorizes scalar; explicit SIMD not beneficial here)
Base64 CN1 encode 72.000 ms
Base64 CN1 decode 92.000 ms
Base64 SIMD encode 113.000 ms
Base64 encode ratio (SIMD/CN1) 1.569x (56.9% slower)
Base64 SIMD decode 98.000 ms
Base64 decode ratio (SIMD/CN1) 1.065x (6.5% slower)
Image encode benchmark iterations 100
Image createMask (SIMD off) 9.000 ms
Image createMask (SIMD on) 4.000 ms
Image createMask ratio (SIMD on/off) 0.444x (55.6% faster)
Image applyMask (SIMD off) 22.000 ms
Image applyMask (SIMD on) 26.000 ms
Image applyMask ratio (SIMD on/off) 1.182x (18.2% slower)
Image modifyAlpha (SIMD off) 29.000 ms
Image modifyAlpha (SIMD on) 19.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.655x (34.5% faster)
Image modifyAlpha removeColor (SIMD off) 34.000 ms
Image modifyAlpha removeColor (SIMD on) 19.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.559x (44.1% faster)

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

✅ Continuous Quality Report

Test & Coverage

Static Analysis

  • SpotBugs [Report archive]
    • ✅ ByteCodeTranslator: 0 findings (no issues)
    • ✅ android: 0 findings (no issues)
    • ✅ backend: 0 findings (no issues)
    • ✅ build-engine: 0 findings (no issues)
    • ✅ build-hint-catalog: 0 findings (no issues)
    • ✅ build-hint-tools: 0 findings (no issues)
    • ✅ codenameone-gradle-plugin: 0 findings (no issues)
    • ✅ codenameone-maven-plugin: 0 findings (no issues)
    • ✅ core-unittests: 0 findings (no issues)
    • ✅ ios: 0 findings (no issues)
    • ✅ project-model: 0 findings (no issues)
  • ✅ PMD: 0 findings (no issues) [Report archive]
  • ✅ Checkstyle: 0 findings (no issues) [Report archive]

Generated automatically by the PR CI workflow.

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 193 screenshots: 193 matched.
✅ JavaScript-port screenshot tests passed.

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 172 screenshots: 172 matched.
Native Windows port (x64 / Intel-AMD): full hellocodenameone screenshot suite rendered offscreen with Direct2D/DirectWrite, plus the real benchmarks (base64 native/CN1/SIMD, image createMask/applyMask/modifyAlpha/PNG/JPEG, SSE2 SIMD kernels). Compared against the in-repo baseline in scripts/windows/screenshots.

Benchmark Results

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 62ms / native 5ms = 12.4x speedup
SIMD float-mul (64K x300) java 57ms / native 4ms = 14.2x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 native bridge unavailable (CN1 + SIMD + image benchmarks only)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path gated to scalar (CPU autovectorizes scalar; explicit SIMD not beneficial here)
Base64 CN1 encode 73.000 ms
Base64 CN1 decode 88.000 ms
Base64 SIMD encode 103.000 ms
Base64 encode ratio (SIMD/CN1) 1.411x (41.1% slower)
Base64 SIMD decode 104.000 ms
Base64 decode ratio (SIMD/CN1) 1.182x (18.2% slower)
Image encode benchmark iterations 100
Image createMask (SIMD off) 12.000 ms
Image createMask (SIMD on) 4.000 ms
Image createMask ratio (SIMD on/off) 0.333x (66.7% faster)
Image applyMask (SIMD off) 26.000 ms
Image applyMask (SIMD on) 25.000 ms
Image applyMask ratio (SIMD on/off) 0.962x (3.8% faster)
Image modifyAlpha (SIMD off) 30.000 ms
Image modifyAlpha (SIMD on) 19.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.633x (36.7% faster)
Image modifyAlpha removeColor (SIMD off) 37.000 ms
Image modifyAlpha removeColor (SIMD on) 23.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.622x (37.8% faster)

ParparVM vs HotSpot (JDK 25): Windows x64

Runner CPU: AMD64 Family 25 Model 1 Stepping 1, AuthenticAMD (baseline windows-x64@amd64-family-25-model-1-authenticamd)

Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in vm/selfhost/perf-baseline/ fails: above it is a regression, below it an improvement that has to be rebaselined (a row whose calibration runs were noisier carries a wider tolerance), and for RAM the change must also exceed 0.05x in absolute terms. Both run unpinned on all of the runner's CPUs, with their own default thread counts.

Benchmark Cores Time RAM Status
hello (WinHelloMain) 4 1.25x (base 1.23x, +1.7%) 0.82x (base 0.80x, +2.1%) ok
translator (self) 4 0.59x (base 0.60x, -3.1%) 0.53x (base 0.54x, -3.0%) ok
intArithmetic 4 1.10x (base 1.10x, -0.0%) 0.06x (base 0.06x, -0.6%) ok
longArithmetic 4 1.08x (base 1.08x, -0.0%) 0.05x (base 0.05x, -0.2%) ok
mathTranscendental 4 0.79x (base 0.79x, -0.3%) 0.06x (base 0.06x, -0.3%) ok
arraySequential 4 1.33x (base 1.36x, -2.2%) 0.38x (base 0.39x, -0.0%) ok
arrayRandom 4 0.94x (base 1.06x, -11.6%) 0.23x (base 0.23x, -0.2%) ok
objectAllocation 4 6.40x (base 4.93x, +29.8%) 0.42x (base 0.40x, +3.9%) ok
valueEscape 4 0.10x (base 0.10x, +0.4%) 0.04x (base 0.04x, -2.5%) ok
hashMapChurn 4 1.81x (base 1.80x, +0.8%) 0.10x (base 0.09x, +3.0%) ok
stringBuilding 4 1.14x (base 1.22x, -6.0%) 0.32x (base 0.32x, -0.3%) ok
recursion 4 1.49x (base 1.49x, -0.2%) 0.06x (base 0.06x, -1.6%) ok
quicksort 4 1.13x (base 1.13x, -0.1%) 0.12x (base 0.12x, -0.7%) ok

Result: no regression

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 172 screenshots: 172 matched.
Native Linux port (x64), GTK3/Cairo/Pango, ParparVM bytecode-to-C (no JVM): the hellocodenameone screenshot suite rendered by a native ELF built + run on the GitHub x64 runner. Baseline: scripts/linux/screenshots.

ParparVM vs HotSpot (JDK 25): Linux x64

Runner CPU: AMD EPYC 7763 64-Core Processor (baseline linux-x64@amd-epyc-7763-64-core-processor)

Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in vm/selfhost/perf-baseline/ fails: above it is a regression, below it an improvement that has to be rebaselined (a row whose calibration runs were noisier carries a wider tolerance), and for RAM the change must also exceed 0.05x in absolute terms. Both run unpinned on all of the runner's CPUs, with their own default thread counts.

Benchmark Cores Time RAM Status
hello (LinuxHelloMain) 4 0.82x (base 0.85x, -2.6%) 0.86x (base 0.84x, +2.3%) ok
translator (self) 4 0.47x (base 0.48x, -2.5%) 0.52x (base 0.52x, -0.5%) ok
intArithmetic 4 1.10x (base 1.12x, -1.2%) 0.08x (base 0.06x, +34.5%) ok
longArithmetic 4 1.08x (base 1.08x, +0.0%) 0.07x (base 0.05x, +35.9%) ok
mathTranscendental 4 1.09x (base 1.08x, +0.4%) 0.08x (base 0.07x, +2.2%) ok
arraySequential 4 2.38x (base 2.00x, +18.8%) 0.36x (base 0.36x, -0.5%) ok
arrayRandom 4 0.95x (base 0.91x, +4.9%) 0.21x (base 0.21x, -0.2%) ok
objectAllocation 4 5.10x (base 5.01x, +1.8%) 0.32x (base 0.37x, -14.5%) ok
valueEscape 4 0.10x (base 0.10x, -0.1%) 0.06x (base 0.04x, +25.6%) ok
hashMapChurn 4 1.44x (base 1.48x, -2.9%) 0.09x (base 0.09x, +0.2%) ok
stringBuilding 4 1.42x (base 1.39x, +2.4%) 0.26x (base 0.26x, +0.3%) ok
recursion 4 1.19x (base 1.25x, -5.4%) 0.06x (base 0.08x, -26.3%) ok
quicksort 4 1.11x (base 1.07x, +3.0%) 0.11x (base 0.11x, -2.2%) ok

Result: no regression

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 172 screenshots: 172 matched.
Native Linux port (arm64), GTK3/Cairo/Pango, ParparVM bytecode-to-C (no JVM): the hellocodenameone screenshot suite rendered by a native ELF built + run on the GitHub arm64 runner. Baseline: scripts/linux/screenshots-arm.

ParparVM vs HotSpot (JDK 25): Linux arm64

Runner CPU: Neoverse-N2 (baseline linux-arm64@neoverse-n2)

Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in vm/selfhost/perf-baseline/ fails: above it is a regression, below it an improvement that has to be rebaselined (a row whose calibration runs were noisier carries a wider tolerance), and for RAM the change must also exceed 0.05x in absolute terms. Both run unpinned on all of the runner's CPUs, with their own default thread counts.

Benchmark Cores Time RAM Status
hello (LinuxHelloMain) 4 0.92x (base 0.93x, -0.7%) 0.87x (base 0.87x, +0.2%) ok
translator (self) 4 0.61x (base 0.62x, -3.1%) 0.54x (base 0.54x, -0.5%) ok
intArithmetic 4 1.04x (base 1.04x, -0.0%) 0.03x (base 0.03x, -1.2%) ok
longArithmetic 4 0.79x (base 0.79x, -0.2%) 0.03x (base 0.03x, +1.1%) ok
mathTranscendental 4 1.10x (base 1.10x, +0.1%) 0.03x (base 0.03x, +0.0%) ok
arraySequential 4 0.35x (base 0.36x, -4.0%) 0.37x (base 0.37x, +0.1%) ok
arrayRandom 4 0.94x (base 0.94x, +0.1%) 0.20x (base 0.20x, +0.4%) ok
objectAllocation 4 3.18x (base 3.04x, +4.4%) 0.24x (base 0.24x, -1.5%) ok
valueEscape 4 0.51x (base 0.51x, -0.1%) 0.03x (base 0.03x, +1.7%) ok
hashMapChurn 4 0.85x (base 0.84x, +1.1%) 0.12x (base 0.12x, -4.9%) ok
stringBuilding 4 1.31x (base 1.36x, -3.5%) 0.26x (base 0.26x, -0.2%) ok
recursion 4 1.42x (base 1.43x, -1.1%) 0.03x (base 0.04x, -1.0%) ok
quicksort 4 1.00x (base 0.99x, +0.6%) 0.09x (base 0.09x, +0.1%) ok

Result: no regression

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 172 screenshots: 172 matched.
Native Windows port (arm64 / Apple Silicon - Arm): full hellocodenameone screenshot suite rendered offscreen with Direct2D/DirectWrite, plus the real benchmarks (base64 native/CN1/SIMD, image createMask/applyMask/modifyAlpha/PNG/JPEG, NEON SIMD kernels). Compared against the in-repo baseline in scripts/windows/screenshots.

Benchmark Results

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 49ms / native 4ms = 12.2x speedup
SIMD float-mul (64K x300) java 48ms / native 3ms = 16.0x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 native bridge unavailable (CN1 + SIMD + image benchmarks only)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path gated to scalar (CPU autovectorizes scalar; explicit SIMD not beneficial here)
Base64 CN1 encode 60.000 ms
Base64 CN1 decode 70.000 ms
Base64 SIMD encode 77.000 ms
Base64 encode ratio (SIMD/CN1) 1.283x (28.3% slower)
Base64 SIMD decode 73.000 ms
Base64 decode ratio (SIMD/CN1) 1.043x (4.3% slower)
Image encode benchmark iterations 100
Image createMask (SIMD off) 5.000 ms
Image createMask (SIMD on) 1.000 ms
Image createMask ratio (SIMD on/off) 0.200x (80.0% faster)
Image applyMask (SIMD off) 17.000 ms
Image applyMask (SIMD on) 17.000 ms
Image applyMask ratio (SIMD on/off) 1.000x (0.0% slower)
Image modifyAlpha (SIMD off) 13.000 ms
Image modifyAlpha (SIMD on) 10.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.769x (23.1% faster)
Image modifyAlpha removeColor (SIMD off) 17.000 ms
Image modifyAlpha removeColor (SIMD on) 9.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.529x (47.1% faster)

ParparVM vs HotSpot (JDK 25): Windows arm64

Runner CPU: ARMv8 (64-bit) Family 8 Model D49 Revision 0, MICROSOFT CORPORATION (baseline windows-arm64@armv8-64-bit-family-8-model-d49-microsoft-corporation)

Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in vm/selfhost/perf-baseline/ fails: above it is a regression, below it an improvement that has to be rebaselined (a row whose calibration runs were noisier carries a wider tolerance), and for RAM the change must also exceed 0.05x in absolute terms. Both run unpinned on all of the runner's CPUs, with their own default thread counts.

Benchmark Cores Time RAM Status
hello (WinHelloMain) 4 1.08x (base 1.11x, -2.2%) 0.92x (base 0.90x, +1.7%) ok
translator (self) 4 0.76x (base 0.75x, +1.4%) 0.54x (base 0.54x, +0.9%) ok
intArithmetic 4 1.04x (base 1.04x, +0.0%) 0.06x (base 0.06x, +0.1%) ok
longArithmetic 4 0.79x (base 0.80x, -0.0%) 0.06x (base 0.06x, +0.5%) ok
mathTranscendental 4 0.72x (base 0.72x, -0.0%) 0.06x (base 0.06x, +0.4%) ok
arraySequential 4 0.38x (base 0.39x, -2.8%) 0.39x (base 0.39x, -0.1%) ok
arrayRandom 4 0.94x (base 0.94x, -0.0%) 0.23x (base 0.23x, +0.1%) ok
objectAllocation 4 2.69x (base 2.81x, -4.1%) 0.49x (base 0.45x, +7.0%) ok
valueEscape 4 0.76x (base 0.76x, +0.0%) 0.05x (base 0.05x, +0.5%) ok
hashMapChurn 4 1.00x (base 0.98x, +1.5%) 0.12x (base 0.12x, -3.3%) ok
stringBuilding 4 1.44x (base 1.43x, +0.7%) 0.27x (base 0.27x, -0.1%) ok
recursion 4 1.46x (base 1.46x, +0.0%) 0.06x (base 0.07x, -0.7%) ok
quicksort 4 1.01x (base 1.00x, +0.5%) 0.12x (base 0.12x, -0.2%) ok

Result: no regression

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

✅ ByteCodeTranslator Quality Report

Test & Coverage

  • ✅ Tests: 740 total, 0 failed, 57 skipped

Benchmark Results

  • Execution Time: 24290 ms

  • Hotspots (Top 20 sampled methods):

    • 7.04% java.util.ArrayList.indexOf (149 samples)
    • 6.19% com.codename1.tools.translator.IteratorEscape.ctorOnlyStoresParamsIntoThis (131 samples)
    • 5.20% java.lang.StringBuilder.append (110 samples)
    • 4.72% com.codename1.tools.translator.BytecodeMethod.equals (100 samples)
    • 3.16% com.codename1.tools.translator.IteratorEscape.mangle (67 samples)
    • 3.02% java.lang.System.identityHashCode (64 samples)
    • 2.74% com.codename1.tools.translator.JavascriptReachability.enqueueResolved (58 samples)
    • 2.65% java.lang.String.equals (56 samples)
    • 2.17% java.util.HashMap.hash (46 samples)
    • 1.70% org.objectweb.asm.tree.analysis.SourceInterpreter.merge (36 samples)
    • 1.61% org.objectweb.asm.tree.analysis.Frame.merge (34 samples)
    • 1.56% com.codename1.tools.translator.IteratorEscape.walk (33 samples)
    • 1.46% com.codename1.tools.translator.ByteCodeClass.hasDeclaredMethod (31 samples)
    • 1.18% java.io.FileOutputStream.open0 (25 samples)
    • 1.18% com.codename1.tools.translator.bytecodes.Invoke.resolveDirectTarget (25 samples)
    • 1.18% com.codename1.tools.translator.BytecodeMethod.optimize (25 samples)
    • 1.09% com.codename1.tools.translator.bytecodes.Invoke.findMethodUp (23 samples)
    • 1.04% com.codename1.tools.translator.BytecodeMethod.appendMethodC (22 samples)
    • 0.99% com.codename1.tools.translator.ByteCodeMethodArg.appendCSig (21 samples)
    • 0.94% com.codename1.tools.translator.BytecodeMethod.appendMethodSignatureSuffixFromDesc (20 samples)
  • ⚠️ Coverage report not generated.

Static Analysis

  • ✅ SpotBugs: no findings (report was not generated by the build).
  • ⚠️ PMD report not generated.
  • ⚠️ Checkstyle report not generated.

Generated automatically by the PR CI workflow.

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 223 screenshots: 223 matched.
✅ Native Apple Watch (watchOS, Core Graphics) screenshot tests passed.

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 155 screenshots: 155 matched.
✅ Native iOS Metal screenshot tests passed.

Benchmark Results

  • VM Translation Time: 0 seconds
  • Compilation Time: 1520 seconds

Build and Run Timing

Metric Duration
Simulator Boot 1000 ms
Simulator Boot (Run) 1000 ms
App Install 14000 ms
App Launch 6000 ms
Test Execution 368000 ms

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 80ms / native 4ms = 20.0x speedup
SIMD float-mul (64K x300) java 92ms / native 4ms = 23.0x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path active (NEON-accelerated)
Base64 CN1 encode 55.000 ms
Base64 CN1 decode 57.000 ms
Base64 native encode 1035.000 ms
Base64 encode ratio (CN1/native) 0.053x (94.7% faster)
Base64 native decode 563.000 ms
Base64 decode ratio (CN1/native) 0.101x (89.9% faster)
Base64 SIMD encode 57.000 ms
Base64 encode ratio (SIMD/CN1) 1.036x (3.6% slower)
Base64 SIMD decode 50.000 ms
Base64 decode ratio (SIMD/CN1) 0.877x (12.3% faster)
Base64 encode ratio (SIMD/native) 0.055x (94.5% faster)
Base64 decode ratio (SIMD/native) 0.089x (91.1% faster)
Image encode benchmark iterations 100
Image createMask (SIMD off) 7.000 ms
Image createMask (SIMD on) 1.000 ms
Image createMask ratio (SIMD on/off) 0.143x (85.7% faster)
Image applyMask (SIMD off) 57.000 ms
Image applyMask (SIMD on) 52.000 ms
Image applyMask ratio (SIMD on/off) 0.912x (8.8% faster)
Image modifyAlpha (SIMD off) 80.000 ms
Image modifyAlpha (SIMD on) 51.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.638x (36.3% faster)
Image modifyAlpha removeColor (SIMD off) 176.000 ms
Image modifyAlpha removeColor (SIMD on) 23.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.131x (86.9% faster)

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 154 screenshots: 154 matched.
✅ Native Mac screenshot tests passed.

Benchmark Results

  • VM Translation Time: 0 seconds
  • Compilation Time: 381 seconds

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 67ms / native 3ms = 22.3x speedup
SIMD float-mul (64K x300) java 69ms / native 3ms = 23.0x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path active (NEON-accelerated)
Base64 CN1 encode 54.000 ms
Base64 CN1 decode 81.000 ms
Base64 native encode 764.000 ms
Base64 encode ratio (CN1/native) 0.071x (92.9% faster)
Base64 native decode 357.000 ms
Base64 decode ratio (CN1/native) 0.227x (77.3% faster)
Base64 SIMD encode 80.000 ms
Base64 encode ratio (SIMD/CN1) 1.481x (48.1% slower)
Base64 SIMD decode 51.000 ms
Base64 decode ratio (SIMD/CN1) 0.630x (37.0% faster)
Base64 encode ratio (SIMD/native) 0.105x (89.5% faster)
Base64 decode ratio (SIMD/native) 0.143x (85.7% faster)
Image encode benchmark iterations 100
Image createMask (SIMD off) 11.000 ms
Image createMask (SIMD on) 4.000 ms
Image createMask ratio (SIMD on/off) 0.364x (63.6% faster)
Image applyMask (SIMD off) 77.000 ms
Image applyMask (SIMD on) 44.000 ms
Image applyMask ratio (SIMD on/off) 0.571x (42.9% faster)
Image modifyAlpha (SIMD off) 37.000 ms
Image modifyAlpha (SIMD on) 31.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.838x (16.2% faster)
Image modifyAlpha removeColor (SIMD off) 70.000 ms
Image modifyAlpha removeColor (SIMD on) 67.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.957x (4.3% faster)

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 166 screenshots: 166 matched.
✅ Native Mac screenshot tests passed.

Benchmark Results

  • VM Translation Time: 0 seconds
  • Compilation Time: 322 seconds

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 60ms / native 3ms = 20.0x speedup
SIMD float-mul (64K x300) java 51ms / native 3ms = 17.0x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 native bridge unavailable (CN1 + SIMD + image benchmarks only)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path active (NEON-accelerated)
Base64 CN1 encode 46.000 ms
Base64 CN1 decode 53.000 ms
Image encode benchmark iterations 100
Image createMask (SIMD off) 8.000 ms
Image createMask (SIMD on) 5.000 ms
Image createMask ratio (SIMD on/off) 0.625x (37.5% faster)
Image applyMask (SIMD off) 56.000 ms
Image applyMask (SIMD on) 51.000 ms
Image applyMask ratio (SIMD on/off) 0.911x (8.9% faster)
Image modifyAlpha (SIMD off) 59.000 ms
Image modifyAlpha (SIMD on) 44.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.746x (25.4% faster)
Image modifyAlpha removeColor (SIMD off) 75.000 ms
Image modifyAlpha removeColor (SIMD on) 40.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 0.533x (46.7% faster)

ParparVM vs HotSpot (JDK 25): macOS arm64

Runner CPU: Apple M1 (Virtual) (baseline macos-arm64)

Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A ratio more than 15% (time) / 15% (RAM) away from its baseline in vm/selfhost/perf-baseline/ fails: above it is a regression, below it an improvement that has to be rebaselined (a row whose calibration runs were noisier carries a wider tolerance), and for RAM the change must also exceed 0.05x in absolute terms. Both run unpinned on all of the runner's CPUs, with their own default thread counts.

Benchmark Cores Time RAM Status
hello (HelloCodenameOne) 3 0.83x (base 0.96x, -13.2%) 1.39x (base 1.34x, +4.0%) ok
translator (self) 3 0.54x (base 0.54x, +0.0%) 0.79x (base 0.74x, +6.4%) ok
intArithmetic 3 1.09x (base 1.03x, +5.3%) 0.12x (base 0.12x, +0.3%) ok
longArithmetic 3 1.07x (base 1.03x, +4.3%) 0.11x (base 0.11x, +0.6%) ok
mathTranscendental 3 0.96x (base 1.01x, -4.9%) 0.12x (base 0.12x, -0.4%) ok
arraySequential 3 0.43x (base 0.42x, +1.6%) 0.53x (base 0.53x, -0.3%) ok
arrayRandom 3 0.96x (base 0.99x, -2.7%) 0.32x (base 0.32x, -0.4%) ok
objectAllocation 3 4.07x (base 3.70x, +10.0%) 0.42x (base 0.48x, -12.6%) ok
valueEscape 3 0.51x (base 0.51x, +0.3%) 0.08x (base 0.09x, -7.4%) ok
hashMapChurn 3 1.00x (base 1.18x, -15.3%) 0.07x (base 0.07x, -0.4%) ok
stringBuilding 3 0.85x (base 0.84x, +1.1%) 0.46x (base 0.46x, -0.3%) ok
recursion 3 1.22x (base 1.24x, -1.1%) 0.12x (base 0.12x, -0.5%) ok
quicksort 3 1.01x (base 1.01x, -0.0%) 0.11x (base 0.11x, -0.1%) ok

Result: no regression

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 150 screenshots: 150 matched.
✅ Native Apple TV (tvOS, Metal) screenshot tests passed.

@shai-almog

shai-almog commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Compared 155 screenshots: 155 matched.
✅ Native iOS Metal screenshot tests passed.

Benchmark Results

  • VM Translation Time: 0 seconds
  • Compilation Time: 1671 seconds

Build and Run Timing

Metric Duration
Simulator Boot 94000 ms
Simulator Boot (Run) 0 ms
App Install 25000 ms
App Launch 1000 ms
Test Execution 437000 ms

Detailed Performance Metrics

Metric Duration
SIMD kernel backend SSE2 (x64) / NEON (arm64) native kernels
SIMD int-add (64K x300) java 63ms / native 3ms = 21.0x speedup
SIMD float-mul (64K x300) java 59ms / native 2ms = 29.5x speedup
SIMD kernel correctness PASS (native result == scalar reference)
Base64 payload size 8192 bytes
Base64 benchmark iterations 6000
Base64 SIMD byte path active (NEON-accelerated)
Base64 CN1 encode 66.000 ms
Base64 CN1 decode 67.000 ms
Base64 native encode 878.000 ms
Base64 encode ratio (CN1/native) 0.075x (92.5% faster)
Base64 native decode 1510.000 ms
Base64 decode ratio (CN1/native) 0.044x (95.6% faster)
Base64 SIMD encode 84.000 ms
Base64 encode ratio (SIMD/CN1) 1.273x (27.3% slower)
Base64 SIMD decode 71.000 ms
Base64 decode ratio (SIMD/CN1) 1.060x (6.0% slower)
Base64 encode ratio (SIMD/native) 0.096x (90.4% faster)
Base64 decode ratio (SIMD/native) 0.047x (95.3% faster)
Image encode benchmark iterations 100
Image createMask (SIMD off) 7.000 ms
Image createMask (SIMD on) 3.000 ms
Image createMask ratio (SIMD on/off) 0.429x (57.1% faster)
Image applyMask (SIMD off) 383.000 ms
Image applyMask (SIMD on) 62.000 ms
Image applyMask ratio (SIMD on/off) 0.162x (83.8% faster)
Image modifyAlpha (SIMD off) 408.000 ms
Image modifyAlpha (SIMD on) 56.000 ms
Image modifyAlpha ratio (SIMD on/off) 0.137x (86.3% faster)
Image modifyAlpha removeColor (SIMD off) 113.000 ms
Image modifyAlpha removeColor (SIMD on) 575.000 ms
Image modifyAlpha removeColor ratio (SIMD on/off) 5.088x (408.8% slower)

@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T10:54:27.093405Z 9526f92 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 709d4e600c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py Outdated
Comment thread vm/selfhost/perf_baseline.py
Comment thread vm/selfhost/README.md Outdated
Comment thread vm/selfhost/perf_baseline.py Outdated
Comment thread vm/selfhost/perf_baseline.py Outdated
Comment thread vm/selfhost/perf_baseline.py
… stricter check

- A pull request may edit or delete an overlay already on its base branch: two
  merged rebaselines of one row otherwise leave master unresolvable with no
  check-compliant repair.
- Combined calibrations get a tolerance covering every contributing row's own
  band around the combined median, so none of the runs they came from fails.
- import-legacy merges a row field by field against the merge base, keeping
  master's changes and refusing fields both sides changed differently.
- policy.json values are type- and range-checked; a change to the retired
  perf-baseline.json fails the check with migration instructions.
- README shows the import command with the --ref flag it requires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0f3252a1b4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py
Comment thread vm/selfhost/perf_baseline.py Outdated
…re stale check

CI on this branch failed macOS and Windows x64 (AMD Family 25 Model 1) as
IMPROVED on stringBuilding RAM. Not an improvement: across 79 CI runs that row
reads 0.267x or 0.322x on that Windows CPU, and 0.389x-0.54x on macOS, with
unchanged VM code, and the baselines were calibrated from runs on the upper
mode. pr/5930.json recalibrates the RAM of those two rows from all their runs;
replaying the 30 runs whose VM and benchmark sources equal master's now gives
0 non-ok verdicts of 780.

- calibrate-perf-baseline.py feeds a run to the row that JUDGED it (macOS reports
  'Apple M1 (Virtual)' but is judged by the plain macos-arm64 row), and gains
  --metric to recalibrate one metric without touching the other.
- A rebaseline's 'from' records the replaced row's tolerance too, so a later
  rebaseline cannot silently restore a tolerance an earlier one changed.
- import-legacy refuses rows the branch deleted (the old --fresh), which an
  overlay cannot express, instead of reporting no change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b2d883fa6a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py
…egacy policy edits

CI failed Windows x64 (AMD Family 25 Model 1) as IMPROVED on translator time,
0.48x against 0.61x, on this branch's unchanged VM code. The row is bimodal:
rounds alternate 0.42x-0.49x and 0.59x-0.64x inside single runs, so a run's median
lands on either mode (this branch's own runs read 0.594x and 0.479x). hello time on
the same CPU is as wide (rounds 0.87x-1.57x) and had already fallen outside on a
master run. pr/5930.json recalibrates the time of both rows from all 11 of that
CPU's runs, leaving RAM alone; the 32 runs whose VM and benchmark sources equal
master's now give 0 non-ok verdicts of 832.

import-legacy now refuses a branch that changed the old file's tolerance or floor,
which an overlay cannot carry, instead of reporting no change and losing the edit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f6138aaeb7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py Outdated
Comment thread vm/selfhost/perf_baseline.py
Comment thread vm/selfhost/perf_baseline.py
Comment thread vm/selfhost/calibrate-perf-baseline.py Outdated
…tricter values

A pull request whose own rebaseline goes stale (another merged change moved the row)
was told to re-measure on top of that change, but the gate then refused to measure
and the calibrator refused to load, so no run existed to re-measure from. The gate
now judges such a run against the baseline without the pull request's own overlay
and fails afterwards, naming it; the calibrator does the same and rewrites every
stale own row from the fresh runs, refusing if the runs do not cover one.

Also: ratios and tolerances must be finite (JSON's 1e999 is infinity), a
rebaseline's 'from' values are validated like the row, and calibrations of one CPU
that combine into an out-of-range tolerance are refused, naming the overlays,
instead of resolving to a row that gates nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7fe248f5c7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/calibrate-perf-baseline.py Outdated
…folded row

When another pull request's calibration of the same CPU was folded into base/
first, this pull request's own calibration of that row is superseded; a fresh run
past the folded row was still written as a calibration, which resolve() ignored
the same way, so the gate could never pass. Such a row is now a rebaseline from
the folded row, and writing it drops the stale calibration.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 071a909e86

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py Outdated
…he fold

A pull request measured on top of a merged but not yet folded rebaseline names
that one's result as its 'from', yet resolve() refused any row two overlays
touched, so it was blocked until the nightly fold. Rebaselines of a row are now
applied as a chain, each from the value the previous one left; two from the same
value (changes that never saw each other) and one whose 'from' the row never
reaches are still refused. The order comes from the 'from' values, not from
which merged first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4ea343750c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/calibrate-perf-baseline.py
Comment thread .github/workflows/perf-baseline.yml
…l rows

baseline_key() selects no row for a runner whose CPU model cannot be decoded on
a platform with per-model rows, so the plain-platform row the calibrator used to
write there could never be read and the gate stayed uncalibrated. It now refuses
and points at perf-gate.cpu_model(). The fold's one-shot site rebuild is kept,
with a comment on why: website-docs.yml rebuilds daily on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1e4a8971da

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/calibrate-perf-baseline.py
Comment thread vm/selfhost/calibrate-perf-baseline.py
Comment thread scripts/website/build.sh
build.sh renders the Port Status JDK 25 table with perf_baseline.py summary, but
website-docs.yml's path filters did not include it, so a change to its output
shape skipped the Hugo build and validate_port_status until a scheduled
production build. The renderer is now a trigger; the baseline data is not (a row
changes numbers, not shape, and the nightly fold dispatches a rebuild).

Comments record why two recovery paths are deliberately absent: recalibrating a
rebaseline another overlay is chained on would make that one stale, and a CPU
calibrated an order of magnitude apart is a broken measurement to delete.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 79c7a685e8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/perf_baseline.py Outdated
Comment thread vm/selfhost/calibrate-perf-baseline.py Outdated
…n the table

The overlay carried ONE reason for the whole file, so once a pull request had
written a rebaseline, any row it rebaselined later passed under that earlier
row's explanation -- a regression accepted for one cause, explained by another.
The reason now belongs to each rebaseline row, like its 'from': required and
validated per row, reused only when that same row is re-measured, never
inherited by a row new to the overlay; a file-level reason is refused.
pr/5930.json gives each of its four rows its own.

The JDK 25 summary recorded one CPU list per platform while each cell's median
could cover fewer models (one calibrated for some benchmarks only), so published
numbers could shift as models were filled in with nothing to show it. Each cell
now records the models it covers, and its tooltip says how many of the
platform's.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ba36b34d61

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vm/selfhost/ci-perf-gate.sh Outdated
… moved a row

The calibrator processes the whole results file, so a run that both fills a
missing row and moves another past its tolerance writes a calibration AND a
rebaseline -- and refuses without a reason for the rebaseline. The verdict step
and the PR comment printed the reason-less command in that case; both now print
the one that works.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shai-almog

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: 9526f92808

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@shai-almog
shai-almog merged commit 2746b62 into master Oct 2, 2026
58 of 59 checks passed
@shai-almog
shai-almog deleted the perf-gate-overlays branch October 2, 2026 12:58
shai-almog added a commit that referenced this pull request Oct 2, 2026
…ilation

Master retired vm/selfhost/perf-baseline.json for per-PR overlays (#5930).
This branch's own edits to it -- the linux-x64 AMD EPYC 9V45 rows measured
with this branch's VM changes, and the Windows AMD Family 26 objectAllocation
RAM row -- are imported into perf-baseline/pr/5883.json with
perf_baseline.py import-legacy (explicit --legacy/--original: run mid-merge,
--ref resolves the merge base to the branch itself and imports nothing).

Conflicts:
- .gitignore: both generated data files.
- port-status.html: both sections, the Flutter comparison and master's
  ParparVM vs JDK 25 table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
shai-almog added a commit that referenced this pull request Oct 3, 2026
… into

A pull_request run checks out GitHub's merge of the pull request into the
current base branch, but the check diffed from pull_request.base.sha, which
can be an older base commit. Every base/ row the nightly fold wrote after it
was then charged to the pull request: this one failed on master's own fold of
#5930 (a8efcce). The merge commit's first parent is the branch it was
merged into, so HEAD^1...HEAD is exactly the pull request's change.
Reproduced on a local merge of this branch into master: the old base reports
both base/ files, HEAD^1 reports OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant