Reduce CI contention with stale-run cancellation and shared Apple builds - #5928
Conversation
|
Developer Guide build artifacts are available for download from this workflow run:
Developer Guide quality checks: |
|
Compared 12 screenshots: 12 matched. |
Native fidelity (Android, Material 3)54 pairs compared -- median 95.6%, worst 91.3% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Cloudflare Preview
|
|
Compared 193 screenshots: 193 matched. |
|
Codex Review: Didn't find any major issues. Already looking forward to the next diff. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
|
|
Compared 157 screenshots: 157 matched. Native Android coverage
✅ Native Android screenshot tests passed. Native Android coverage
Benchmark ResultsDetailed Performance Metrics
|
✅ Continuous Quality ReportTest & Coverage
Static Analysis
Generated automatically by the PR CI workflow. |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows x64Runner CPU: AMD64 Family 25 Model 1 Stepping 1, AuthenticAMD (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
Native fidelity (ios-27-metal)68 pairs compared -- median 94.6%, worst 74.5% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux x64Runner CPU: AMD EPYC 7763 64-Core Processor (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux arm64Runner CPU: Neoverse-N2 (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows arm64Runner CPU: ARMv8 (64-bit) Family 8 Model D49 Revision 0, MICROSOFT CORPORATION (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
✅ ByteCodeTranslator Quality ReportTest & Coverage
Benchmark Results
Static Analysis
Generated automatically by the PR CI workflow. |
Native fidelity (iOS Modern, Metal)68 pairs compared -- median 95.0%, worst 83.5% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
Compared 166 screenshots: 166 matched. Benchmark Results
Detailed Performance Metrics
ParparVM vs HotSpot (JDK 25): macOS arm64Runner CPU: Apple M1 (Virtual) (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
Compared 223 screenshots: 223 matched. |
|
Compared 154 screenshots: 154 matched. Benchmark Results
Detailed Performance Metrics
|
|
Compared 150 screenshots: 150 matched. |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
@codex review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
@codex review |
|
Codex Review: Didn't find any major issues. 🎉 Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |




















































































































































































































































































































































































Mac runner queueing dominates turnaround: sampled September 30 runs spent 147 minutes waiting for a 43-second accessibility check, 118 minutes waiting for 43 minutes of Catalyst work, and 255 minutes waiting along an iOS run's 59-minute final dependency chain. Superseded revisions, repeated port preparation, and cn1lib matrix fan-out increase that contention.
This change cancels whole obsolete validation runs and reduces Mac allocations while retaining the performance gate and native coverage:
always()conditions with!cancelled()so failed runs retain diagnostics but obsolete runs can stop.pull_request_targetcleanup on synchronization and closure. It cancels obsolete PR checks even when path filters no longer trigger a replacement, rechecks current PR/run state, distinguishes forks, and never targets push, release, manual, or scheduled runs.Validation:
The ten cn1lib allocations and three duplicate iOS producers are configuration reductions when their full suites are selected, not measured wall-time savings. Hosted exact-head execution and post-merge timings remain to be observed. The cleanup workflow becomes active from the trusted default branch after merge.