Skip to content

Accelerate CI image materialization and dependency reuse - #74173

Open
zozo123 wants to merge 17 commits into
apache:mainfrom
zozo123:ci/ci-image-data-root-snapshot
Open

zozo123 wants to merge 17 commits into
apache:mainfrom
zozo123:ci/ci-image-data-root-snapshot

Conversation

@zozo123

@zozo123 zozo123 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR reduces repeated CI-image preparation in two independent, conservative layers:

  1. Image materialization: publish an optional Docker data-directory snapshot from the existing build-ci-images builder for consumers in the same pull-request workflow run. Consumers restore it when compatible and retain the existing image-stash/load path as fallback.
  2. Dependency installation: cache locked third-party dependency installation separately from workspace source changes. Source-only edits can reuse the dependency layer; the authoritative full workspace installation still runs afterward.

Neither optimization skips image builds or tests, and neither introduces a Bazel dependency.

Image snapshot changes

  • Create and upload the snapshot at the end of build-ci-images, after image and mount-cache publication, using the already-built daemon.
  • Remove the separate producer job, its image reload, and its stash download from the snapshot path.
  • Restrict snapshots to pull requests using the workflow revision; use ci-image-snapshot-v2-* artifacts with two-day retention.
  • Keep snapshot creation in trusted workflow commands rather than executing code from a configurable checkout.
  • Write the archive to runner temporary storage so archive creation does not read and write the AMD builder's Docker data disk simultaneously.
  • Validate checkout revision, daemon fingerprint, image ID, empty image/container inventories, and supported overlay2 storage.
  • Recover Docker after failures, remove partially restored state, and emit warnings for unsupported storage or fingerprint mismatches.

Dependency-layer changes

  • Add a BuildKit manifest stage that copies every regular workspace pyproject.toml and uv.lock before the full checkout.
  • Preinstall only the locked third-party dependencies with the existing architecture-specific UV cache when this is not an upgrade/highest-resolution build.
  • Preserve the existing full source installation, workspace package installation, packaging-tool setup, dependency-resolution fallback, and pip check path.
  • Treat optional preinstallation failure, missing lockfiles, upgrade builds, and unsupported paths as misses/fallbacks rather than changing the current installation contract.
  • Use the existing source checkout and cache mechanisms; no trusted cross-run environment artifact is introduced.

Safety and fallback

Both optimizations are optional. Missing artifacts, malformed or stale snapshot metadata, incompatible daemon state, unsupported storage, failed extraction, failed Docker recovery, failed image startup, missing lockfiles, upgrade builds, or failed dependency preinstallation fall back to the existing path. The PR does not reduce test coverage.

Measurements and limitations

The earlier separate-producer snapshot design measured consumer preparation at a median of 102 seconds versus 189 seconds for the baseline, but delayed consumers by approximately 236 seconds. The current design creates the snapshot in the builder after cache publication, removing the separate producer reload; the net workflow benefit still requires measurement.

The first current-design run created the snapshot in 74 seconds and uploaded it in 14 seconds. The latest revision writes the archive off the Docker data disk. No guaranteed cache-hit rate, minute saving, or whole-workflow speedup is claimed.

The dependency-layer miniature probe demonstrated the intended cache boundary: source-only changes reused the dependency layer, while manifest and lockfile changes invalidated it. A full Airflow image equivalence/performance run remains environment/CI dependent.

Validation

  • Combined focused regression tests: 30 passed.
  • Full scripts suite on the combined branch: 1,664 passed, 1 skipped.
  • Dependency tests cover manifest extraction with and without uv.lock, nested paths including spaces, upgrade bypass, and successful/failed optional preinstallation.
  • Snapshot tests cover creation, restore, compatibility checks, warnings, cleanup, and fallback behavior.
  • The first combined CI image attempt exposed and fixed a real Dockerfile issue: the manifest stage now explicitly uses Bash rather than Docker's default /bin/sh. CI is rerunning for head 490da23a65.
  • Full Docker equivalence and end-to-end performance measurements are not claimed from the local environment because its Docker daemon was unavailable.

Scope

This PR optimizes image materialization and dependency-layer reuse. It does not implement semantic CI-environment identity, skip tests, prune coverage, or make reusable cross-run/fork-produced environment artifacts trusted.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 5.5) and Codex (GPT-6)

Generated-by: Claude Code (Opus 5.5) and Codex (GPT-6), following the Airflow guidelines.

Comment thread .github/workflows/ci-image-build.yml Fixed
@zozo123
zozo123 force-pushed the ci/ci-image-data-root-snapshot branch from f58c30f to 80d5cb7 Compare October 3, 2026 22:39
Comment thread .github/workflows/ci-image-build.yml Fixed
@zozo123
zozo123 force-pushed the ci/ci-image-data-root-snapshot branch from 1d49fda to bf067ab Compare October 4, 2026 04:49
Comment thread .github/workflows/ci-image-build.yml Fixed

@shahar1 shahar1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Measured on this PR's own CI run: each CI-image consumer saves ~87 s of image prep, but every consumer now starts ~4 min later, because snapshot-ci-images lives inside the reusable workflow those consumers wait on. Net wall-clock per PR goes up, while runner-minutes go down. One structural change (create the snapshot inside build-ci-images) should turn this into a pure win, so I'd like that resolved before merging.

Snapshot job extends the critical path of every consumer (.github/workflows/ci-image-build.yml:407)

Numbers from this PR's run 37206844037 and a baseline PR run from the same afternoon on the same runner pool, 37210552741, both ubuntu-22.04:

baseline this PR
Build CI linux/amd64 image 3.10 finished 14:54:13 13:55:59
first consumer job started 14:54:13 (+0 s) 13:59:55 (+3 min 56 s)
consumer CI-image prep, median 189 s (stash 67 s + load 106 s, n=29) 102 s (download 59 s + unpack 31 s, n=79)

Every job in ci-amd.yml that has needs: build-ci-images waits for the whole called workflow, and snapshot-ci-images is part of it; continue-on-error: true changes the conclusion, not the wait. The job spent 178 s of its 229 s restoring and docker image load-ing the stash that build-ci-images had exported seconds earlier, then 25 s snapshotting and 15 s uploading. So the ~87 s saved per consumer is paid for with ~236 s added before any consumer can start. For a PR the end-to-end duration gets roughly 2.5 min longer; what improves is total runner time (79 consumers × 87 s ≈ 115 runner-minutes in this run). The description lists "end-to-end workflow duration" among the metrics to evaluate but does not report it.

Suggestion: create the snapshot as the last step of build-ci-images instead of in a separate job. The image is already in that daemon, so the 178 s reload disappears and the extra critical path shrinks to roughly the 40 s of create + upload. move_docker_to_mnt.sh bind-mounts /mnt/var-lib-docker onto /var/lib/docker, so DockerRootDir still reads /var/lib/docker and check_supported_daemon passes there. Keep the same gate (github.event_name == 'pull_request' && inputs.upload-image-artifact == 'true' && inputs.image-stash-ref == '') and continue-on-error on the step, and place it after "Stash cache mount": create runs docker builder prune --all, which would otherwise wipe the mount cache before it is exported. If the separate job was chosen for isolation reasons, please spell them out in the job comment. In pull_request context the token is read-only either way, and build-ci-images already runs the PR's own sources.

Smaller observations

  • scripts/ci/docker_data_root_snapshot.sh:53 — check_supported_daemon accepts only overlay2. When GitHub's runner images move to the containerd image store, every snapshot will be rejected, the producer keeps spending its minutes, and nothing surfaces it because the create step is continue-on-error. An echo "::warning::..." on the fingerprint-mismatch and unsupported-store exits would make that visible in the run summary.
  • The description still says the key is ci-image-snapshot-v1-*; the code uses ci-image-snapshot-v2-.
  • The script tests drive the real bash through shimmed docker/sudo/zstd/git and cover every bail-out path. All 23 pass locally in under a second. Nice approach.

Worth a second look from

This change touches the CI image stash/restore flow; folks with the most context here:

  • @potiuk — authored 10 of the last 25 commits on ci-image-build.yml and prepare_breeze_and_image/action.yml (the stash and image-stash-ref design)
  • @jscheffl — 2 of the last 25 commits on the same files

None of them have been notified — asking any of them for an extra pass is the maintainer's call, and optional.


This review was drafted by an AI-assisted tool and
confirmed by an Apache Airflow maintainer. The findings
below are observations, not blockers; an Apache Airflow
maintainer — a real person — will take the next look at the
PR. If you think a finding is mis-applied, please reply on
the PR and a maintainer will weigh in.

More on how Apache Airflow handles maintainer review:
Contributing guide.

Comment thread .github/workflows/ci-image-build.yml Outdated
Comment thread .github/workflows/ci-image-build.yml Outdated
Comment thread scripts/ci/docker_data_root_snapshot.sh
zozo123 and others added 12 commits October 4, 2026 18:42
Every job that prepares the CI image runs `docker image load` on the stash,
which unpacks and checksums every layer of an 8 GB image again, about two
minutes per job. Extracting a copy of the image store into a stopped daemon
gives the same image in about half a minute on the same runners. The image
stash stays as it is, for other consumers and as the fallback when a
snapshot does not fit the daemon a job runs on.
Snapshot failures must not prevent the authoritative stash from loading, and an older branch snapshot must not substitute an image for the current checkout.
Raw daemon snapshots are job handoffs, so publishing them into branch caches unnecessarily exposes later runs to arbitrary checkout refs.
The snapshot script runs checkout code and prunes Docker state. Existing cache publications must finish before that optional execution.
The optional snapshot producer executes checkout code without publishing branch caches; its artifacts are available only to consumers in the same workflow run.
Snapshot producers must not accept arbitrary checkout refs in privileged workflow contexts, and a partially failed Docker stop must still trigger restart cleanup.
Preparing the authoritative image in a second runner adds nearly three minutes before consumers can start. All branch publications must finish before the optional snapshot code prunes Docker state.
The builder can check out a configurable ref before publishing caches. Snapshot execution must read the workflow commit itself rather than trust files left in that working tree.
Executing snapshot code inside the configurable-ref builder introduces a cache-poisoning finding. A separate job preserves the security boundary; its latency cost must remain explicit when measuring the runner-time benefit.
An unavailable daemon must not be mistaken for an empty image store before destructive snapshot materialization.
Reloading the freshly built image in a second job delays every consumer by nearly four minutes. Trusted workflow commands reuse the builder daemon without executing code from the configurable checkout.
@zozo123
zozo123 force-pushed the ci/ci-image-data-root-snapshot branch from d3c3f01 to 6e8fd38 Compare October 4, 2026 18:44
The builder bind-mounts Docker from /mnt. Reading its image store and writing the compressed archive to that same disk competes for I/O; runner temp keeps the output on the root disk.
@zozo123 zozo123 changed the title Restore the CI image from a snapshot of Docker's data directory Restore CI images from Docker data-directory snapshots Oct 4, 2026
zozo123 and others added 3 commits October 5, 2026 00:01
Ordinary source edits invalidate the dependency installation layer because it follows the full checkout copy. Reuse locked third-party dependencies while retaining the complete source installation and its resolution fallback.
Use portable path-preserving shell commands for the BuildKit manifest stage and cover nested paths containing spaces.\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use the shell required by the manifest extraction stage so the BuildKit RUN command works with Docker's default shell.\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@zozo123 zozo123 changed the title Restore CI images from Docker data-directory snapshots Accelerate CI image materialization and dependency reuse Oct 4, 2026
@zozo123
zozo123 requested a review from shahar1 October 4, 2026 21:25
Restore the current UV and prek versions, propagate manifest discovery and copy failures, and cover failure injection in the extraction tests.\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@zozo123

zozo123 commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

@shahar1 — measured the current f53486a4 design against a successful baseline with the same active job names. Verdict: this run demonstrates useful CI savings, and supports keeping this optimization. Both total runner time and workflow wall time improved.

Metric Matched baseline This PR Improvement
Median CI-image preparation per consumer 210s 127s 83s faster (39.5%)
Total runner execution time 1,306.0 min 1,236.4 min 69.6 runner-minutes saved (5.3%)
AMD workflow wall time 51m 14s 49m 47s 87s faster (2.8%)
Median gap from CI builder completion to consumer start 3s 3s No increase

All 79 active CI-image consumers successfully restored the snapshot and skipped the legacy image-load action. Snapshot creation took 73s and upload took 23s; the total workflow numbers above already include that 96s overhead.

The critical-path tradeoff is also favorable for the median consumer in this comparison: consumers started 69s later relative to workflow creation, but their faster image preparation meant the median paired job reached the end of preparation 16s earlier. Across all 79 paired consumers, preparation used 110.6 fewer runner-minutes. The total workflow saving is 69.6 minutes after accounting for every executed job, including the builder.

Sources: current-head run, attempt 1 · matched baseline run, attempt 1. Measurements use GitHub Actions job/step timestamps and consumer action outcomes. Wall time is workflow creation to the last executed job completing; runner time is the sum of executed job durations.

This is one treatment run and one exact-matrix baseline, so the percentages describe this comparison rather than a guaranteed result for every PR. The dependency-layer cache saving is separate and not included as a demonstrated benefit here. Additional repetitions would establish consistency, but the measured result is already positive: faster preparation, less runner time, and a faster completed workflow.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants