Split out of #10318 (point 3), refs #8476.
Problem
With a corrupt chunks index and no pack errors, a full borg check --repair rebuilds the index from all packs twice:
Repository.check() rebuilds the index from every pack and stores it
(repository.py:1700, slow_rebuild=True, write_immediately=True).
ArchiveChecker.check() then invalidates that index and rebuilds it from every pack again
(archive.py:2324, slow_rebuild=repair), and finish() stores it again.
Both walks validate every object (metadata slot read + decryption per object), so the packs are read and
decrypted twice and the index is written twice.
Proposal
In a full check, leave the rebuild to the archives phase: Repository.check() already gets repo_only
from check_cmd.py, so it can rebuild the index only if repo_only is set.
This is simpler than letting the archives phase reuse the index the repository check built: with
--repair, the archives phase rebuilds from the packs even if the stored index is intact, so reusing it
would need an extra "index was just rebuilt" signal between the two phases.
Consequences:
- In a full check, the "Repository index was corrupted and has been rebuilt from the packs." message and the
count of skipped byte ranges move to the archives phase (note_dropped_objects).
- If the archives phase gets interrupted, the index stays corrupt and is rebuilt on next use, as with
today's "rebuild interrupted" path.
Related question
If some packs are corrupt, Repository.check() does not rebuild the index (refs #8572), but in a full
--repair the archives phase rebuilds it from all packs anyway (objects failing validation are left out).
So that guard only has an effect for --repository-only --repair. Decide whether that is intended.
Split out of #10318 (point 3), refs #8476.
Problem
With a corrupt chunks index and no pack errors, a full
borg check --repairrebuilds the index from all packs twice:Repository.check()rebuilds the index from every pack and stores it(
repository.py:1700,slow_rebuild=True, write_immediately=True).ArchiveChecker.check()then invalidates that index and rebuilds it from every pack again(
archive.py:2324,slow_rebuild=repair), andfinish()stores it again.Both walks validate every object (metadata slot read + decryption per object), so the packs are read and
decrypted twice and the index is written twice.
Proposal
In a full check, leave the rebuild to the archives phase:
Repository.check()already getsrepo_onlyfrom
check_cmd.py, so it can rebuild the index only ifrepo_onlyis set.This is simpler than letting the archives phase reuse the index the repository check built: with
--repair, the archives phase rebuilds from the packs even if the stored index is intact, so reusing itwould need an extra "index was just rebuilt" signal between the two phases.
Consequences:
count of skipped byte ranges move to the archives phase (
note_dropped_objects).today's "rebuild interrupted" path.
Related question
If some packs are corrupt,
Repository.check()does not rebuild the index (refs #8572), but in a full--repairthe archives phase rebuilds it from all packs anyway (objects failing validation are left out).So that guard only has an effect for
--repository-only --repair. Decide whether that is intended.