Skip to content

borg extract: only newer/non-existing files #6857

Description

@unilynx

Have you checked borgbackup docs, FAQ, and open Github issues?

Yes

Is this a BUG / ISSUE report or a QUESTION?

Question

Describe the problem you're observing.

Is there a way to have 'borg extract' ignore already existing/extracted files ? This would be useful if the connection is lost during an extract, or if halfway a large restore you decide it's better to abort and add an additional --exclude mask.

Activity

  1. ThomasWaldmann commented on Jul 13, 2022

    @ThomasWaldmann
    Member

    I agree it would be useful to optimize such a case (and still making sure everything is correct in the end), but that's not implemented yet.

    IIRC, there is a more generic ticket about this already, maybe you can find it.

    The general case is extracting merging into some (arbitrary) existing directory tree and maybe even efficiently updating existing big files.

  2. ThomasWaldmann commented on Jul 13, 2022

    @ThomasWaldmann
    Member

    What you can try is borg mount and rsync from the mount to the target dir.

    Not sure if it is really faster, guess it depends...

    But be aware that borg mount does not support ACLs and (bsd / filesystem) flags.

  3. xwst commented on Aug 10, 2022

    @xwst

    IIRC, there is a more generic ticket about this already, maybe you can find it.

    I am also interested in this feature, so I looked for related issues. Are you referring to #1986?

  4. ThomasWaldmann commented on Aug 10, 2022

    @ThomasWaldmann
    Member

    @xwst Yes, guess that was the one.

  5. mbunkus commented on Nov 23, 2022

    @mbunkus

    I'd like to ask for this feature as well. I've just experienced extraction abort in the middle of the ~10TB restore, which had already been running for two days, and having to re-do all of those hurts.

    I'll look into mounting & rsyncing for now.

    Apart from that: borg is an amazing piece of software, and I'm very, very grateful for all your work!

  6. ThomasWaldmann commented on Nov 23, 2022

    @ThomasWaldmann
    Member

    OK, maybe it is worth implementing the simplest usecase (not the most generic usecase of "i have something, bring it in sync"):

    • we start from an empty extraction base directory D
    • an extraction of some archive A is attempted, but interrupted
    • nothing inside D is modified (especially: nothing added or renamed)
    • the extraction attempt of A shall be efficiently repeated without re-extracting what we already have in D
    • expectation: have a full, valid extraction of A, no more, no less

    So we have these cases for some file Fa (in archive) and Ff (in filesystem):

    • Ff is not present: extract Fa
    • Ff is already present, but there is a mismatch in size or mtime compared to Fa: delete Ff, extract Fa
    • Ff is already present, its size and mtime matches what we have in Fa: nothing to do

    For some directory Da (in archive) and Df (in filesystem):

    • Df is not present: extract Da (== create directory, set metadata). Note: if we write files into Df, we modify the timestamps of Df by doing that and need to update timestamps of Df again at the end.
    • Df is already present, but there is a mismatch in mtime: update timestamps of Df again at the end
    • Df is already present, its mtime matches Da: nothing to do

    TODO: consider xattrs, acls and other metadata.

  7. added this to the 2.1 aka N+1 milestone on Nov 23, 2022
  8. ThomasWaldmann commented on Nov 24, 2022

    @ThomasWaldmann
    Member

    The current code restores metadata in this order:

    • uid/gid
    • mode
    • atime/birthtime
    • atime/mtime
    • acls
    • xattrs
    • flags (includes immutable flag, thus must be done at the end)

    Note: if metadata restoration gets interrupted somewhere after mtime, the fs item would have a "correct" (matching) mtime, but would not have complete acls or xattrs.

    Thus, I guess this would need to change to:

    • uid/gid
    • mode
    • acls
    • xattrs
    • atime/birthtime
    • atime/mtime
    • flags (includes immutable flag, thus must be done at the end)

    That way (doing mtime as late as possible), having a matching mtime (archive vs. filesystem) would imply that metadata restoration was finished for that fs item (with a small remaining risk concerning the flags, which aren't used that much).

    Comments?

  9. thebalaa commented on Mar 3, 2023

    @thebalaa

    would love to see this functionality, if someone has the time to handhold with me on how this should be implemented I am wiling to give it a shot, end goal for me is #1986

  10. ThomasWaldmann commented on Mar 3, 2023

    @ThomasWaldmann
    Member

    "mtime (2nd) last" was already implemented in master and 1.2-maint branches.

    see archive.py -> restore_attrs.

    @thebalaa if you want to help, just ask (e.g. on IRC) and open a PR.

  11. ThomasWaldmann commented on Jun 8, 2023

    @ThomasWaldmann
    Member

    borg extract --continue (master branch) does some of this. #1356

  12. ThomasWaldmann commented on Oct 6, 2026

    @ThomasWaldmann
    Member

    State in master (borg 2):

    borg extract --continue implements the simple use case described above: extracting into an empty directory, getting interrupted, then running the same extraction again into the same directory.

    • An existing regular file with the archived type, mode, size and mtime is considered complete and skipped. Any other existing file is removed and extracted again.
    • restore_attrs sets the mtime after ACLs and xattrs, so a matching mtime means the metadata restoration was complete (with the small remaining risk about flags mentioned above).
    • Existing directories are reused, never removed.
    • Groups of hard links stay together when some of their members are skipped (extract --continue: keep groups of hard links together #10387).
    • Without --continue, borg extract refuses to extract into a non-empty directory (extract: shall borg warn if cwd is not empty? #10057), so merging into existing content always has to be asked for explicitly.

    #10506 fixes two gaps:

    • Directories with the archived mode and mtime were supposed to be skipped, but their metadata was still restored again (the code for that was not wired up).
    • --progress did not count the skipped files, so the percentage stayed far below the real one.

    Still not covered:

    • A partially extracted file is extracted again from the beginning, so an interrupted extraction of a big file starts over. Resuming inside a file, or updating it in place, is Restore remote data inplace with deduplication or rsync algorithm #95 (Feature Request: Differential Extraction  #1986 was closed as a duplicate of it).
    • mtimes are compared with ns precision and modes are compared too. On filesystems with coarse timestamps or without POSIX modes (HFS+, FAT/exFAT, some SMB mounts), and when extracting e.g. a Linux archive on Windows, nothing ever matches, so everything gets extracted again. The result is still correct, just without the speedup.
    • borg 1.x does not have --continue. There, the workaround is still borg mount + rsync, but note that borg mount does not provide ACLs and flags.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions