mirror of https://github.com/Zygo/bees.git synced 2025-08-23 14:32:20 +02:00

Go to file

Zygo Blaxell ce0367dafe scan_one_extent: reduce the number of LOGICAL_INO calls before finding a duplicate block range

When we have multiple possible matches for a block, we proceed in three
phases:

1.  retrieve each match's extent refs and put them in a list,
2.  iterate over the list converting viable block matches into range matches,
3.  sort and flatten the list of range matches into a non-overlapping
list of ranges that cover all duplicate blocks exactly once.

The separation of phase 1 and 2 creates a performance issue when there
are many block matches in phase 1, and all the range matches in phase
2 are the same length.  Even though we might quickly find the longest
possible matching range early in phase 2, we first extract all of the
extent refs from every possible matching block in phase 1, even though
most of those refs will never be used.

Fix this by moving the extent ref retrieval in phase 1 into a single
loop in phase 2, and stop looping over matching blocks as soon as any
dedupe range is created.  This avoids iterating over a large list of
blocks with expensive `LOGICAL_INO` ioctls in an attempt to improve the
match when there is no hope of improvement, e.g. when all match ranges
are 4K and the content is extremely prevalent in the data.

If we find a matched block that is part of a short matching range,
we can replace it with a block that is part of a long matching range,
because there is a good chance we will find a matching hash block in
the long range by looking up hashes after the end of the short range.
In that case, overlapping dedupe ranges covering both blocks in the
target extent will be inserted into the dedupe list, and the longest
matches will be selected at phase 3.  This usually provides a similar
result to that of the loop in phase 1, but _much_ more efficiently.

Some operations are left in phase 1, but they are all using internal
functions, not ioctls.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>

2024-11-30 23:30:33 -05:00

bin

bees: remove local cruft, throw at github

2016-11-17 12:12:13 -05:00

docs

docs: event counter updates after fixing counter names and scan_one_extent improvements

2024-11-30 23:30:33 -05:00

include/crucible

multilock: allow turning it off

2024-11-30 23:30:33 -05:00

lib

multilock: allow turning it off

2024-11-30 23:30:33 -05:00

scripts

Merge github PR #148

2022-12-23 00:26:33 -05:00

src

scan_one_extent: reduce the number of LOGICAL_INO calls before finding a duplicate block range

2024-11-30 23:30:33 -05:00

test

test: GCC 13 fix for limits.cc

2023-05-07 21:24:21 -04:00

.gitignore

gitignore: clang creates a lot of *.tmp files

2021-11-29 21:27:48 -05:00

COPYING

GPL-3: license it

2016-11-17 12:12:15 -05:00

Defines.mk

beesd: Honor DESTDIR on installation.

2022-12-23 11:10:17 +08:00

Makefile

Makefile: also drop fiemap and fiewalk from main Makefile

2023-01-28 11:21:51 +01:00

makeflags

lib: deprecate memset_zero template, use C99 compound literals instead

2021-11-29 21:27:48 -05:00

README.md

docs: working around btrfs send issues isn't really a feature

2023-03-07 10:25:51 -05:00

README.md

BEES

Best-Effort Extent-Same, a btrfs deduplication agent.

About bees

bees is a block-oriented userspace deduplication agent designed for large btrfs filesystems. It is an offline dedupe combined with an incremental data scan capability to minimize time data spends on disk from write to dedupe.

Strengths

Space-efficient hash table and matching algorithms - can use as little as 1 GB hash table per 10 TB unique data (0.1GB/TB)
Daemon incrementally dedupes new data using btrfs tree search
Works with btrfs compression - dedupe any combination of compressed and uncompressed files
Works around btrfs filesystem structure to free more disk space
Persistent hash table for rapid restart after shutdown
Whole-filesystem dedupe - including snapshots
Constant hash table size - no increased RAM usage if data set becomes larger
Works on live data - no scheduled downtime required
Automatic self-throttling based on system load

Weaknesses

Whole-filesystem dedupe - has no include/exclude filters, does not accept file lists
Requires root privilege (or CAP_SYS_ADMIN)
First run may require temporary disk space for extent reorganization
First run may increase metadata space usage if many snapshots exist
Constant hash table size - no decreased RAM usage if data set becomes smaller
btrfs only

Installation and Usage

More Information

Bug Reports and Contributions

Email bug reports and patches to Zygo Blaxell bees@furryterror.org.

You can also use Github:

    https://github.com/Zygo/bees

Copyright & License

GPL (version 3 or later).

README.md

BEES

About bees

Strengths

Weaknesses

Installation and Usage

Recommended Reading

More Information

Bug Reports and Contributions

Copyright & License