GGLinnk/bees - bees - Virtual World Git

mirror of https://github.com/Zygo/bees.git synced 2025-07-15 22:33:27 +02:00

Author	SHA1	Message	Date
Zygo Blaxell	8d08a3c06f	readahead: inject some sanity at the foundation of an insane architecture This solves some of the worst problems with bees reads: 1. The kernel readahead doesn't work. More precisely, it's much better adapted for a very different use case: a single thread alternating between reading a file sequentially and processing the data that was read. bees has multiple threads which compete for access to IO and then issue reads in random order immediately after the call to readahead. The kernel uses idle ioprio scheduling for the readaheads, so the readaheads get preempted by the random reads, or cancels the readaheads because the data access pattern isn't sequential after the readahead was issued. 2. Seeking drives perform terribly with multiple competing readers, especially with btrfs striped profiles where the iops are broken into tiny stripe-sized pieces. At one point I intended to read the btrfs device map and figure out which devices can be read in parallel, but to make that useful, the user needs to have an array with multiple drives in single profile, or 4+ drives in raid1 profile. In all other cases, the elaborate calculations always return the same result: there can be only one reader at a time. This commit fixes both problems: 1. Don't use the kernel readahead. Use normal reads into a dummy buffer instead. 2. Allow only one thread to readahead at any time. Once the read is completed, the data is in the page cache, and all the random-order small reads that bees does will hit the page cache, not a spinning disk. In some cases we need to read two things close together, so add a `bees_readahead_pair` which holds one lock across both reads. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	cdcdf8e218	hash: use kernel readahead instead of bees_readahead to prefetch hash table The hash table is read sequentially and from a single thread, so the kernel's implementation of readahead is appropriate here. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	37f5b1bfa8	docs: add allocator regression in 6.0+ kernels Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	abe2afaeb2	context: when a task fails to acquire an extent lock, don't go ahead and scan the extent anyway Commit `c3b664fea5` ("context: don't forget to retry locked extents") removed the critical return that prevents a Task from processing an extent that is locked. Put the return back. Fixes: `c3b664fea5` ("context: don't forget to retry locked extents") Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	792fdbbb13	fs: get rid of 16 MiB limit on dedupe requests The kernel has not required a 16 MiB limit on dedupe requests since v4.18-rc1 b67287682688 ("Btrfs: dedupe_file_range ioctl: remove 16MiB restriction"). Kernels before v4.18 would truncate the request and return the size actually deduped in `bytes_deduped`. Kernel v4.18 and later will loop in the kernel until the entire request is satisfied (although still in 16 MiB chunks, so larger extents will be split). Modify the loop in userspace to measure the size the kernel actually deduped, instead of assuming the kernel will only accept 16 MiB. On current kernels this will always loop exactly once. Since we now rely on `bytes_deduped`, make sure it has a sane value. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	30a4fb52cb	Revert "context: add experimental code for avoiding tiny extents" because this problem is better solved elsewhere. This reverts commit `11fabd66a8`.	2024-11-30 23:30:33 -05:00
Zygo Blaxell	90d7075358	usage: the default scan mode is 3 (recent) The code and docs were changed some time ago, but not the usage message. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	faac895568	docs: add the 6.10..6.12 delayed refs bug Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	a7baa565e4	crawl: rename next_transid() to avoid confusion with BeesScanMode::next_transid() Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:30:33 -05:00
Zygo Blaxell	b408eac98e	trace: add file and line numbers all the way up the stack These were added to crucible all the way back in 2018 (`1beb61fb78` "crucible: error: record location of exception in what() message") but it's even more useful in the stack tracer in bees. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-11-30 23:27:24 -05:00
Zygo Blaxell	75131f396f	context: reduce the size of LOGICAL_INO buffers Since we'll never process more than BEES_MAX_EXTENT_REF_COUNT extent references by definition, it follows that we should not allocate buffer space for them when we perform the LOGICAL_INO ioctl. There is some evidence (particularly https://github.com/Zygo/bees/issues/260#issuecomment-1627598058) that the kernel is subjecting the page cache to a lot of disruption when trying allocate large buffers for LOGICAL_INO. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-04-17 23:14:35 -04:00
Zygo Blaxell	cfb7592859	usage: the default scan mode is 1 (independent) Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-04-17 23:14:35 -04:00
Zygo Blaxell	3839690ba3	lib: fix btrfs_data_container pointer casts for 32-bit userspace on 64-bit kernels Apparently reinterpret_cast<uint64_t> sign-extends 32-bit pointers. This is OK when running on a 32-bit kernel that will truncate the pointer to 32 bits, but when running on a 64-bit kernel, the extra bits are interpreted as part of the (now very invalid) address. Use <uintptr_t> instead, which is unsigned, integer, and the same word size as the arch's pointer type. Ordinary numeric conversion can take it from there, filling the rest of the word with zeros. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2024-04-17 23:07:41 -04:00
Zygo Blaxell	124507232f	docs: add vmalloc bug to kernel bugs list The bug is: v6.3-rc6: f349b15e183d mm: vmalloc: avoid warn_alloc noise caused by fatal signal The fixes are: v6.4: 95a301eefa82 mm/vmalloc: do not output a spurious warning when huge vmalloc() fails v6.3.10: c189994b5dd3 mm/vmalloc: do not output a spurious warning when huge vmalloc() fails The bug has been backported to LTS, but the fix has not: v6.2.11: 61334bc29781 mm: vmalloc: avoid warn_alloc noise caused by fatal signal v6.1.24: ef6bd8f64ce0 mm: vmalloc: avoid warn_alloc noise caused by fatal signal v5.15.107: a184df0de132 mm: vmalloc: avoid warn_alloc noise caused by fatal signal Signed-off-by: Zygo Blaxell <bees@furryterror.org> v0.10	2023-07-06 13:50:12 -04:00
Zygo Blaxell	3c5e13c885	context: log when LOGICAL_INO returns 0 refs There was a bug in kernel 6.3 where LOGICAL_INO with IGNORE_OFFSET sometimes fails to ignore the offset. That bug is now fixed, but LOGICAL_INO still returns 0 refs much more often than seems appropriate. This is most likely because bees frequently deletes extents while there is still work waiting for them in Task queues. In this case, LOGICAL_INO correctly returns an empty list, because every reference to some extent is deleted, but the new extent tree with that extent removed is not yet committed in btrfs. Add a DEBUG-level log message and an event counter to track these events. In the absence of a kernel bug, the debug message may indicate CPU time was wasted performing a search whose outcome could have been predicted. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-07-06 12:54:33 -04:00
Zygo Blaxell	a6ca2fa2f6	docs: add IGNORE_OFFSET regression in 6.2..6.3 to kernel bugs list This doesn't impact the current bees master, but it does break bees-next. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-07-06 12:49:36 -04:00
Zygo Blaxell	3f23a0c73f	context: downgrade toxic extent workaround message Toxic extents are much less of a problem now than they were in kernels before 5.7. Downgrade the log message level to reflect their lesser importance. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-07-06 12:49:36 -04:00
Zygo Blaxell	d6732c58e2	test: GCC 13 fix for limits.cc GCC complains that #include <cstdint> is missing, so add that. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-05-07 21:24:21 -04:00
Zygo Blaxell	75b2067cef	btrfs-tree: fix build on clang++16 The "loops" variable isn't read (only set) if not built with extra debug code. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-05-07 21:23:27 -04:00
Zygo Blaxell	da3ef216b1	docs: working around `btrfs send` issues isn't really a feature The critical kernel bugs in send have been fixed for years. The limitations that remain aren't bugs, and bees has no sustainable workaround for them. Also update copyright year range. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-03-07 10:25:51 -05:00
Zygo Blaxell	b7665d49d9	docs: fill in missing LTS backports for "1119a72e223f btrfs: tree-checker: do not error out if extent ref hash doesn't match" Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-03-07 10:17:44 -05:00
Zygo Blaxell	717bdf5eb5	roots: make sure transid_max's computed value isn't max We check the result of transid_max_nocache(), but not the result of transid_max(). The latter is a computed result that is even more likely to be wrong[citation needed]. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:45:29 -05:00
Zygo Blaxell	9b60f2b94d	docs: add "missing" features that have been in development for some time already Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:42:42 -05:00
Zygo Blaxell	8978d63e75	docs: update GCC versions list and clarify markdown statement I don't know if anyone else is testing GCC versions before 8.0 any more, but I'm not. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:39:55 -05:00
Zygo Blaxell	82474b4ef4	docs: update front page At least one user was significantly confused by "designed for large filesystems". The btrfs send workarounds aren't new any more. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:38:50 -05:00
Zygo Blaxell	73834beb5a	docs: minor changes to how-it-works based on past user questions Clarify that "too large" and "too small" are some distance away from each other. The Goldilocks zone is _wide_. The interval between cache drops is now shorter. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:37:37 -05:00
Zygo Blaxell	c92ba117d8	docs: various gotcha updates Fixing the obviously wrong and out of date stuff. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:37:23 -05:00
Zygo Blaxell	c354e77634	docs: simplify the exit-with-SIGTERM description The description now matches the code again. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:36:44 -05:00
Zygo Blaxell	f21569e88c	docs: update the feature interactions page Fixing the obviously out-of-date and no-longer-tested things. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:34:22 -05:00
Zygo Blaxell	3d5ebe4d40	docs: update kernel bugs and workarounds list for 6.2.0 Remove some of the repetition to make the document easier to edit. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-25 03:32:52 -05:00
Zygo Blaxell	3430f16998	context: create a Pool of BtrfsIoctlLogicalInoArgs objects Each object contains a 16 MiB buffer, which is very heavy for some malloc implementations. Keep the objects in a Pool so that their buffers are only allocated and deallocated once in the process lifetime. Signed-off-by: Zygo Blaxell <bees@furryterror.org> v0.9.3	2023-02-23 22:45:31 -05:00
Zygo Blaxell	7c764a73c8	fs: allow BtrfsIoctlLogicalInoArgs to be reused, remove virtual methods Some malloc implementations will try to mmap() and munmap() large buffers every time they are used, causing a severe loss of performance. Nothing ever overrode the virtual methods, and there was no virtual destructor, so they cause compiler warnings at build time when used with a template that tries to delete pointers to them. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-23 22:40:12 -05:00
Zygo Blaxell	a9a5cd03a5	ProgressTracker: reduce memory usage with long-running work items ProgressTracker was only freeing memory for work items when they reach the head of the work tracking queue. If the first work item takes hours to complete, and thousands of items are processed every second, this leads to millions of completed items tracked in memory at a time, wasting gigabytes of system RAM. Rewrite ProgressHolderState methods to keep only incomplete work items in memory, regardless of the order in which they are added or removed. Also fix the unit tests which were relying on the memory leak to work, and add test cases for code coverage. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-23 22:33:35 -05:00
Zygo Blaxell	299509ce32	seeker: fix the test for ILP32 platforms Not sure what I was thinking, but the argument here should clearly be uint64_t. Fixes: https://github.com/Zygo/bees/issues/248 Signed-off-by: Zygo Blaxell <bees@furryterror.org> v0.9.2	2023-02-20 11:30:56 -05:00
Zygo Blaxell	d5a99c2f5e	roots: don't share a RootFetcher between threads If the send workaround is enabled, it is possible for two threads (a thread running the crawl_new task, and a thread attempting to apply the send workaround) to access the same RootFetcher object at the same time. That never ends well. Give each function its own BtrfsRootFetcher object. Fixes: https://github.com/Zygo/bees/issues/250 Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-02-20 11:14:34 -05:00
Kai Krakow	fd6c3b3769	Makefile: also drop fiemap and fiewalk from main Makefile Fixes: `ccd8dcd43f` Signed-off-by: Kai Krakow <kai@kaishome.de> v0.9.1	2023-01-28 11:21:51 +01:00
Zygo Blaxell	849c071146	hash: flush the table more slowly With SIGTERM and fast exit, the trickle writeback is less important. We don't want to flood people's IO subsystems with continuous writes. This really should be configurable at runtime. Signed-off-by: Zygo Blaxell <bees@furryterror.org> v0.9	2023-01-27 22:16:02 -05:00
Zygo Blaxell	85ff543695	test: simplify Makefile Make can build dependencies in parallel, so let Make do that. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	8147f80a5a	src: bees-version.cc cleanups Do rebuild bees-version.cc if libcrucible changes. Don't rebuild bees-version.cc if it doesn't change. Also use the standard suffix for new files. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	cbde237f79	src: simplify Makefile Make can build dependencies in parallel, so let Make do that. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	3b85fc8bc7	lib: drop version.cc entirely crucible::VERSION doesn't make much sense now that libcrucible no longer exists as a shared library. Nothing ever referenced it, so it can go away. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	4df1b2c834	lib: simplify dependency generation We don't need to run all the dependencies first, Make can do those in parallel. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	495218104a	fd: FS_IOC_SETFLAGS takes an int* argument not a long* According to ioctl_iflags(2): The type of the argument given to the FS_IOC_GETFLAGS and FS_IOC_SETFLAGS operations is int , notwithstanding the implication in the kernel source file include/uapi/linux/fs.h that the argument is long . So this code doesn't work on be64 machines. Also, Valgrind complains about it. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	e82ce3c06e	fd: pwrite returns ssize_t not int A subtle distinction, and not one that is particularly relevant to bees, but it does make toolchains complain. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	bd336e81a6	fs: get rid of base class btrfs_ioctl_logical_ino_args Another instance of the pattern where we derived a crucible class from a btrfs struct. Make it an automatic variable instead. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	ea17c89165	fs: remove duplicate BTRFS_COMPRESS_ definitions This was fixed in `7f660f50b` lib: fs: stop using libbtrfs-dev helper functions to re-enable buffer length checks but apparently some copies live on. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-27 22:16:02 -05:00
Zygo Blaxell	ccd8dcd43f	fiemap, fiewalk: drop dead example/test code These tools are obsolete. fiemap was a thin wrapper around FIEMAP, but FIEMAP is not useful on btrfs. fiewalk was a thin wrapper around BtrfsExtentWalker, but development on BtrfsExtentWalker has been abandoned. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-23 00:09:26 -05:00
Zygo Blaxell	facf4121a6	context: remove the one call to operator vector<> method in BtrfsIoctlLogicalInoArgs There's only one user of this method. Open-code it so we can kill the method in libcrucible. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-23 00:09:26 -05:00
Zygo Blaxell	cbc76a7457	hash: don't spin when writes fail When a hash table write fails, we skip over the write throttling because we didn't report that we successfully wrote an extent. This can be bad if the filesystem is full and the allocations for writes are burning a lot of CPU time searching for free space. We also don't retry the write later on since we assume the extent is clean after a write attempt whether it was successful or not, so the extent might not be written out later when writes are possible again. Check whether a hash extent is dirty, and always throttle after attempting the write. If a write fails, leave the extent dirty so we attempt to write it out the next time flush cycles through the hash table. During shutdown this will reattempt each failing write once, after that the updated hash table data will be dropped. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-23 00:09:26 -05:00
Zygo Blaxell	28ee2ae1a8	docs: fix broken link in options.md Links in docs/ are relative to docs/, not the top level. Signed-off-by: Zygo Blaxell <bees@furryterror.org>	2023-01-23 00:08:54 -05:00

... 2 3 4 5 6 ...

824 Commits