Replying back with a large update.
After moving some datasets I had from dedup=on checksum=sha256 compression=zstd to dedup=off checksum=blake3 compression=zstd, and deleting the original source dataset, and moving one large dataset to a subdirectory under an existing dataset (collapsing 2 datasets into 1), I’ve managed to free up 9.2 terabytes of capacity.
Yes, that’s a real number. I thought I accidentally deleted a large dataset, but nothing I have is that large in a single dataset. There were no snapshots of these datasets, that was purely savings freed up from some pending backlog of async destroys finally freeing blocks on the pool.
Those must have been backlogged for the better part of 2 years. Even with multiple dozens of scrubs in that time, my capacity had been diminishing down to 11TB total free, now I have 21TB total free. Apparently the bptree was so backed up, it could never process the frees, and the rename old/create new/rsync old to new/destroy old, managed to shake it loose.
Instead of using dedup, I’m now going to use hardlink and jdupes inside each dataset to shed some of the duplicate files (there are millions of them) and see if that gives me even more capacity back.
I started with something that looked like this:
dedup: DDT entries 88810794, size 28.4G on disk, 12.6G in core
bucket allocated referenced
______ ______________________________ ______________________________
refcnt blocks LSIZE PSIZE DSIZE blocks LSIZE PSIZE DSIZE
------ ------ ----- ----- ----- ------ ----- ----- -----
1 77.6M 10.5T 10.2T 10.2T 77.6M 10.5T 10.2T 10.2T
2 6.38M 985G 954G 957G 13.3M 1.98T 1.92T 1.92T
4 458K 48.0G 42.0G 42.9G 2.14M 224G 195G 199G
8 136K 10.6G 7.55G 7.92G 1.38M 109G 76.7G 80.6G
16 55.1K 4.17G 2.48G 2.65G 1.16M 89.7G 53.2G 56.8G
32 29.9K 2.47G 1.07G 1.17G 1.26M 106G 44.2G 48.8G
64 16.0K 1.50G 460M 517M 1.48M 147G 44.1G 49.4G
128 5.07K 426M 140M 160M 902K 74.5G 24.5G 28.0G
256 2.25K 171M 54.6M 63.0M 720K 49.3G 15.7G 18.4G
512 466 14.5M 4.03M 6.17M 311K 9.84G 2.72G 4.16G
1K 192 12.8M 2.46M 3.13M 293K 21.5G 4.00G 4.97G
2K 74 4.04M 842K 1.08M 209K 11.4G 2.34G 3.09G
4K 28 638K 210K 348K 150K 3.43G 1.14G 1.86G
8K 16 500K 152K 227K 180K 5.80G 1.75G 2.58G
16K 2 10K 7K 14.2K 41.4K 209M 145M 294M
32K 3 128K 29K 42.6K 135K 6.11G 1.28G 1.91G
64K 1 1.50K 1.50K 7.10K 120K 180M 180M 852M
128K 1 11.5K 8K 14.2K 213K 2.40G 1.67G 2.96G
Total 84.7M 11.5T 11.2T 11.2T 102M 13.3T 12.6T 12.6T
And right now, it’s still pruning, but I see this:
dedup: DDT entries 83668274, size 27.2G on disk, 12.7G in core
bucket allocated referenced
______ ______________________________ ______________________________
refcnt blocks LSIZE PSIZE DSIZE blocks LSIZE PSIZE DSIZE
------ ------ ----- ----- ----- ------ ----- ----- -----
1 73.7M 10.1T 9.87T 9.87T 73.7M 10.1T 9.87T 9.87T
2 5.78M 961G 935G 935G 11.9M 1.93T 1.88T 1.88T
4 283K 44.1G 40.5G 40.6G 1.29M 203G 187G 188G
8 60.4K 7.28G 6.60G 6.65G 612K 73.8G 66.5G 67.0G
16 16.3K 1.91G 1.84G 1.85G 348K 40.9G 39.1G 39.4G
32 4.19K 509M 454M 457M 156K 18.5G 16.4G 16.6G
64 233 24.9M 24.3M 24.5M 20.3K 2.18G 2.13G 2.15G
128 193 22.1M 21.7M 21.8M 34.1K 3.92G 3.84G 3.86G
256 45 5.08M 4.59M 4.64M 14.0K 1.56G 1.39G 1.41G
512 4 2.50K 2.50K 28.4K 2.65K 1.71M 1.71M 18.8M
1K 2 128K 4.50K 14.2K 2.31K 130M 4.69M 16.4M
Total 79.8M 11.1T 10.8T 10.8T 88.0M 12.4T 12.1T 12.1T
I learned quite a bit in this process about ARC, tuning, and why fast_dedup and dedup are basically the same, and neither are helpful at all.