Space is being lost every time I move data between datasets

Hello,

I have a 4 x 4TB disks Nas running with the TrueNAS service on it. The ZFS disk storage configuration is RAIDZ1 so the available space is about 75% of the physical storage. The service is up to date so right now I am running the TrueNAS 25.10.3.1 version.

Some days ago, my NAS has reached the 90% of the used space (It’s not related with the reasons but maybe it’s important). While the ZFS space was above the 90%, I have started to move some films from one dataset to another, and during that movement I have noticed that the used space has reached the 99.1%.

The first time I used the mv command to move the content of a folder to the new dataset but the films were not deleted until all were “copied” first, so I decide to change the way and I used a for loop to move one by one, and also I have used the command rsync with the --remove-source-files for the same reason (all the commands were executed directly on the NAS, not by SMB or NFS). No matter the way I have used, I have noticed that the free space is dissapearing even when the source file is deleted.

Here I have started to ask me where is the free space, because a film uses 30GB of space regardless the compression (they cannot be compressed), and the files weren’t duplicated. Also there are no snapshots on the pool, the datasets or similar:

% zfs list -t snapshot -r nas                                                 
no datasets available

also the freeing is almost 0 compared with the missing space:

%  zpool get freeing nas
NAME  PROPERTY  VALUE    SOURCE
nas   freeing   17.8G    -

The ZFS system was freeing some space but very far away from the lost space.

Thinking that maybe was something related with the fragmentation or the lack of free space because the pool was in a critical state staying above the 90% of the usage, I decided to free some space and move again the films. I deleted about 500GB of non critical videos, moved out to some old external HDD about 2.6TB freeing up to 3.1TB of space and leaving the ZFS near the much more confortable 70% of usage. After freeing all that space I have started to move the films again between datasets and for my surprise, another 500GB have been dissapeared from the pool and now I am unable to move the data that I have on my external drives back to their datasets. I am moving another 2TB of downloads to another dataset and I am getting the same, loosing again a lot of GB. I have not generated any new data in that NAS and from the 3.1TB I had, now I have 2.35TB and going down.

My pool and datasets distribution is:

  • pool “nas”:
    • nas/Media: source (Películas folder inside the dataset)
    • nas/Media/Peliculas: destination

I am desesperated because I don’t know where is my space, what to do, what to check… I only know that my space is dissapearing without grown the containing data.

  • I have checked if there are snapshots, and no, there are no snapshots. I have never created any but I have checked it anyway.
  • The data is the same and it’s not duplicated. I have not seen any duplicated file.
  • The configuration between the datasets is the same. Only changes the blocksize that was changed to 1MB (change to the original 128k blocksize later with the same result).

I have done a simple test moving one file:

% du 10\ minutos\ menos\ \(2019\)/10\ minutos\ menos\ \(2019\)\ Bluray-480p.avi --apparent-size
1797950 10 minutos menos (2019)/10 minutos menos (2019) Bluray-480p.avi

% du 10\ minutos\ menos\ \(2019\)/10\ minutos\ menos\ \(2019\)\ Bluray-480p.avi                
1742350 10 minutos menos (2019)/10 minutos menos (2019) Bluray-480p.avi


mv 10\ minutos\ menos\ \(2019\)/10\ minutos\ menos\ \(2019\)\ Bluray-480p.avi ../

% du ../10\ minutos\ menos\ \(2019\)\ Bluray-480p.avi --apparent-size
1797950 ../10 minutos menos (2019) Bluray-480p.avi

% du ../10\ minutos\ menos\ \(2019\)\ Bluray-480p.avi 
1795034 ../10 minutos menos (2019) Bluray-480p.avi

This is normal, the first uses a 1mb blocksize while the 2nd uses a 128k blocksize, so the compression of the data is a bit better and then the 2nd is bigger, but right now I am mooving the files from the 128k blocksize dataset to the 1mb blocksize dataset and I am loosing space. The compression cannot be because a mkv file contains a compressed video that at most you can compress a little, but not about a 23%.

Anyone can help my with this problem?, because for now the only solution I see is to destroy everything and start again, but I have no enough space outside the NAS to move all the data out and in again. Maybe I am missing something.

Thanks!!

Weird - any chance zfs list -t all gives you something of interest?

ZFS space reporting is a bit hard. You may also be using block cloning. Try searching forum for block cloning and space and see if one explains how to view it in the command line.

The important thing is to stop moving any data until you sort out how much space you really have as you will crash the system and lose the pool.

Yeah, is very weird but I am not able to find the reason… I am fighting agains this for about 4 days trying to search something similar in Google and asking the IA, but nothing helps. The numbers provided by the zfs commands matches, but the free space is going out without a reason. I have passed from to have about 1.1TB of space, to a debt of about 500GB. Considering the about 600GB of deleted data that is 2.2TB of free space that is nowere.

Can the fragmentation provoke a higher usage? because the fragmentation was about 58%, after the cleanup and moving data out the pool was reduced to about 28%, and now is again at about 41%.

Looking at the command I don’t see anything special, just the datasets that I have created with their size.

nas/Media                                                                                            6.85T  2.39T  1004G  /mnt/nas/Media
nas/Media/Downloads                                                                                  1.24T  2.39T  1.24T  /mnt/nas/Media/Downloads
nas/Media/Fotos                                                                                      62.3G  2.39T  62.3G  /mnt/nas/Media/Fotos
nas/Media/Others                                                                                      384K  2.39T   384K  /mnt/nas/Media/Others
nas/Media/Peliculas                                                                                  2.53T  2.39T  2.53T  /mnt/nas/Media/Peliculas
nas/Media/Series                                                                                     2.03T  2.39T  2.03T  /mnt/nas/Media/Series

The deduplication option can be kept at folder level even after disabling it?. Doesn’t makes sense for me, but it’s the only reason I am able to find. The downloads data is copied to the film folder once the download is finished and then after some time deleted. Only a deduplication can save enough space, but the deduplication was disabled a little time after enabling it and much time before the download of those files.

Also another command shows a different free space:

% zpool get freeing,fragmentation,leaked,allocated,free nas
NAME  PROPERTY       VALUE    SOURCE
nas   freeing        0        -
nas   fragmentation  41%      -
nas   leaked         0        -
nas   allocated      11.1T    -
nas   free           3.46T    -

Which is near the point I have started. Anther one:

% zfs get recordsize,compression,used,logicalused nas 
NAME  PROPERTY     VALUE           SOURCE
nas   recordsize   128K            local
nas   compression  zstd-6          local
nas   used         8.07T           -
nas   logicalused  8.46T           -

I’ll take a look and I’ll try to search about the block cloning. Maybe it’s related. The configuration of copies in the dataset is just 1 (another thing that I have checked), but maybe is something hidden in the interface.

Thanks!

I think taking a look at the following thread / post will help you find the issue or explain it a bit better.

Different blocksizes prevents block cloning, so your moves are real copies.
With high fragmentation and a nearly full pool, you may also be triggering the alternative allocation logic.

:index_pointing_up:
I suspect that this is the safest path to sanity…

Addiitional headache… Data written while dedup was enabled relies on the DDT and need to retain it; the only way to fully remove dedup is to replicate the data to a dataset that is not deduplicated.
tl;dr Don’t mess with dedup.

Block cloning can be checked as below. I’ve used my desktop’s root pool as an example, which has block cloning disabled.

rpool    bcloneused                     0                          -
rpool    bclonesaved                    0                          -
rpool    bcloneratio                    1.00x                      -
rpool    feature@block_cloning          disabled                   local

On the subject of fragmentation causing reduction in storage efficiency, yes, that can happen. Especially with RAID-Zx. Basically if their are no full width stripes available, then shorter width stripes have to be used. Thus, more storage used because more parity is in use, (and more metadata pointing to those stripes, will be used).

However, unless your pool had a lot of churn / massive changes, I would doubt that is the whole answer.

Thanks to all!!,

The BCLONE makes a lot of sense in my case. On the same Dataset I had Downloads, Peliculas and Series. Everything in Downloads is also in Peliculas and Series so having the same blocksize and the same configuration (same dataset), will lead to the use of reflinks and the bclone feature. Now Is time to try to revert the problem which will be hard…

% zpool get feature@block_cloning                                                                                    
NAME       PROPERTY               VALUE                  SOURCE
apps       feature@block_cloning  active                 local
boot-pool  feature@block_cloning  active                 local
nas        feature@block_cloning  enabled                local

It’s enabled but right now it’s 1.00x:

% zpool list -o name,bcloneused,bclonesaved,bclonerati
NAME       BCLONE_USED  BCLONE_SAVED  BCLONE_RATIO
apps              244K          244K         2.00x
boot-pool          52K           52K         2.00x
nas                  0             0         1.00x

Maybe a python script with a database of the files which performs a reflink copy from one to another once a match is detected.

Greetings!!

Block cloning operates at block level (who would have guessed?). So “file matches” do not matter. ZFS can only effect block cloning during the initial copy; if different setting prevented block cloning while copying from Downloads to another dataset you have two different files which will remain so forever and take twice the space.

Time to create a bigger pool, and rethink your storage layout.

Yes, that is what I am doing. For now I am reverting the changes and joining all together in the same dataset which makes easier to manage the reflinks. Luckily I have knowledge about programming and I have designed a way to revert the problem. For sure I will use a python script with a sqlite database to store the data and I will analize all the files in the dataset to detect duplicates, Once a duplicate is found, I will delete one of them and I’ll perform a reflink copy from the other. Nothing special, just like the zfs dedup do but I’ll perform it at file level :smile:. Maybe I will revert the state to the original one and even better.

Greetings!