So while planning my data layout I came across the realisation that these 2 datasets inherit their recordsize from the root dataset which for me is currently set at 128KiB.
.truenas_containers
ix-apps (and this does not seem a period)
So these are my thoughts
So for app and container writes 128KiB blocks is very large.
The default 128KiB is the recommended default rcordsize for datasets so this is probably a good value to leave for the root dataset where the rest inherit from.
I will have more than apps and containers on this vdev. Unless this is a really bad idea.
The issue
Without using the command line, which I should not have to because of how TrueNAS is supposed to work, how can I give these datasets a better recordsize
Can someone recommend me a suitable block size?
Is this a design flaw that needs to be flagged to the TN team?
I would say when you first initialise the apps and containers you have the options to set a recordsize. perhaps even change it later.
Only possible through the command-line, unless they are exposed in the GUI like how the “iocage” dataset was with Core.
Why do you think 128-KiB is very large? Even if those datasets were configured with a 16-MiB recordsize, an incompressible 4-KiB file will only consume 4-KiB on the disk. An imcompressible 3.25-MiB file will only consume 3.25-MiB on the disk.
Unless you’re regularly doing in-place modifications on your files, such as with DB software, the smaller recordsizes will not really benefit you, and may in fact add more overhead.
Not familiar with SCALE/CE. Are the VM images placed under “.truenas_containers” or “.ix-apps”? What about “ix-virt”? I thought that the user is meant to create or specify the volume or image to use, which will be its own “zvol”. The “zvol” creation wizard exposes “blocksize” options, which work similar to a dataset’s recordsize. No need to use the command-line.
As for in-place modifications, perhaps there’s some truth to what you said about Apps having their own small databases. I doubt their size is enough to warrant overriding the default 128-KiB recordsize for those hidden datasets.
It doesn’t matter for files smaller than 128 KiB. If the file is 8 KiB, it will only consume 8 KiB on the disk as a single block. To read its contents into RAM requires pulling this single block of 8 KiB.
If a file’s size goes beyond the recordsize (i.e, 128 KiB), the subsequent “raw” blocks will each be 128 KiB, including the very last block, even if the last block only contains a few kilobytes of file data. Why isn’t this a problem? It’s because any compression will shrink the padding of null bytes into nothing. After compression and encryption is where you get the actual size and form of the block on the disk and in the ARC.
When it comes to reading and writing, you won’t notice reading a 128-KiB block from disk, even if it is truly 128 KiB. Same thing with 1-MiB blocks. It happens fast. In fact, larger recordsizes yield better speeds and less overhead for most use-cases, since each block must be checksummed, compressed, and encrypted. Imagine doing that eight times for 1 MiB of data split into 8 x 128-KiB blocks, whereas it only needs to be done once for a single 1-MiB block if the recordsize was 1-MiB.
Where smaller recordsizes matter is for VMs, databases, and cases where a lot of true in-place modifications are being done to the data, since it reduces write-amplification and inefficient space usage for small changes inside a file.
My default recordsize is 1-MiB for all my datasets, unless I need to override a particular dataset. I have considered going to 4-MiB. While ZFS allows up to a 16-MiB recordsize, you start to see performance drop after 4-MiB, since you lose out on parallel processing.
Not the point I was making, I think this is a missing feature and should be in the gui. I am following the mantra that true as is an appliance and is locked down.
It always is. It iterates in binary: 4K → 8K → 16K → all the way to the recordsize, where it then starts a new block. Then compression can further shrink it if the data inside is compressible. (Only the first block is truly variable if there is no compression involved.)
All subsequent blocks will use the full recordsize, but compression will deal with the padding so that no space is wasted pointlessly.
I agree. I think the “iocage” (jails) datasets exposed in the GUI in FreeNAS and TrueNAS Core was an oversight, but ironically ended up being a useful feature.
Most importantly it allowed creating snapshot and replication tasks for your jails. I never assumed it was an oversight. That was a central part of my backup strategy.
The “app” system really needs improvement in that regard.