`ix-apps` and `.truenas_containers` are system and hidden, how should i configure the recordsize

background

So while planning my data layout I came across the realisation that these 2 datasets inherit their recordsize from the root dataset which for me is currently set at 128KiB.

  • .truenas_containers
  • ix-apps (and this does not seem a period)

So these are my thoughts

  • So for app and container writes 128KiB blocks is very large.
  • The default 128KiB is the recommended default rcordsize for datasets so this is probably a good value to leave for the root dataset where the rest inherit from.
  • I will have more than apps and containers on this vdev. Unless this is a really bad idea.

The issue

  • Without using the command line, which I should not have to because of how TrueNAS is supposed to work, how can I give these datasets a better recordsize
  • Can someone recommend me a suitable block size?
  • Is this a design flaw that needs to be flagged to the TN team?
    • I would say when you first initialise the apps and containers you have the options to set a recordsize. perhaps even change it later.

Any thoughts would be appreciated.

Only possible through the command-line, unless they are exposed in the GUI like how the “iocage” dataset was with Core.


Why do you think 128-KiB is very large? Even if those datasets were configured with a 16-MiB recordsize, an incompressible 4-KiB file will only consume 4-KiB on the disk. An imcompressible 3.25-MiB file will only consume 3.25-MiB on the disk.

Unless you’re regularly doing in-place modifications on your files, such as with DB software, the smaller recordsizes will not really benefit you, and may in fact add more overhead.

if you do not use hostpaths then all the writes for apps go into ix-apps an I assume containers are not all write only.

Apps will have small writes to databases, configs. A lot of apps are web-based and will accept very small PHP POSTs with things like

Containers are by their very nature a large collection of small files.

Is it not the case that in a dataset your cannot have more than one file in a block and files can be made up of multiple blocks.

A lot of home-labbers will only have one pool so this becomes and issue. 16KiB is the recommended records size for databases I believe.

I am planning ahead. I will have a mix of VMs, Apps and media on the one pool. I might have some containers once V26 is live.

You realise that recordsize is a maximum value, not a minimum or a fixed size, right?

can a single block hold more than one file ?

Not familiar with SCALE/CE. Are the VM images placed under “.truenas_containers” or “.ix-apps”? What about “ix-virt”? I thought that the user is meant to create or specify the volume or image to use, which will be its own “zvol”. The “zvol” creation wizard exposes “blocksize” options, which work similar to a dataset’s recordsize. No need to use the command-line.

V26 containers are in .truenas_containers

It’s a moot point because of this:

As for in-place modifications, perhaps there’s some truth to what you said about Apps having their own small databases. I doubt their size is enough to warrant overriding the default 128-KiB recordsize for those hidden datasets.

I thought the dataset record size denotes the block size that was read and written for each I/o operation

I thought I bookmarked it. I laid out the layman’s explanation of recordsizes and blocks. I will try to find it.

EDIT: Found it.

I will have a read. Thanks

What is the point in the record size?

Does zfs not read and write in these 128k blocks rather than variable sizes.

Does Zfs uses the recordsize to calculate checksums?

A lot of tutorial say to tune your recordsize to match that datasets purpose.

I am confused :confused:

It doesn’t matter for files smaller than 128 KiB. If the file is 8 KiB, it will only consume 8 KiB on the disk as a single block. To read its contents into RAM requires pulling this single block of 8 KiB.

If a file’s size goes beyond the recordsize (i.e, 128 KiB), the subsequent “raw” blocks will each be 128 KiB, including the very last block, even if the last block only contains a few kilobytes of file data. Why isn’t this a problem? It’s because any compression will shrink the padding of null bytes into nothing. After compression and encryption is where you get the actual size and form of the block on the disk and in the ARC.

When it comes to reading and writing, you won’t notice reading a 128-KiB block from disk, even if it is truly 128 KiB. Same thing with 1-MiB blocks. It happens fast. In fact, larger recordsizes yield better speeds and less overhead for most use-cases, since each block must be checksummed, compressed, and encrypted. Imagine doing that eight times for 1 MiB of data split into 8 x 128-KiB blocks, whereas it only needs to be done once for a single 1-MiB block if the recordsize was 1-MiB.

Where smaller recordsizes matter is for VMs, databases, and cases where a lot of true in-place modifications are being done to the data, since it reduces write-amplification and inefficient space usage for small changes inside a file.

My default recordsize is 1-MiB for all my datasets, unless I need to override a particular dataset. I have considered going to 4-MiB. While ZFS allows up to a 16-MiB recordsize, you start to see performance drop after 4-MiB, since you lose out on parallel processing.

So I am right it matters. :blush:

Most apps have databases and from what I have been reading you should set the record size to match your database page size

Most stuff above has assumed just storage and not heavy oops interactions.

So I think my initial post has some merit.

So the first block can be of variable length?

If you’re comfortable with the command-line and you know what to set for optimal performance for your apps databases:

zfs set recordsize=16k mypool/ix-apps/somedataset

Change 16k to whatever.

Not the point I was making, I think this is a missing feature and should be in the gui. I am following the mantra that true as is an appliance and is locked down.

It always is. It iterates in binary: 4K → 8K → 16K → all the way to the recordsize, where it then starts a new block. Then compression can further shrink it if the data inside is compressible. (Only the first block is truly variable if there is no compression involved.)

All subsequent blocks will use the full recordsize, but compression will deal with the padding so that no space is wasted pointlessly.

I agree. I think the “iocage” (jails) datasets exposed in the GUI in FreeNAS and TrueNAS Core was an oversight, but ironically ended up being a useful feature. :laughing:

Most importantly it allowed creating snapshot and replication tasks for your jails. I never assumed it was an oversight. That was a central part of my backup strategy.

The “app” system really needs improvement in that regard.