ZFS pool unmountable: invalid label

The hexdumps from sdj and sdk seem to be entirely missing their labels from the expected offset of 0000c000 onwards - they’re not just “corrupt”, they’re zeroed all the way to 00019000 - normally, I’d think this is a case of the wrong partition offset, and the labels would just be elsewhere.

What does the last 32MB of your partition show?

Your partition layout seems to be identical to sdc and sdd which say they’re members of zpool - but you said you expect your pool name to be zpool4t (for the 4T size of the member disks I assume) which isn’t showing anywhere in the label either.

Was this pool made initially from these five disks, or was there a replacement/swap for either expansion (replacing smaller with larger) or failure reasons?

tail -c 32M /dev/sdx | hexdump -C > sdx.txt

I uploaded it on
https://drive.g :grin:oogle.com/drive/folders/1Yrd2e2dEOhB78c0 :grin:cGHkUot2ch82SFamx?usp=sharing
I added emojis to prevent search engines from crawling.

I forgot how I formatted the drives at that time. Did I use some feature of the web UI? Back then, I formatted the hard drive because there was an issue with support for CJK filenames. After formatting, I used the following command to move the files back into the pool.
This is what I found from history which created my pool.

zpool create -f -o ashift=12 -O atime=off -o autotrim=on -O compression=on
-O normalization=formC -O utf8only=on zpool4t raidz
ata-ST4000VN008-2DR166_ZDH8HKB0 ata-ST4000VN008-2DR166_ZDH8J4R4
ata-ST4000VN008-2DR166_ZDH8J4Y4 ata-ST4000VN008-2DR166_ZDH9AX8E
ata-ST4000VN008-2DR166_ZDH9C7HR

At first, its name was “zpool”. When I created it for the second time, I thought “zpool4t” suited it better, so I used “zpool4t” in the command.

I seriously suspect I was being stupid at the time. I unassigned the ZFS pool from the web UI, but doing so didn’t actually destroy the pool. Then I directly created a new pool on top of the old pool. To be honest, I can’t remember what I did at the time.

Just getting back to this now.

I haven’t ever messed with the ZFS support in UnRAID, so I don’t know what the expected method of creation through the webUI or CLI is - but from your zpool create command:

zpool create -f -o ashift=12 -O atime=off -o autotrim=on -O compression=on
-O normalization=formC -O utf8only=on zpool4t raidz
ata-ST4000VN008-2DR166_ZDH8HKB0 ata-ST4000VN008-2DR166_ZDH8J4R4
ata-ST4000VN008-2DR166_ZDH8J4Y4 ata-ST4000VN008-2DR166_ZDH9AX8E
ata-ST4000VN008-2DR166_ZDH9C7HR

You appear to have passed full disks as opposed to partitions, and picking apart the dump of the tail end of the disks there, I’m guessing that stomped all over whatever automatic partitioning scheme was used by UnRAID in the GUI. Using the -f flag cemented it into existence as well rather than giving ZFS a chance to throw up an error along the lines of “I see existing labels, are you sure you want to do this?” so that is likely further contributing to the issues here.

You definitely created a pool on top of a pool. I can see overlapping labels with different GUIDs on sdc, see the two segments below and the different hex after the pool_guid line:

016ae030  00 00 00 24 00 00 00 20  00 00 00 04 6e 61 6d 65  |...$... ....name|
016ae040  00 00 00 09 00 00 00 01  00 00 00 07 7a 70 6f 6f  |............zpoo|
016ae050  6c 34 74 00 00 00 00 24  00 00 00 20 00 00 00 05  |l4t....$... ....|
016ae060  73 74 61 74 65 00 00 00  00 00 00 08 00 00 00 01  |state...........|
016ae070  00 00 00 00 00 00 00 01  00 00 00 20 00 00 00 20  |........... ... |
016ae080  00 00 00 03 74 78 67 00  00 00 00 08 00 00 00 01  |....txg.........|
016ae090  00 00 00 00 00 20 a2 60  00 00 00 28 00 00 00 28  |..... .`...(...(|
016ae0a0  00 00 00 09 70 6f 6f 6c  5f 67 75 69 64 00 00 00  |....pool_guid...|
016ae0b0  00 00 00 08 00 00 00 01  9e 63 c2 8e d9 ed 92 00  |.........c......|
...
01fb6030  00 00 00 24 00 00 00 20  00 00 00 04 6e 61 6d 65  |...$... ....name|
01fb6040  00 00 00 09 00 00 00 01  00 00 00 05 7a 70 6f 6f  |............zpoo|
01fb6050  6c 00 00 00 00 00 00 24  00 00 00 20 00 00 00 05  |l......$... ....|
01fb6060  73 74 61 74 65 00 00 00  00 00 00 08 00 00 00 01  |state...........|
01fb6070  00 00 00 00 00 00 00 01  00 00 00 20 00 00 00 20  |........... ... |
01fb6080  00 00 00 03 74 78 67 00  00 00 00 08 00 00 00 01  |....txg.........|
01fb6090  00 00 00 00 00 2f 84 21  00 00 00 28 00 00 00 28  |...../.!...(...(|
01fb60a0  00 00 00 09 70 6f 6f 6c  5f 67 75 69 64 00 00 00  |....pool_guid...|
01fb60b0  00 00 00 08 00 00 00 01  f9 73 fb 4e 32 db 2c 34  |.........s.N2.,4|

On sdd I see the remnants of a Windows EFI boot manager as well at the tail so this disk didn’t get cleared properly by whatever added it to a pool in the first place.

01ffbd80  74 09 b4 0e bb 07 00 cd  10 eb f2 c3 0d 0a 41 20  |t.............A |
01ffbd90  64 69 73 6b 20 72 65 61  64 20 65 72 72 6f 72 20  |disk read error |
01ffbda0  6f 63 63 75 72 72 65 64  00 0d 0a 42 4f 4f 54 4d  |occurred...BOOTM|
01ffbdb0  47 52 20 69 73 20 6d 69  73 73 69 6e 67 00 0d 0a  |GR is missing...|
01ffbdc0  42 4f 4f 54 4d 47 52 20  69 73 20 63 6f 6d 70 72  |BOOTMGR is compr|
01ffbdd0  65 73 73 65 64 00 0d 0a  50 72 65 73 73 20 43 74  |essed...Press Ct|
01ffbde0  72 6c 2b 41 6c 74 2b 44  65 6c 20 74 6f 20 72 65  |rl+Alt+Del to re|
01ffbdf0  73 74 61 72 74 0d 0a 00  8c a9 be d6 00 00 55 aa  |start.........U.|

The command that followed this line was the zpool create - just so we’re clear, did you move the files off the pool first, or did you just try to create a new pool on top of the old one? Because if the aim was just to rename the pool, that doesn’t involve create - you can just import with oldname newname and change it that way.

I wanted to recreate the pool because I wanted to add the two parameters: “-O normalization=formC -O utf8only=on”. As far as I remember, I cut all the files to another location before performing the operation I mentioned earlier. I really regret doing this through the web UI. But what’s strange is that I had restarted several times before, and there were no errors.

I don’t know why there exists a Windows EFI boot manager. I haven’t installed an operating system on it; maybe it’s the installation image of the operating system?

With the inconsistent/ambiguous labels you have, it’s likely that you were basically doing a “coin toss” on each boot, and previously you were “winning” the coin toss.

I’m not sure if there’s a way to explicitly force zpool import to look at a specific label for import - but what happens if you ask zdb -lu /dev/sdc directly, without specifying a partition? Does it show the zpool4t pool correctly there, I wonder?

Since you have full backups of the drives (bit-for-bit copies, I assume?) we might be able to take riskier steps - such as zeroing the zpool related labels on a single disk and seeing if ZFS will grab the zpool4t ones automatically.

Yes, I used dd.

losetup -f --show /mnt/disk6/MyShare/sdc_backup.img → /dev/loop2
Then zdb -lu /dev/loop2
→

failed to unpack label 0
------------------------------------
LABEL 0 (Bad label cksum)
------------------------------------
    Uberblock[32]
        magic = 0000000000bab10c
        version = 5000
        txg = 3113984
        guid_sum = 18232135290920727704
        timestamp = 1740539289 UTC = Wed Feb 26 11:08:09 2025
        bp = DVA[0]=<0:7000064000:2000> DVA[1]=<0:740005a000:2000> DVA[2]=<0:7800298000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3113984L/3113984P fill=1224 cksum=00000003034553a2:00000b9e7e717ac2:00166e1b94236fb3:1ce61945a516ed63
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
    Uberblock[36]
        magic = 0000000000bab10c
        version = 5000
        txg = 3114017
        guid_sum = 18232135290920727704
        timestamp = 1740539408 UTC = Wed Feb 26 11:10:08 2025
        bp = DVA[0]=<0:a000050000:2000> DVA[1]=<0:a800046000:2000> DVA[2]=<0:b00025e000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3114017L/3114017P fill=1215 cksum=000000042c9a71a9:000010187afde551:001f10fa19275bf4:2804bb7809e38790
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
    Uberblock[40]
        magic = 0000000000bab10c
        version = 5000
        txg = 3113890
        guid_sum = 18232135290920727704
        timestamp = 1740538051 UTC = Wed Feb 26 10:47:31 2025
        bp = DVA[0]=<0:2d8cbc74000:2000> DVA[1]=<0:43883d66000:2000> DVA[2]=<0:1c04f54000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3113890L/3113890P fill=1446 cksum=000000039fa7fe53:00000dfe274aa193:001b0bf84edbc635:22e432f2ff006ee1
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
failed to unpack label 1
failed to unpack label 2
failed to unpack label 3

Then sdi(loop3) →

failed to unpack label 0
------------------------------------
LABEL 0 (Bad label cksum)
------------------------------------
    Uberblock[32]
        magic = 0000000000bab10c
        version = 5000
        txg = 3113984
        guid_sum = 18232135290920727704
        timestamp = 1740539289 UTC = Wed Feb 26 11:08:09 2025
        bp = DVA[0]=<0:7000064000:2000> DVA[1]=<0:740005a000:2000> DVA[2]=<0:7800298000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3113984L/3113984P fill=1224 cksum=00000003034553a2:00000b9e7e717ac2:00166e1b94236fb3:1ce61945a516ed63
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
    Uberblock[36]
        magic = 0000000000bab10c
        version = 5000
        txg = 3114017
        guid_sum = 18232135290920727704
        timestamp = 1740539408 UTC = Wed Feb 26 11:10:08 2025
        bp = DVA[0]=<0:a000050000:2000> DVA[1]=<0:a800046000:2000> DVA[2]=<0:b00025e000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3114017L/3114017P fill=1215 cksum=000000042c9a71a9:000010187afde551:001f10fa19275bf4:2804bb7809e38790
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
    Uberblock[40]
        magic = 0000000000bab10c
        version = 5000
        txg = 3113890
        guid_sum = 18232135290920727704
        timestamp = 1740538051 UTC = Wed Feb 26 10:47:31 2025
        bp = DVA[0]=<0:2d8cbc74000:2000> DVA[1]=<0:43883d66000:2000> DVA[2]=<0:1c04f54000:2000> [L0 DMU objset] fletcher4 uncompressed unencrypted LE contiguous unique triple size=1000L/1000P birth=3113890L/3113890P fill=1446 cksum=000000039fa7fe53:00000dfe274aa193:001b0bf84edbc635:22e432f2ff006ee1
        mmp_magic = 00000000a11cea11
        mmp_delay = 0
        mmp_valid = 0
        checkpoint_txg = 0
        raidz_reflow state=0 off=0
        labels = 0 
failed to unpack label 1
failed to unpack label 2
failed to unpack label 3

Please tell me how to do that.

You’ll have to quite directly dd if=/dev/zero using the oseek option to skip a certain number of bs-sized blocks on the target.

I’m going to want to test this out myself to see just how it will react to a completely absent label, but you’re basically going to obliterate the ambiguous data and hope that it will pick up the correct one. That’s assuming the remaining label is intact and points to a functional/active pool.

Could you tell me the details? To be honest, I’m not familiar with file system. How can I determine exactly where I should overwrite? Is there any way to figure it out? Thank u really much.

Can you tell me where I can find documentation about the binary-level structure of ZFS? I’ve been searching for a long time but haven’t found anything.