Pool disappears and disks are unattached

How odd.

We will have to try with a command line. This can be done from Clonezilla at a prompt, once you are absolutely certain you have identified the correct source and destination disks

dd if=/dev/sourcedisk of=/dev/destinationdisk bs=32M status=progress

This should clone the drive. Once you’ve done that, remove the USB/portable SSD, and then try the zhack label repair command - again, ensuring you are targeting the correct disk only.

Get stuck at this point

I tried testing with bs=4M but it also failed at the same point

Maybe creating a partition of the same size and copying one partition and then another in other partiotn of the same size will make it work? Sorry if I don’t explain myself.

EDIT: recibe an error when try to copy from /dev/sdf2 to /dev/sde2

dd: error reading '/dev/sdf2': Input/output error

i think i have a problem with that specific partiton but idk why and im not sure if of software or hardware error

Is it possible to recover both of these from this disk with the other two disks? As if it were a raid5

The drive may have permanent damage at a specific position.
So you’re deep into “Hail Mary” territory. I suppose you could try zhack on the drive without a cloned copy to go back. Or maybe, if you’re potentially willing to spend $399 to recover some data, first plug the drives into a Windows computer and scan with Klennet ZFS Recovery to see what could be salvaged (scanning is free; actual recovery will require the paid licence).
Wait for @HoneyBadger 's advice.

You have a stripe: No redundancy, no recovery from other drives.

You may need to use a program such as ddrescue that can handle automatically skipping over bad or damaged sectors - but as this is a stripe you may have permanent data loss regardless.

Checking with Klennet’s ZFS Recovery may be an option (as it’s free to check) - if you are able to ddrescue the drive then definitely attempt the zhack label repair - but if you are not, then you have a bit of a dilemma I’m afraid.

i recibe this result with ZFS Recovery, dont look bad, r?

Is there an alternative to this program that doesn’t cost 400€? free or at least much cheaper?

1.2 TB out of a 2*0.5+1 TB stripe? Not bad indeed.

Not as far as I know. Klennet is pretty much alone on this very specific market.

Assuming that the maximum usage of the stripe was 80% (and it could have been lower) that would be a maximum of 1600GB of data (or less if Klennet is reporting GiB but saying GB), and it could easily be only 60% which would be 1200GB, so indeed Klennet may be able to recover almost all (or even actually all) the data.

Sorry for bumping this thread, but I’ve made good progress. Ever since this issue occurred, I felt very unmotivated and was busy with my studies. Using the tool ReclaiMe Pro, I was able to export the image from the problematic Toshiba disk to a new hard drive, which the server now recognizes. When I reconnected it to the server, it detected the disk and I managed to import the pool without any major issues (except that a few files from some apps were corrupted, but none of my personal files were affected).

However, after rebooting the server and trying to fix the Kubernetes error, the server no longer detected the pool when I turned it back on. It won’t let me import it, although it does let me add the existing disks to another pool. I haven’t touched anything yet, because ChatGPT told me that I would lose the data (though I doubt it can help or give correct information in this case).

So I ask here: if I import the disks, will I lose the data on them?
Thanks

You should probably run the commands from the second forum post and post them back using Preformatted Text (Ctrl+e) </> on the tool bar. We need to start over if anything has changed.

Ah yes, ChatGPT that well meaning but bumbling idiot with a hallucination problem - everyone’s favourite self-proclaimed expert. Except in this case, it is correct that if you do the wrong thing then you will indeed lose all your data.

My guess is that if you are able to import the pool then you probably won’t lose all your data, but if the UI is reporting that your disks are available to create a new pool then I doubt very much that an import will work. But worth a try.

I sincerely hope that having recovered your data the first time you learned your lesson and migrated from a simple stripe to a redundant configuration.

So, you will need to provide us with a new set of diagnostic information and hope that genuine experts here (of which I am NOT one) can assist you with getting you access to your data yet again.

Please run the following commands and post the output here, with each output in a separate </> box:

  • lsblk -bo NAME,LABEL,MAJ:MIN,TRAN,ROTA,ZONED,VENDOR,MODEL,SERIAL,PARTUUID,START,SIZE,PARTTYPENAME
  • sudo zpool status -vsc upath,media,lsblk,serial,smartx,smart
  • sudo zpool import
  • lspci
  • sudo sas2flash -list
  • sudo sas3flash -list
  • sudo storcli show all
  • for disk in /dev/sd?1; do; sudo zdb -l $disk; done
  • for disk in /dev/sd?; do; sudo smartctl -x $disk; done

It is to create a new pool, or add them to an existing one, so I don’t know if maybe adding them would work, because the first time when I connected the failed disk it let me import the pool from import pools, but now it doesn’t, I understand that the pool is already imported, it just doesn’t detect the disks, I still await your response before doing anything

There are all commands, and the outputs, but someone they give an error

root@truenas:~# lsblk -bo NAME,LABEL,MAJ:MIN,TRAN,ROTA,ZONED,VENDOR,MODEL,SERIAL,PARTUUID,START,SIZE,PARTTYPENAME
NAME     LABEL     MAJ:MIN TRAN   ROTA ZONED VENDOR   MODEL                   SERIAL          PARTUUID                                START          SIZE PARTTYPENAME
sda      data        8:0   sata      1 none  ATA      ST2000DM008-2UB102      WFL8RZ7L                                                      2000398934016 
sdb                  8:16  sata      0 none  ATA      EMTEC X250 256GB        A2205CW03290                                                   256060514304 
├─sdb1               8:17            0 none                                                   3bcb9985-ce6f-11ee-b719-d45d64208664       40     272629760 EFI System
├─sdb2   boot-pool   8:18            0 none                                                   3bd3eff9-ce6f-11ee-b719-d45d64208664 34086952  238605565952 FreeBSD ZFS
└─sdb3               8:19            0 none                                                   3bd0a529-ce6f-11ee-b719-d45d64208664   532520   17179869184 FreeBSD swap
  └─sdb3           253:0             0 none                                                                                                   17179869184 
sdc                  8:32  sata      1 none  ATA      Hitachi HDS721050CLA362 JPB530HA06UJ1B                                                 500107862016 
├─sdc1               8:33            1 none                                                   b2112368-ce76-11ee-9fb8-d45d64208664      128    2147483648 FreeBSD swap
└─sdc2   data        8:34            1 none                                                   b24ce596-ce76-11ee-9fb8-d45d64208664  4194432  497960292352 FreeBSD ZFS
sdd                  8:48  sata      1 none  ATA      WDC WD5000AAKX-22ERMA0  WD-WCC2E0LH1EET                                                500107862016 
├─sdd1               8:49            1 none                                                   b226e7fc-ce76-11ee-9fb8-d45d64208664      128    2147483648 FreeBSD swap
└─sdd2   data        8:50            1 none                                                   b25df2bc-ce76-11ee-9fb8-d45d64208664  4194432  497960292352 FreeBSD ZFS
root@truenas:~# sudo zpool status -vsc upath,media,lsblk,serial,smartx,smart
Can't run -c with root privileges unless ZPOOL_SCRIPTS_AS_ROOT is set.
root@truenas:~# sudo zpool import
   pool: data
     id: 15103091714514370022
  state: FAULTED
status: The pool metadata is corrupted.
 action: The pool cannot be imported due to damaged devices or data.
        The pool may be active on another system, but can be imported using
        the '-f' flag.
   see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-72
 config:

        data                                    FAULTED  corrupted data
          sda                                   ONLINE
          b25df2bc-ce76-11ee-9fb8-d45d64208664  ONLINE
          b24ce596-ce76-11ee-9fb8-d45d64208664  ONLINE
root@truenas:~# lspci
00:00.0 Host bridge: Intel Corporation 8th Gen Core Processor Host Bridge/DRAM Registers (rev 0d)
00:01.0 PCI bridge: Intel Corporation 6th-10th Gen Core Processor PCIe Controller (x16) (rev 0d)
00:14.0 USB controller: Intel Corporation 200 Series/Z370 Chipset Family USB 3.0 xHCI Controller
00:16.0 Communication controller: Intel Corporation 200 Series PCH CSME HECI #1
00:17.0 SATA controller: Intel Corporation 200 Series PCH SATA controller [AHCI mode]
00:1c.0 PCI bridge: Intel Corporation 200 Series PCH PCI Express Root Port #5 (rev f0)
00:1c.7 PCI bridge: Intel Corporation 200 Series PCH PCI Express Root Port #8 (rev f0)
00:1d.0 PCI bridge: Intel Corporation 200 Series PCH PCI Express Root Port #11 (rev f0)
00:1f.0 ISA bridge: Intel Corporation Device a2ca
00:1f.2 Memory controller: Intel Corporation 200 Series/Z370 Chipset Family Power Management Controller
00:1f.3 Audio device: Intel Corporation 200 Series PCH HD Audio
00:1f.4 SMBus: Intel Corporation 200 Series/Z370 Chipset Family SMBus Controller
01:00.0 VGA compatible controller: NVIDIA Corporation GP108 [GeForce GT 1030] (rev a1)
01:00.1 Audio device: NVIDIA Corporation GP108 High Definition Audio Controller (rev a1)
03:00.0 Ethernet controller: Realtek Semiconductor Co., Ltd. RTL8111/8168/8411 PCI Express Gigabit Ethernet Controller (rev 15)
root@truenas:~# sudo sas2flash -list
LSI Corporation SAS2 Flash Utility
Version 20.00.00.00 (2014.09.18) 
Copyright (c) 2008-2014 LSI Corporation. All rights reserved 

        No LSI SAS adapters found! Limited Command Set Available!
        ERROR: Command Not allowed without an adapter!
        ERROR: Couldn't Create Command -list
        Exiting Program
root@truenas:~# sudo sas3flash -list
Avago Technologies SAS3 Flash Utility
Version 16.00.00.00 (2017.05.02) 
Copyright 2008-2017 Avago Technologies. All rights reserved.

        No Avago SAS adapters found! Limited Command Set Available!
        ERROR: Command Not allowed without an adapter!
        ERROR: Couldn't Create Command -list
        Exiting Program.
root@truenas:~# sudo storcli show all
CLI Version = 007.1504.0000.0000 June 22, 2020
Operating system = Linux 6.1.74-production+truenas
Status Code = 0
Status = Success
Description = None

Number of Controllers = 0
Host Name = truenas
Operating System  = Linux 6.1.74-production+truenas
root@truenas:~# for disk in /dev/sd?1; do; sudo zdb -l $disk; done
-bash: syntax error near unexpected token `;'
root@truenas:~# for disk in /dev/sd?; do; sudo smartctl -x $disk; done
-bash: syntax error near unexpected token `;'

Jumping in late, the core issue is that the labels at the START of sda2 are MISSING:

I do not know off the top of my head how to re-create a missing disk label. Perhaps @HoneyBadger or @NickF1227 do?

Apologies. Try:

  • /sbin/zpool status -vsc upath,media,lsblk,serial,smartx,smart

I have no idea why the last two didn’t work - they work fine on my own system running EE.

However try:

  • sudo zdb -l /sda
  • sudo zdb -l /sdc2
  • sudo zdb -l /sdd2

It should be noted that when you managed to recover the data and replace the previous disk you used a full disk and not a partition - hence the mixture of disks and partitions in the above list of zdb commands.

And when you have your system running again I would suggest that you immediately rebuild it with redundancy.

This is ancient information from an old, already solved issue. Hence my request for fresh diagnostic info.

Im running this on the web shell, maybe that the problem? im not sure btw i recibe the same error

root@truenas:~# /sbin/zpool status -vsc upath,media,lsblk,serial,smartx,smart
Can't run -c with root privileges unless ZPOOL_SCRIPTS_AS_ROOT is set.
root@truenas:~# sudo zdb -l /sda
cannot open '/sda': No such file or directory
root@truenas:~# sudo zdb -l /sdc2
cannot open '/sdc2': No such file or directory
root@truenas:~# sudo zdb -l /sdd2
cannot open '/sdd2': No such file or directory

if helps, there is the fdisk -l output

Disk /dev/sda: 1.82 TiB, 2000398934016 bytes, 3907029168 sectors
Disk model: ST2000DM008-2UB1
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 4096 bytes
I/O size (minimum/optimal): 4096 bytes / 4096 bytes


Disk /dev/sdc: 465.76 GiB, 500107862016 bytes, 976773168 sectors
Disk model: Hitachi HDS72105
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disklabel type: gpt
Disk identifier: B1C8748B-CE76-11EE-9FB8-D45D64208664

Device       Start       End   Sectors   Size Type
/dev/sdc1      128   4194431   4194304     2G FreeBSD swap
/dev/sdc2  4194432 976773127 972578696 463.8G FreeBSD ZFS


Disk /dev/sdb: 238.47 GiB, 256060514304 bytes, 500118192 sectors
Disk model: EMTEC X250 256GB
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disklabel type: gpt
Disk identifier: 3BCAEEDA-CE6F-11EE-B719-D45D64208664

Device        Start       End   Sectors   Size Type
/dev/sdb1        40    532519    532480   260M EFI System
/dev/sdb2  34086952 500113447 466026496 222.2G FreeBSD ZFS
/dev/sdb3    532520  34086951  33554432    16G FreeBSD swap

Partition table entries are not in disk order.


Disk /dev/sdd: 465.76 GiB, 500107862016 bytes, 976773168 sectors
Disk model: WDC WD5000AAKX-2
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disklabel type: gpt
Disk identifier: B1FAE7F8-CE76-11EE-9FB8-D45D64208664

Device       Start       End   Sectors   Size Type
/dev/sdd1      128   4194431   4194304     2G FreeBSD swap
/dev/sdd2  4194432 976773127 972578696 463.8G FreeBSD ZFS


Disk /dev/mapper/sdb3: 16 GiB, 17179869184 bytes, 33554432 sectors
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes

ofc, my new plan its build an raid 5 with 3 disk of 4tb each one

D’oh. My stupidity!!!

Try:

  • sudo zdb -l /dev/sda
  • sudo zdb -l /dev/sdc2
  • sudo zdb -l /dev/sdd2

And sorry - the fdisk output is more like the lsblk than zdb in terms of diagnostic information.

No such thing in TrueNAS as RAID5 - I think you mean RAIDZ1.

I’m not sure if this matters, but I only cloned the data partition — the swap partition wasn’t cloned to the new disk. I’m not sure if this affects anything, because the first time I did it, there was no problem detecting the pool and importing it.

I think there are no longer any problems with the disk, apart from the errors that were already there from the beginning (even with those, I was able to import the pool). That makes me think that the new error is no longer on the disk, but rather in TrueNAS or (hopefully not) in the other disks.

It seems the pool is actually imported, but it’s not detecting the disks. I’m not sure if I’m right or not, but I just mention it in case it helps in finding a new solution or not. Thank you very much.

Here are the commands:

root@truenas:~# sudo zdb -l /dev/sda
------------------------------------
LABEL 0 
------------------------------------
    version: 5000
    name: 'data'
    state: 0
    txg: 3303346
    pool_guid: 15103091714514370022
    errata: 0
    hostid: 859060325
    hostname: 'truenas'
    top_guid: 12879759490428939429
    guid: 12879759490428939429
    vdev_children: 3
    vdev_tree:
        type: 'disk'
        id: 0
        guid: 12879759490428939429
        path: '/dev/disk/by-partuuid/b2639d11-ce76-11ee-9fb8-d45d64208664'
        phys_path: 'id1,enc@n3061686369656d30/type@0/slot@1/elmdesc@Slot_00/p2'
        metaslab_array: 270
        metaslab_shift: 33
        ashift: 12
        asize: 998052462592
        is_log: 0
        DTL: 29176
        create_txg: 4
    features_for_read:
        com.delphix:hole_birth
        com.delphix:embedded_data
    labels = 0 1 
failed to unpack label 2
failed to unpack label 3
root@truenas:~# sudo zdb -l /dev/sdc2
------------------------------------
LABEL 0 
------------------------------------
    version: 5000
    name: 'data'
    state: 0
    txg: 3382283
    pool_guid: 15103091714514370022
    errata: 0
    hostid: 859060325
    hostname: 'truenas'
    top_guid: 5094482715523507685
    guid: 5094482715523507685
    vdev_children: 3
    vdev_tree:
        type: 'disk'
        id: 2
        guid: 5094482715523507685
        path: '/dev/disk/by-partuuid/b24ce596-ce76-11ee-9fb8-d45d64208664'
        phys_path: 'id1,enc@n3061686369656d30/type@0/slot@4/elmdesc@Slot_03/p2'
        metaslab_array: 256
        metaslab_shift: 32
        ashift: 12
        asize: 497955373056
        is_log: 0
        DTL: 29175
        create_txg: 4
    features_for_read:
        com.delphix:hole_birth
        com.delphix:embedded_data
    labels = 0 1 2 3
root@truenas:~# sudo zdb -l /dev/sdd2
------------------------------------
LABEL 0 
------------------------------------
    version: 5000
    name: 'data'
    state: 0
    txg: 3382283
    pool_guid: 15103091714514370022
    errata: 0
    hostid: 859060325
    hostname: 'truenas'
    top_guid: 7372822762195821153
    guid: 7372822762195821153
    vdev_children: 3
    vdev_tree:
        type: 'disk'
        id: 1
        guid: 7372822762195821153
        path: '/dev/disk/by-partuuid/b25df2bc-ce76-11ee-9fb8-d45d64208664'
        phys_path: 'id1,enc@n3061686369656d30/type@0/slot@3/elmdesc@Slot_02/p2'
        metaslab_array: 264
        metaslab_shift: 32
        ashift: 12
        asize: 497955373056
        is_log: 0
        DTL: 29174
        create_txg: 4
    features_for_read:
        com.delphix:hole_birth
        com.delphix:embedded_data
    labels = 0 1 2 3

This likely isn’t helping anything, as it’s possible that the cloning tool didn’t capture the entirety of the ZFS labels at the start/end of the partition. A full disk clone would have been a better option here.

But let’s attack the root of the problem here. You have a striped pool with no redundancy that’s expecting three valid disks for import.

    vdev_children: 3

You have one disk, /dev/sda with the partition-only-clone, that is way out of alignment in terms of transaction group number

That disk/member is roughly 80,000 transactions behind the other two. As such, it’s considered invalid by ZFS - and as such, you have a missing disk from a stripe, making your pool un-importable.

I’m not sure why it would have let you import the pool successfully in the prior case - perhaps there’s a better, valid label elsewhere on the disk?

Your /dev/sda appears to have an empty partition table:

And I’m just now noticing that you’ve chosen an SMR Barracuda to replace it which is not going to help you any, and could even be part of the issue if it’s had a sudden power cut. Do you have a non-SMR drive you can re-clone the entire Toshiba to?

I tried to clone the disk with BalenaEtcher directly, but after a few seconds of starting to work, the program stopped working, remained grayed out and the task manager did not show any usage of both disks.

I used the same method to export the disk image, although it’s true that the second time I tried to write the image to the new drive, an error popped up. I fixed it by reinstalling the program. I don’t remember exactly what it said, but it was something related to metadata or something like that.

I don’t have any other disk, I could try using the SSD disk that I tried to use a while ago, but since it has exactly the same size, there might be problems.