Resilver fails/hangs but all drives pass SMART long test

I have a 7x18TB raidz2 array connected to a Dell PERC H310 HBA, originally created under Truenas CORE, currently running under SCALE 25.10.2.1
recently 2 of the drives started reporting errors, so I disconnected the failing drives, RMA’d them, installed replacements, and resilver started automatically.
Initially everything seemed fine, but then lots of checksum errors (~30-40k) started cropping up across most of the disks, followed by a handful (3 to ~700) read/write errors on 2 or 3 of the disks, and resilver seems to be hung at ~9.6% completion. After reboot disk I/O is initially high (200MB/s) but once the resilver stalls disk I/O drops to near zero. Completion time initially shows a reasonable number (~32 hours) but once it stalls the time seems to increase indefinitely (currently showing 17 days 16 hours, with resilver progress hung at 9.6%)
I was able to export the pool and ran long SMART tests on all the drives and they all passed.
When I first installed the new drives I was able to access the pool and backup the little bit of data that would be difficult to replace, but after subsequent reboots accessing the pool through SMB shares shows lots of folder/files missing entirely, and of the ones that remain, very few seem to still be accessible, the vast majority give me an error on attempting to open/copy them.
I had resilvered this pool several times before under truenas CORE without issue, this is my first time attempting a drive replacement under SCALE.
So far my leading theories are:

  1. drive failure that SMART doesn’t detect
  2. failing HBA/cables
  3. array somehow got into an inconsistent state

Any ideas on how I might save this? Data is replaceable it would just be inconvenient, and at this point I’m assuming I’ve lost everything anyway.
I can try replacing the HBA/cables but it will take me probably a week to get new parts.
Alternatively I’ve wondered if installing truenas CORE on another boot drive and attempting resilver through that might make a difference, though I doubt it.

Welcome to the forums!

It is always helpful to show the output of zpool status in these cases. Plus, for CE / SCALE, the output of

lsblk -o NAME,MODEL,SERIAL,LABEL,UUID,PARTUUID,TYPE

Both in CODE tags.

Next, during resilvers the HBA could be getting quite hot. Thus, throwing a lot of checksum errors. Is your H310 HBA adequately cooled? And is it in IT mode?

Thanks for the quick reply, here’s more info:

zpool status:

root@truenas[~]# zpool status
  pool: CLOSED
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
  scan: resilver in progress since Wed Aug 12 12:45:07 2026
        9.92T / 81.3T scanned at 543M/s, 7.80T / 81.3T issued at 38.3M/s
        2.14T resilvered, 9.60% done, 23 days 06:59:59 to go
config:

        NAME                                        STATE     READ WRITE CKSUM
        CLOSED                                      DEGRADED     0     0     0
          raidz2-0                                  DEGRADED   737    60     0
            replacing-0                             DEGRADED 32.5K 32.0K    16
              sdc2                                  OFFLINE      0     0     0
              c298d8c6-de53-4f07-ae70-cedad352a4b2  DEGRADED     3 33.2K     0  too many errors
            6e431bf6-b111-11ee-80df-00155d011501    ONLINE       0     0     0
            8a006ac7-9881-11ef-a106-1cfd0876a4af    ONLINE       0     0     0
            f0bc3ab2-2d63-11f1-bf17-7085c2aafbee    ONLINE       0     0     0  (resilvering)
            replacing-4                             DEGRADED   844   173    16
              sdb2                                  OFFLINE      0     0     0
              cb031d0b-7773-4da2-8e3a-a49e769848d6  ONLINE     662   173     0  (resilvering)
            f708bd25-a790-11f0-9bc1-c1f323032d4a    DEGRADED    17 32.6K     0  too many errors
            sdf2                                    ONLINE       0     0     0  (resilvering)
lsblk -o NAME,MODEL,SERIAL,LABEL,UUID,PARTUUID,TYPE
NAME        MODEL                 SERIAL              LABEL     UUID                                 PARTUUID                             TYPE
sda         OOS18000G             0008EKD3                                                                                                disk
└─sda1                                                CLOSED    4785475302194226398                  cb031d0b-7773-4da2-8e3a-a49e769848d6 part
sdb         OOS18000G             0003QQ7Z                                                                                                disk
├─sdb1                                                                                               f6f6e353-a790-11f0-9bc1-c1f323032d4a part
└─sdb2                                                CLOSED    4785475302194226398                  f708bd25-a790-11f0-9bc1-c1f323032d4a part
sdc         OOS18000G             0008GH6D                                                                                                disk
└─sdc1                                                CLOSED    4785475302194226398                  c298d8c6-de53-4f07-ae70-cedad352a4b2 part
sdd         OOS18000G             0008R7ZC                                                                                                disk
├─sdd1                                                                                               6d93de1a-b111-11ee-80df-00155d011501 part
└─sdd2                                                CLOSED    4785475302194226398                  6e431bf6-b111-11ee-80df-00155d011501 part
sde         TOSHIBA MD09ACA18TR   Y5K2A00BTK2H                                                                                            disk
├─sde1                                                                                               f0a357cb-2d63-11f1-bf17-7085c2aafbee part
└─sde2                                                CLOSED    4785475302194226398                  f0bc3ab2-2d63-11f1-bf17-7085c2aafbee part
sdf         WDC WD180EDGZ-11B2DA0 3FHZD5UT                                                                                                disk
├─sdf1                                                                                               53a8eba6-e74f-11f0-97f6-c377a88a3dbd part
└─sdf2                                                CLOSED    4785475302194226398                  53c35297-e74f-11f0-97f6-c377a88a3dbd part
sdg         SATA SSD              19013110240205                                                                                          disk
├─sdg1                                                EFI       4224-8A00                            0608012f-b1bc-11ee-bbcc-a4badb2c639f part
├─sdg2                                                boot-pool 4269388194816694075                  061016ed-b1bc-11ee-bbcc-a4badb2c639f part
└─sdg3                                                swap0     fee1e0bc-5a06-90b4-146a-7154df06a5fc 060be06a-b1bc-11ee-bbcc-a4badb2c639f part
sdh         OOS18000G             00014KMG                                                                                                disk
├─sdh1                                                                                               89e5b034-9881-11ef-a106-1cfd0876a4af part
└─sdh2                                                CLOSED    4785475302194226398                  8a006ac7-9881-11ef-a106-1cfd0876a4af part
sdi         SATA SSD              19013110240670                                                                                          disk
├─sdi1                                                EFI       4216-02F4                            05fc17a6-b1bc-11ee-bbcc-a4badb2c639f part
├─sdi2                                                boot-pool 4269388194816694075                  060337a0-b1bc-11ee-bbcc-a4badb2c639f part
└─sdi3                                                swap0     fee1e0bc-5a06-90b4-146a-7154df06a5fc 06005270-b1bc-11ee-bbcc-a4badb2c639f part
nvme0n1     JAJP600M4TB           C24ED55011141129001                                                                                     disk
├─nvme0n1p1                                                                                          80105f51-b1ec-11ee-9cda-a4badb2c639f part
└─nvme0n1p2                                           nvme      18235672659811483595                 80146394-b1ec-11ee-9cda-a4badb2c639f part

as for HBA cooling it’s hard to say, the HBA itself does not have a heatsink but I’ve cranked the case fans to max through the BIOS just in case (it didn’t seem to make any difference) and yes its in IT mode

if I’m understanding these error counts correctly, I’m guessing f708bd25-a790-11f0-9bc1-c1f323032d4a and the 2 new drives are bad

if that’s the case, can I create a new pool with 4 drives in raidz2 and expand it later once I get replacements? I’ve read somewhere doing this might give less pool capacity than if I were to create a new pool with all 7 drives present.

Welcome to the forums

Is this a used HBA? The Dell Perc 310 should have a heatsink on the main chip at a minimum. Do an internet search for “Dell Perc 310” and look at some of the images.

Standard case fans alone typically do not offer enough cooling. Cooling is measured in how many linear feet of air pass by the card, and it is a lot. And I’m not saying this is your problem, BUT we have seen this many times before.

A question of my own here:
You stated that you started with TrueNAS CORE, but you did not provide any duration that the system worked.
Q: How long has the system operated before failures started to occur? Days, Weeks, Months?
Q: Was it stable on CORE? Also the duration it was running CORE please.

I ask because if this is a new installation, it changes how troubleshooting is addressed. We will start with the common things but if your system was just built, then you may have bad components.

You said SMART passes on all the drives, that is a good data point.
Your zpool status output shows you have a lot going on. In the future, replace only ONE drive, let it RESILVER, then replace the next failing drive. This is part of the cause for the very long resilvering time.

Do you have a backup of all your important data? Hopefully the answer is Yes.

Good luck, hope to see some positive news.

Ah, you’re correct, the HBA actually does have a small heatsink on the main chip, I misremembered.

I built the system in Jan or Feb ‘24, originally in an old Dell Poweredge t310 system. CORE was very stable in terms of uptime, but I did suffer 4 drive failures over the 2 years since, prior to this event. Drives are refurbished so I assume I just got unlucky. Thankfully the vendor sold them with 5 year warranty and has been very responsive in replacing them without issue.
In march of this year the PSU in the t310 died and I decided consolidate the system with components from my jellyfin machine in a Fractal Design Define 7 case. I also switched to truenas SCALE at this time as I had to run jellyfin as a SCALE app. It was fairly stable since, though I did have a few occasions where the system hanged (which seemingly never happened in CORE).

In the future, replace only ONE drive, let it RESILVER, then replace the next failing drive. This is part of the cause for the very long resilvering time.

I will keep this in mind, I had thought it would be better to install both simultaneously, but oh well.
I do have the most important data backed up, otherwise I’ve largely used this as a giant pool for torrents (which I can fairly easily replace) and backups of my bluray collection (which I still have and can re-rip).

Your Poweredge T310 may have had better cooling for your HBA over the Define 7 case.

you may be correct. I didn’t think it was getting much airflow in the old system but I just checked and it was pretty hot to the touch.
I attempted the tried and true method of “take the side panel off and point a box fan at it” (/s) and after reboot my files are all back and accessible again, so I was able to backup a little bit more.
resilver time still keeps going up and one of the new drives is still showing errors though. Is there any way to disconnect that drive and continue resilver with the rest?

I don’t think you can remove any drives with your current pool status. Removing a drive may cause pool failure.
You can post back a zpool status -v CLOSED if you still want to try. It will allow others to comment on what it returned.

actually I was able to just disconnect the new drive showing errors, and resilver appears to be proceeding normally. Disk I/O looks good and time remaining shows ~34 hours

zpool status -v CLOSED shows some permanent errors on some torrent files but assuming the resilver eventually completes I should be able to repair those with a recheck:

root@truenas[~]# zpool status -v CLOSED
  pool: CLOSED
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
  scan: resilver in progress since Fri Aug 14 11:58:55 2026
        3.20T / 81.3T scanned at 4.70G/s, 494G / 81.3T issued at 727M/s
        66.6G resilvered, 0.59% done, 1 days 08:22:52 to go
config:

        NAME                                        STATE     READ WRITE CKSUM
        CLOSED                                      DEGRADED     0     0     0
          raidz2-0                                  DEGRADED     0     0     0
            replacing-0                             UNAVAIL      0     0     0  insufficient replicas
              sdc2                                  OFFLINE      0     0     0
              10567549841703865306                  UNAVAIL      0     0     0  was /dev/disk/by-partuuid/c298d8c6-de53-4f07-ae70-cedad352a4b2
            6e431bf6-b111-11ee-80df-00155d011501    ONLINE       0     0     0
            8a006ac7-9881-11ef-a106-1cfd0876a4af    ONLINE       0     0     0
            f0bc3ab2-2d63-11f1-bf17-7085c2aafbee    ONLINE       0     0     0
            replacing-4                             DEGRADED     0     0   244
              sdb2                                  OFFLINE      0     0     0
              cb031d0b-7773-4da2-8e3a-a49e769848d6  ONLINE       0     0     0  (resilvering)
            f708bd25-a790-11f0-9bc1-c1f323032d4a    ONLINE       0     0     0
            sdg2                                    ONLINE       0     0     0

errors: Permanent errors have been detected in the following files:

        CLOSED/fixed_storage/torrents:<0x0>
        (whole bunch of files with errors, all torrents)

This puts both HBA cooling and insufficient PSU power on the HDD rails as possible suspects.

This is metadata corruption. You might resolve it by deleting the whole torrents directory—or not even.
Raidz2 is not an ideal geometry for the many small writes that active torrents receive.

I’ve been concerned about the PSU as well, my old dell poweredge was only rated for like 350w and seemed to work OK, while the new system is 650w. neither had enough sata power connectors for all the drives so I just used an expander to connect them all. I know that’s not ideal but I wasn’t able to find a PSU with so many SATA power connectors. I have the system connected to a kill-a-watt meter and the whole system is currently drawing 95w even with the resilver running.

As for the write issue, I’ve been aware of this (downloading directly to the pool killed performance) so I have qbittorrent set to download files to a 4TB nvme drive on a pcie card and then move them after completion.

EDIT: the specific PSU I’m using is my old OCZ-ZS650W, which appears to be rated for 22A / 130W on 5V rail, so hopefully OK but definitely worth upgrading. the system peaks at ~103w on bootup according to the kill-a-watt.