I have a 7x18TB raidz2 array connected to a Dell PERC H310 HBA, originally created under Truenas CORE, currently running under SCALE 25.10.2.1
recently 2 of the drives started reporting errors, so I disconnected the failing drives, RMA’d them, installed replacements, and resilver started automatically.
Initially everything seemed fine, but then lots of checksum errors (~30-40k) started cropping up across most of the disks, followed by a handful (3 to ~700) read/write errors on 2 or 3 of the disks, and resilver seems to be hung at ~9.6% completion. After reboot disk I/O is initially high (200MB/s) but once the resilver stalls disk I/O drops to near zero. Completion time initially shows a reasonable number (~32 hours) but once it stalls the time seems to increase indefinitely (currently showing 17 days 16 hours, with resilver progress hung at 9.6%)
I was able to export the pool and ran long SMART tests on all the drives and they all passed.
When I first installed the new drives I was able to access the pool and backup the little bit of data that would be difficult to replace, but after subsequent reboots accessing the pool through SMB shares shows lots of folder/files missing entirely, and of the ones that remain, very few seem to still be accessible, the vast majority give me an error on attempting to open/copy them.
I had resilvered this pool several times before under truenas CORE without issue, this is my first time attempting a drive replacement under SCALE.
So far my leading theories are:
drive failure that SMART doesn’t detect
failing HBA/cables
array somehow got into an inconsistent state
Any ideas on how I might save this? Data is replaceable it would just be inconvenient, and at this point I’m assuming I’ve lost everything anyway.
I can try replacing the HBA/cables but it will take me probably a week to get new parts.
Alternatively I’ve wondered if installing truenas CORE on another boot drive and attempting resilver through that might make a difference, though I doubt it.
Next, during resilvers the HBA could be getting quite hot. Thus, throwing a lot of checksum errors. Is your H310 HBA adequately cooled? And is it in IT mode?
root@truenas[~]# zpool status
pool: CLOSED
state: DEGRADED
status: One or more devices is currently being resilvered. The pool will
continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
scan: resilver in progress since Wed Aug 12 12:45:07 2026
9.92T / 81.3T scanned at 543M/s, 7.80T / 81.3T issued at 38.3M/s
2.14T resilvered, 9.60% done, 23 days 06:59:59 to go
config:
NAME STATE READ WRITE CKSUM
CLOSED DEGRADED 0 0 0
raidz2-0 DEGRADED 737 60 0
replacing-0 DEGRADED 32.5K 32.0K 16
sdc2 OFFLINE 0 0 0
c298d8c6-de53-4f07-ae70-cedad352a4b2 DEGRADED 3 33.2K 0 too many errors
6e431bf6-b111-11ee-80df-00155d011501 ONLINE 0 0 0
8a006ac7-9881-11ef-a106-1cfd0876a4af ONLINE 0 0 0
f0bc3ab2-2d63-11f1-bf17-7085c2aafbee ONLINE 0 0 0 (resilvering)
replacing-4 DEGRADED 844 173 16
sdb2 OFFLINE 0 0 0
cb031d0b-7773-4da2-8e3a-a49e769848d6 ONLINE 662 173 0 (resilvering)
f708bd25-a790-11f0-9bc1-c1f323032d4a DEGRADED 17 32.6K 0 too many errors
sdf2 ONLINE 0 0 0 (resilvering)
NAME MODEL SERIAL LABEL UUID PARTUUID TYPE
sda OOS18000G 0008EKD3 disk
└─sda1 CLOSED 4785475302194226398 cb031d0b-7773-4da2-8e3a-a49e769848d6 part
sdb OOS18000G 0003QQ7Z disk
├─sdb1 f6f6e353-a790-11f0-9bc1-c1f323032d4a part
└─sdb2 CLOSED 4785475302194226398 f708bd25-a790-11f0-9bc1-c1f323032d4a part
sdc OOS18000G 0008GH6D disk
└─sdc1 CLOSED 4785475302194226398 c298d8c6-de53-4f07-ae70-cedad352a4b2 part
sdd OOS18000G 0008R7ZC disk
├─sdd1 6d93de1a-b111-11ee-80df-00155d011501 part
└─sdd2 CLOSED 4785475302194226398 6e431bf6-b111-11ee-80df-00155d011501 part
sde TOSHIBA MD09ACA18TR Y5K2A00BTK2H disk
├─sde1 f0a357cb-2d63-11f1-bf17-7085c2aafbee part
└─sde2 CLOSED 4785475302194226398 f0bc3ab2-2d63-11f1-bf17-7085c2aafbee part
sdf WDC WD180EDGZ-11B2DA0 3FHZD5UT disk
├─sdf1 53a8eba6-e74f-11f0-97f6-c377a88a3dbd part
└─sdf2 CLOSED 4785475302194226398 53c35297-e74f-11f0-97f6-c377a88a3dbd part
sdg SATA SSD 19013110240205 disk
├─sdg1 EFI 4224-8A00 0608012f-b1bc-11ee-bbcc-a4badb2c639f part
├─sdg2 boot-pool 4269388194816694075 061016ed-b1bc-11ee-bbcc-a4badb2c639f part
└─sdg3 swap0 fee1e0bc-5a06-90b4-146a-7154df06a5fc 060be06a-b1bc-11ee-bbcc-a4badb2c639f part
sdh OOS18000G 00014KMG disk
├─sdh1 89e5b034-9881-11ef-a106-1cfd0876a4af part
└─sdh2 CLOSED 4785475302194226398 8a006ac7-9881-11ef-a106-1cfd0876a4af part
sdi SATA SSD 19013110240670 disk
├─sdi1 EFI 4216-02F4 05fc17a6-b1bc-11ee-bbcc-a4badb2c639f part
├─sdi2 boot-pool 4269388194816694075 060337a0-b1bc-11ee-bbcc-a4badb2c639f part
└─sdi3 swap0 fee1e0bc-5a06-90b4-146a-7154df06a5fc 06005270-b1bc-11ee-bbcc-a4badb2c639f part
nvme0n1 JAJP600M4TB C24ED55011141129001 disk
├─nvme0n1p1 80105f51-b1ec-11ee-9cda-a4badb2c639f part
└─nvme0n1p2 nvme 18235672659811483595 80146394-b1ec-11ee-9cda-a4badb2c639f part
as for HBA cooling it’s hard to say, the HBA itself does not have a heatsink but I’ve cranked the case fans to max through the BIOS just in case (it didn’t seem to make any difference) and yes its in IT mode
if I’m understanding these error counts correctly, I’m guessing f708bd25-a790-11f0-9bc1-c1f323032d4a and the 2 new drives are bad
if that’s the case, can I create a new pool with 4 drives in raidz2 and expand it later once I get replacements? I’ve read somewhere doing this might give less pool capacity than if I were to create a new pool with all 7 drives present.
Is this a used HBA? The Dell Perc 310 should have a heatsink on the main chip at a minimum. Do an internet search for “Dell Perc 310” and look at some of the images.
Standard case fans alone typically do not offer enough cooling. Cooling is measured in how many linear feet of air pass by the card, and it is a lot. And I’m not saying this is your problem, BUT we have seen this many times before.
A question of my own here:
You stated that you started with TrueNAS CORE, but you did not provide any duration that the system worked.
Q: How long has the system operated before failures started to occur? Days, Weeks, Months?
Q: Was it stable on CORE? Also the duration it was running CORE please.
I ask because if this is a new installation, it changes how troubleshooting is addressed. We will start with the common things but if your system was just built, then you may have bad components.
You said SMART passes on all the drives, that is a good data point.
Your zpool status output shows you have a lot going on. In the future, replace only ONE drive, let it RESILVER, then replace the next failing drive. This is part of the cause for the very long resilvering time.
Do you have a backup of all your important data? Hopefully the answer is Yes.
Ah, you’re correct, the HBA actually does have a small heatsink on the main chip, I misremembered.
I built the system in Jan or Feb ‘24, originally in an old Dell Poweredge t310 system. CORE was very stable in terms of uptime, but I did suffer 4 drive failures over the 2 years since, prior to this event. Drives are refurbished so I assume I just got unlucky. Thankfully the vendor sold them with 5 year warranty and has been very responsive in replacing them without issue.
In march of this year the PSU in the t310 died and I decided consolidate the system with components from my jellyfin machine in a Fractal Design Define 7 case. I also switched to truenas SCALE at this time as I had to run jellyfin as a SCALE app. It was fairly stable since, though I did have a few occasions where the system hanged (which seemingly never happened in CORE).
In the future, replace only ONE drive, let it RESILVER, then replace the next failing drive. This is part of the cause for the very long resilvering time.
I will keep this in mind, I had thought it would be better to install both simultaneously, but oh well.
I do have the most important data backed up, otherwise I’ve largely used this as a giant pool for torrents (which I can fairly easily replace) and backups of my bluray collection (which I still have and can re-rip).
you may be correct. I didn’t think it was getting much airflow in the old system but I just checked and it was pretty hot to the touch.
I attempted the tried and true method of “take the side panel off and point a box fan at it” (/s) and after reboot my files are all back and accessible again, so I was able to backup a little bit more.
resilver time still keeps going up and one of the new drives is still showing errors though. Is there any way to disconnect that drive and continue resilver with the rest?
I don’t think you can remove any drives with your current pool status. Removing a drive may cause pool failure.
You can post back a zpool status -v CLOSED if you still want to try. It will allow others to comment on what it returned.
actually I was able to just disconnect the new drive showing errors, and resilver appears to be proceeding normally. Disk I/O looks good and time remaining shows ~34 hours
zpool status -v CLOSED shows some permanent errors on some torrent files but assuming the resilver eventually completes I should be able to repair those with a recheck:
root@truenas[~]# zpool status -v CLOSED
pool: CLOSED
state: DEGRADED
status: One or more devices is currently being resilvered. The pool will
continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
scan: resilver in progress since Fri Aug 14 11:58:55 2026
3.20T / 81.3T scanned at 4.70G/s, 494G / 81.3T issued at 727M/s
66.6G resilvered, 0.59% done, 1 days 08:22:52 to go
config:
NAME STATE READ WRITE CKSUM
CLOSED DEGRADED 0 0 0
raidz2-0 DEGRADED 0 0 0
replacing-0 UNAVAIL 0 0 0 insufficient replicas
sdc2 OFFLINE 0 0 0
10567549841703865306 UNAVAIL 0 0 0 was /dev/disk/by-partuuid/c298d8c6-de53-4f07-ae70-cedad352a4b2
6e431bf6-b111-11ee-80df-00155d011501 ONLINE 0 0 0
8a006ac7-9881-11ef-a106-1cfd0876a4af ONLINE 0 0 0
f0bc3ab2-2d63-11f1-bf17-7085c2aafbee ONLINE 0 0 0
replacing-4 DEGRADED 0 0 244
sdb2 OFFLINE 0 0 0
cb031d0b-7773-4da2-8e3a-a49e769848d6 ONLINE 0 0 0 (resilvering)
f708bd25-a790-11f0-9bc1-c1f323032d4a ONLINE 0 0 0
sdg2 ONLINE 0 0 0
errors: Permanent errors have been detected in the following files:
CLOSED/fixed_storage/torrents:<0x0>
(whole bunch of files with errors, all torrents)
This puts both HBA cooling and insufficient PSU power on the HDD rails as possible suspects.
This is metadata corruption. You might resolve it by deleting the whole torrents directory—or not even.
Raidz2 is not an ideal geometry for the many small writes that active torrents receive.
I’ve been concerned about the PSU as well, my old dell poweredge was only rated for like 350w and seemed to work OK, while the new system is 650w. neither had enough sata power connectors for all the drives so I just used an expander to connect them all. I know that’s not ideal but I wasn’t able to find a PSU with so many SATA power connectors. I have the system connected to a kill-a-watt meter and the whole system is currently drawing 95w even with the resilver running.
As for the write issue, I’ve been aware of this (downloading directly to the pool killed performance) so I have qbittorrent set to download files to a 4TB nvme drive on a pcie card and then move them after completion.
EDIT: the specific PSU I’m using is my old OCZ-ZS650W, which appears to be rated for 22A / 130W on 5V rail, so hopefully OK but definitely worth upgrading. the system peaks at ~103w on bootup according to the kill-a-watt.