I’m having a somewhat bizarre issue, or perhaps not. For years, I’ve had these 2 x 6TB Seagates & 3 x 6TB WD Red Pluses in a 5-disk RAIDZ2 pool. In addition, I had a spare 6TB WD Red. This pool has existed for about 4.5 years, and has migrated across a carpload of hardware. It’s moved across two chassis, three different controllers, multiple CPUs/boards and now the controlling Truenas appliance is virtualized. Well, barely any sooner do I get this virtualized TrueNAS running that it is reports that one of the WD Reds is dying. I’d noticed some read errors on it in the past, so I decided to swap it out. Within a couple weeks, it reports that ALL of the WDs are dying. But the Seagates are just fine. The Seagates are about a year younger, and the WD Reds were all replacements surrounding that ugly SMR debacle–so same age, same batch, same origin.
Before this current incarnation of TrueNAS, life happened and my older NAS was down for about a year and a half. I’d moved twice, and surely the disks got a little jostled. I may have even dropped one about a foot onto carpeting while migrating into the new NAS. But the Seagates are just fine. What are the odds that all the WDs would go within the same week, or did they just sustain something environmental that the Seagates withstood?
I changed out the controller from an LSI 93xx series, which has a reputation for running hotter than the surface of the sun, to its lower powered 9400 cousin. No change. I rechecked my wiring and moved drives around to different power cables & sata cables. Nada.
I’m not devastated, I rescued the data and have other drives. But these bricks aren’t exactly cheap nowadays and I don’t want to exclude them based on a false positive (though this seems very,v very real). Coincidence or no?
Current setup:
Truenas on Proxmox VE 9.1
Broadcom 9400-16i with HBA passthrough
WD Red 6TB WD60EFRX, Seagate Ironwolf 6TB
Intel D-1557
Have you run SMART Long tests on the drives and looked at the results? Does the HBA have additional cooling or is it in a true server chassis?
I mean you seem experienced enough - any chance you can spin a bare metal system & run at minimum smart longs on the ‘failed’ drives?
Also any errors, logs, or anything else could go a long way… btw you did blacklist & passthrough the HBA, not the individual drives, right?
May have? That is so very bad for the drive, but stuff happens.
When you say the drives died, you need to be very specific with what error you are talking about. Was a ZFS Error? or a Drive Error? These are two distinctly different things. While a failing drive could cause a ZFS error, that isn’t always the case.
Take a look at my Drive Troubleshooting Flowcharts link in my signature. This is likely to get you headed down the correct path. I suspect you have a ZFS error.
Also, since your running on Proxmox, I would recommend that you create a bootable USB drive and install TrueNAS to it, then boot your system from the USB drive. You should now be running on bare metal, assuming you did not create virtual disks for your VDEVs.
@joeschmuck May have? That is so very bad for the drive, but stuff happens.
I admit to nothing, sir.
When you say the drives died, you need to be very specific with what error you are talking about. Was a ZFS Error?
I got some closure on this; maybe it can help someone else.
Originally, the problems manifested as ZFS read errors. TrueNAS eventually began flagging drives as DEGRADED then subsequently FAILED.
I’d been running smartctl long tests and scrubs to try to validate drive health, and there were indeed some errors, but eventually I learned I knew less about HDD error-handling than I thought. Misplaced confidence will always get ya.
This video from Level1Techs (Mechanical Drive Bad Sectors: The Drive is Not Always Dead (A Primer) taught me I needed to write to those bad sectors, and the whole drive, to validate the health. I removed the drives and connected them directly a board on a Linux system to bypass any HBA or virtualization weirdness, so I took your advice @Fleshmauler (and yes, I’d made sure mpt3sas.sys was blacklisted in Proxmox). I then used fio to write the whole drives.
The outcome?
- The WD RED drive that was flagged with smart errors was indeed bad.
- The other two WD RED drives were REALLY bad.
- One WD RED drive that I’d removed from service for recurring ZFS read errors was in perfect health. I can attribute this error to problems elsewhere in the pool perhaps, which was clearly not healthy.
- Interestingly both Seagates in the pool were humming along without issue, but they had both registered a single G-Shock event. Not good. I don’t know when that happened or whether it explains the fate of the WDs, but it’s quite possible that something happened and the Seagates were robust enough to survive while the WDs were not.
The good news is the virtualized appliance is not to blame and everything else checks out fine.
@SmallBarky I wrote a rather long-winded post about Sliger CX3701 case and the power, cooling and connection issues it has, but I fixed a Corsair RS120MAX 30mm-thick wind turbine 1" off the face of that HBA, and it is sucking air through it with tornadic force–so that’s one less thing to worry about at this point.
Sucks about the drives, but glad you got it figured out
Earlier this year I dropped a small 500GB laptop hard drive on the carpet. The drop was about 12", at the most. The carpet is well padded, BUT the drive never worked again. I’ve done worse before but the drive must have hit a corner that broke it.
So, these things happen and we want to kick ourselves in the butt for doing it, but it happens to us all at some point in time. Maybe not dropping a hard drive but it would be something terrible or expensive.
@joeschmuck A laptop disk’s plater substrate is glass based where as desktop drives use metal.Howver you probably know this 
As far as I am aware, there are a few different types of substrate. While 2.5" drives are typically Glass, the 3.5" drives have a variety of methods.
I don’t like to make a blanket statement as factual data, if it isn’t completely factual. In many of my postings I may say “as far as I know” or similar, indicating I do not know it for a fact.
But from what I have read, it appears 2.5" platters are made of glass, but I don’t know how accurate the Wiki page is for this information.
But the point was, it is bad to drop a drive at all. Sometimes they survive, sometimes the will never spin up again.