Coming from a single disk?
Yes, from the disk that was replaced, as always. No surprises here… -.-
Have you tried replacing that drive?
I do very much appreciate you trying to help but please read the thread.
Yes, three times actually to be absolutely sure that it’s not a faulty drive.
Sorry I got confused a bit around the drive replacements you had already done. So the corruption is happening during rebuilds. It does sound very much like hardware atm.
This part is also interesting and makes me wonder about cooling elsewhere in the system.
those are Samsung PM1735 MZPLJ1T6HBJR.
If we use all four they’re right in neigbouring PCIe slots so barely any airflow is left.
The HBA is currently at 57°C while the RAIDZ3 is being scrubbed and the temporary pool is resilvering. I do not think it has thermal issues. All drives are monitored and their temps are good.
Fair enough. Perhaps cabling or backplane/expander issues then.
That’s what I thought too in the beginning but that absolutely does not explain the symptoms.
Four HDDs in four different drive bays that all showed the same issues in the old pool.
Three are now showing no issues at all in the temporary pool even after a scrub and now during a resilver (which is already at ~50%) while the fourth is again throwing checksum errors that no “old” disk has.
Just as an update: as nothing else seemed to help I’ve now replicated and then destroyed the pool.
I added two more SSDs for special vDevs so the redundancy now matches and I am currently copying the data back.
So far everything (ZFS errors, disk temps, controller temps) looks fine.
If the replication and a scrub afterwards finish without errors I’ll mark this as the solution - even if it’s not a good one.