Weird CHKSUM errors after drive replacement

Just a small update: after a third HDD shows the exact same errors I have now updated system firmware, BIOS und TrueNAS to 25.04.2.1.

Errors still showing up.

Failed a new ticket.

Is this controller in IT mode and what firmware version is it running?

sas3flash -listall

Yes it is in IT-mode…or I think so. The disks just get passed through.
sas3flash did not work (no adapters found).
storcli /c0 show says following:

FW Package Build = 18.00.01.00
FW Version = 18.00.01.00
BIOS Version = 09.35.00.00_18.00.00.00
NVDATA Version = 18.00.00.24
PSOC FW Version = 0x0001
PSOC Part Number = 05689
Driver Name = mpt3sas
Driver Version = 48.100.00.00
Bus Number = 175
Device Number = 0
Function Number = 0
Domain ID = 0
Vendor Id = 0x1000
Device Id = 0xAF
SubVendor Id = 0x1000
SubDevice Id = 0x3010
Board Name = HBA 9400-8i
Board Assembly = 05-50008-01A

Cool. Yeah my bad 9400 onwards is storcli not sas3flash.

Just to update you guys here:
By now I’ve tested four disks - all show the same symptoms but only the new ones do.

My new Ticket was now closed with this message:

Sorry but this is for reporting bugs in software of TrueNAS. We do not have the time or resources to investigate every variation of hardware system and so we’re closing this ticket. Please note, before claiming it’s not hardware understand we have many 10’s of thousands of users running TrueNAS without reporting similar issues. The probability is that this is, in fact, some sort of hardware/software issue. Maybe reaching out to upstream zfs would be next step or trying to run truenas on another server to see if the same errors follow. Tough situation to be in but there isn’t anything actionable for us here because we don’t suspect this to be a bug.

I do understand that this is not likely to be a bug in TrueNAS’s software but rather in ZFS itself or even in my hardware - although I have no idea what kind of hardware defect could cause something like this - but I think it very frustrating, that the answer to my first ticket was to update to the latest TrueNAS version and file a new ticket only for this to be closed like that.

Hardware issues are often very frustrating and difficult to identify. Do you have some other hardware you can use to move things around? I’m thinking a different HBA and ideally a different chassis.

What’s the history to this? How long have you been running this system for and on what version/s?

Sounds like the issue only started when you went to replace a failing disk if I’ve read correctly?

PS: What is the physical hardware here?

that’s my next “move”. Currently I have created a temporary pool with three of the new drives that all produced checksum errors in the existing pool. I am moving data there and will do a scrub afterwards. I this new pool then has no errors it basically cannot be a hardware issue.

If the new pool produces errors I’ll move the disks into another computer and test there.

The system has been running for a few years. Afaik it was first installed with SCALE Angelfish and then upgraded over the years.

That’s correct.

So is this a new pool within the same system? What is the hardware being used here?

Resilver is very IO intensive so I wouldn’t assume because you can add data with no errors it’s not hardware. For example your HBA could be overheating thus causing errors.

The existing pool is probably there since the system was first installed.
There is a precise hardware description in the ticket.

I’ll read the HBA temperature next time a scrub or resilver is running. While writing this it reports 56°C and that’s while it’s replicating a ~5TB dataset from the existing pool to the temporary one.

I still don’t see, why an HBA temperature issue should create checksum errors only for one disk in a pool though (especially during scrubs after the resilvering, that should stress all disks in a similar way),

@TheColin21 listing the hardware only in the Jira ticket really sucks as anyone who wants to help must have a Jira account. Please post your system configuration in the forum. Should be an easy cut and paste.

Sorry, didn’t think about that.

Here is a description of the hardware:

  • Thomas-Krenn 4HE Intel-Dual-CPU RI2424 Server
  • 2x Intel Xeon Silver 4215R
  • 128GB DDR4-ECC
  • 2x 240GB SATA SSD - Boot-pool
  • 3x 1,92TB SAS SSD - pool “local”
  • 2x 1,6TB NVMe SSD - special metadata device for pool “pool”
  • 18x 18TB SAS HDD ST18000NM004J - RAIDZ3 data vdev of pool “pool”
  • all SCSI disks connected to Broadcom / LSI HBA 9400-8i

The new temporary pool consisting of three of the drives that showed errors in the original pool is now filled with ~7TB of data and, so far, shows no checksum errors.
I am now running a scrub there but if that completes without any errors as well I’d basically rule out any hardware error.

The question is: in what way could a ZFS pool be erroneous, to cause these symptoms?

The new temporary pool was scrubbed without error. I actually thin, the existing pool might be borked in some way now.

So the original error happened when you replaced a drive correct? Then perhaps try replacing a drive on your new pool and see what happens.

I am now doing two things as a last resort before destryoing and restoring the pool.

  • I made a checkpoint on the existing pool, upgraded it, cleared errors and am now running a scrub for the last time. I honestly don’t think it will help but my the ZFS upgrade changed…something.
  • I removed one of the disks from the temporary pool, wiped and added it again - it is now resilvering and I’ll check for errors there.

If the first idea doesn’t help and the temprary pool resilvers and scrubs without errors we will completely rebuild the pool.

What’s the current state of the original pool? Can we see zpool status for it?

It’s currently healthy as I just cleared all errors to see if new ones will pop up during the scrub.
Here is the zpool status:

  pool: pool
 state: ONLINE
  scan: scrub in progress since Tue Aug 19 11:23:30 2025
        10.9T / 64.6T scanned at 8.69G/s, 1.48T / 64.6T issued at 1.18G/s
        0B repaired, 2.29% done, 15:12:58 to go
    scan warning: skipping blocks that are only referenced by the checkpoint.
checkpoint: created Tue Aug 19 11:22:49 2025, consumes 1.07G
config:

        NAME                                      STATE     READ WRITE CKSUM
        pool                                      ONLINE       0     0     0
          raidz3-0                                ONLINE       0     0     0
            79466275-6640-49ce-b8eb-92f814efa9fc  ONLINE       0     0     0
            40640f17-ac71-4645-9ad1-5d574335ca4d  ONLINE       0     0     0
            0be792f1-7511-4c14-8a49-fd8a91367467  ONLINE       0     0     0
            d9a9c9d9-1334-4fbb-b558-2ccd599fc67c  ONLINE       0     0     0
            9f35f45e-4e27-4584-b13b-737f9ba0d10e  ONLINE       0     0     0
            eff3ea9b-060c-4854-97d2-de4b9c31f87e  ONLINE       0     0     0
            de24582d-ae26-4739-af0f-e518521fc313  ONLINE       0     0     0
            5d22eeae-e54f-4afe-88dc-080c3561fc39  ONLINE       0     0     0
            9dffed98-acdc-4af8-8375-5a89af26c726  ONLINE       0     0     0
            e5605e5d-c91e-4311-9738-22e7221f5679  ONLINE       0     0     0
            c0142394-4fcf-4150-a3db-d3a4d53be78f  ONLINE       0     0     0
            8aa6762f-84df-4362-8f33-0cb7dd35e360  ONLINE       0     0     0
            972cb2ac-237e-4636-bb0f-4c6d347c8cc5  ONLINE       0     0     0
            321aa17f-6349-449e-8224-f5d26200ab11  ONLINE       0     0     0
            1fdbec9b-c388-45be-bbde-255df8b680de  ONLINE       0     0     0
            ac070634-3a00-4e88-82b1-fd7a6d5dd955  ONLINE       0     0     0
            5501523d-f02b-46a1-9923-085e2d7ec452  ONLINE       0     0     0
            57ae22bd-c261-4ef1-a204-cba55ea98452  ONLINE       0     0     0
        dedup
          mirror-2                                ONLINE       0     0     0
            1f5ec529-b6d7-0a49-960f-5cd27b1c5f44  ONLINE       0     0     0
            94babd1c-a682-b544-8b8f-d7d69be65618  ONLINE       0     0     0
        special
          mirror-3                                ONLINE       0     0     0
            9f7b71b4-76a2-e646-a108-9583643e1eec  ONLINE       0     0     0
            563aa614-47e2-0d4c-8aa4-ef33fa5acbf6  ONLINE       0     0     0
        logs
          mirror-1                                ONLINE       0     0     0
            9b230d2d-9ba7-bf40-be65-2952d79a37e3  ONLINE       0     0     0
            6bfa6e18-a089-544e-a72a-edf3e05fa708  ONLINE       0     0     0
        cache
          nvme0n1p2                               ONLINE       0     0     0
          nvme1n1p2                               ONLINE       0     0     0

errors: No known data errors

Ok. I’ve only just noticed this now and not relevant to the current situation but you have a mirror special vdev and a mirror dedup vdev however your data vdev is an 18 wide Z3. For info if you lost your two drives in your special or dedup vdevs you would lose the whole pool. I wouldn’t recommend it but having an 18 wide Z1 would give you the same fault tolerance as the other vdevs. To have the same tolerance of a Z3 each special and dedup vdev would need to be a four-way mirror.

Just something to consider if/when reconstructing your pool.

Thanks, we’re aware of that.
We do have four SSDs but when all were installed they overheated.
We’re going to try, if we can get all four in with an extra fan but we’ll also discuss, if the special vdev is even necessary for this pool.

The first 3 checksum errors have appeared on the original pool again after 10% of the scrub.

The temporary pool is still resilvering without issues.