I’ve got TrueNAS Scale 25.10.3 running in a Proxmox VM. HDD controller card passed through, 4x 16tb pooled in a raid 10 style setup. I’ve been having issues with storage pool randomly disappearing. If I restart TrueNAS, the storage pool re-appears, but almost always with errors. I’ve scrubbed, removed error files, cleared, rescrubbed - had it where there was nothing but green lights across the board but then it eventually disappears again. It usually stays up for a couple of days before it goes.
I’ve run smart long tests on the 4 hdds and the bootpool nvme and attached the results in notepad files. They all say “passed” but if someone in the know can please scan over them and let me know if any are on the way out?
The pool is currently in a degraded state. I’m running another scrub atm. That won’t finish for another about 9 hours, provided it doesn’t disappear during that process.
I have 2 other vm’s running (Win11 and HAOS), which seem to be fine. I’m not sure what’s going on, but I’m not convinced its resource based.
If anyone has some suggestions that I could try, I’m all ears.
Your SMART results show incredible numbers of raw read errors, seek errors, and communications errors between all the drives and the controller. This smells like either the controller is overheating and resetting constantly, or a controller that’s not long for this world. Your data is probably not valid either.
I second @Farout: Is the controller an HBA in IT mode that just needs more cooling, or is it a cheap PCIe-to-SATA port multiplier? If the latter, it’s not engineered for server loads. If the former, it’s expecting cubic yards per minute of airflow.
Didn’t know about the blacklisted. Will have a hunt for that.
It’s got a fan close by that is blowing directly across the heatsink, but you’ve raised a fair point. It’s been a while since I’ve pulled the server out of the rack and had a look. I’ll do that tonight - make sure everything is how it should be.
I checked the fan and general cooling. Everything was as I left it, but I’ve repositioned the fan closer to really focus on the heatsink. You’ve got to wonder if these things run so hot that they fail, why the manufacturer didn’t send it with an attached fan already.
I’ll post back again when I get a chance to upgrade the firmware, hopefully in the next couple of days.
See my original post: …it’s expecting cubic yards per minute of airflow.
There’s no single heatsink fan that can generate that much airflow. HBAs are engineered for a server chassis built like a wind tunnel. In a home-brew chassis, the best and most affordable option is a water-cooling block.
Actually, I’ve got a old Thermaltake cpu water cooler that would be good for it. I had a look but I’d need to frankenstein a mount to secure the block on to the chip. Waiting for a 3d printer to arrive, so that might be a project.
Do you know of a water cooler that fits these?
So, apart from a hiccup last night (pretty sure that was my error) it’s been running really well. I pulled the heatsink off and re did the thermal paste again, got a better suited fan and positioned it as good as it’s going to get and flashed the card with 16.00.16.00. All of that has seemed to fix my problems.