… and at the same time, once management decides that it’s time to make changes to some of the features that are cornerstones to said operational results, it behooves a very careful rollout. This is not libvert vs. incus, or Docker vs. kubernetes.
SMART as a “industry standard” has tremendous improvement potential thanks to OEMs taking significant liberties re: how and what to report, making arbitrary changes, etc. One would think that with only three competitors left that the engineering oligopoly could at least agree on SMART interoperability?
And yet, we have negative disk temperatures and other bugs that may be as annoying as a stuck glove compartment in a car but not as dangerous as cooked brakes. Hence, SMART alarms have always required parsing, usually via experts who had lived with similar drives and know which symptoms were bad (“click of death”) and which could be likely safely ignored.
If ixsystems wants to get into the parsing game by only selectively alarming when a subset of SMART errors crop up, I suggest a fully-fleshed out documentation page that lays out what will and what will not cause a SMART warning with default settings. I’d also document the default cadence of scans, etc. This is not for legal reasons but primarily for customer goodwill.
The team doesn’t have to lay out why it’s made its decisions (ie those may be based on proprietary data) but the end user can then make a informed decision re: whether to solely rely on the ixsystems suggested set of SMART parameters getting monitored (with likely fewer false positives) or whether to rely instead on more inclusive packages like multi-report (that will likely result in more false positives, additional research requirements but also a fuller picture of alleged drive health).
Instead of eliminating GUI scheduling of SMART scans, i suggest starting the likely multi-year process of building and updating a better SMART error parser.
The official support of this SMART error parser could be limited to the series of enterprise drives that ixsystems ships with its server hardware, further aligning engineering efforts with paid customers. WD used to be ixsystems official supplier, and for all I know would be happy to lend engineering help involving their drives.
CE users have to understand that drives that were never shipped by ixsystems may not enjoy 100% parser compatibility. But, CE users can rely on multi-report instead and scrutiny now has a new lifeline (thank you, @pmh for that update).
IIRC, similar efforts were made way back when to bring some bespoke Marvell support to FreeBSD / TrueNAS in order to make the Asrock Rack C2750D4I in the mini series fully reliable. Other chips didn’t get bespoke drivers and hence remained buggy. C’est la vie.
But putting together a better SMART error parser is worlds away from removing a GUI feature and replacing it with an cron job. I’m flummoxed and saddened by the damage left in the wake of this decision, including the banning of at least one user.