Resilvering before adding replacement drive?

I had a drive start throwing SMART errors before vaction. I turned off the nas, bought a replacement drive and figured I would deal with it when I got back.

Now I’m back and the replacement drive is here.

Booting up truenas, however, shows that it has started its resilvering process despite me not adding the replacement disk.

What’s more confusing: I can see the critical warnings in the the alert log but inspecting that specific drive’s SMART test results reports no errors. Pics below:

In summary:

  • truenas alerts tells me there are several critical issues with sde, but…
  • truenas dashboard says there’s no problem, yet…
  • truenas has started resilvering my pool, despite…
  • I haven’t installed the replacement HDD for the ‘bad’ drive.

Can anyone shine a light on whether this is a drive problem vs. something else?

Separately:

  • I’ve read that it’s highly discouraged to swap in a drive during a resilvering process.
  • I don’t know how this process started when I didn’t tell truenas to do so and there are no spare disks in this system to resilver onto.

Please let me know if this is expected behavior, This is my first drive replacement process!

Happy to provide more information as needed

Might have been a little glib with the dashboard comment: it does say pool1 is not healthy

Small suggestion because I missed this the first few times: If there is an error, it should highlight that text so it is visually distinct from the non-error reports

OK, I think I got myself on the right track. Not sure why Truenas indicated it was resilvering on that first boot - whatever it tried to do must have been temporary. It’s no longer running (wish I got a screenshot…) but I went ahead with off lining and replacing sde as described here

The resilvering process has started again (currently scanning) and this time it makes sense.

I’ll follow up if I have any more questions.

Sometimes ZFS will off-line a drive it thinks is failing. But, with your shutdown, on boot it tried the drive and it appeared reasonably healthy. Thus, re-silvered the difference between when it was offlined and current state. These resilvers are shorter, (and thus faster). All perfectly normal.

While ZFS is not perfect, nor it’s management, it is mostly state of the art today.