16 day scrub time?

Maybe I’m misremembering, but I recall reading that restarts clear ZFS error counters.

As you’ve noticed, the device names (sdX) can change during restarts. Instead of using the device name, take note of the unique serial number when tracking failing drives.

2 Likes

That’s too much extra work. :face_with_head_bandage: There is a simpler solution.

2 Likes

The error seems to be on the same serial number drive, I am pretty sure. Working out which drive that is in there might be a bit hit and miss but I’ll work it out. Not looking forward to chasing up warranty on it though.

Current actions: I have installed a fresh drive in there ready to go. Thankfully I bought 2 spares right before prices went insane. Also currently have all 6 drives running a long test to see what shows up. After that, I guess I will need to work out how to add the fresh drive into the pool and remove the iffy one.

Also, as a side note, super appreciate the help. Just a really nice community which is a change from how jaded a lot of places have become. You all rock :slight_smile:

4 Likes

Good that you already have spares!
As long as you do not reboot the problem drive will keep its letter. Write down its serial number for when you will remove it.

Keep the iffy one in there, as you have enough slots, and click on the “Replace” button that is clearly visible in the bottom right corner of your screenshot.
You can do that already, or wait for the long SMART tests to complete first.

I’m afraid you’ll have to: Checksum errors could be the cable, the controller, or a flaky PSU (but the last two would probably show up on multiple drives), but read errors are most likely from the drive. We can help you decyphering the SMART report when the long test has finished to confirm.
Better go through RMA to get a new spare (or at the very least a refund) than writing down a net loss.

3 Likes

I have a wide variety of files from video to text. Scrub takes about 24hrs.

  • Data Topology: 2 x MIRROR | 2 wide | 12.73 TiB
  • Usable Capacity: 25.32 TiB
1 Like

SMART tests on all drives have finished wis I believe no errors:

Still have this though:

So I am wondering if I still replace it, or clear and see if it comes back?

In your screenshot I only see 4 extended tests and 3 of them aborted before they managed to complete.

A short test is effectively useless at catching all but the most obvious and total fault.

What you do in terms of replacing the failing drive is up to you. If you changed nothing in terms of how you connect it, I imagine it will fail again soon.

It’s your time and your data.

2 Likes

With the tests stopping, I was still messing around with the box and needed to turn it off to wire new drive in. Will give the plugs at the back of the hot bays an extra push now and then hit replace on that drive and see how we go. Also need to work out which one it actually is in the big stack.

So I am thinking replace that one either way, and contact east digital to see how painful replacing it is gonna be?

Wish me luck.

Seems like it’s done, got a message saying the error had cleared and the dodgy drive serial is no longer on the list.


I marked the drive last night so now I just pull it out I guess? Also, just a quick one, can I move the other drives around so there isn’t an empty one in the middle, or will TrueNAS care about the drives staying in the same SATA ports?

You can reshuffle drives as you please.

But I suggest to have a full long SMART test of the removed drive to confirm its condition, and have grounds for RMA if needed.

1 Like

Did a smart test and it came back entirely clean. I’ll maybe plug it into the desktop and run a scan through that and see what comes back though. The worst screenshot of it right now is this one:

“SUCCESS” does not mean that the drive is fine, especially for a short test.
What I’d like to see is the full output of
smartctl -x /dev/sdd (or whatever the drive letter is…)
preferably after a completed long test.
Of course, you can run the test in a desktop.

2 Likes

Yep, that’s my plan. About to plug the now pulled drive into the desktop and test it on there. Just need to pull the back off the case and run some cables. I remember when I installed these I ran a full surface test through hard disk sentinel. Should I do something like that again or just run tests through something like Seatools?

As it turns out, X870 ProArt only has 4 ports and all are in use, but as luck would have it, had one of those USB adapters laying around.

Current info but long test running in Seatools.

Well then, kinda light on details. Maybe it’s ok after all and just something weird happened in a file transfer? Thinking it might be ok to just keep this one one hand as a spare now?

Details would come from smartctl -x (or -a) on the command line, or at least some utility like the one you used to provide a screenshot in your previous post.

It’s plugged into the PC now and not in the NAS anymore.

SMART info:

And the Seatools long test doesn’t give me any details other than it passed.

Your Seatools image looks to have a Results button at the bottom. Maybe that is how you view?

2 Likes

That screenshot I took is the results page, quality application.

1 Like

In CrystalDiskInfo, go to Function, Advanced, Raw Values, 10 DEC for much more human readable information in the bottom half of the screen. (This is from memory so i may have missed one step in the menu, but you’ll find it.)

1 Like