Unbuffered ECC RAM Going Bad?

I upgraded from Core and it went smooth. I’m running a Ryzen 5 Pro 5650GE on a Gigabyte X570S Aero G with 4 x 32GB 3200MHz unbuffered ECC from Nemix on Scale 25.10.4.

I noticed dmesg would spam every 30 minutes or so lines like:

Logs

[Mon Jun 15 21:45:56 2026] mce: [Hardware Error]: Machine check events logged
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Corrected error, no action required.
[Mon Jun 15 21:45:56 2026] [Hardware Error]: CPU:0 (19:50:0) MC18_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-]: 0x9c2040000000011b
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Error Addr: 0x00000004b724c2e0
[Mon Jun 15 21:45:56 2026] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000014b40a401202
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[Mon Jun 15 21:45:56 2026] EDAC MC0: 1 CE on mc#0csrow#2channel#1 (csrow:2 channel:1 page:0x9ae498 offset:0x5e0 grain:64 syndrome:0x14b4)
[Mon Jun 15 21:45:56 2026] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
2026 Jun 15 21:45:57 truenas [Hardware Error]: Corrected error, no action required.
2026 Jun 15 21:45:57 truenas [Hardware Error]: CPU:0 (19:50:0) MC18_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-]: 0x9c2040000000011b
2026 Jun 15 21:45:57 truenas [Hardware Error]: Error Addr: 0x00000004b724c2e0
2026 Jun 15 21:45:57 truenas [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000014b40a401202
2026 Jun 15 21:45:57 truenas [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
2026 Jun 15 21:45:57 truenas [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD

It was always that channel mc#0csrow#2channel#1 so I researched and determined that would be my 4th slot (furthest from CPU). I’m pretty confident that stick was bad based on the frequency, so I removed that stick and booted with 3 DIMMs, giving me 96GB available.

I fired up a debian-slim docker container and ran memtester with 85GB (leaving a little for some services and ZFS cache to breathe) and it passed. No more errors. Until roughly 24 hours later, I got another message exactly like that but for the 2nd DIMM slot mc#0csrow#2channel#0

New Logs

[76271.776673] mce: [Hardware Error]: Machine check events logged
[76271.776678] [Hardware Error]: Corrected error, no action required.
[76271.777023] [Hardware Error]: CPU:0 (19:50:0) MC17_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[76271.777385] [Hardware Error]: Error Addr: 0x000000002754b000
[76271.777743] [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000066820a401402
[76271.778069] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[76271.778082] EDAC MC0: 1 CE on mc#0csrow#2channel#0 (csrow:2 channel:0 page:0x2754b offset:0x0 grain:64 syndrome:0x6682)
[76271.778687] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD

This error has only happened once. I’m sitting at about 85GB used (ZFS cache + some misc services) and it’s been a few more hours and I haven’t seen that error again. I plan on letting it sit for a few more days to see if it logs anything else. Now chances are I had a bad stick for so many years on Core, but apparently it couldn’t log it (Linux “talks” better with the Ryzen integrated memory controller since this is a consumer board with no IPMI).

But is it likely this other stick is bad too? Should I run a specific test other than memtester? Or is just a one off correction like this somewhat typical? Hopefully some of the Ryzen ECC users can chime in.

That’s an odd way to run a memtest…

I would only trust full passes of a memtest if booted directly from a USB on the actual system itself. No virtualization, no docker, no “allocated” memory.

I would use Passmark’s software in this case, since Memtest86+ doesn’t have full support for reporting ECC corrections, AFAIK.

4 Likes

Yeah I’ll have to do that. I just didn’t want to take her offline but I guess no other choice.

Not worth testing that 4th stick though because of the frequency of errors?

That’s the cost of knowing for certain.

I would run a few passes with all sticks inside and at the settings you normally use. If it fails, then you can start a process of elimination.

It’s not off the table that your sticks might be good and this could be a matter of a BIOS setting.

I would use Passmark’s software in this case, since Memtest86+ doesn’t have full support for reporting ECC corrections, AFAIK.

1 Like

I felt like chiming in here as well.

Everyone have stated the proper way to use MemTest86+. Version 8.10 is the current version and I would recommend you use that version on a bare metal USB boot drive, not in a VM.

Populate all the RAM you want to have in the system, and I’d also install the stick you think might be bad. You can’t trust the test when run as a VM. Also, the errors you were seeing, it looks like they were being corrected.

Make sure you leave all your drives in the system to put the same power draw on the system as you normally run the system.

Now this is very important and many people don’t take this advice to heart…
Run Memtest86+ for “At Least 4 complete passed” and I personally prefer 5 passes minimum. Before i sell a MB/CPU/RAM, I let Memtest86+ run for a minimum of 10 passes. One pass is not enough unless you have failures, then you need to try and isolate the problem. Sometimes a failure will show up on pass #4, seen it here in the forums before. If you are able, run the test longer. With that much RAM, it will take a while.

Other thoughts… Your RAM may be perfectly good, it could be the CPU or the MB, or even (dare I say it) a power supply. I actually do not see many sticks of RAM go bad, nor the CPU, but power supplies… Yes. I’m only saying these things so you can keep an open mind.

If you have a Memtest86+ error, write down the error, or better yet, snap a photo of it.
Next, rotate your RAM around (slot 1 to slot 2, slot 2 to slot 3, slot 3 to slot 4, slot 4 to slot 1) and this will move the RAM to a different interleave channel.
Run Memtest86+ again, does the failure happen at the same place or does it completely move?

These are the little things to help you diagnose the problem. With RAM prices being what they are today, I’d rather replace a power supply! or better yet, have the test pass 10 times consecutively.

Don’t cut any corners, you will regret it later if you miss a problem and down the road it causes you some serious grief. At least right now you can do it on your own schedule, which is much better than when it fails and you are busy with other things.

Oh yes, a few questions:

  1. What are the models of the RAM sticks?
  2. Have you checked to see if the MB BIOS is current? Often the BIOS will be updated to fix timing and voltage issues to fix stability problems.

Best of luck to you.

5 Likes

In this case he probably should consider Passmark’s Memtest instead of Memtest86+, since Passmark’s free version reports ECC corrections, which will not show up as “errors”.

I’m not sure how far Memtest86+ has progressed on this feature. (AFAIK they don’t fully support it.) If it doesn’t report ECC corrections, then it will “pass”, even if the RAM has to constantly fix flipped bits, which could be a sign of future failure.

2 Likes

So yeah I think the passmark version does support ECC better so that’s what I’m running. I started the test with 3 sticks, it just rebooted like 10 minutes into the test. I pulled the second stick, and now I’m testing with two sticks installed (I moved them to 2nd and 4th slot for proper slot population). It’s been running almost 7 hours with no errors yet. I’ll let it finish then either test the other two sticks at the same time or individually.

Now why it rebooted with 3 sticks? Maybe memory controller in “Flex” mode is no good and should only use single channel or dual channel. I also don’t know if the 5650GE integrated memory controller is super strong to just run 4 dual rank DIMMs of 3200MHz- like I think it should be but I don’t know? Testing two at a time will help me with this. Maybe I need to up the CPU SOC voltage slightly or I’m even willing to downclock to 2933/2666MHz to get it stable, but I’ll try that after testing.

They are from Nemix and Micron chips. I did the FBGA decoder on the ICs and it is proper Micron F Die genuinely rated for unbuffered ECC 3200MHz, 1.2V @ CAS 22 (which is what default SPD runs them at). Motherboard BIOS is on F8a which is newest. BIOS is mostly on default other than changing ECC from Auto to Enabled but I think ECC would’ve picked up right regardless.

Were you using “flex mode” with all 4 sticks? I assume all 4 sticks are the same model, which means you should be using dual-channel.

If you’re going to test 3 sticks, maybe set it to single channel to take “flex mode” out of the troubleshooting.

What about running memtest passes with all 4 sticks in the standard configuration? You shouldn’t start off only testing 2 or 3 of them or figuring out which slots might be at fault. All 4 sticks, dual-channel, let it run multiple passes and see if the system reboots, shows memory errors, or reports ECC corrections.

Right now we don’t know if it’s a bad stick, multiple bad sticks, a bad slot, a configuration in the BIOS, or an issue with the CPU.

1 Like

You are 100% correct, sort of. I know Memtest86+ does support various AMD CPUs with ECC testing as verified by AMD. BUT, I do not see a switch to enable ECC testing, unless it is automatic. I’m not holding my breath but I did submit a question about the ECC support as it started being implemented in version 7, and we are at 8.10 now.

Thanks for the correction and when/if I find out something about Memtest86+ and ECC testing, I will make sure Everyone Knows. Maybe a Meme.

2 Likes

No. 4 sticks is dual channel.

Unfortunately can’t set this in BIOS. I think you just automatically get single channel (1 stick), dual channel (2 or 4 sticks) or the weird flex mode (3 sticks).

I’m already down the path of testing two at a time. I’ll make a follow up post in a few days once I run memtest on all the sticks.

I personally don’t think it’s a bad slot or “issue” with the CPU. My working theory right now is either bad stick(s) or weak IMC on CPU. If all sticks test good in dual channel- this severely lessens the load on the IMC compared to running 4 sticks- my solution would be to run 4 sticks and increase SOC voltage or drop frequency (e.g. 2933MHz or 2666MHz instead of 3200MHz.) If some sticks fail, then RMA is in order.

Funny to me I probably had this issue for years but TrueNAS Core / FreeBSD never logged it.

I don’t understand what version is better to use in this case :smirking_face:

come back serious :smile:

as far you push both mainboard and CPU IMC to the max (128gb of ram, for sure 2Rx8), in your place i would really prefer to drop frequency instead of increase voltages to reach stability. Less performance (it is really appreciable on a nas run at 3200 instead of safer 2933-2666?) but less heat, less stress on components… just my 2 cents

2 Likes

Ok so one stick is 100% bad. It causes memtest to just black screen. This happened when it was by itself in the second slot, and also with a good module in the second slot and having it in the 4th slot. I’m not sure when it exactly went black screen, but within 30 minutes of running the test for sure. System was unresponsive.

Two stick pass

Other stick pass

Nemix approved RMA of one stick so I’m going to proceed with that. In the meantime I’m only running 2 sticks in 2nd and 4th slot for proper dual channel.

When I was running 3 sticks dmesg only had one error, but I’m chalking that up to flex mode. I think Ryzen doesn’t like that. Once I get the replacement stick in I’ll be back up to all 4 DIMMs populated.

4 Likes