I upgraded from Core and it went smooth. I’m running a Ryzen 5 Pro 5650GE on a Gigabyte X570S Aero G with 4 x 32GB 3200MHz unbuffered ECC from Nemix on Scale 25.10.4.
I noticed dmesg would spam every 30 minutes or so lines like:
Logs
[Mon Jun 15 21:45:56 2026] mce: [Hardware Error]: Machine check events logged
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Corrected error, no action required.
[Mon Jun 15 21:45:56 2026] [Hardware Error]: CPU:0 (19:50:0) MC18_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-]: 0x9c2040000000011b
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Error Addr: 0x00000004b724c2e0
[Mon Jun 15 21:45:56 2026] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000014b40a401202
[Mon Jun 15 21:45:56 2026] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[Mon Jun 15 21:45:56 2026] EDAC MC0: 1 CE on mc#0csrow#2channel#1 (csrow:2 channel:1 page:0x9ae498 offset:0x5e0 grain:64 syndrome:0x14b4)
[Mon Jun 15 21:45:56 2026] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
2026 Jun 15 21:45:57 truenas [Hardware Error]: Corrected error, no action required.
2026 Jun 15 21:45:57 truenas [Hardware Error]: CPU:0 (19:50:0) MC18_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-]: 0x9c2040000000011b
2026 Jun 15 21:45:57 truenas [Hardware Error]: Error Addr: 0x00000004b724c2e0
2026 Jun 15 21:45:57 truenas [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000014b40a401202
2026 Jun 15 21:45:57 truenas [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
2026 Jun 15 21:45:57 truenas [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
It was always that channel mc#0csrow#2channel#1 so I researched and determined that would be my 4th slot (furthest from CPU). I’m pretty confident that stick was bad based on the frequency, so I removed that stick and booted with 3 DIMMs, giving me 96GB available.
I fired up a debian-slim docker container and ran memtester with 85GB (leaving a little for some services and ZFS cache to breathe) and it passed. No more errors. Until roughly 24 hours later, I got another message exactly like that but for the 2nd DIMM slot mc#0csrow#2channel#0
New Logs
[76271.776673] mce: [Hardware Error]: Machine check events logged
[76271.776678] [Hardware Error]: Corrected error, no action required.
[76271.777023] [Hardware Error]: CPU:0 (19:50:0) MC17_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[76271.777385] [Hardware Error]: Error Addr: 0x000000002754b000
[76271.777743] [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000066820a401402
[76271.778069] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[76271.778082] EDAC MC0: 1 CE on mc#0csrow#2channel#0 (csrow:2 channel:0 page:0x2754b offset:0x0 grain:64 syndrome:0x6682)
[76271.778687] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
This error has only happened once. I’m sitting at about 85GB used (ZFS cache + some misc services) and it’s been a few more hours and I haven’t seen that error again. I plan on letting it sit for a few more days to see if it logs anything else. Now chances are I had a bad stick for so many years on Core, but apparently it couldn’t log it (Linux “talks” better with the Ryzen integrated memory controller since this is a consumer board with no IPMI).
But is it likely this other stick is bad too? Should I run a specific test other than memtester? Or is just a one off correction like this somewhat typical? Hopefully some of the Ryzen ECC users can chime in.


