Hi everyone,
I’m running into a frustrating issue with my TrueNAS Scale box freezing completely, randomly — often during the night. When it happens, NFS, SMB, Web UI, and SSH all become unreachable, and I have to force-reboot the machine via the power button.
Hardware Configuration:
- Motherboard: ASUS Prime B350-PLUS (AM4)
- CPU: AMD Ryzen 5 1600X
- RAM: 32 GB DDR4
- Storage Devices:
- 1x 10 TB HDD
- 2x 3 TB HDD
- 3x 2 TB HDD
- 5x 1 TB HDD
- 1x 4 TB SAS HDD
- 1x 128 GB SATA SSD (boot)
- 1x 500 GB SATA SSD
- 1x 512 GB NVMe SSD
- HBA: Dell PERC H310 flashed in IT mode (8-port SAS)
- Network: Chelsio T520-CR 10GbE SFP+ card
- PSU: Generic 600 W (non-certified)
ZFS Pool Layout
I currently manage 4 separate pools, each designed for specific purposes:
Pool_1 – Main Storage (Films + Series
- ** Layout**: 1 x 10Tb SATA
- Purpose: Multimedia library, long-term data
Pool_2 – Virtual Machine NFS Target
- Layout: 4x 2 TB SAS & SATA HDD in RAID-Z1
- L2ARC: 500 GB SATA SSD (read cache)
- Purpose: Multimedia library, long-term data
Pool_3 – Archive Storage
- Layout: 3x 3 TB SAS, SATA HDD in RAID-Z1
- SLOG: NVMe SSD
- Used as: High-throughput NFS target for Proxmox VMs
Pool_4 – Test/Temp/Frigate
- Layout: 1x 4 TB SATA
Symptom Summary:
- System works well for several hours, but freezes silently (no panic, no shutdown messages)
/var/log/messagesends abruptly before the freeze- After reboot, system works again for a while
- Error detected when power draw before freeze: ~120–130 W constant (via smart plug)
- A forced reboot brings it back up cleanly
- Here is logs error & failed last day
truenas_admin@truenas[~]$ sudo cat /var/log/messages | grep failed
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.0: VF BAR 2 [mem size 0x00080000 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.1: VF BAR 2 [mem size 0x00080000 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.3: VF BAR 2 [mem size 0x00080000 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.3: VF BAR 4 [mem 0x00010000-0x00020fff 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.3: VF BAR 0 [mem size 0x00100000 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: pci 0000:08:00.3: VF BAR 6 [mem size 0x00100000 64bit]: failed to assign
Jun 8 15:59:21 truenas kernel: gpio_generic: module verification failed: signature and/or required key missing - tainting kernel
Jun 8 15:59:21 truenas kernel: cxgb4 0000:08:00.4: Direct firmware load for cxgb4/t5-config.txt failed with error -2
truenas_admin@truenas[~]$ sudo cat /var/log/messages | grep error
Jun 8 15:59:21 truenas kernel: cxgb4 0000:08:00.4: Direct firmware load for cxgb4/t5-config.txt failed with error -2
truenas_admin@truenas[~]$ sudo cat /var/log/messages | grep critical
truenas_admin@truenas[~]$
What I’ve Tried:
Removed cron reboot at 4:00 AM (initially suspected an intentional reboot)
Confirmed: No OOM, no kernel panic, no scrub/snapshot conflict
Chelsio T520-CR: proper cxgb4driver loaded; SR-IOV enabled on Proxmox host
truenas kernel: cxgb4 0000:08:00.4: Direct firmware load for cxgb4/t5-config.txt failed with error -2
Switched from Proxmox LXC to full VM for NFS client to eliminate container-layer issues
Main Suspects:
- PSU Instability: Low-end 600 W noname PSU may not deliver stable voltage under multi-disk spin-up + 10GbE + Ryzen draw
- Network stack / driver conflict with Chelsio T520-CR and heavy I/O
- Possibly thermal or chipset instability on the aging AM4 board
- RAM issue ==>MemTest needed?
Planned Fixes:
- Replacing PSU with be quiet! System Power 10 450W 80+ Bronze (or higher)==> Alim changed and behavior still present
- Consider migrating from Chelsio to Intel X520 if issues persist ==> Can I charge new drivers?
My questions:
- Has anyone experienced similar hard freezes with TrueNAS and Chelsio T520 cards?
- Any way to dump more debug logs (like kernel ring buffers or hang detection)?
Thanks in advance for any help. I’m happy to share logs or test scripts if needed.
Cheers,
Pierre