Pool i/o suspended

Hey everyone I have a huge issue with my truenas server that I can’t seem to resolve. My server is technically online and I can technically login to the webui, but when i go to the shell and use zpool status -v it throws me the pool i/o is suspended error. I know there’s a drive in my pool that needs to be replaced but I’m unable to replace it because i get a "errno(6): No such device or address" error. I don’t know if the failing drive is related to this but it’s been running for months with that problem and it seems fine.

My issue now is that the server is functionally not working due to suspended i/o activity. How exactly am I supposed to fix this? I rebooted the server and got stuck on a boot loop of “power on or device reset occurred”.

I was able to fix this boot loop by replacing the hba card. However when booting up it still threw that error but it was able to boot regardless. But now I’m still stuck as to why my i/o is suspended. The only thing I can think of that may be causing this problem is some kind of pcie bus/bandwidth issue. Im running truenas off of a mini itx lga1700 motherboard and I recently added a 5gbe nic that slots into the m.2 slot thats typically used for bluetooth and wifi modules. My thinking is that they share a lane with the hba cards pcie slot and it’s causing throughput issues? Otherwise I’m totally baffled why this is happening. Any suggestions because I am totally confused. Thanks!

I have 7 cmr 16tb hdds in raidz1, but 10 drives total. 2 of those drives are currently not in the pool, and the other drive is an ssd used for the boot pool. I have an 850w psu. So I don’t think its a power supply issue.

Update: removed the 5gbe nic and used the onboard lan but still getting i/o errors. I am completely at a loss right now :confused:

i/o suspended is zfs locking the pool after a disk path died. errno 6 means that device is already gone off the bus.

-zpool status -v
-ls -l /dev/disk/by-id/
-dmesg | grep -iE ‘reset|offline|mpt|mps|sd[a-z]’

if something is FAULTED/UNAVAIL pull that drive, then:
-zpool clear YOURPOOL
-still stuck: zpool export YOURPOOL && zpool import -f YOURPOOL

download a config backup under System → General → Manage Config b4 replace. device-reset boot loop after the old hba fits the controller dying, so match the pool gptids to what Storage → Disks still shows.

Thank you for the response. I’m not really sure what those last two commands do or how to parse the information, but these are the results. I also don’t know what you mean by matching the pool gptids. Sorry I’m fairly new to all of this.

truenas_admin@truenas[~]$ ls -l /dev/disk/by-id/
total 0
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-KINGSTON_SA400S37240G_50026B76855761C4 -> ../../sdd
lrwxrwxrwx 1 root root 10 Sep 27 15:18 ata-KINGSTON_SA400S37240G_50026B76855761C4-part1 -> ../../sdd1
lrwxrwxrwx 1 root root 10 Sep 27 15:18 ata-KINGSTON_SA400S37240G_50026B76855761C4-part2 -> ../../sdd2
lrwxrwxrwx 1 root root 10 Sep 27 15:18 ata-KINGSTON_SA400S37240G_50026B76855761C4-part3 -> ../../sdd3
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-MG08ACP16TE_21K0A362FWXG -> ../../sdi
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-MG08ACP16TE_21K0A362FWXG-part1 -> ../../sdi1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-MG08ACP16TE_3130A8F2FWXG -> ../../sdb
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-MG08ACP16TE_3130A8F2FWXG-part1 -> ../../sdb1
lrwxrwxrwx 1 root root  9 Sep 27 15:41 ata-MG08ACP16TE_31A0A226FWXG -> ../../sdc
lrwxrwxrwx 1 root root 10 Sep 27 15:30 ata-MG08ACP16TE_31A0A226FWXG-part1 -> ../../sdc1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-ST16000NM001G-2KK103_ZL2B5NLL -> ../../sdj
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-ST16000NM001G-2KK103_ZL2B5NLL-part1 -> ../../sdj1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-ST16000NM001G-2KK103_ZL2BEYVZ -> ../../sdg
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-ST16000NM001G-2KK103_ZL2BEYVZ-part1 -> ../../sdg1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-TOSHIBA_MG08ACA16TE_1170A081FVGG -> ../../sde
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-TOSHIBA_MG08ACA16TE_1170A081FVGG-part1 -> ../../sde1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-WDC_WD160EDFZ-11AFWA0_3WGXUGUJ -> ../../sdf
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-WDC_WD160EDFZ-11AFWA0_3WGXUGUJ-part1 -> ../../sdf1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 ata-WUH721816ALE6L4_2PGDDPGT -> ../../sda
lrwxrwxrwx 1 root root 10 Sep 27 15:29 ata-WUH721816ALE6L4_2PGDDPGT-part1 -> ../../sda1
lrwxrwxrwx 1 root root 13 Sep 27 15:17 nvme-PNY_CS1030_500GB_SSD_PNY251125031201002C6 -> ../../nvme0n1
lrwxrwxrwx 1 root root 15 Sep 27 15:18 nvme-PNY_CS1030_500GB_SSD_PNY251125031201002C6-part1 -> ../../nvme0n1p1
lrwxrwxrwx 1 root root 13 Sep 27 15:17 nvme-PNY_CS1030_500GB_SSD_PNY251125031201002C6_1 -> ../../nvme0n1
lrwxrwxrwx 1 root root 15 Sep 27 15:18 nvme-PNY_CS1030_500GB_SSD_PNY251125031201002C6_1-part1 -> ../../nvme0n1p1
lrwxrwxrwx 1 root root 13 Sep 27 15:17 nvme-eui.6479a7a49ac01df9 -> ../../nvme0n1
lrwxrwxrwx 1 root root 15 Sep 27 15:18 nvme-eui.6479a7a49ac01df9-part1 -> ../../nvme0n1p1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000039aa8c897c7 -> ../../sde
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000039aa8c897c7-part1 -> ../../sde1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000039ab8d397f0 -> ../../sdi
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000039ab8d397f0-part1 -> ../../sdi1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000039ac8c8e83b -> ../../sdb
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000039ac8c8e83b-part1 -> ../../sdb1
lrwxrwxrwx 1 root root  9 Sep 27 15:41 wwn-0x5000039ac8cb4d3e -> ../../sdc
lrwxrwxrwx 1 root root 10 Sep 27 15:30 wwn-0x5000039ac8cb4d3e-part1 -> ../../sdc1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000c500c7f3a2d6 -> ../../sdj
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000c500c7f3a2d6-part1 -> ../../sdj1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000c500c8168ff1 -> ../../sdg
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000c500c8168ff1-part1 -> ../../sdg1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000cca284cd1a08 -> ../../sdf
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000cca284cd1a08-part1 -> ../../sdf1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x5000cca2c1c5a419 -> ../../sda
lrwxrwxrwx 1 root root 10 Sep 27 15:29 wwn-0x5000cca2c1c5a419-part1 -> ../../sda1
lrwxrwxrwx 1 root root  9 Sep 27 15:17 wwn-0x50026b76855761c4 -> ../../sdd
lrwxrwxrwx 1 root root 10 Sep 27 15:18 wwn-0x50026b76855761c4-part1 -> ../../sdd1
lrwxrwxrwx 1 root root 10 Sep 27 15:18 wwn-0x50026b76855761c4-part2 -> ../../sdd2
lrwxrwxrwx 1 root root 10 Sep 27 15:18 wwn-0x50026b76855761c4-part3 -> ../../sdd3
truenas_admin@truenas[~]$ sudo dmesg | grep -iE ‘reset|offline|mpt|mps|sd[a-z]’
zsh: command not found: offline
zsh: command not found: mpt
zsh: no matches found: sd[a-z]’
truenas_admin@truenas[~]$ zsh: command not found: mps

[1]    done       sudo dmesg |
exit 1     grep -iE ‘reset |
exit 127   offline | mpt | mps

The two disks I don’t have in my pool, one of them has been in vdev expansion but stalled for months, and the other one is the one that gives me the errno error. I would like to use that one to replace the drive I think is the problem but it literally just won’t let me. They’re both plugged into the hba so I’m not sure what’s wrong.

action: The pool can be imported despite missing or damaged devices.  The
fault tolerance of the pool may be compromised if imported.
config:

    tank                                      DEGRADED
      raidz1-0                                DEGRADED
        141e82c4-40f0-4af4-8f54-08ef32fe831b  UNAVAIL
        30f144c0-333a-497e-b224-b059728210b6  ONLINE
        5d3fdde3-1177-4492-af9a-587128ccdf3f  ONLINE
        1f69e3db-742c-4fe2-a4ce-6751b43dc242  ONLINE
        7a2dbc18-c97f-4d82-9f3a-f16dfda1c6c0  ONLINE
        b4095240-c068-409d-b781-76cf4942658d  ONLINE
        28c97425-b8b3-4eb9-9566-cdeaead5506f  ONLINE
        b54e2b3d-0a61-4992-abd8-62369e36d4ac  ONLINE

truenas_admin@truenas[~]$ sudo zpool import -f tank
cannot import 'tank': insufficient replicas
Destroy and re-create the pool from
a backup source.

I pulled the faulty drive and when I try to import I get this message. So I’m assuming I just need to add another drive to the array and I can import? Or is there something else I need to do? I couldn’t offline the disk before I pulled it because of the i/o suspension. But I still have 2 drives in the server that aren’t part actually part of the server yet (one of them being in the process of vdev expansion) (and the other giving me the errno error ) so I don’t know what is going on.

insufficient replicas on raidz1 usually means zfs thinks 2 disks are missing. a blank spare wont make import work.

-plug the drive u just pulled back in
-zpool import
-if it shows tank DEGRADED with only 1 UNAVAIL: zpool import tank
-still insufficient: zpool import -nF tank
-then zpool import -F tank

spare disks only after the pool is imported and u replace the missing gptid

cannot import 'tank': one or more devices is currently unavailable

I plugged the drive back in but its not being recognized by truenas so the commands gave me this error above. Do you think I’m fucked here?

I shut it down, unplugged and replugged in the cables to that drive and its being recognized again. I just did sudo zpool import -F tank and it seems to be working as I haven’t gotten an error yet. Can you walk me through how to replace the missing gptid?

The order of operations here seems to be, finish waiting for import>replace gptid>use the disk that was previously giving me the errno error as the replacement for the failing drive>remove failing drive from system? Does that seem correct or am I missing something? Thank you for all the help you have no idea how much it means to me truly.

Zpool import failed and now I’m stuck in an endless bootloop.

bootloop after -F usually means the box is wedged trying to bring the pool up on every boot.

boot the truenas install usb → Shell (dont install):
-zpool import -o readonly=on tank
-zpool status -v tank

if readonly import works u still have the data. then:
-zpool export tank
-reboot off the usb into the installed system
-at grub pick an older boot environment if u have one

if installed system still loops b4 any login, from the usb again:
-zpool import tank
-zpool status -v
-zpool replace tank OLDGPTID /dev/disk/by-id/NEWERDRIVE
dont yank the failing drive till replace finishes. paste status -v after the readonly import.

root@truenas-installer: # zpool import -o readonly=on tank cannot import 'tank": pool was previously in use from another system. Last accessed by truenas (hostid=3655dfd1) at Sun Sep 27 23:18:18 2026 The pool can be imported, use 'pool import -f' to import the pool. roototruenas-installer: # zpool status -v tank cannot open 'tank" : no such pool

Hey thanks for the response! So this is the message I get from the shell inside the truenas usb. Should I try the export command now?

yeah that line just wants -f. export does nothing till the pool is imported.

from the usb shell:
-zpool import -f -o readonly=on tank
-zpool status -v tank

if status looks sane (raidz1 w/ at most one UNAVAIL) ur data is still there. paste that status b4 any rw import or replace. dont export yet.

It says state online for all the drives. Tank and raidz1-0 also listed as online.

It also shows permanent errors in a few files but theyre just media files so I don’t really care about those.

It says a resilver is in progress at 0% and that drive that has been stuck expanding the vdev is still listed as well.

online + resilver running means ur past the scary part. leave every drive plugged and let the resilver finish.

-zpool status -v tank every so often
-dont replace / offline / export while its still resilvering
-permanent errors on a few media files are whatever, note the paths and delete or restore those later

after resilver hits 100%:
-zpool status -v tank
-if that expanding / stuck member is still weird, then replace with a known good disk by-id

keep the box on wall power til that bar is done.

Thanks! The resilvering has been at 0% and 0b scanned for hours though. It seems like something is wrong. I’ll check back tomorrow in the morning to see if there’s any progress but this resilvering behavior is how it was acting before I pulled the failing drive and still had the pool show up in the web ui.

Are you checking in the command-line or UI?

Command line via zpool status -v tank. I’m in the shell via a usb boot drive of truenas

Isn’t this pool mounted Read Only at this point and wouldn’t that lock out a resilver or am I missing something?

Well the last thing command before zpool status -v that i typed was zpool import -f -o readonly=on tank but I have no idea if that prevents resilvering from progressing. Something is causing it to not resilver though it has not changed at all sadly.

So I had a backup configuration from a couple of days ago that allowed me to boot into my truenas instance without getting stuck in the bootloop but my pool is still not there. Zpool import shows just one drive as unavailable and tank and raidz as degraded. Whereas truenas from the usb installer from before showed that drive as online and tank and raidz as online. It feels like I’m running in circles here.

The disk that was giving me that errno 6 error turned out to be completely broken, windows couldn’t initialize it, testdisk couldn’t clear it, same with macos. So I’m thinking the only thing I can even try is buying a new drive and doing zpool replace to replace the degraded drive and then try to import?

Do NOT try to “initialise” a drive. If possible, you want it back in the pool.
You can’t replace a drive without first importing the pool.

zpool status above shows an 8-wide raidz1. Would you care to describe your hardware in detail and how the drives are attached?