NetBSD-Bugs archive

[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]

Re: PR/59663 CVS commit: src/sys/dev



> Module Name:	src
> Committed By:	hannken
> Date:		Wed Jul 29 17:39:23 UTC 2026
> 
> Modified Files:
> 	src/sys/dev: fss.c
> 
> Log Message:
> Bracket I/O from or to the backing store with fstrans_start_lazy()
> so the fss device doesn't block before the file system is suspended.
> 
> Should fix some deadlocks with revokation of ttys.
> 
> PR kern/59663 (ffs_snapshot_read -> uvm_fault (or pool page empty)

Do you hypothesize that this will fix the deadlocks reported in
https://mail-index.netbsd.org/tech-kern/2025/06/29/msg030580.html
and/or this PR?

I'm trying to figure out how it could help, but I don't see any direct
connection to the fss I/O path in the stack traces in question:

- In this PR, dd(1) on /dev/fssN hasn't reached it yet -- it's still
  waiting to busy the buffer in specfs:

  cv_timedwait() at netbsd:cv_timedwait+0xd4
  bbusy() at netbsd:bbusy+0xe3
  getblk() at netbsd:getblk+0x68
  bio_doread() at netbsd:bio_doread+0x1d
  bread() at netbsd:bread+0x18
  spec_read() at netbsd:spec_read+0x1e6
  VOP_READ() at netbsd:VOP_READ+0x42
  vn_read() at netbsd:vn_read+0x136
  dofileread() at netbsd:dofileread+0x79
  sys_read() at netbsd:sys_read+0x49
  syscall() at netbsd:syscall+0x9d

  I'm not sure who has the buffer busy; it's probably _not_ the
  process that's exiting and revoking its controlling terminal because
  I don't see anything in that path that could hold any buffers busy
  for any reason:

  cv_wait_sig() at netbsd:cv_wait_sig+0xd2
  fstrans_setstate() at netbsd:fstrans_setstate+0x10c
  genfs_suspendctl() at netbsd:genfs_suspendctl+0x32
  VFS_SUSPENDCTL() at netbsd:VFS_SUSPENDCTL+0x25
  vfs_suspend() at netbsd:vfs_suspend+0x8d
  vrevoke_suspend_next() at netbsd:vrevoke_suspend_next+0x2a
  vrevoke() at netbsd:vrevoke+0x44
  genfs_revoke() at netbsd:genfs_revoke+0x13
  VOP_REVOKE() at netbsd:VOP_REVOKE+0x3b
  exit1() at netbsd:exit1+0x7bd
  sys_exit() at netbsd:sys_exit+0x3b

  Having stack traces of other threads might help to find the cycle.

- In https://mail-index.netbsd.org/tech-kern/2025/06/29/msg030580.html
  buhrow never followed up with stack traces, but csh was waiting in
  either open, close, or revoke (most likely revoke, I'd guess) on
  specfs_iocv for another thread in the middle of spec_io_enter/exit
  or spec_node_close.  Also not sure how that path could hold any
  buffers busy.

I note that the VCHR path for spec_read/write drops the vnode lock and
enters spec_io_enter/exit transaction instead while it works (or fails
if the vnode was revoked between VOP_UNLOCK and spec_io_enter), in
order to avoid blocking indefinitely with the vnode lock held on,
e.g., tty devices:

   1123 		VOP_UNLOCK(vp);
   1124 		error = spec_io_enter(vp, &sn, &dev);
   1125 		if (error)
   1126 			goto out;
   1127 		error = cdev_read(dev, uio, ap->a_ioflag);
   1128 		spec_io_exit(vp, sn);
   1129 out:		/* XXX What if the caller held an exclusive lock?  */
   1130 		vn_lock(vp, LK_SHARED | LK_RETRY);
   1131 		return error;

https://nxr.netbsd.org/xref/src/sys/miscfs/specfs/spec_vnops.c?r=1.219#1108

I wonder whether the VBLK case should do the same, instead of blocking
indefinitely on bbusy while a snapshot is taken and the file system is
being synced, involving bwrite?

But that's probably not it, because if we reach spec_read, we should
already be in an fstrans(9) transaction, so any concurrent
snapshot-taking has to wait for us to complete anyway -- including
waiting for bbusy.

Side note: I wonder whether we need a path for revoking a VBLK-type
vnode to interrupt concurrent bbusy and make it fail.



Home | Main Index | Thread Index | Old Index