NetBSD-Bugs archive

[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]

Re: kern/60782: SIGIO says socket ready for send before send does



The following reply was made to PR kern/60782; it has been noted by GNATS.

From: mlelstv%serpens.de@localhost (Michael van Elst)
To: gnats-bugs%netbsd.org@localhost
Cc: 
Subject: Re: kern/60782: SIGIO says socket ready for send before send does
Date: Thu, 24 Sep 2026 21:17:18 -0000 (UTC)

 gnats-admin%NetBSD.org@localhost ("campbell+netbsd%mumble.net@localhost via gnats") writes:
 
 >>Number:         60782
 >>Category:       kern
 >>Synopsis:       SIGIO says socket ready for send before send does
 >>Confidential:   no
 >>Severity:       serious
 >>Priority:       medium
 >>Responsible:    kern-bug-people
 >>State:          open
 >>Class:          sw-bug
 >>Submitter-Id:   net
 >>Arrival-Date:   Thu Sep 24 15:00:00 +0000 2026
 >>Originator:     Taylor R Campbell
 >>Release:        current, 11, 10, ...
 >>Organization:
 >The NetBSD Foun*** I/O available
 >>Environment:
 >>Description:
 
 >	Consider a local socket pair,connecting processes A and B,
 >	whose buffer is currently full in the direction from A (sender)
 >	to B (receiver).
 
 >	If process A has subscribed to SIGIO on one socket, and process
 >	B receives a single byte on the peer socket, the system will
 >	send SIGIO with si_code=POLL_OUT si_band=POLLOUT to process A.
 
 >	But if process A then tries to send anything, even a single
 >	byte, it will block, because the default low water mark of the
 >	socket is >>1, and send will block until the space available in
 >	the buffer is at least the low water mark.
 
 >	Perhaps we should avoid sending SIGIO notifications claiming
 >	that the socket is writable when it is not, in fact, writable.
 
 
 
 In unp_rcvd the receiver updates two hiwater marks for the sender:
 
 mbmax is increased for the memory used by mbufs consumed by the reader:
 
      snd->sb_mbmax += unp->unp_mbcnt - rcv->sb_mbcnt;
      unp->unp_mbcnt = rcv->sb_mbcnt;
 
 hiwat is increased for the number of bytes consumed by the reader:
 
      newhiwat = snd->sb_hiwat + unp->unp_cc - rcv->sb_cc;
      (void)chgsbsize(so2->so_uidinfo,
          &snd->sb_hiwat, newhiwat, RLIM_INFINITY);
      unp->unp_cc = rcv->sb_cc;
 
 
 The sender blocks when it tries to write more than there is free space
 in the buffer. But the free space calculation takes both hiwater marks
 into account in sbspace():
 
         if (sb->sb_hiwat <= sb->sb_cc || sb->sb_mbmax <= sb->sb_mbcnt) 
                 return 0;
 
 For the test program, the hiwat value doesn't matter. cc on the sender
 side is 0 as previous writes are already in the receive buffer and
 hiwat has been bumped by the read to 1. But the mbuf highwater value
 stays at 0 until the first mbuf has been consumed.
 
 The first mbuf in the chain contains a header. That's why on amd64 with
 MSIZE=512 you need to read at least 400 bytes (MHLEN) to free the first mbuf.
 
 
 Sending SIGIO and waking up a blocked writer is done at the same time,
 but a blocked writer just blocks again without noticable issue.
 
 poll(2) checks sbspace(sb) >= sb->sb_lowat. So it signals a writable
 descriptor at a different time. If sbspace(sb) > 0 a send smaller
 than sb_lowat may even succeed when, according to poll(2) it's not
 yet possible.
 
 
 Maybe changing the test in sbspace to:
 
 	if (sb->sb_hiwat <= sb->sb_cc)
 		return 0;
 	if (sb->sb_mbmax <= sb->sb_mbcnt) 
 		return 1;
 
 already helps.
 



Home | Main Index | Thread Index | Old Index