NetBSD-Bugs archive
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]
kern/60794: uvm: pagedaemon can deadlock waiting for vmmpepl pool growth
>Number: 60794
>Category: kern
>Synopsis: uvm: pagedaemon can deadlock waiting for vmmpepl pool growth
>Confidential: no
>Severity: serious
>Priority: medium
>Responsible: kern-bug-people
>State: open
>Class: sw-bug
>Submitter-Id: net
>Arrival-Date: Fri Sep 25 21:25:00 +0000 2026
>Originator: Izumi Tsutsui
>Release: NetBSD 11.99.8 around 20260922
>Organization:
>Environment:
System: NetBSD 11.99.8 (GENERIC) #41: Tue Sep 22 08:39:31 JST 2026
tsutsui@mirage:/s/tsutsui/netbsd/src/sys/arch/news68k/compile/GENERIC
Architecture: m68k
Machine: news68k
>Description:
I got a system hang on a real Sony NWS-1750 running
NetBSD/news68k -current (11.99.8 around 20260922).
The machine was still partly alive. DDB could be entered with
CTRL-ALT-ESC and ping still worked, but rlogin did not respond.
A cat(1) on the framebuffer console also did not respond to
CTRL-D, CTRL-C, or CTRL-Z.
"ps /w" showed many processes waiting on "vmmpepl".
One cron process was waiting on "plpg".
The vmmpepl pool was completely in use:
db> show pool/c 2c99c0
POOLCACHE vmmpepl: itemsize 80, totalmem 98304 ...
minitems 0, minpages 0, maxpages 4294967295, npages 12
itemsperpage 102, nitems 0, nout 1224
nget 1881, nfail 0, nput 657
npagealloc 12, npagefree 0, hiwat 12, nidle 0
12 * 102 is 1224, so there was no free vm_map_entry left in the
underlying pool.
cron (PID 3569 per `ps` on DDB) was trying to grow this pool and
was waiting for a physical page:
db> trace/t 0t3569
mi_switch
sleepq_block
mtsleep
uvm_wait
uvm_km_kmem_alloc
pool_page_alloc
pool_grow
pool_get
pool_cache_get_slow
...
uvm_map
uvm_pagermapin
genfs_getpages
VOP_GETPAGES
ubc_fault
uvm_fault_internal
...
execve1
sys_execve
Its wait message was "plpg".
The pagedaemon itself was waiting for the same vmmpepl pool grow:
db> trace/a 6236c0
trace: pid 0 lid 96
mi_switch
sleepq_block
cv_wait
pool_grow
pool_get
pool_cache_get_slow
...
uvm_map_clip_start
uvm_unmap_remove
uvm_unmap1
...
pmap_page_protect
genfs_do_putpages
genfs_putpages
VOP_PUTPAGES
uvm_pageout
My reading is that the deadlock looks like this:
normal LWP (cron)
vmmpepl pool_grow
PR_GROWING is set
needs another physical page
uvm_wait("plpg")
waits for pagedaemon
|
*
pagedaemon
reclaims pages
-> genfs_do_putpages
-> uvm_unmap_remove
-> uvm_map_clip_start
-> needs another vm_map_entry
-> vmmpepl is empty
-> pool_grow sees PR_GROWING
-> waits for cron's pool_grow
|
*deadlock*
The UVM state when I entered DDB was:
pagesize=8192
1605 VM pages
585 active
179 inactive
23 free
freemin=16
free-target=21
resv-pg=1
resv-kernel=5
poolpages=970
paging=0
swpages=40499
swpginuse=1784
swpgonly=1629
These values are from after the system had already hung.
The UVM waiter state was:
uvm_pagedaemon_waiters = 1
uvm_waiter_wakeup = 0
uvmpd_sleeping = 0
uvmpd_wakeup = 1
Note this is a uniprocessor m68k, with only 16 MB RAM.
I have left the machine stopped in DDB in this state for now, so I
can collect more information if needed.
>How-To-Repeat:
I do not have a reproducer for now.
The machine had been left mostly idle for about six hours before
I noticed the hang. A cat(1) process had been left running on
the framebuffer console, but there was no interactive or remote
activity during that time.
It is possible that a daily cron job triggered the condition.
There were several cron-related processes when I entered DDB,
and one cron process was the thread blocked in uvm_wait("plpg"),
but I have not confirmed that the daily job was the trigger.
>Fix:
I do not know the right fix.
It looks like the pagedaemon reclaim path must not end up waiting for
a vmmpepl pool grow which can itself be waiting for the pagedaemon.
I have no idea whether this should be handled in the VM map code,
the pool allocator, or elsewhere.
---
Izumi Tsutsui
Home |
Main Index |
Thread Index |
Old Index