NetBSD-Bugs archive

[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]

kern/60794: uvm: pagedaemon can deadlock waiting for vmmpepl pool growth



>Number:         60794
>Category:       kern
>Synopsis:       uvm: pagedaemon can deadlock waiting for vmmpepl pool growth
>Confidential:   no
>Severity:       serious
>Priority:       medium
>Responsible:    kern-bug-people
>State:          open
>Class:          sw-bug
>Submitter-Id:   net
>Arrival-Date:   Fri Sep 25 21:25:00 +0000 2026
>Originator:     Izumi Tsutsui
>Release:        NetBSD 11.99.8 around 20260922
>Organization:
>Environment:
System: NetBSD 11.99.8 (GENERIC) #41: Tue Sep 22 08:39:31 JST 2026
       tsutsui@mirage:/s/tsutsui/netbsd/src/sys/arch/news68k/compile/GENERIC
Architecture: m68k
Machine: news68k
>Description:

I got a system hang on a real Sony NWS-1750 running
NetBSD/news68k -current (11.99.8 around 20260922).

The machine was still partly alive.  DDB could be entered with
CTRL-ALT-ESC and ping still worked, but rlogin did not respond.
A cat(1) on the framebuffer console also did not respond to
CTRL-D, CTRL-C, or CTRL-Z.

"ps /w" showed many processes waiting on "vmmpepl".
One cron process was waiting on "plpg".

The vmmpepl pool was completely in use:

        db> show pool/c 2c99c0
        POOLCACHE vmmpepl: itemsize 80, totalmem 98304 ...
                minitems 0, minpages 0, maxpages 4294967295, npages 12
                itemsperpage 102, nitems 0, nout 1224
                nget 1881, nfail 0, nput 657
                npagealloc 12, npagefree 0, hiwat 12, nidle 0

12 * 102 is 1224, so there was no free vm_map_entry left in the
underlying pool.

cron (PID 3569 per `ps` on DDB) was trying to grow this pool and
was waiting for a physical page:

        db> trace/t 0t3569
        mi_switch
        sleepq_block
        mtsleep
        uvm_wait
        uvm_km_kmem_alloc
        pool_page_alloc
        pool_grow
        pool_get
        pool_cache_get_slow
        ...
        uvm_map
        uvm_pagermapin
        genfs_getpages
        VOP_GETPAGES
        ubc_fault
        uvm_fault_internal
        ...
        execve1
        sys_execve

Its wait message was "plpg".

The pagedaemon itself was waiting for the same vmmpepl pool grow:

        db> trace/a 6236c0
        trace: pid 0 lid 96
        mi_switch
        sleepq_block
        cv_wait
        pool_grow
        pool_get
        pool_cache_get_slow
        ...
        uvm_map_clip_start
        uvm_unmap_remove
        uvm_unmap1
        ...
        pmap_page_protect
        genfs_do_putpages
        genfs_putpages
        VOP_PUTPAGES
        uvm_pageout

My reading is that the deadlock looks like this:

        normal LWP (cron)
            vmmpepl pool_grow
                PR_GROWING is set
                needs another physical page
                    uvm_wait("plpg")
                        waits for pagedaemon
                                 |
                                 *
        pagedaemon
            reclaims pages
              -> genfs_do_putpages
                -> uvm_unmap_remove
                  -> uvm_map_clip_start
                    -> needs another vm_map_entry
                      -> vmmpepl is empty
                      -> pool_grow sees PR_GROWING
                        -> waits for cron's pool_grow
                                 |
                             *deadlock*

The UVM state when I entered DDB was:

        pagesize=8192
        1605 VM pages
        585 active
        179 inactive
        23 free
        freemin=16
        free-target=21
        resv-pg=1
        resv-kernel=5
        poolpages=970
        paging=0
        swpages=40499
        swpginuse=1784
        swpgonly=1629

These values are from after the system had already hung.

The UVM waiter state was:

        uvm_pagedaemon_waiters = 1
        uvm_waiter_wakeup      = 0
        uvmpd_sleeping         = 0
        uvmpd_wakeup           = 1

Note this is a uniprocessor m68k, with only 16 MB RAM.

I have left the machine stopped in DDB in this state for now, so I
can collect more information if needed.


>How-To-Repeat:

I do not have a reproducer for now.

The machine had been left mostly idle for about six hours before
I noticed the hang.  A cat(1) process had been left running on
the framebuffer console, but there was no interactive or remote
activity during that time.

It is possible that a daily cron job triggered the condition.
There were several cron-related processes when I entered DDB,
and one cron process was the thread blocked in uvm_wait("plpg"),
but I have not confirmed that the daily job was the trigger.

>Fix:

I do not know the right fix.

It looks like the pagedaemon reclaim path must not end up waiting for
a vmmpepl pool grow which can itself be waiting for the pagedaemon.
I have no idea whether this should be handled in the VM map code,
the pool allocator, or elsewhere.

---
Izumi Tsutsui




Home | Main Index | Thread Index | Old Index