Port-mips archive
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]
Re: NFS weirdness on EdgeRouter Lite (aka ERLite3)
On Tue, Sep 22, 2026 at 10:04:06AM +0100, Nick Hudson wrote:
>
>
> > On 21 Sep 2026, at 21:46, Alexander Schreiber <als%thangorodrim.ch@localhost> wrote:
> >
> > On Sat, Sep 12, 2026 at 11:37:37AM +0200, Alexander Schreiber wrote:
> >>
> >> Hi,
> >>
> >> On Fri, Sep 04, 2026 at 07:26:34AM +0100, Nick Hudson wrote:
> >>> On 03/09/2026 20:07, Jiaxun Yang wrote:
> >>>> On Thu, 3 Sep 2026, at 7:47 PM, Alexander Schreiber wrote:
> >>>> [...]
> >>>>> Does anybody have any ideas what might be going on here? Am I just holding
> >>>>> it wrong in a particularly weird way?
> >>>> Hi Alex,
> >>>> I was hunting another swap issue on 11.0 which end up being a pmap dirty tracking
> >>>> issue. Haven't got a chance to send-pr as I'm not 100% sure about my conclution.
> >>>> Your symptom seems to be a dirty tracking issue as well.
> >>>> Can you try to pull-up those two commits and see if it fix your problem:
> >>>> https://github.com/NetBSD/src/commit/f3dc7bd2ff5ac5ae49ed4bd4869fad0802d7c6a0
> >>>> https://github.com/NetBSD/src/commit/020b2df9f010c1882cd3ac0a4eb142d377ea02e
> >>>
> >>> I created a diff (and a draft pull-up) for this (refmod.diff)
> >>
> >>
> >> TLDR: Thank you, Jiaxun Yang your changes seem to have fixed it and
> >> thanks to Nick Hudson for a simple to apply patch. Waiting for tests
> >> to complete.
> >>
> >> Just a short interim report: I'm currently waiting on
> >> "./build.sh distribution" with /usr/(src,obj) on NFS to complete. So far
> >> it looks good.
> >
> > And another report: running "./build.sh distribution" with /usr/(src,obj)
> > on NFS for stress testing successfully completed after about 11d.
> >
> > So I call the NFS write problems fixed. This also fixed another issue
> > I noticed: writes to msdos FS (/boot on OCTEON) produced garbage files:
> > - file size correct in ls (but that is in the directory entry)
> > - reading the file ends quickly, producing zeros
> > This no longer happens and writes to msdos FS now work too.
> >
> > The test above was run with:
> > - 11.0 RELEASE kernel
> > - the patch above
> > - configured for uniprocessor
>
> Thanks for confirming. The pull-up for the patch was processed.
>
> https://releng.netbsd.org/cgi-bin/req-11.cgi?show=472
>
>
> >
> > When I changed the kernel config for that to multiprocessor, the kernel
> > exploded during init of cpu1:
> >
> > ------------ cut here ----------------
> > [ 1.0000030] cpu1: Cavium CN5020-500 (0xd0601) Rev. 1 with software emulated floating point
> > [ 1.0000030] cpu1: 64 TLB entriesReader / writer lock error: rw_vector_enter,298: assertion failed: !cpu_intr_p()
> >
> > [ 1.0200030] lock address : netbsd:kernel_map_store+0x8
> > [ 1.0200030] current cpu : 1
> > [ 1.0200030] current lwp : 0x980000041fd68080
> > [ 1.0300030] owner/count : 000000000000000000 flags : 000000000000000000
> >
> > [ 1.0400030] panic: lock error: Reader / writer lock: rw_vector_enter,298: assertion failed: !cpu_intr_p(): lock 0xffffffff808b03f0 cpu 1 lwp 0x980000041fd68080
> > [ 1.0500030] cpu1: Begin traceback...
> > [ 1.0500030] pid 268678176 not found
> > [ 1.0600030] cpu1: End traceback...
> > [ 1.0600030] kernel: breakpoint trap
> > Stopped in pid 0.16 (system) at netbsd:cpu_Debugger+0x4: jr ra
> > bdslot: nop
> > db{1}> bt
> > 0x980000041003b440: cpu_Debugger+0x4 (0x4,0x7000,0x10,0xffffffff80001580) ra 0xf
> > fffffff804fc0c0 sz 0
> > 0x980000041003b440: vpanic+0x1e8 (0x4,0x7000,0x10,0xffffffff80001580) ra 0xfffff
> > fff804fc158 sz 64
> > 0x980000041003b480: panic+0x30 (0x4,0x980000041003b4a8,0xffffffff8066f140,0x12a)
> > ra 0xffffffff804ef0f4 sz 96
> > 0x980000041003b4e0: lockdebug_abort+0xf4 (0x4,0x980000041003b4a8,0xffffffff8066f
> > 140,0x12a) ra 0xffffffff804b48cc sz 176
> > 0x980000041003b590: rw_enter+0x3c4 (0x4,0x980000041003b4a8,0xffffffff8066f140,0x
> > 12a) ra 0xffffffff8043e20c sz 112
> > 0x980000041003b600: uvm_fault_internal+0x14c (0,0x980000041003b4a8,0x1,0x12a) ra
> > 0xffffffff802572e4 sz 576
> > 0x980000041003b840: trap+0x954 (0,0x980000041003b4a8,0x1,0xffffffff804d050c) ra
> > 0xffffffff80245af4 sz 528
> > 0x980000041003ba50: mips64r2_kern_gen_exception+0xd8 (0x70a3,0xffffffff80000008,
> > 0x6,0xffffffff804d050c) ra 0xffffffff804d0514 sz 352
>
> This exception shouldn’t happen. I think interrupts should be really be disabled at this point to prevent this.
>
> > 0x980000041003bbb0: callout_hardclock+0x24 (0x70a3,0xffffffff80000008,0x6,0xffff
> > ffff804d050c) ra 0xffffffff80201324 sz 32
> > 0x980000041003bbd0: cpu_intr+0xfc (0x70a3,0xffffffff80000008,0x6,0xffffffff804d0
> > 50c) ra 0xffffffff80245d40 sz 112
> > 0x980000041003bc40: mips64r2_kern_intr+0xc0 (0,0xff00,0,0xffffffff80001580) ra 0
> > xffffffff804fb72c sz 352
> > 0x980000041003bda0: aprint_normal_internal+0xa4 (0,0xff00,0,0xffffffff80001580)
> > ra 0xffffffff804fb81c sz 48
> > 0x980000041003bdd0: aprint_normal_dev+0x54 (0,0xff00,0x40,0xffffffff80001580) ra
> > 0xffffffff8025a360 sz 96
> > 0x980000041003be30: cpu_identify+0x288 (0,0xff00,0x40,0xffffffff80001580) ra 0xf
> > fffffff80248cb0 sz 80
> > 0x980000041003be80: cpu_hatch+0x128 (0,0xff00,0x40,0xffffffff80001580) ra 0 sz 6
> > 4
> > User-level: pid 0.16
> > db{1}>
> > ------------ cut here ----------------
>
> Please raise a PR.
pr/60778
> >
> > I also tried the HEAD/20260919024529Z snapshot from
> > https://nycdn.netbsd.org/pub/NetBSD-daily/HEAD
> > configured it for multiprocessor. It booted, found both CPUs and ran
> > apparently fine, so I started the same stress test as above (distribution
> > build) and after a few minutes of that it panic()ed:
> >
> > ------------ cut here ----------------
> > login: [ 1836.7526691] panic: kernel diagnostic assertion "lwp_eprio(l) >= ts->ts_eprio" failed: file "/usr/build/HEAD_20260919024529Z/usr/src/sys/kern/kern_turnstile.c", line 356
> > [ 1836.7526691] cpu0: Begin traceback...
> > [ 1836.7526691] pid 280145376 not found
> > [ 1836.7526691] cpu0: End traceback...
> > [ 1836.7526691] rebooting...
> > ------------ cut here ----------------
>
> Please raise a PR.
pr/60781 and was deduped to pr/60776
> >
> > So I turned on all the debug/diag/etc things I could find and that were
> > supported, which gave me on booting:
> >
> > ---------------- cut here ---------------
> > Bootloader: Done loading app on coremask: 0x3
> > [ 1.0000000] MIPS32/64 params: cpu arch: 256
> > [ 1.0000000] MIPS32/64 params: TLB entries: 64
> > [ 1.0000000] MIPS32/64 params: Icache: line=128, total=32768, ways=4, sets=64, colors=1
> > [ 1.0000000] MIPS32/64 params: Dcache: line=128, total=16384, ways=64, sets=2, colors=0
> > [ 1.0000000] MIPS32/64 params: SDcache: line=128, total=131072, ways=8, sets=128, colors=2
> > [ 1.0000000] Dcache is coherent
> > [ 1.0000000] UBSan: Undefined Behavior in /usr/build/HEAD_20260919024529Z/usr/src/sys/arch/mips/mips/mips_machdep.c:1117:20, left shift of 55 by 26 places cannot be represented in type 'int'
> > [ 1.0000000] UBSan: Undefined Behavior in /usr/build/HEAD_20260919024529Z/usr/src/sys/arch/mips/mips/mips_fixup.c:481:2, left shift of negative value -8
> > [ 1.0000000] UBSan: Undefined Behavior in /usr/build/HEAD_20260919024529Z/usr/src/sys/arch/mips/mips/mips_fixup.c:491:46, left shift of negative value -8
> > [ 1.0000000] mips_fixup_addr: unexpected mips insn 0x67bdffe0 at 0xffffffff813c0118
> > [ 1.0000000] mips_fixup_addr: unexpected mips insn 0x67bdfff0 at 0xffffffff813c05f8
> > ---------------- cut here ---------------
> >
> > with a lot of repeats of the "unexpected mips insn" complaint, which
> > https://www.eg.bucknell.edu/~csci320/mips_web/ decoded to:
> > DADDIU - Doubleword Add Immediate Unsigned
> > so I'm guessing that is just the kernel trying to figure out the CPU
> > details.
> >
>
>
> Think there are two PRs here. The UB stuff and the mips_fixup_addr errors.
pr/60779 and pr/60780
Kind regards,
Alex.
--
"Opportunity is missed by most people because it is dressed in overalls and
looks like work." -- Thomas A. Edison
Home |
Main Index |
Thread Index |
Old Index