Unprivileged local DoS via u_int truncation of iov_len in /dev/null and /dev/zero write (infinite kernel loop)
| Field | Value |
|---|---|
| ID | DF-0079 |
| Status | new |
| Severity | Medium |
| CVSS 3.1 | CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H |
| CWE | CWE-835 Loop with Unreachable Exit Condition; CWE-197 Integer Truncation |
| File | sys/kern/kern_memio.c |
| Lines | 292-382 |
| Area | kern (/dev/mem, /dev/null, /dev/zero drivers) |
| Confidence | certain |
| Discovered | 2026-06-30 |
| Reported | pending |
Summary
In mmrw(), the per-iteration byte count c is declared u_int (32-bit)
(kern_memio.c:225), while iov->iov_len is size_t (64-bit on amd64). The
/dev/null write path assigns c = iov->iov_len; directly with no clamp
(:298); /dev/zero does the same (:364).
When the caller supplies a length whose low 32 bits are zero (e.g. exactly
2Β³Β²), c truncates to 0. After the switch, the bookkeeping at :379-382
performs iov->iov_len -= c (subtracting 0) and uio->uio_resid -= c
(subtracting 0), leaving the loop state unchanged. The while (uio->uio_resid > 0
&& error == 0) predicate at :232 is still true, so mmrw spins forever in
kernel context.
The upper-layer sys_write path does not clamp nbyte below 2Β³Β²: it only
rejects (ssize_t)nbyte < 0 (sys_generic.c:336-337), so iov_len = 2Β³Β² passes
through. Worse, on the /dev/null write path no uiomove/copyin is ever
issued, so the user buffer pointer is never validated β the attacker can pass an
arbitrary (even unmapped) address.
/dev/null is created mode 0666 (:842) and /dev/zero mode 0666
(:848), so this is reachable by any local unprivileged user. A single
syscall wedges one kernel thread indefinitely; N syscalls wedge N CPUs.
Recommended fix
Widen c to size_t throughout mmrw (and adjust the (int)c casts passed to
uiomove), or clamp c to a bounded value in the /dev/null and /dev/zero
write cases:
--- a/sys/kern/kern_memio.c
+++ b/sys/kern/kern_memio.c
@@ case 2: /* /dev/null */
if (uio->uio_rw == UIO_READ)
return (0);
- c = iov->iov_len;
+ c = min(iov->iov_len, PAGE_SIZE);
break;
Apply the same clamp to case 12 (/dev/zero write, :363-366) and case 1
(/dev/kmem, :264). The cleanest fix is to make c size_t everywhere.
Proof of concept
See findings/poc/DF-0079/. A one-line write(fd, buf, 0x100000000ULL) to
/dev/null pegs a CPU forever.
Timeline
- 2026-06-30 Discovered during automated file-by-file audit of
sys/kern/kern_memio.c. - pending Reported to DragonFlyBSD security contact.
Discussion (0)
PoC verification
Evidence pack
findings/poc/DF-0079 Β· 15 files| File | Type | Description | Size | |
|---|---|---|---|---|
| df0079.c | trigger-source | Minimal unprivileged trigger: write(/dev/null, (void*)0x1, 0x100000000). Fork-N mode to wedge N CPUs. | 2.7 KB | view raw |
| watch_df0079.sh | observer | Serial-console watcher: polls pgrep df0079 and writes ps/top to /dev/ttyd0 (survives reset). | 730 B | view raw |
| build.sh | build-script | cc -o df0079 df0079.c | 196 B | view raw |
| run.sh | run-script | Runs ./df0079 [N]; documents the DoS and observation guidance. | 1.2 KB | view raw |
| build.log | build-log | Clean build output + /dev/null & /dev/zero perms (0666) + vulnerable source excerpt (unpatched baseline). | 567 B | view raw |
| run.log | run-log | Host-side baseline DoS timeline (ssh unreachable t+1s) + ps excerpt from decisive run on #0. | 2.6 KB | view raw |
| serial_wedge_capture.txt | run-log | Per-iteration ps/top from serial watcher on #0: pid 852, UID 1001, STAT R0, cputime 0.50s -> 20.56s, 50% system (one core wedged). | 4.6 KB | view raw |
| env.txt | environment | uname, cc, ncpu, device perms, vulnerable line citations. | 1.8 KB | view raw |
| fix.diff | suggested-fix | One-line root-cause fix: widen u_int c -> size_t c at kern_memio.c:225 (closes /dev/null, /dev/zero AND /dev/kmem). Validated on a built+booted single-fix kernel. | 268 B | view raw |
| fix_build.log | build-log | Full nativekernel build output after applying fix.diff (35510 lines, rc=0, no errors). kern_memio.o rebuilt at 15:50. | 5.6 MB | β download |
| fix_run.log | run-log | Patched #1 kernel PoC re-run: 5 runs (single /dev/null x3, /dev/zero x1, fork-2 x1) all return write()=4294967296 in 0.00s, EXIT=0, guest stays UP; + baseline contrast narrative. | 5.3 KB | view raw |
| VERDICT.md | verdict | Full narrative: mechanism (path:line each hop), reproduced evidence, exploit ceiling, fix rationale, Phase-8 fix-validation before/after. | 8.7 KB | β raw |
| README.md | readme | Human summary: status, root cause, build/run, expected result, evidence index. | 3.7 KB | β raw |
| ../fix_build_combined.log | build-log | Combined 41-finding kernel build (rc=0, -Werror clean) | 5.6 MB | β download |
| ../fix_build_summary.txt | build-summary | Summary of the combined 41-finding kernel build | 826 B | view raw |
DF-0079 PoC β /dev/null (and /dev/zero) infinite kernel loop DoS
Status: REPRODUCED (trivial unprivileged local full-system DoS)
A single write(fd, buf, (size_t)1<<32) to world-writable /dev/null (mode
0666) by an unprivileged user pegs one CPU at 100% in kernel context
forever. The write() syscall never returns; the process is unkillable
from userspace; only a reboot recovers. Forking N copies (one per CPU) wedges
all cores.
Root cause (verified in audited master DEV sys/kern/kern_memio.c)
In mmrw(), the per-iteration byte count c is declared u_int (32-bit)
at kern_memio.c:225, while iov->iov_len is size_t (64-bit). The
/dev/null write path assigns c = iov->iov_len; directly with no clamp
(:298); /dev/zero write does the same (:364).
When the caller passes a length whose low 32 bits are zero (e.g. exactly
2Β³Β² = 0x100000000), c truncates to 0. After the switch, the bookkeeping
at :379-382 performs iov->iov_len -= c (subtracting 0) and
uio->uio_resid -= c (subtracting 0), leaving the loop state unchanged. The
while (uio->uio_resid > 0 && error == 0) predicate at :232 is still true,
so mmrw spins forever in kernel context.
The early if (iov->iov_len == 0) { ... continue; } guard at :234 does NOT
trip, because it compares the full 64-bit iov_len (= 2Β³Β², not 0).
The upper-layer sys_write does not clamp nbyte below 2Β³Β²: it only
rejects (ssize_t)nbyte < 0 (sys_generic.c:336), so iov_len = 2Β³Β² passes
through. Worse, on the /dev/null write path no uiomove/copyin is ever
issued, so the user buffer pointer is never validated β the attacker can pass
an arbitrary (even unmapped) address; the PoC passes (void *)0x1.
Build & run
cc -o df0079 df0079.c # or: ./build.sh
./df0079 # pegs 1 CPU forever; or ./run.sh
./df0079 4 # fork 4 copies to wedge 4 CPUs
Run as any local user β /dev/null and /dev/zero are mode 0666
(kern_memio.c:842/:847). No privilege required.
Expected result (on a vulnerable kernel)
Each invocation calls write(/dev/null, buf, 0x100000000) and never
returns. The kernel thread spins in mmrw at :232. Observation (serial
console, since ssh itself gets starved) shows the wedged process in state
R0 (running on CPU, not blocked) with cputime climbing ~1.18 s per wall
second (= 100% of one core) indefinitely, and top reporting one CPU fully
in sys. On a 2-CPU guest a single wedge typically makes the box
unresponsive to ssh within ~1 s (the wedged CPU also services the network IRQ).
Recovery is a hard reset only.
Verified evidence (in this folder)
run.logβ host-side DoS timeline (ssh unreachable t+1 s) + ps excerpt.serial_wedge_capture.txtβ per-iterationps/topfrom a serial-console watcher (pid 852, UID 1001, STATR0, cputime 0.50 s β 20.56 s).build.logβ clean build +/dev/null//dev/zeroperms + source excerpt.env.txtβ guest uname, cc, ncpu, device perms, vulnerable lines.VERDICT.mdβ full narrative + path:line mechanism + fix.fix.diffβ one-line root-cause fix: widenu_int cβsize_t c.watch_df0079.shβ serial-console observer (writes ps/top to/dev/ttyd0).manifest.jsonβ artifact catalog.
Notes
- The same bug affects
/dev/zerowrite (:364) and/dev/kmem(:264, root-only). The fix infix.diff(widenctosize_t) closes all three. - The wedged process cannot be killed (SIGKILL/SIGTERM are never delivered: the thread is in an unyielding kernel loop with no signal-check point), so the only recovery is a reboot.
DF-0079 β VERDICT
Verdict: REPRODUCED β trivial unprivileged local full-system DoS
A single write(/dev/null, buf, (size_t)1<<32) by an unprivileged user
(uid 1001, not in wheel) wedges one CPU at 100% in kernel context forever.
The syscall never returns, the process is unkillable from userspace, and on a
small guest the whole machine becomes unreachable within ~1 second. Forking N
copies wedges N cores. Recovery is a hard reset only.
Mechanism (every hop cited path:line, confirmed in audited master DEV)
-
sys_write(sys/kern/sys_generic.c:336) only rejects(ssize_t)nbyte < 0.nbyte = 2Β³Β²is positive as a 64-bitssize_t, so it is accepted.sys_generic.c:340setsaiov.iov_len = uap->nbyte(= 2Β³Β²) and:344setsauio.uio_resid = uap->nbyte(= 2Β³Β²). No clamping anywhere. -
The write reaches
mmwriteβmmrw(sys/kern/kern_memio.c:222). The per-iteration byte count is declaredu_int c;(32-bit) at:225. -
mmrwenterswhile (uio->uio_resid > 0 && error == 0)at:232. The early guardif (iov->iov_len == 0) { ... continue; }at:234does NOT fire because the full 64-bitiov_lenis 2Β³Β², not 0. -
For
/dev/null(minor 2),case 2:at:292returns early on read (:296-297) and on write executesc = iov->iov_len;at:298β a directsize_tβu_inttruncation. Withiov_len = 0x100000000the low 32 bits are zero, soc = 0. There is nouiomove/copyinon this path (the user buffer pointer(void*)0x1is never dereferenced), so noEFAULTrescues us. -
After
breakat:299, control falls to the bookkeeping at:377-382:iov->iov_base += c(+= 0),iov->iov_len -= c(-= 0 β still 2Β³Β²),uio->uio_offset += c(+= 0),uio->uio_resid -= c(-= 0 β still 2Β³Β²). -
The
whilepredicate at:232is still true (uio_resid == 2Β³Β² > 0,error == 0). Steps 3β5 repeat with zero net change to the loop state.mmrwspins forever in kernel context on the calling CPU. -
The mem cdev is
D_MPSAFE | D_QUICK(:85) andmmrwholds no lock while spinning, so the wedge is a pure unyielding tight kernel loop β not a lockup. It never blocks, never callslwkt_yield/uiomove/tsleep, and never reaches a signal-check point, so: - the thread is stuck in stateR(running) consuming ~100% of one CPU; - it cannot be preempted or signalled (SIGKILL never takes effect); - if that CPU also services the network IRQ (vtnet), sshd is starved and the guest becomes unreachable within ~1 s.
/dev/null and /dev/zero are created mode 0666 (:842/:847), so the
attack is reachable by any local user. /dev/zero write (case 12,
:363-365) has the identical c = iov->iov_len truncation and the same
infinite loop. (/dev/kmem :264 has it too but is root-only.)
Reproduced evidence
serial_wedge_capture.txtβ a serial-console watcher (writingps/topto/dev/ttyd0, which lands in the hostboot.logand survives reset) captured the wedged process across 18 iterations. Decisive excerpt:PID PPID STAT UID %CPU TIME COMMAND 852 1 R0 1001 0.0 0:00.50 ./df0079 (t+0.0s) 852 1 R0 1001 0.0 0:01.68 ./df0079 (t+0.7s) ... 852 1 R0 1001 0.0 0:20.56 ./df0079 (t+19.5s) CPU states: 0.0% user, 0.0% nice, 50.0% system, 0.0% interrupt, 50.0% idleSTATR0= running on CPU (not blocked);UID 1001= unprivilegedmaxx; cputime grows ~1.18 s per wall-second = 100% of one of two CPUs; residual never drains. (50% system / 50% idle = one CPU fully wedged; the fork-N variant takes the rest.)run.logβ host-side timeline: a single detached./df0079(one process) made the guest unresponsive to ssh within 1 second; ssh stayed unreachable for the full 10 s observation window. No kernel panic (clean starvation, perboot.log). Onlyvm.sh resetrecovered the guest. Reproduced across three independent runs (each requiring a reset).build.log,env.txtβ clean build;/dev/null&/dev/zerobothcrw-rw-rw-(0666); the vulnerableu_int c;decl confirmed at:225.
Exploit chain / weaponization
Not a memory-corruption primitive β there is no corruption, only an infinite
loop β so there is no escalation chain. The realistic impact ceiling is
reliable unprivileged local full-system denial of service:
- One write() per CPU pegs every core at 100% in-kernel and never returns.
- The wedged threads are unkillable (no signal-check point), so even an
administrator cannot recover short of a reboot.
- Trivially scriptable: a one-line attacker (python -c 'import os; os.write(os.open("/dev/null",1),b"x"*(1<<32))' or the C PoC) run by any user, or auto-started at login by a compromised low-priv account, takes the whole machine down indefinitely.
Fix (in fix.diff)
Root-cause, one-line fix: widen the per-iteration byte count from 32-bit to
64-bit so c = iov->iov_len cannot truncate:
--- a/sys/kern/kern_memio.c
+++ b/sys/kern/kern_memio.c
@@ -222,7 +222,7 @@
mmrw(cdev_t dev, struct uio *uio, int flags)
{
int o;
- u_int c;
+ size_t c;
With c being size_t, c = iov->iov_len (= 2Β³Β²) no longer truncates; the
bookkeeping at :380/:382 subtracts the full 2Β³Β², draining uio_resid to 0
in a single iteration, so the while predicate at :232 exits and the write
returns normally. This closes /dev/null (:298), /dev/zero (:364) and
/dev/kmem (:264) in one change. All existing function-argument uses of c
(uiomove/read_random/add_buffer_randomness_src at :253/:289/:309/:314/:319/:328/:354/:372)
operate on values already bounded by min(β¦, PAGE_SIZE), so widening c to
size_t compiles cleanly and changes no bounded-path behaviour.
This supersedes the finding markdown's per-case c = min(iov->iov_len, PAGE_SIZE)
clamp proposal: that also works for /dev/null and /dev/zero but would need
separate edits at three sites and would not fix the latent /dev/kmem (:264)
truncation. Widening c is the single-change root-cause fix the finding itself
names as "cleanest".
FIX VALIDATION (Phase 8 β built & booted a single-fix kernel)
The fix.diff was validated end-to-end on the DragonFly master DEV guest:
Baseline (unpatched #0, kern.version 6.5-DEVELOPMENT #0: Thu Jul 2 06:02:54 UTC 2026):
A single detached ./df0079 (uid 1001, write(/dev/null, 0x1, 2^32)) wedged
the box: within ~1 s ssh to both dfbsd-maxx and dfbsd returned
Connection reset by peer / kex_exchange_identification: read: Connection
reset; a follow-up root ssh hit timeout 124 after 8 s; vm.sh status
reported down. No kernel panic (clean starvation, no new output in
boot.log) β consistent with a tight unyielding in-kernel loop in mmrw.
Only vm.sh reset with-src recovered the guest. Baseline DoS reproduced
deterministically.
Patched (#1, kern.version 6.5-DEVELOPMENT #1: Thu Jul 2 15:52:03 UTC 2026,
kernel.stripped sha256 7f49550e41a1fba71b19d49f44f4d5f6144e720657a68c4a215b233a04d9b677):
Built by applying ONLY fix.diff to /usr/src and running
make -j6 nativekernel KERNCONF=X86_64_GENERIC (rc=0, no errors β see
fix_build.log). The SAME PoC was re-run 5Γ on the patched kernel:
| run | variant | result | real time | exit | guest after |
|---|---|---|---|---|---|
| 1 | /dev/null single |
write returned 4294967296 |
0.00 s | 0 | UP, load 0.1 |
| 2 | /dev/null single |
write returned 4294967296 |
0.00 s | 0 | UP |
| 3 | /dev/null single |
write returned 4294967296 |
0.00 s | 0 | UP |
| 4 | /dev/zero single (case 12) |
write returned 4294967296 |
0.00 s | 0 | UP |
| 5 | /dev/null fork-2 (2 children) |
both write returned 4294967296 |
0.00 s | 0 | UP, load 0.12 |
write() now returns the full 2^32 = 4294967296 (drained in one iteration
because c is 64-bit), in sub-millisecond time. The guest stays fully
responsive; load stays ~0.1; no CPU is pegged; no process lingers. The DoS is
gone on the patched kernel and present on the unpatched baseline β a
clean before/after. fix_status: fixed.
The full patched-kernel run output (all 5 runs + guest state) is in
fix_run.log; the full nativekernel build output (35510 lines, rc=0) is in
fix_build.log.
PoC changes made during verification
df0079.c: rewritten with an explanatory header citing every relevantkern_memio.cline, atrigger()helper, and a fork-N mode. The core trigger is unchanged from the original PoC:write(fd, (void*)0x1, 0x100000000ULL).- Added
watch_df0079.sh(serial-console observer),build.sh,run.sh,env.txt,VERDICT.md, thisREADME.md,fix.diff,manifest.json, and the full logs (build.log,run.log,serial_wedge_capture.txt).
Fix verification
fixedVALIDATED the fix: ./df0079 (write /dev/null 2^32) on the unpatched 6.5-DEVELOPMENT #0 baseline wedged the box (ssh to both maxx and root starved within ~1s, vm.sh status => down, only reset recovered β baseline DoS reproduced) and does NOT wedge on the single-fix #1 kernel (write returns 4294967296 in 0.00s across 5 runs incl. /dev/zero and fork-2 variants; guest stays UP, load ~0.1) => fix.diff closes the bug. nativekernel build rc=0, no errors.
BASELINE #0: LAUNCHED pid=926 -> maxx ssh 'Connection reset by peer' at t+1.5s -> root ssh timeout 124 at t+8s -> vm.sh status 'down' -> only vm.sh reset recovered. PATCHED #1: ./df0079 -> 'write returned 4294967296' / '0.00 real 0.00 user 0.00 sys' / EXIT=0 (x5 runs: /dev/null x3, /dev/zero x1, fork-2 x1); guest UP, load 0.12, no CPU pegged. nativekernel rc=0 (fix_build.log).
Confirmed kernel references
- sys/kern/kern_memio.c:225
- sys/kern/kern_memio.c:232
- sys/kern/kern_memio.c:234
- sys/kern/kern_memio.c:264
- sys/kern/kern_memio.c:298
- sys/kern/kern_memio.c:364
- sys/kern/kern_memio.c:379
- sys/kern/kern_memio.c:380
- sys/kern/kern_memio.c:382
- sys/kern/kern_memio.c:842
- sys/kern/kern_memio.c:847
- sys/kern/sys_generic.c:336
- sys/kern/sys_generic.c:340
- sys/kern/sys_generic.c:344
Detail
Exploit chain
Unprivileged local full-system DoS via u_int truncation of iov_len in /dev/null (and /dev/zero) write -> infinite kernel loop in mmrw; the wedged thread is unkillable (no signal-check point in the tight loop) so only a reboot recovers; N writes wedge N cores. Not a memory-corruption primitive, so no escalation chain β the realistic impact ceiling is reliable unprivileged local denial of service. Trivially scriptable (one-line write() to a mode-0666 device).
Evidence (decisive lines)
BASELINE (#0, unpatched): LAUNCHED pid=926 -> sleep 1.5s -> 'kex_exchange_identification: read: Connection reset by peer' (maxx ssh dead); follow-up root ssh: SSH_RC=124 (timeout); vm.sh status => 'down'; only vm.sh reset with-src recovered. No panic in boot.log (tight-loop starvation). PATCHED (#1, fix.diff applied): [15:58:13] ./df0079 -> 'write returned 4294967296 (unexpected!)' / '0.00 real 0.00 user 0.00 sys' / TRIGGER_EXIT=0; repeated 5x (/dev/null x3, /dev/zero x1, fork-2 x1) all identical; guest UP, load 0.12, no CPU pegged. nativekernel build rc=0 (fix_build.log, 35510 lines), kern_memio.o rebuilt 15:50, mmrw/mmwrite symbols present.
PoC changes
No source changes (df0079.c/build.sh/run.sh unchanged from prior verified session β the trigger was already correct). fix.diff unchanged (already the verified one-line root-cause fix). Added this session: fix_build.log (full nativekernel output, rc=0), fix_run.log (5 patched-kernel runs + baseline contrast), and appended a Phase-8 fix-validation section to VERDICT.md; refreshed manifest.json with fix_kernel_uname/sha256 and the new artifacts.
Verified recommended fix
Widen the per-iteration byte count in mmrw() from 32-bit to 64-bit at sys/kern/kern_memio.c:225: change 'u_int c;' to 'size_t c;'. With c being size_t, 'c = iov->iov_len' (= 2^32) no longer truncates to 0; the bookkeeping at :380/:382 subtracts the full 2^32, draining uio_resid to 0 in one iteration so the while(uio_resid>0) predicate at :232 exits and the write returns. This single change closes /dev/null (:298), /dev/zero (:364) AND /dev/kmem (:264, root-only). All other in-function uses of c operate on values already bounded by min(..., PAGE_SIZE), so the widening compiles cleanly (verified rc=0) and changes no bounded-path behaviour. SUPERSIDES the finding markdown's per-case 'c = min(iov->iov_len, PAGE_SIZE)' clamp proposal (which would need 3 separate edits and would not fix the latent /dev/kmem truncation). Full git-apply-able diff in findings/poc/DF-0079/fix.diff; validated on a built+booted single-fix kernel.
Verdict
REPRODUCED + FIX VALIDATED. On the unpatched #0 master DEV kernel, a single detached ./df0079 (uid 1001, write(/dev/null, 0x1, 2^32)) wedged the box: within ~1 s ssh to BOTH dfbsd-maxx and dfbsd returned 'Connection reset by peer' / 'kex_exchange_identification: read: Connection reset', a follow-up root ssh hit timeout 124 after 8 s, and vm.sh status reported 'down'. No kernel panic (clean starvation, no new boot.log output) β consistent with a tight unyielding in-kernel loop in mmrw (kern_memio.c:232) where the u_int c at :225 truncates 2^32 -> 0, so the bookkeeping at :380/:382 subtracts 0 and the while(uio_resid>0) predicate never becomes false. Only vm.sh reset recovered the guest. The single-fix #1 kernel (same source + fix.diff: u_int c -> size_t c at :225) makes the same write return 4294967296 (=2^32) in 0.00 s real across 5 runs (single /dev/null x3, /dev/zero case-12 x1, fork-2 x1), guest stays UP with load ~0.1. Clean before/after => fix closes the bug.
No comments yet.