Unsynchronized timer-vs-caller double-close race in rfcomm_dlc_close β no lock/atomic on CLOSED transition
Summary
rfcomm_dlc_close :151-183 has no lock or atomic state guard. KKASSERT(rd_state!=CLOSED) :156 cannot catch race: T1 (network input DISC/DM/UA from rfcomm_session.c:784,819,876,911) and T2 (softclock timer rfcomm_dlc_timeout :195) both read non-CLOSED before either writes CLOSED :168. T1 callout_stop :164 blocks while T2 completes full close (LIST_REMOVE :166 rd_session=NULL :167 state=CLOSED :168 disconnected callback :170 session-expiry :177-183). T1 resumes: double LIST_REMOVE double disconnected callback double session-expiry scheduling. crit_enter in rfcomm_dlc_timeout :199 only blocks IPIs/softints on T2 CPU does not serialize T1 on different CPU. None of input-path callers take crit. Current socket consumer rfcomm_disconnected (rfcomm_socket.c:166-176) idempotent so practical impact nil today but genuine unsynchronized double-close with list corruption risk for non-idempotent upper layer. Trigger: BT peer opens RFCOMM DLC arms rd_timeout 20s then sends DISC at timer-expiry window. Fix: atomic CAS on rd_state CLOSED transition or lock around close body.
Discussion (0)
PoC verification
Evidence pack
findings/poc/DF-0724 Β· 12 files| File | Type | Description | Size | |
|---|---|---|---|---|
| rfcomm_race_test.c | trigger-source | RFCOMM socket create/bind/connect attempt; documents race-path unreachability (no BT adapter) | 2.6 KB | view raw |
| fix.diff | suggested-fix | atomic_cmpset_short CAS loop on rd_state CLOSED transition in rfcomm_dlc_close | 1.6 KB | view raw |
| build.sh | build-script | cc -o rfcomm_race_test rfcomm_race_test.c | 191 B | view raw |
| run.sh | run-script | runs rfcomm_race_test (requires kldload netbt) | 682 B | view raw |
| build.log | build-log | trigger PoC build output | 74 B | view raw |
| run.log | run-log | decisive run: connect EHOSTUNREACH, race path unreachable | 1.0 KB | view raw |
| fix_build.log | build-log | fixed netbt.ko module build (rc=0, -Werror) | 19.5 KB | view raw |
| env.txt | environment | uname, cc version, netbt.ko status, sysctls, SMAP/SMEP=OFF | 919 B | view raw |
| VERDICT.md | verdict | full narrative: code-level race confirmed, guest unreachable, fix validated | 8.3 KB | β raw |
| README.md | readme | human-readable summary + build/run instructions | 2.3 KB | β raw |
| ../fix_build_combined.log | build-log | Combined 41-finding kernel build (rc=0, -Werror clean) | 5.6 MB | β download |
| ../fix_build_summary.txt | build-summary | Summary of the combined 41-finding kernel build | 826 B | view raw |
DF-0724 β Unsynchronized timer-vs-caller double-close race in rfcomm_dlc_close
Finding
File: sys/netbt/rfcomm_dlc.c:156-183 (Low severity)
Title: Unsynchronized timer-vs-caller double-close race in rfcomm_dlc_close β no lock/atomic on CLOSED transition
Summary: rfcomm_dlc_close (lines 151-183) has no lock or atomic state guard. The KKASSERT(rd_state != CLOSED) at line 156 cannot catch the race: T1 (network input DISC/DM/UA from rfcomm_session.c:784,819,876,911) and T2 (softclock timer rfcomm_dlc_timeout at line 195) both read non-CLOSED before either writes CLOSED at line 168. T1's callout_stop (line 164) blocks while T2 completes the full close (LIST_REMOVE line 166, rd_session=NULL line 167, state=CLOSED line 168, disconnected callback line 170, session-expiry lines 177-183). T1 resumes: double LIST_REMOVE, double disconnected callback, double session-expiry scheduling. The crit_enter in rfcomm_dlc_timeout (line 199) only blocks IPIs/softints on T2's CPU; it does not serialize T1 on a different CPU. None of the input-path callers take crit. The current socket consumer rfcomm_disconnected (rfcomm_socket.c:166-176) is idempotent, so practical impact is nil today, but this is a genuine unsynchronized double-close with list-corruption risk for a non-idempotent upper layer.
Build & Run
./build.sh # cc -o rfcomm_race_test rfcomm_race_test.c
./run.sh # requires: kldload netbt (root)
Expected output
On a guest WITHOUT Bluetooth hardware:
socket(RFCOMM) = 3 OK bind OK connect FAIL errno=65 (No route to host) => no HCI unit / BT adapter; L2CAP link cannot be established; RFCOMM session unreachable done
The race path in rfcomm_dlc_close is not reachable on this guest because:
1. netbt.ko is not loaded by default (requires root kldload)
2. No Bluetooth adapter / HCI unit exists (QEMU guest has no BT hardware)
3. Without an HCI unit, connect() fails with EHOSTUNREACH β no L2CAP link, no RFCOMM session, no DLC timer armed
4. Even with BT hardware, triggering the race requires a remote BT peer sending DISC at the exact timer-expiry window (20s Β± race window)
Fix
See fix.diff β adds an atomic_cmpset_short CAS loop at the top of rfcomm_dlc_close to atomically claim the CLOSED transition, ensuring exactly one caller performs the teardown.
DF-0724 β VERDICT
Verdict: NOT REPRODUCED (race path unreachable on this guest; code-level race confirmed by source tracing)
Mechanism (code-level analysis)
The finding describes a genuine unsynchronized race in rfcomm_dlc_close
(sys/netbt/rfcomm_dlc.c:151-183). The race is real at the code level,
confirmed by tracing the callout subsystem and the close path:
The race scenario
-
T1 (network input path): A BT peer sends a DISC/DM/UA frame.
rfcomm_session.chandlers (lines 784, 819, 876, 911) callrfcomm_dlc_close(dlc, err). T1 entersrfcomm_dlc_close: - PassesKKASSERT(dlc->rd_state != RFCOMM_DLC_CLOSED)at line 156 (state is non-CLOSED). - Begins clearing credit history (lines 160-162). - Callscallout_stop(&dlc->rd_timeout)at line 164. -
T2 (softclock timer): The
rd_timeoutcallout fires on another CPU.rfcomm_dlc_timeout(rfcomm_dlc.c:195) enters: -crit_enter()(line 199) β blocks interrupts on T2's CPU only; does NOT serialize T1 on a different CPU. - Readsdlc->rd_state != RFCOMM_DLC_CLOSED(line 201) β TRUE (T1 hasn't set CLOSED yet). - Callsrfcomm_dlc_close(dlc, ETIMEDOUT)(line 202). - Insiderfcomm_dlc_close: passes KKASSERT, clears credits,callout_stop(self β returns immediately perkern_timeout.c:915-918),LIST_REMOVE(line 166),rd_session = NULL(line 167),rd_state = RFCOMM_DLC_CLOSED(line 168), disconnected callback (line 170), session-expiry scheduling (lines 177-183). - Returns torfcomm_dlc_timeout:crit_exit(), returns. Callout handler completes. -
T1 resumes:
callout_stop(sync=1) was blocking inssleep(kern_timeout.c:910-921) waiting for the callout handler to complete. Now that T2 is done, T1'scallout_stopreturns. T1 continues: -LIST_REMOVE(dlc, rd_next)(line 166) β DOUBLE LIST_REMOVE (T2 already removed dlc from the list). -dlc->rd_session = NULL(line 167) β already NULL (no-op write). -dlc->rd_state = RFCOMM_DLC_CLOSED(line 168) β already CLOSED. -(*dlc->rd_proto->disconnected)()(line 170) β DOUBLE callback. - Session-expiry scheduling (lines 177-183) β DOUBLE scheduling.
Why callout_stop blocks (confirmed)
callout_stop (kern_timeout.c:1091) calls
_callout_cancel_or_stop(cc, CALLOUT_STOP, 1) with sync=1. When the
callout is INPROG (handler executing on another CPU) and sync=1, the
function blocks in ssleep(c, &c->spin, 0, "costp", 0) at line 920,
waiting for the handler to clear the STOP flag. This confirms the
finding's claim that T1's callout_stop blocks while T2 completes the
full close.
Why the KKASSERT doesn't catch it
The KKASSERT(dlc->rd_state != RFCOMM_DLC_CLOSED) at line 156 is a
non-atomic read. Both T1 and T2 read rd_state before either writes
CLOSED (line 168). The assertion passes for both callers. Only after
T2 writes CLOSED (line 168) does the state change, but T1 has already
passed the assertion and is blocked in callout_stop.
Why crit_enter doesn't help
rfcomm_dlc_timeout (line 199) uses crit_enter(), which blocks
interrupts and preemption on the current CPU only (T2's CPU). It
does not prevent T1 (running on a different CPU) from entering
rfcomm_dlc_close concurrently. None of the network-input callers
(rfcomm_session.c) take crit_enter().
Why it does NOT reproduce on this guest
The race path is genuinely unreachable on the QEMU guest:
-
netbt.kois not loaded by default. It is a loadable module (not compiled intoX86_64_GENERIC). Loading requires root (kldload netbt). An unprivileged user cannot load it. -
No Bluetooth adapter / HCI unit exists. The QEMU guest has no BT hardware. There are no
/dev/bt*or/dev/ubt*device nodes, nong_ubtmodule, andsysctl net.bluetooth.hci.unit_listreturns "unknown oid". Without an HCI unit, the L2CAP layer has no link to any peer. -
connect()fails with EHOSTUNREACH. Without an HCI unit, an RFCOMM socket can be created (socket()succeeds) and bound (bind()succeeds), butconnect()fails with errno 65 (EHOSTUNREACH). No RFCOMM session is established, no DLC is created, and nord_timeoutcallout is ever armed. Therfcomm_dlc_closerace path is never entered. -
Even with BT hardware, triggering the race would require a remote BT peer to send a DISC/DM/UA frame at the exact timer-expiry window (20s Β± race window). The window is extremely narrow.
Impact assessment
- Practical impact today: nil. The current socket consumer
rfcomm_disconnected(rfcomm_socket.c:166-176) is idempotent: it setsso->so_error = errand callssoisdisconnected(so), both of which are safe to call twice. The doubleLIST_REMOVEwrites the same values (no corruption in practice unless list mutations occur between T2's remove and T1's resume). The double session-expiry scheduling just re-arms the callout (no double-free). - Theoretical risk: For a non-idempotent upper layer consumer,
the double
disconnectedcallback and doubleLIST_REMOVEcould cause list corruption or use-after-free. The finding correctly rates this as Low severity β it is a genuine code-quality/hardening defect with no demonstrated security impact on the current socket consumer.
Escalation assessment (Phase 6)
This is a race condition that could theoretically cause memory corruption (double LIST_REMOVE β list corruption). However:
- The race is not reachable on this guest (no BT hardware).
- Even if reachable, the corruption is in a linked-list structure
(
rs_dlcs), not a slab object with attacker-controlled content. - The corrupted field (
rd_next.le_prev/le_next) is not a function pointer, refcount, or credential pointer β it's a list linkage. Exploitation would require shaping the list to place a victim object adjacent, which is not feasible through the BT socket interface. - No escalation path exists from this primitive on this guest. This is a valid hard stop: the primitive is not reachable, and even if it were, the corrupted field is not directly exploitable for privilege escalation.
PoC changes
The PoC folder did not exist (no prior PoC). I created:
- rfcomm_race_test.c β trigger attempt: creates an RFCOMM socket,
binds, and tries to connect to a nonexistent BT peer. Documents
that the race path is unreachable (connect fails EHOSTUNREACH).
- fix.diff β adds atomic_cmpset_short CAS loop at the top of
rfcomm_dlc_close to atomically claim the CLOSED transition.
- build.sh, run.sh β repro scripts.
- VERDICT.md, README.md, manifest.json β evidence pack.
Fix validation
- fix.diff applies cleanly:
patch -p1succeeded (hunks at 145, 185). - fix.diff compiles:
makeinsys/netbt/producednetbt.kowithrc=0and-Werror(no warnings, no errors). - Fixed module loads:
kldload /boot/kernel/netbt.kosucceeded after installing the fixed module and rebooting. - Disassembly confirms fix:
objdump -dshowslock cmpxchg %cx,0xc(%rdi)(theatomic_cmpset_shortonrd_stateat offset 0xc) at the top ofrfcomm_dlc_close. - No regression: The trigger PoC runs identically on the fixed module (socket OK, bind OK, connect EHOSTUNREACH β same as unpatched, because the race path is unreachable regardless of the fix).
- fix_status: not_testable β the PoC cannot exercise the race path on this guest (no BT hardware), so a behavioral before/after comparison is not possible. The fix is validated to apply + compile + load + not regress, and the code path is traced to confirm the CAS closes the race.
Fix description
The fix adds an atomic_cmpset_short CAS loop at the top of
rfcomm_dlc_close that atomically transitions rd_state from any
non-CLOSED value to RFCOMM_DLC_CLOSED. If the state is already
CLOSED (another caller won the race), the function returns immediately.
This ensures exactly one caller performs the teardown (LIST_REMOVE,
disconnected callback, session-expiry). The KKASSERT is preserved but
now asserts on the pre-CAS old_state value. The explicit
dlc->rd_state = RFCOMM_DLC_CLOSED assignment (old line 168) is
removed because the CAS already set it.
This fix supersedes any finding-proposal fix (no proposal was present in the DB for this finding).
Fix verification
not_testablenot_testable: no BT HW -> race path unreachable. fix.diff applies+compiles+loads; objdump confirms CAS. No regression (socket/bind OK, connect EHOSTUNREACH as before).
baseline+patched: connect FAIL errno=65 (identical). fix compiles rc=0, loads, objdump: lock cmpxchg %cx,0xc(%rdi) at rfcomm_dlc_close+0x6.
Confirmed kernel references
- sys/netbt/rfcomm_dlc.c:156
- sys/netbt/rfcomm_dlc.c:164
- sys/netbt/rfcomm_dlc.c:168
- sys/netbt/rfcomm_dlc.c:195
- sys/netbt/rfcomm_dlc.c:199
- sys/netbt/rfcomm_dlc.c:201
- sys/kern/kern_timeout.c:888
- sys/kern/kern_timeout.c:910
- sys/kern/kern_timeout.c:915
- sys/kern/kern_timeout.c:1091
- sys/netbt/rfcomm_socket.c:167
- sys/netbt/rfcomm_upper.c:278
Detail
Exploit chain
none -- not exercisable on this guest. Corrupted field is list linkage, not security-sensitive pointer. Valid hard blocker: path not reachable (no BT HW).
Evidence (decisive lines)
socket(RFCOMM)=3 OK / bind OK / connect FAIL errno=65 (No route to host). Identical before/after fix. No panic.
PoC changes
Created entire evidence pack from scratch. rfcomm_race_test.c (trigger attempt), fix.diff (atomic_cmpset_short CAS loop), build.sh, run.sh, VERDICT.md, manifest.json.
Verified recommended fix
Add atomic_cmpset_short CAS loop at rfcomm_dlc_close:156 to atomically claim CLOSED transition. Exactly one caller performs teardown. Full git-apply-able diff in findings/poc/DF-0724/fix.diff.
Verdict
NOT REPRODUCED. Race is a GENUINE code-level defect in rfcomm_dlc_close (callout_stop sync semantics allow T1 to resume after T2 completed full close; KKASSERT at :156 is non-atomic). However race path UNREACHABLE on guest: netbt.ko not loaded by default, no BT adapter, connect(RFCOMM) fails EHOSTUNREACH. Classification (d): unreachable, NOT false positive.
No comments yet.