β¬’ DragonFlyBSD Kernel Audit
← triage Β· dashboard
DF-1366

Heap OOB write in vegam_populate_smc_vce_level via unbounded VBIOS mm_dep_table->count

Summary

vegam_populate_smc_vce_level at vegam_smumgr.c:1217: VceLevelCount=(uint8_t)(mm_table->count) from VBIOS ucNumEntries (u8 0-255, no clamp). Loop writes VceLevel[SMU75_MAX_LEVELS_VCE=8]. count>8 -> ~3KB heap overflow past DpmTable into vegam_smumgr fields. Sibling of DF-1271/1296/1346. Crafted VBIOS. Fix: clamp to SMU75_MAX_LEVELS_VCE.

Discussion (0)

No comments yet.

PoC verification

Evidence pack

findings/poc/DF-1366 Β· 9 files
FileTypeDescriptionSize
fix.diff suggested-fix git-apply-able fix; compile-validated under -Werror in the amdgpu build 595 B view raw
VERDICT.md verdict full source trace + reachability analysis + compile-validation 3.1 KB ↓ raw
README.md readme summary, mechanism, trigger conditions, fix 3.1 KB ↓ raw
build.sh build script that applies fix.diff and rebuilds the amdgpu 383 B view raw
run.sh run guest reachability probe 659 B view raw
build_fix.log build-log build log slice showing the patched file compiles cleanly (rc=0) 118 B view raw
env.txt environment uname, cc, PCI/kldstat/device-node state proving no HW 1.4 KB view raw
../fix_build_combined.log build-log Combined 41-finding kernel build (rc=0, -Werror clean) 5.6 MB ↓ download
../fix_build_summary.txt build-summary Summary of the combined 41-finding kernel build 826 B view raw
README.md readme summary, mechanism, trigger conditions, fix
↓ download raw

DF-1366 -- Heap OOB write in vegam_populate_smc_vce_level (unclamped VceLevelCount)

File: sys/dev/drm/amd/powerplay/smumgr/vegam_smumgr.c:1217 Class: memory corruption (heap OOB / OOB write / OOB read)

Status: INCONCLUSIVE at runtime -- CONFIRMED source bug, hardware-gated on this guest

The vulnerable code path was traced line-by-line in sys/ and confirmed to be a genuine bug (missing bounds check / integer overflow / unvalidated HBA-supplied index). However it is not exercisable on the DragonFly audit guest because the guest has no amdgpu hardware:

  • PCI shows only QEMU stdvga (vgapci0 class=0x030000 chip=0x11111234), virtio_net and virtio_blk -- no AMD GPU, no Intel iGPU, no LSI SAS HBA, no floppy controller, no TI ThunderLAN NIC, no Emulex OneConnect NIC, no BusLogic SCSI HBA.
  • The amdgpu driver (whether a loadable .ko or compiled-in) never attaches.
  • /dev/fd0, /dev/dri, /dev/dsp* do not exist on this guest.

This is the valid hard-blocker "vulnerable code path unreachable at runtime on this guest AND no harness can exercise it" -- the bug is a real latent defect that would manifest on a system with the relevant hardware (or, for VBIOS-driven bugs, a crafted VBIOS via passthrough/hotplug).

Mechanism (confirmed by source trace)

vegam_populate_smc_vce_level() sets table->VceLevelCount = (uint8_t)(mm_table->count) then loops count < VceLevelCount writing table->VceLevel[count]. The VceLevel[] array is fixed at SMU75_MAX_LEVELS_VCE = 8 (smu75_discrete.h:293). mm_table->count comes from VBIOS ucNumEntries (u8, 0-255) with no clamp, so count>8 overflows into the following fields of SMU75_Discrete_DpmTable (and into vegam_smumgr fields such as power_tune_defaults). Reached during powerplay init.

Live trigger conditions

Requires the amdgpu hardware (and the driver loaded). For VBIOS-driven bugs, requires a crafted VBIOS via PCI passthrough or hotplug. The audit QEMU guest has none of this hardware, so the bug is unreachable here.

Fix

A standalone, git apply-able fix is in fix.diff. Compile-validated: applied to in-guest /usr/src and built with the kernel's -Werror flags (rc=0, no warnings/errors in the patched translation unit). See build_fix.log.

Clamped the count using the same ternary pattern already used elsewhere in the file for pcie-link levels: VceLevelCount = (uint8_t)((SMU75_MAX_LEVELS_VCE < mm_table->count) ? SMU75_MAX_LEVELS_VCE : mm_table->count). (matches finding proposal.)

Reproduce / validate

# 1. Confirm the bug site (read-only source trace):
grep -n ... sys/dev/drm/amd/powerplay/smumgr/vegam_smumgr.c

# 2. Compile-validate the fix on the audit guest:
scp -F dfbsd-qemu/config findings/poc/DF-1366/fix.diff dfbsd:/root/DF-1366.fix.diff
./dfbsd-qemu/vm.sh run_root 'cd /usr/src && patch -p1 --forward < /root/DF-1366.fix.diff'
# Then either:
#   cd /usr/src && make -j6 nativekernel KERNCONF=X86_64_GENERIC        # kernel-internal drivers
# OR
#   cd /usr/src/sys/dev/drm/<module> && KERNCONF=X86_64_GENERIC SYSDIR=/usr/src/sys make -m /usr/src/share/mk  # GPU modules

# 3. (requires real hardware) Exercise the bug: attach the HW and trigger.
VERDICT.md verdict full source trace + reachability analysis + compile-validation
↓ download raw

VERDICT -- DF-1366

Verdict: INCONCLUSIVE at runtime; source bug CONFIRMED; fix COMPILE-VALIDATED

Citations confirmed: - sys/dev/drm/amd/powerplay/smumgr/vegam_smumgr.c:1217 - sys/dev/drm/amd/powerplay/smumgr/vegam_smumgr.c:1220 - sys/dev/drm/amd/powerplay/smumgr/vegam_smumgr.c:1221

Is the bug real? -- YES (source trace)

vegam_populate_smc_vce_level() sets table->VceLevelCount = (uint8_t)(mm_table->count) then loops count < VceLevelCount writing table->VceLevel[count]. The VceLevel[] array is fixed at SMU75_MAX_LEVELS_VCE = 8 (smu75_discrete.h:293). mm_table->count comes from VBIOS ucNumEntries (u8, 0-255) with no clamp, so count>8 overflows into the following fields of SMU75_Discrete_DpmTable (and into vegam_smumgr fields such as power_tune_defaults). Reached during powerplay init.

Can it be reproduced on this guest? -- NO (hardware-gated)

No AMD GPU in PCI list (only QEMU stdvga 0x11111234); amdgpu.ko is not loaded and would not attach without the HW.

The amdgpu driver is a loadable module only (NOT in X86_64_GENERIC), is not loaded, and cannot be kldload'd by an unprivileged user (kldload is root-only). Even loaded, it would not attach without the hardware. Therefore the vulnerable code is unreachable at runtime here. Because the sinks are device-integrated parsers / DRM ioctls / DMA-supplied indices / hardware-dependent paths, no userspace harness on this guest can exercise them. This is the documented valid hard-blocker "unreachable at runtime + no feasible harness"; the bug is a real latent defect with the live trigger conditions noted above.

No escalation chain (and why that is correct here)

There is no memory-corruption primitive to escalate on this guest: the corruption sinks live entirely inside the not-attached driver behind hardware that is absent. The escalation work the audit expects (slab groom -> victim -> uid0) presupposes a reachable write primitive; here there is none on the guest. The deliverable is therefore the confirmed root-cause + a compile-validated fix.

Fix (fix.diff) -- authored and COMPILE-VALIDATED

Clamped the count using the same ternary pattern already used elsewhere in the file for pcie-link levels: VceLevelCount = (uint8_t)((SMU75_MAX_LEVELS_VCE < mm_table->count) ? SMU75_MAX_LEVELS_VCE : mm_table->count). (matches finding proposal.)

The fix was applied to in-guest /usr/src (all hunks applied cleanly) and the module was rebuilt with the kernel's -Werror flags: cd /usr/src/sys/dev/drm/amdgpu && KERNCONF=X86_64_GENERIC SYSDIR=/usr/src/sys make -m /usr/src/share/mk => rc=0 (see build_fix.log). No warnings or errors in the patched translation unit. The runtime before/after of the bug cannot be tested on this guest (no hardware), so fix_status is not_testable (diff applies + compiles; code path traced closed).

Why not not_reproduced (false-positive)?

This is NOT a false positive. The cited sys/ code is genuinely missing the guard / has the overflow / has the unclamped loop -- verified by reading the source. It is a real bug that is simply out of reach of this particular (driverless) QEMU guest.

Confirmed kernel references

Detail

Exploit chain

none (not a memory-corruption primitive on this guest): the bug is in the AMD powerplay SMU manager; amdgpu is not in X86_64_GENERIC, is not kldload'd, and cannot attach without an AMD GPU which is absent.

Evidence (decisive lines)

Source trace: vegam_smumgr.c:1217,1220,1221 + smu75.h:60 + smu75_discrete.h:293 confirmed. Guest PCI: only vgapci0 0x11111234. kldstat: amdgpu not loaded. /dev/dri absent. Fix compile-validated: amdgpu module rebuilt rc=0 under -Werror, vegam_smumgr.o produced, no warnings.

PoC changes

Wrote fresh evidence pack under findings/poc/DF-1366/ plus git-apply-able fix.diff that clamps VceLevelCount to SMU75_MAX_LEVELS_VCE using the same ternary pattern already used for pcie-link levels in this file.

Verified recommended fix

At vegam_smumgr.c:1217 change table->VceLevelCount = (uint8_t)(mm_table->count); to table->VceLevelCount = (uint8_t)((SMU75_MAX_LEVELS_VCE < mm_table->count) ? SMU75_MAX_LEVELS_VCE : mm_table->count);. matches finding proposal.

Verdict

Source-confirmed. vegam_populate_smc_vce_level() at vegam_smumgr.c:1217 sets table->VceLevelCount = (uint8_t)(mm_table->count) from VBIOS ucNumEntries (u8, 0-255, no clamp) then loops count8 -> ~3KB heap overflow past DpmTable into vegam_smumgr fields. Real heap OOB write. Not runtime-reachable: no AMD GPU on guest (only QEMU stdvga chip=0x11111234), amdgpu.ko is a loadable module NOT in GENERIC, not loaded, would not attach without HW.