| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
IB/mlx4: Fix use-after-free on pkey sysfs registration failure
register_pkey_tree() ignores errors from register_one_pkey_tree() and
continues registering the remaining slaves. The per-slave error path has
already released the pkey parent kobjects, but their pointers remain
stored in the device. A later device cleanup therefore passes the stale
pointers to kobject_put(), causing a use-after-free.
Clear the parent pointers after releasing a failed slave tree and skip
unregistered trees during device cleanup. This preserves the existing
best-effort registration behavior while preventing a second cleanup of
the failed tree. |
| In the Linux kernel, the following vulnerability has been resolved:
drop_monitor: use timer_shutdown_sync() to prevent timer rearming during teardown
In drop_monitor teardown paths (net_dm_trace_off_set(),
net_dm_hw_monitor_stop(), and error unwind paths in net_dm_trace_on_set()
and net_dm_hw_monitor_start()), per-CPU timers are stopped using
timer_delete_sync() followed by cancel_work_sync().
However, there is a circular dependency between send_timer and
dm_alert_work:
1) sched_send_work() (timer callback) schedules dm_alert_work.
2) send_dm_alert() / net_dm_hw_summary_work() calls reset_per_cpu_data()
or net_dm_hw_reset_per_cpu_data().
3) If memory allocation fails under memory pressure in the reset
function, it re-arms the timer via mod_timer(&data->send_timer, ...).
If dm_alert_work is running concurrently while timer_delete_sync()
executes on another CPU, an allocation failure in the worker will
re-arm the timer after timer_delete_sync() has already returned.
Once cancel_work_sync() completes and module_put() is called, the timer
remains active in the timer wheel. If the module is then unloaded, the
timer will fire and execute sched_send_work() in freed memory,
triggering a kernel panic / use-after-free.
Switch from timer_delete_sync() to timer_shutdown_sync(). This guarantees
that any in-flight timer handler has finished and prevents subsequent
re-arming attempts from running workers from succeeding. When monitoring
is restarted later, timer_setup() is invoked, which cleanly
re-initializes the timer. |
| In the Linux kernel, the following vulnerability has been resolved:
pppoatm: ensure a writable skb header and linear data
In pppoatm_send(), LLC encapsulation checks whether there is sufficient
headroom for the 4-byte LLC header, but does not ensure that the skb header
is writable.
Normal transmit packets passing through ppp_start_xmit() have their header
unshared via skb_cow_head(). However, packets can also reach pppoatm_send()
via PPP channel bridging (PPPIOCBRIDGECHAN) without going through
ppp_start_xmit().
Use skb_cow_head() to ensure both sufficient headroom and a writable
header before pushing the LLC header.
While at it:
- Call pskb_may_pull(skb, 1) before inspecting skb->data[0] to prevent
out-of-bounds reads on zero-length or non-linear frames (e.g. from
bridging).
- Defer SC_COMP_PROT protocol compression until after pppoatm_may_send()
succeeds. This eliminates the temporary skb allocation on admission failure
and completely removes the fragile "undo" heuristic at the nospace label,
avoiding any risk of reading uninitialized headroom or performing an
unbalanced skb_push(). |
| In the Linux kernel, the following vulnerability has been resolved:
smb: client: fix next_buffer UAF and NextCommand bounds in compound PDUs
Fix several related bounds checking and pointer lifecycle issues in
receive_encrypted_standard()'s handling of compound encrypted frames:
- Clear next_buffer after assigning it to server->bigbuf. A stale
next_buffer pointer can lead to a use-after-free on subsequent
error paths.
- Update pdu_length to the decrypted plaintext size (buf_size). Using
the pre-decryption length allows NextCommand to point into stale
ciphertext residue.
- Reject next_cmd values smaller than MID_HEADER_SIZE(server).
- Fix an integer overflow in the upper bound check by verifying
pdu_length - next_cmd < MID_HEADER_SIZE(server), ensuring the
trailing slice is large enough for a header. |
| In the Linux kernel, the following vulnerability has been resolved:
sched/rt,dl: Skip migrate-disabled tasks when picking a push candidate
A migrate_disable()'d RT task cannot be moved to another CPU, but the
scheduler still keeps such a task on that CPU's pushable list
(rq->rt.pushable_tasks) and still marks the runqueue RT-overloaded
(rq->rt.overloaded = 1). So the RT balancer keeps treating this CPU as
having a task to move away, and keeps trying to move the task, but the
push can never succeed. When the head is pinned, push_rt_task() does not
give up either. It falls back to pushing rq->curr instead, using the
per-CPU stopper, as added by commit a7c81556ec4d ("sched: Fix
migrate_disable() vs rt/dl balancing").
The CPU spends tens of milliseconds in this retry loop. The core is
isolated for real-time work, but during the loop nearly half of its time
is consumed by pushes that cannot succeed.
An ftrace capture of the affected CPU, with sched_switch enabled and
commit 94894c9c477e ("sched/rt: Skip currently executing CPU in
rto_next_cpu()") applied, shows where the CPU time went. Two SCHED_FIFO
tasks at equal priority shared the CPU, taskA migrate_disable()'d and
queued, taskB as rq->curr. In one 89 ms window, taskB got only 52 ms of
CPU. The other 37 ms went to the stopper thread.
The scheduler kept trying to push taskA, the pinned head of the pushable
list, fell back to pushing taskB instead, and woke the stopper 5204
times. Every one of those pushes failed and no task was moved. taskA
stayed runnable and queued the whole time, and never ran.
Pushing taskB fails on a re-check. find_lock_lowest_rq() drops the rq
lock to take the target rq lock, then checks again with
"task != pick_next_pushable_task(rq)".
The task being pushed is taskB, but the pick returns taskA, the head of
the pushable list. taskB is rq->curr, and set_next_task_rt() removes the
running task from that list, so taskB can never be the head. The check
expects a candidate taken from the pushable list, but the fallback
pushes rq->curr, which is never on that list. So the check fails every
time.
.--> push-IPI arrives
| |
| v
| pushable head = taskA -> pinned, cannot be pushed
| |
| v
| so push taskB instead -> wake migration/N, a stop-class
| | thread, so it preempts taskB
| v
| re-check compares taskB against the pushable head,
| which is still taskA -> give up
| |
| v
| nothing moved, taskA still queued, rq still overloaded
| |
'----------'
repeats every ~17 us, 5204 times, for 89 ms
The loop cannot stop itself. Every round leaves the runqueue
exactly as it was, so the next push-IPI does the same thing. In
the capture it ended only when taskB went to sleep on its own.
taskA was then picked locally and left the pushable list.
CPU time per task in the window, from sched_switch:
taskB 51.95 ms real work
migration/N 37.18 ms nothing moved
taskA 0.00 ms queued the whole time, never picked
idle 0.01 ms
Counts over the same window:
7667 push-IPIs handled on this CPU
17481 pick_next_pushable_task() returned taskA, still pinned
5204 find_lock_lowest_rq() gave up on the re-check
1 push that actually completed
0 migrations of taskA
The CPU times and the window length come from the standard
sched_switch tracepoint. The counts needed tracepoints added inside
the RT balancer for this investigation.
The self-IPI path is closed by the rto_next_cpu() fix above, and that
part works. But the runqueue is still marked overloaded, because the
pinned task is still advertised as pushable. Other CPUs now send the
push-IPIs during their own RT balancing, and the same loop runs again.
Closing the self-IPI path did not stop a pinn
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
iptfs_skb_reset_frag_walk() advances to the fragment containing @offset
with an unbounded loop:
while (offset >= walk->past + walk->frags[walk->fragi].len)
walk->past += walk->frags[walk->fragi++].len;
walk->fragi is advanced and walk->frags[walk->fragi] is dereferenced
without ever checking fragi against walk->nr_frags. When the requested
offset is at or beyond the total length spanned by the walk's fragments,
fragi runs past nr_frags and off the end of the fixed-size on-stack
frags[MAX_SKB_FRAGS + 1] array, reading out-of-bounds stack memory.
The two callers behave differently: iptfs_skb_add_frags() already guards
against this with
if (!walk->nr_frags ||
offset >= walk->total + walk->initial_offset)
return len;
but iptfs_skb_can_add_frags() has no such guard and calls
iptfs_skb_reset_frag_walk() unconditionally, so it performs the
out-of-range walk. Its own "fragi < walk->nr_frags" bound check runs only
afterwards, too late to prevent the read.
This is reachable from the receive path: a crafted IP-TFS (AGGFRAG)
payload delivered to an IPTFS SA drives iptfs_reassem_cont() ->
iptfs_skb_can_add_frags() with an offset past the fragment total, e.g.:
BUG: KASAN: stack-out-of-bounds in iptfs_skb_reset_frag_walk+0x235/0x250
Read of size 4 at addr ffff888008ad7210 by task repro/345
iptfs_skb_reset_frag_walk+0x235/0x250 net/xfrm/xfrm_iptfs.c:392
iptfs_skb_can_add_frags+0x155/0x310 net/xfrm/xfrm_iptfs.c:420
iptfs_reassem_cont+0xcf8/0x1140 net/xfrm/xfrm_iptfs.c:902
iptfs_input_ordered+0x552/0x670 net/xfrm/xfrm_iptfs.c:1280
iptfs_input+0x3d6/0xde0 net/xfrm/xfrm_iptfs.c:1741
xfrm_input+0x282f/0x6140 net/xfrm/xfrm_input.c:700
xfrm4_esp_rcv+0x93/0x120 net/ipv4/xfrm4_protocol.c:104
ip_rcv+0x278/0x2d0 net/ipv4/ip_input.c:612
Give iptfs_skb_can_add_frags() the same up-front guard that
iptfs_skb_add_frags() already has, so the walk is never entered with an
out-of-range offset. When it triggers, the caller falls back to the
existing linearize-and-copy path, which is safe. |
| In the Linux kernel, the following vulnerability has been resolved:
drm/amdkfd: Avoid integer underflow in EOP ring size calculation.
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as
order_base_2(q->eop_ring_buffer_size / 4) - 1
But order_base_2 can in theory return 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).
Change this to
order_base_2(q->eop_ring_buffer_size / 8)
using properties of logarithms.
Also add to the above comment to make the mathematics more clear.
(cherry picked from commit f0f43fcf8b2b3a924cad9444340921c96ed5f634) |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: mac80211: don't offload TC setup on AP_VLAN interfaces
AP_VLAN interfaces are purely virtual, so don't try to offload
TC setup to drivers. We can't really use the AP interface either
since we may not know it all the time, and it could technically
even change.
Just reject the TC offload so things get done in software. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: mac80211: don't start a ROC while scanning
The ROC work can be pending when a scan starts (which requires
ROC list to be empty, but that's possible), and then a new ROC
can be added to the list and the work will pick it up.
Avoid starting that ROC if a scan made it between things, as
otherwise we'll hit a warning later:
WARNING: net/mac80211/offchannel.c:404 at ieee80211_start_next_roc+0x256/0x2d0
Workqueue: events_unbound cfg80211_wiphy_work
Call Trace:
__ieee80211_scan_completed+0x4fd/0xe40 net/mac80211/scan.c:537
ieee80211_scan_work+0x472/0x1ff0 net/mac80211/scan.c:1193
cfg80211_wiphy_work+0x410/0x570 net/wireless/core.c:513 |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: cfg80211: don't get the radio mask for netdev-less wdevs
cfg80211_calculate_bi_data() calls rdev_get_radio_mask() with
wdev->netdev, which can be NULL and then crashes in mac80211.
To avoid that, invert the order of checks since wdev->netdev
is always valid for beaconing interfaces. |
| In the Linux kernel, the following vulnerability has been resolved:
btrfs: do not force reloc root creation during qgroup_account_snapshot()
[BUG]
When running btrfs/252 with quota enabled through MKFS_OPTIONS="-O quota",
it has a high chance to trigger the following kernel warning and flips
the fs RO:
BTRFS info (device dm-2): relocating block group 30408704 flags metadata|dup
------------[ cut here ]------------
WARNING: fs/btrfs/extent-tree.c:879 at lookup_inline_extent_backref+0x74b/0x960 [btrfs], CPU#4: btrfs/2173
CPU: 4 UID: 0 PID: 2173 Comm: btrfs Not tainted 7.2.0-rc6-custom+ #457 PREEMPT(full) 3adc6528fb66f7a55fe1095385818e742f200aab
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS unknown 02/02/2022
RIP: 0010:lookup_inline_extent_backref+0x74b/0x960 [btrfs]
Call Trace:
<TASK>
insert_inline_extent_backref+0x7c/0x160 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
__btrfs_inc_extent_ref+0xa9/0x270 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
__btrfs_run_delayed_refs+0x4af/0x11c0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_run_delayed_refs+0x9d/0xf0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
create_pending_snapshot+0x39d/0xf00 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
create_pending_snapshots+0x9b/0xc0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_commit_transaction+0x280/0xeb0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
prepare_to_relocate+0x147/0x200 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
relocate_block_group+0x6b/0x5e0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_relocate_block_group+0x92c/0x2380 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_relocate_chunk+0x3f/0x1a0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_balance+0xa2c/0x19c0 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
btrfs_ioctl+0x2839/0x2d30 [btrfs 32f09462c54d9c922fca74a3e4866f4aa7737b72]
__x64_sys_ioctl+0x416/0x9a0
do_syscall_64+0xe1/0x790
entry_SYSCALL_64_after_hwframe+0x4b/0x53
</TASK>
---[ end trace 0000000000000000 ]---
BTRFS info (device dm-2): leaf 4593991680 gen 233 total ptrs 175 free space 5953 owner 2
BTRFS info (device dm-2): refs 3 lock_owner 2173 current 2173
item 0 key (166772736 METADATA_ITEM 1) itemoff 16250 itemsize 33
extent refs 1 gen 222 flags 2
ref#0: tree block backref root 266
[ Skip the tree dump ]
item 174 key (263225344 METADATA_ITEM 0) itemoff 10328 itemsize 33
extent refs 1 gen 162 flags 258
ref#0: tree block backref root 267
BTRFS error (device dm-2): extent item not found for insert, bytenr 179847168 num_bytes 16384 parent 4594335744 root_objectid 273 owner 0 offset 0
BTRFS error (device dm-2): failed to run delayed ref for logical 179847168 num_bytes 16384 type 182 action 1 ref_mod 1: -117
[CAUSE]
The above error is showing that there is a tree reference to a metadata
extent that is no longer there.
With "ref_verify" mount option (requires CONFIG_BTRFS_DEBUG), there is
some extra debug output:
BTRFS error (device dm-2): dumping block entry [180961280 16384], num_refs 0, metadata 1, from disk 0
BTRFS error (device dm-2): root entry 256, num_refs 18446744073709551615
BTRFS error (device dm-2): root entry 273, num_refs 18446744073709551615
BTRFS error (device dm-2): Ref action 3, root 273, ref_root 273, parent 0, owner 0, offset 0, num_refs 1
btrfs_force_cow_block+0x129/0x7d0 [btrfs]
btrfs_cow_block+0x10a/0x250 [btrfs]
btrfs_search_slot+0x5eb/0xf40 [btrfs]
btrfs_insert_empty_items+0x3a/0x70 [btrfs]
insert_with_overflow+0x53/0x130 [btrfs]
btrfs_insert_dir_item+0x125/0x290 [btrfs]
btrfs_add_link+0xaa/0x410 [btrfs]
btrfs_rename+0x5ea/0xcd0 [btrfs]
btrfs_rename2+0x28/0x60 [btrfs]
vfs_rename+0x5b2/0xe10
filename_renameat2+0x244/0x430
__x64_sys_rename+0x48/0x70
do_syscall_64+0xe1/0x790
entry_SYSCALL_64_after_hwframe+0x4b/0x53
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: core: Fix potential UAF after asynchronous card release
Usually a sound driver releases the resources assigned to the card via
snd_card_free(), and it synchronizes with the whole release procedure.
However, when the card is released asynchronously via
snd_card_free_when_closed() like USB-audio driver, the situation is
slightly different; although the snd_card_disconnect() call at the
disconnection guarantees that any newer accesses will be gated, the
in-flight tasks might be still accessing to the underlying card->dev
device even after the disconnection, which would cause a
use-after-free in the end, as reported by fuzzers.
For addressing the bug above, this patch takes the refcount of
card->dev at initialization of the card object, and releases at its
destructor. This assures the availability of the card->dev in its
whole lifecycle. |
| In the Linux kernel, the following vulnerability has been resolved:
drm/amdkfd: Avoid integer underflow with ffs in EOP ring size calc
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as
ffs(q->eop_ring_buffer_size / sizeof(unsigned int)) - 1 - 1
But ffs can in theory return 1 or 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).
Change this to
ffs(q->eop_ring_buffer_size / sizeof(unsigned int) / 4)
using properties of logarithms.
(cherry picked from commit 4f18c56630383c14bfc6b2d65f88f2f895d2121a) |
| In the Linux kernel, the following vulnerability has been resolved:
drm/ttm: fix swapped-out resources never leaving their bulk_move range
ttm_tt_swapout() returns the number of pages swapped out on success and
a negative error code on failure; for a populated ttm it never returns
zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
walk on swapout failure") moved the bulk_move bookkeeping in
ttm_bo_swapout_cb() under "if (!ret)", so the
ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
pair is now skipped on every successful swapout. The equivalent change
for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
infinite LRU walk on backup failure") tests "lret > 0", which is what
was intended here as well.
Before b2ed01e7ad3d the resource was taken off the bulk_move before the
swapout; since then a swapped-out resource stays inside its BO's
bulk_move range (and on the manager LRU) although it is unevictable.
When it is later freed or the BO leaves the bulk_move
(ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
ttm_resource_del_bulk_move() skips it because of its
!ttm_resource_unevictable() guard, so a range endpoint in pos->first /
pos->last is left pointing at freed memory. The next
ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
"list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
dereference in ttm_resource_manager_next() -- minutes to hours after a
hibernation, or at process exit / reboot following one. Samuel
Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
dangling cursor; the missing removal at swapout time is the reason it
dangles.
Testing the condition for success restores the removal. On an AMD
Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
crashed 5 of 18 hibernation cycles; a function profile of one
hibernation showed 336 ttm_tt_swapout() calls and zero
ttm_resource_del_bulk_move_unevictable() calls. With this change the
removal happens for every swapped-out resource and 12 further cycles
were clean. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: sdhci-of-aspeed: Remove children before releasing SDC resources
Probe failure and removal leave SDHCI child devices registered after the
parent clock and managed resources are released.
Unregister the OF children in reverse order before disabling the parent
clock on both paths. Use of_platform_device_destroy() because manual
child creation does not set the flag required by of_platform_depopulate().
This issue was identified during our ongoing static-analysis research
while reviewing kernel code. |
| In the Linux kernel, the following vulnerability has been resolved:
net: dsa: mxl862xx: disable the stats poll on teardown
mxl862xx_setup() arms the stats poll before mxl862xx_setup_mdio(), and
nothing stops it until dsa_register_switch() has returned an error to
mxl862xx_probe(). DSA frees the dsa_port list before it returns, so a
poll that fires once .setup or a later step of dsa_tree_setup() has
failed walks freed ports. On shutdown the user ports stay registered,
and the WORK_STOPPED flag test in mxl862xx_get_stats64() is not atomic
with the cancel in mxl862xx_shutdown(), so a re-arm that read the flag
before it was set queues the poll after cancel_delayed_work_sync() has
returned.
Arm the poll once .setup has succeeded and stop it from a .teardown op,
which DSA calls on unregister and after a failed registration, in both
cases before it frees the ports. Use disable_delayed_work_sync() there
and in shutdown(): it drains a running poll as the cancel did and turns
every later attempt to queue the work into a no-op, so the re-arm
cannot bring the poll back. remove() and the probe error path only set
WORK_STOPPED, which crc_err_work tests before it walks the ports. |
| In the Linux kernel, the following vulnerability has been resolved:
smb: client: fix potential OOB read in smb3_enum_snapshots()
If snapshot_array_size is smaller than GMT_TOKEN_SIZE,
smb3_enum_snapshots() sets ret_data_len to
sizeof(struct smb_snapshot_array) without verifying the actual length
of the server's reply.
Because SMB2_ioctl() places no lower bound on the server-supplied
OutputCount and allocates retbuf to exactly that length, a short reply
results in ret_data_len exceeding the size of retbuf. The subsequent
copy_to_user() then reads past the end of retbuf, leaking adjacent slab
memory to userspace. The subsequent clamp check is ineffective as it
only reduces ret_data_len.
Fix this by rejecting replies shorter than
sizeof(struct smb_snapshot_array) with -EIO. Note that the bound is set
to the 12-byte struct size rather than the 16-byte
MIN_SNAPSHOT_ARRAY_SIZE defined in MS-SMB2 3.3.5.15.1, because 12 bytes
is exactly what copy_to_user() attempts to read. |
| In the Linux kernel, the following vulnerability has been resolved:
KVM: x86/mmu: Check write tracking in all address spaces
kvm_gfn_is_write_tracked() checks only the supplied memslot, but page
tracking is per-address-space and shadow pages are shared across all
address spaces. With SMM, a GFN can therefore be write-tracked in one
address space and appear untracked through the other.
Check the supplied slot first, then the slot for the other address space.
This ensures all callers honor write tracking regardless of the active
address space. In particular, it prevents mmu_try_to_unsync_pages() from
marking an upper-level shadow page unsync and eventually triggering the
BUG in pte_list_remove().
[invert direction of the conditional. - Paolo] |
| In the Linux kernel, the following vulnerability has been resolved:
cgroup: Avoid iteration of dying tasks with zero refcount
The commit 260fbcb92bbea ("cgroup: Move dying_tasks cleanup from
cgroup_task_release() to cgroup_task_free()") extended the lifetime of
tasks on the dying_tasks list.
The iterators have provision to go through dying_tasks because of
dying threadgroup leaders or explicit CSS_TASK_ITER_WITH_DEAD, however,
it was expected that such tasks can obtain a new reference (that is
possible before cgroup_task_release()/put_task_struct_rcu_user()).
The tasks after cgroup_task_release() and before cgroup_task_free()
are subject to race when they may or may not have ->usage count > 0.
The race window is between css_task_iter_next() invocations
when css_set_lock is released and we may arrive at a new ->task_pos.
The iterator should not attempt to resurrect tasks whose ->usage count
dropped to zero. (When that happens, __put_task_struct_rcu_cb() is
already imminent and the returned task_struct would could be used
after free.)
As for the fix, we cannot simply check the signal->live count of a task
on the dying list because that won't distinguish regular zombies waiting
to be reaped from RCU remnant tasks that are going to be free'd.
Therefore add an extra check to rule out ->usage==0 tasks from any
iteration.
The repeat: loop in css_task_iter_advance() doesn't consider ->usage
count, so add a new loop to css_task_iter_next() to skip de-used tasks
on the dying_list.
Rough illustration of the possible race
R (reader of cgroup.procs) T (thread) L (group leader)
--------------------------------- -------------------------------- --------------------------------
L exits, signal->live > 0
cgroup_task_dead(L)
css_set_skip_task_iters() // skips only cset->tasks
list_add_tail(&L->cg_list, &cset->dying_tasks)
css_task_iter_next()
take css_set_lock
css_task_iter_advance()
leader && signal->live != 0
=> it->task_pos = &L->cg_list
release css_set_lock
T exits
--signal->live == 0
cgroup_task_dead(T) // css_set_lock
release_task(T)
cgroup_task_release(T)
release_task(L) // zap_leader
cgroup_task_release(L)
put_task_struct_rcu_user(L)
...RCU...
put_task_struct(L)
L->usage = 0
/* L still on dying_tasks */
...RCU...
__put_task_struct(L)
css_task_iter_next() // another iteration
take css_set_lock
it->task_pos = &L->cg_list
get_task_struct(L)
=> addition on 0
drop css_set_lock
cgroup_task_free(L)
css_set_skip_task_iters() // dying skip comes too late
free_task(L)
cgroup_procs_show()
task_pid_vnr(L) |
| In the Linux kernel, the following vulnerability has been resolved:
sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
When the root scheduler has sub-scheds attached, the COMPAT kfunc
wrappers scx_bpf_select_cpu_and() and scx_bpf_dsq_insert_vtime() refuse
the call and report to @p's scheduler:
scx_error(scx_task_sched(p), "... must be used");
The wrappers are reachable with tasks that have no scheduler.
scx_bpf_select_cpu_and() is in the select_cpu kfunc group, which
scx_kfunc_context_filter() opens to BPF_PROG_TYPE_SYSCALL programs;
scx_bpf_dsq_insert_vtime() is in the enqueue_dispatch group, which
ops.enqueue() and ops.dispatch() may call with any KF_RCU task -- the
group has no kf_tasks validation, and scx_dsq_insert_preamble() checks
task ownership with scx_task_on_sched() precisely because @p may be an
arbitrary task.
scx_task_sched(p) is p->scx.sched, which is NULL for tasks past
sched_ext_dead() -- which clears it via scx_disable_and_exit_task() on
exit -- and for idle tasks, which the enable paths skip as they are
never scheduled through SCX. It is also an rcu_dereference_protected()
that expects @p's pi_lock or rq lock, which neither wrapper holds.
Passing NULL to scx_error() reaches scx_vexit(), which dereferences
sch->exit_info, oopsing the kernel.
One concrete trigger exercised while developing the fix: a
BPF_PROG_TYPE_SYSCALL program calling the select_cpu_and wrapper on an
exited-but-not-reaped task while a sub-scheduler was attached (its pid
stays findable while the zombie is unreaped; faulting instruction is
the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
the offset of sch->exit_info):
sched_ext: BPF scheduler "kfunc_subsched_null" enabled
sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
sched_ext: Unassociated program run_select_cpu_ (id 76)
BUG: kernel NULL pointer dereference, address: 0000000000000398
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
RIP: 0010:scx_vexit+0x25/0xa0
Code: ... <4c> 8b bf 98 03 00 00 ...
CR2: 0000000000000398
Call Trace:
<TASK>
__scx_exit+0x4f/0x70
scx_bpf_select_cpu_and+0xab/0xb0
bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
? __x64_sys_bpf+0x2c/0x40
bpf_prog_test_run_syscall+0x130/0x2f0
__sys_bpf+0x930/0x10d0
? __x64_sys_bpf+0x2c/0x40
__x64_sys_bpf+0x2c/0x40
do_syscall_64+0xbc/0x460
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Read @p's scheduler under RCU instead, which the wrappers can do from
their guard(rcu)(): fault it when it can be determined, and when it
can't be determined -- @p is a task past sched_ext_dead() or an idle
task -- there is nothing obviously wrong to report, so just refuse the
call as before without faulting any scheduler.
These COMPAT wrappers are scheduled for eventual removal once the
deprecation grace period elapses, but until then -- and regardless of
their removal timeline -- they must not oops the kernel on a task they
are handed. |