<feed xmlns='http://www.w3.org/2005/Atom'>
<title>kernel.git/kernel/sched/core.c, branch linux-rolling-stable</title>
<subtitle>Hosts the 0x221E linux distro kernel.
</subtitle>
<id>https://git.0xinfinity.dev/distro/kernel.git/atom?h=linux-rolling-stable</id>
<link rel='self' href='https://git.0xinfinity.dev/distro/kernel.git/atom?h=linux-rolling-stable'/>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/'/>
<updated>2026-03-19T15:15:08Z</updated>
<entry>
<title>sched/mmcid: Avoid full tasklist walks</title>
<updated>2026-03-19T15:15:08Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-03-10T20:29:09Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=81f70f0ee9eae29cd06830261a996c39f3bdd818'/>
<id>urn:sha1:81f70f0ee9eae29cd06830261a996c39f3bdd818</id>
<content type='text'>
[ Upstream commit 192d852129b1b7c4f0ddbab95d0de1efd5ee1405 ]

Chasing vfork()'ed tasks on a CID ownership mode switch requires a full
task list walk, which is obviously expensive on large systems.

Avoid that by keeping a list of tasks using a mm MMCID entity in mm::mm_cid
and walk this list instead. This removes the proven to be flaky counting
logic and avoids a full task list walk in the case of vfork()'ed tasks.

Fixes: fbd0e71dc370 ("sched/mmcid: Provide CID ownership mode fixup functions")
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Matthieu Baerts (NGI0) &lt;matttbe@kernel.org&gt;
Link: https://patch.msgid.link/20260310202526.183824481@kernel.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/mmcid: Remove pointless preempt guard</title>
<updated>2026-03-19T15:15:08Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-03-10T20:29:04Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=a4787eecf756b3d5e0c8b08e0dfe87471f83fbf8'/>
<id>urn:sha1:a4787eecf756b3d5e0c8b08e0dfe87471f83fbf8</id>
<content type='text'>
[ Upstream commit 7574ac6e49789ddee1b1be9b2afb42b4a1b4b1f4 ]

This is a leftover from the early versions of this function where it could
be invoked without mm::mm_cid::lock held.

Remove it and add lockdep asserts instead.

Fixes: 653fda7ae73d ("sched/mmcid: Switch over to the new mechanism")
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Matthieu Baerts (NGI0) &lt;matttbe@kernel.org&gt;
Link: https://patch.msgid.link/20260310202526.116363613@kernel.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/mmcid: Handle vfork()/CLONE_VM correctly</title>
<updated>2026-03-19T15:15:08Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-03-10T20:28:58Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=e6761cdce78a8919a537989afb6aaf6881469f83'/>
<id>urn:sha1:e6761cdce78a8919a537989afb6aaf6881469f83</id>
<content type='text'>
[ Upstream commit 28b5a1395036d6c7a6c8034d85ad3d7d365f192c ]

Matthieu and Jiri reported stalls where a task endlessly loops in
mm_get_cid() when scheduling in.

It turned out that the logic which handles vfork()'ed tasks is broken. It
is invoked when the number of tasks associated to a process is smaller than
the number of MMCID users. It then walks the task list to find the
vfork()'ed task, but accounts all the already processed tasks as well.

If that double processing brings the number of to be handled tasks to 0,
the walk stops and the vfork()'ed task's CID is not fixed up. As a
consequence a subsequent schedule in fails to acquire a (transitional) CID
and the machine stalls.

Cure this by removing the accounting condition and make the fixup always
walk the full task list if it could not find the exact number of users in
the process' thread list.

Fixes: fbd0e71dc370 ("sched/mmcid: Provide CID ownership mode fixup functions")
Closes: https://lore.kernel.org/b24ffcb3-09d5-4e48-9070-0b69bc654281@kernel.org
Reported-by: Matthieu Baerts &lt;matttbe@kernel.org&gt;
Reported-by: Jiri Slaby &lt;jirislaby@kernel.org&gt;
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Matthieu Baerts (NGI0) &lt;matttbe@kernel.org&gt;
Link: https://patch.msgid.link/20260310202526.048657665@kernel.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/mmcid: Prevent CID stalls due to concurrent forks</title>
<updated>2026-03-19T15:15:08Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-03-10T20:28:53Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=f0189d49282e0458f3a737bd486c1ec048148f66'/>
<id>urn:sha1:f0189d49282e0458f3a737bd486c1ec048148f66</id>
<content type='text'>
[ Upstream commit b2e48c429ec54715d16fefa719dd2fbded2e65be ]

A newly forked task is accounted as MMCID user before the task is visible
in the process' thread list and the global task list. This creates the
following problem:

 CPU1			CPU2
 fork()
   sched_mm_cid_fork(tnew1)
     tnew1-&gt;mm.mm_cid_users++;
     tnew1-&gt;mm_cid.cid = getcid()
-&gt; preemption
			fork()
			  sched_mm_cid_fork(tnew2)
			    tnew2-&gt;mm.mm_cid_users++;
                            // Reaches the per CPU threshold
			    mm_cid_fixup_tasks_to_cpus()
			    for_each_other(current, p)
			         ....

As tnew1 is not visible yet, this fails to fix up the already allocated CID
of tnew1. As a consequence a subsequent schedule in might fail to acquire a
(transitional) CID and the machine stalls.

Move the invocation of sched_mm_cid_fork() after the new task becomes
visible in the thread and the task list to prevent this.

This also makes it symmetrical vs. exit() where the task is removed as CID
user before the task is removed from the thread and task lists.

Fixes: fbd0e71dc370 ("sched/mmcid: Provide CID ownership mode fixup functions")
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Matthieu Baerts (NGI0) &lt;matttbe@kernel.org&gt;
Link: https://patch.msgid.link/20260310202525.969061974@kernel.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched: Re-evaluate scheduling when migrating queued tasks out of throttled cgroups</title>
<updated>2026-02-26T23:00:49Z</updated>
<author>
<name>Zicheng Qu</name>
</author>
<published>2026-01-30T08:34:38Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=9c801bf15791ff4092e0473fca0b9eb40c8b16e5'/>
<id>urn:sha1:9c801bf15791ff4092e0473fca0b9eb40c8b16e5</id>
<content type='text'>
[ Upstream commit e34881c84c255bc300f24d9fe685324be20da3d1 ]

Consider the following sequence on a CPU configured with nohz_full:

1) A task P runs in cgroup A, and cgroup A becomes throttled due to CFS
   bandwidth control. The gse (cgroup A) where the task P attached is
dequeued and the CPU switches to idle.

2) Before cgroup A is unthrottled, task P is migrated from cgroup A to
   another cgroup B (not throttled).

   During sched_move_task(), the task P is observed as queued but not
running, and therefore no resched_curr() is triggered.

3) Since the CPU is nohz_full, it remains in do_idle() waiting for an
   explicit scheduling event, i.e., resched_curr().

4) For kernel &lt;= 5.10: Later, cgroup A is unthrottled. However, the task
   P has already been migrated out of cgroup A, so unthrottle_cfs_rq()
may observe load_weight == 0 and return early without resched_curr()
called. For kernel &gt;= 6.6: The unthrottling path normally triggers
`resched_curr()` almost cases even when no runnable tasks remain in the
unthrottled cgroup, preventing the idle stall described above. However,
if cgroup A is removed before it gets unthrottled, the unthrottling path
for cgroup A is never executed. In a result, no `resched_curr()` can be
called.

5) At this point, the task P is runnable in cgroup B (not throttled), but
the CPU remains in do_idle() with no pending reschedule point. The
system stays in this state until an unrelated event (e.g. a new task
wakeup or any cases) that can trigger a resched_curr() breaks the
nohz_full idle state, and then the task P finally gets scheduled.

The root cause is that sched_move_task() may classify the task as only
queued, not running, and therefore fails to trigger a resched_curr(),
while the later unthrottling path no longer has visibility of the
migrated task.

Preserve the existing behavior for running tasks by issuing
resched_curr(), and explicitly invoke check_preempt_curr() for tasks
that were queued at the time of migration. This ensures that runnable
tasks are reconsidered for scheduling even when nohz_full suppresses
periodic ticks.

Fixes: 29f59db3a74b ("sched: group-scheduler core")
Signed-off-by: Zicheng Qu &lt;quzicheng@huawei.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Reviewed-by: Aaron Lu &lt;ziqianlu@bytedance.com&gt;
Tested-by: Aaron Lu &lt;ziqianlu@bytedance.com&gt;
Link: https://patch.msgid.link/20260130083438.1122457-1-quzicheng@huawei.com
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched: Fix build for modules using set_tsk_need_resched()</title>
<updated>2026-02-26T23:00:45Z</updated>
<author>
<name>Gabriele Monaco</name>
</author>
<published>2026-01-12T14:04:13Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=fd0ed8b66c7d12563335dd370acd76cd7e1611e5'/>
<id>urn:sha1:fd0ed8b66c7d12563335dd370acd76cd7e1611e5</id>
<content type='text'>
[ Upstream commit 8d737320166bd145af70a3133a9964b00ca81cba ]

Commit adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
added a tracepoint to the need_resched action that can be triggered also
by set_tsk_need_resched.
This function was previously accessible from out-of-tree modules but
it's no longer available because the __trace_set_need_resched() symbol
is not exported (together with the tracepoint itself, which was exported
in a separate patch) and building such modules fails.

Export __trace_set_need_resched to modules to fix those build issues.

Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
Signed-off-by: Gabriele Monaco &lt;gmonaco@redhat.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Phil Auld &lt;pauld@redhat.com&gt;
Link: https://patch.msgid.link/20260112140413.362202-1-gmonaco@redhat.com
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched: Export hidden tracepoints to modules</title>
<updated>2026-02-26T23:00:45Z</updated>
<author>
<name>Gabriele Monaco</name>
</author>
<published>2025-12-05T13:16:16Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=82962381896fd08918e47feba655ae5680962083'/>
<id>urn:sha1:82962381896fd08918e47feba655ae5680962083</id>
<content type='text'>
[ Upstream commit 6c125b85f3c87b4bf7dba91af6f27d9600b9dba0 ]

The tracepoints sched_entry, sched_exit and sched_set_need_resched
are not exported to tracefs as trace events, this allows only kernel
code to access them. Helper modules like [1] can be used to still have
the tracepoints available to ftrace for debugging purposes, but they do
rely on the tracepoints being exported.

Export the 3 not exported tracepoints.
Note that sched_set_state is already exported as the macro is called
from modules.

[1] - https://github.com/qais-yousef/sched_tp.git

Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
Signed-off-by: Gabriele Monaco &lt;gmonaco@redhat.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Phil Auld &lt;pauld@redhat.com&gt;
Link: https://patch.msgid.link/20251205131621.135513-9-gmonaco@redhat.com
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/mmcid: Don't assume CID is CPU owned on mode switch</title>
<updated>2026-02-16T09:13:33Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-02-10T16:20:51Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=81f29975631db8a78651b3140ecd0f88ffafc476'/>
<id>urn:sha1:81f29975631db8a78651b3140ecd0f88ffafc476</id>
<content type='text'>
commit 1e83ccd5921a610ef409a7d4e56db27822b4ea39 upstream.

Shinichiro reported a KASAN UAF, which is actually an out of bounds access
in the MMCID management code.

   CPU0						CPU1
   						T1 runs in userspace
   T0: fork(T4) -&gt; Switch to per CPU CID mode
         fixup() set MM_CID_TRANSIT on T1/CPU1
   T4 exit()
   T3 exit()
   T2 exit()
						T1 exit() switch to per task mode
						 ---&gt; Out of bounds access.

As T1 has not scheduled after T0 set the TRANSIT bit, it exits with the
TRANSIT bit set. sched_mm_cid_remove_user() clears the TRANSIT bit in
the task and drops the CID, but it does not touch the per CPU storage.
That's functionally correct because a CID is only owned by the CPU when
the ONCPU bit is set, which is mutually exclusive with the TRANSIT flag.

Now sched_mm_cid_exit() assumes that the CID is CPU owned because the
prior mode was per CPU. It invokes mm_drop_cid_on_cpu() which clears the
not set ONCPU bit and then invokes clear_bit() with an insanely large
bit number because TRANSIT is set (bit 29).

Prevent that by actually validating that the CID is CPU owned in
mm_drop_cid_on_cpu().

Fixes: 007d84287c74 ("sched/mmcid: Drop per CPU CID immediately when switching to per task mode")
Reported-by: Shinichiro Kawasaki &lt;shinichiro.kawasaki@wdc.com&gt;
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Tested-by: Shinichiro Kawasaki &lt;shinichiro.kawasaki@wdc.com&gt;
Cc: stable@vger.kernel.org
Closes: https://lore.kernel.org/aYsZrixn9b6s_2zL@shinmob
Reviewed-by: Mathieu Desnoyers &lt;mathieu.desnoyers@efficios.com&gt;
Signed-off-by: Linus Torvalds &lt;torvalds@linux-foundation.org&gt;
Signed-off-by: Greg Kroah-Hartman &lt;gregkh@linuxfoundation.org&gt;
</content>
</entry>
<entry>
<title>sched/mmcid: Drop per CPU CID immediately when switching to per task mode</title>
<updated>2026-02-04T11:21:12Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-02-02T09:39:50Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=007d84287c7466ca68a5809b616338214dc5b77b'/>
<id>urn:sha1:007d84287c7466ca68a5809b616338214dc5b77b</id>
<content type='text'>
When a exiting task initiates the switch from per CPU back to per task
mode, it has already dropped its CID and marked itself inactive. But a
leftover from an earlier iteration of the rework then reassigns the per
CPU CID to the exiting task with the transition bit set.

That's wrong as the task is already marked CID inactive, which means it is
inconsistent state. It's harmless because the CID is marked in transit and
therefore dropped back into the pool when the exiting task schedules out
either through preemption or the final schedule().

Simply drop the per CPU CID when the exiting task triggered the transition.

Fixes: fbd0e71dc370 ("sched/mmcid: Provide CID ownership mode fixup functions")
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Mathieu Desnoyers &lt;mathieu.desnoyers@efficios.com&gt;
Link: https://patch.msgid.link/20260201192835.032221009@kernel.org
</content>
</entry>
<entry>
<title>sched/mmcid: Protect transition on weakly ordered systems</title>
<updated>2026-02-04T11:21:12Z</updated>
<author>
<name>Thomas Gleixner</name>
</author>
<published>2026-02-02T09:39:45Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=47ee94efccf6732e4ef1a815c451aacaf1464757'/>
<id>urn:sha1:47ee94efccf6732e4ef1a815c451aacaf1464757</id>
<content type='text'>
Shrikanth reported a hard lockup which he observed once. The stack trace
shows the following CID related participants:

  watchdog: CPU 23 self-detected hard LOCKUP @ mm_get_cid+0xe8/0x188
  NIP: mm_get_cid+0xe8/0x188
  LR:  mm_get_cid+0x108/0x188
   mm_cid_switch_to+0x3c4/0x52c
   __schedule+0x47c/0x700
   schedule_idle+0x3c/0x64
   do_idle+0x160/0x1b0
   cpu_startup_entry+0x48/0x50
   start_secondary+0x284/0x288
   start_secondary_prolog+0x10/0x14

  watchdog: CPU 11 self-detected hard LOCKUP @ plpar_hcall_norets_notrace+0x18/0x2c
  NIP: plpar_hcall_norets_notrace+0x18/0x2c
  LR:  queued_spin_lock_slowpath+0xd88/0x15d0
   _raw_spin_lock+0x80/0xa0
   raw_spin_rq_lock_nested+0x3c/0xf8
   mm_cid_fixup_cpus_to_tasks+0xc8/0x28c
   sched_mm_cid_exit+0x108/0x22c
   do_exit+0xf4/0x5d0
   make_task_dead+0x0/0x178
   system_call_exception+0x128/0x390
   system_call_vectored_common+0x15c/0x2ec

The task on CPU11 is running the CID ownership mode change fixup function
and is stuck on a runqueue lock. The task on CPU23 is trying to get a CID
from the pool with the same runqueue lock held, but the pool is empty.

After decoding a similar issue in the opposite direction switching from per
task to per CPU mode the tool which models the possible scenarios failed to
come up with a similar loop hole.

This showed up only once, was not reproducible and according to tooling not
related to a overlooked scheduling scenario permutation. But the fact that
it was observed on a PowerPC system gave the right hint: PowerPC is a
weakly ordered architecture.

The transition mechanism does:

    WRITE_ONCE(mm-&gt;mm_cid.transit, MM_CID_TRANSIT);
    WRITE_ONCE(mm-&gt;mm_cid.percpu, new_mode);

    fixup()

    WRITE_ONCE(mm-&gt;mm_cid.transit, 0);

mm_cid_schedin() does:

    if (!READ_ONCE(mm-&gt;mm_cid.percpu))
       ...
       cid |= READ_ONCE(mm-&gt;mm_cid.transit);

so weakly ordered systems can observe percpu == false and transit == 0 even
if the fixup function has not yet completed. As a consequence the task will
not drop the CID when scheduling out before the fixup is completed, which
means the CID space can be exhausted and the next task scheduling in will
loop in mm_get_cid() and the fixup thread can livelock on the held runqueue
lock as above.

This could obviously be solved by using:
     smp_store_release(&amp;mm-&gt;mm_cid.percpu, true);
and
     smp_load_acquire(&amp;mm-&gt;mm_cid.percpu);

but that brings a memory barrier back into the scheduler hotpath, which was
just designed out by the CID rewrite.

That can be completely avoided by combining the per CPU mode and the
transit storage into a single mm_cid::mode member and ordering the stores
against the fixup functions to prevent the CPU from reordering them.

That makes the update of both states atomic and a concurrent read observes
always consistent state.

The price is an additional AND operation in mm_cid_schedin() to evaluate
the per CPU or the per task path, but that's in the noise even on strongly
ordered architectures as the actual load can be significantly more
expensive and the conditional branch evaluation is there anyway.

Fixes: fbd0e71dc370 ("sched/mmcid: Provide CID ownership mode fixup functions")
Closes: https://lore.kernel.org/bdfea828-4585-40e8-8835-247c6a8a76b0@linux.ibm.com
Reported-by: Shrikanth Hegde &lt;sshegde@linux.ibm.com&gt;
Signed-off-by: Thomas Gleixner &lt;tglx@kernel.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Mathieu Desnoyers &lt;mathieu.desnoyers@efficios.com&gt;
Link: https://patch.msgid.link/20260201192834.965217106@kernel.org
</content>
</entry>
</feed>
