<feed xmlns='http://www.w3.org/2005/Atom'>
<title>kernel.git/kernel/sched/fair.c, branch linux-rolling-lts</title>
<subtitle>Hosts the 0x221E linux distro kernel.
</subtitle>
<id>https://git.0xinfinity.dev/distro/kernel.git/atom?h=linux-rolling-lts</id>
<link rel='self' href='https://git.0xinfinity.dev/distro/kernel.git/atom?h=linux-rolling-lts'/>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/'/>
<updated>2026-03-12T11:09:13Z</updated>
<entry>
<title>sched/fair: Fix lag clamp</title>
<updated>2026-03-12T11:09:13Z</updated>
<author>
<name>Peter Zijlstra</name>
</author>
<published>2025-04-22T10:16:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=22a1536b3f5acfb3a5c186653a235716368e3375'/>
<id>urn:sha1:22a1536b3f5acfb3a5c186653a235716368e3375</id>
<content type='text'>
[ Upstream commit 6e3c0a4e1ad1e0455b7880fad02b3ee179f56c09 ]

Vincent reported that he was seeing undue lag clamping in a mixed
slice workload. Implement the max_slice tracking as per the todo
comment.

Fixes: 147f3efaa241 ("sched/fair: Implement an EEVDF-like scheduling policy")
Reported-off-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Tested-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Tested-by: Shubhang Kaushik &lt;shubhang@os.amperecomputing.com&gt;
Link: https://patch.msgid.link/20250422101628.GA33555@noisy.programming.kicks-ass.net
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/eevdf: Update se-&gt;vprot in reweight_entity()</title>
<updated>2026-03-12T11:09:13Z</updated>
<author>
<name>Wang Tao</name>
</author>
<published>2026-01-20T12:31:13Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=c1591343440ef1fd81c06281b8d1f9be45e90a9f'/>
<id>urn:sha1:c1591343440ef1fd81c06281b8d1f9be45e90a9f</id>
<content type='text'>
[ Upstream commit ff38424030f98976150e42ca35f4b00e6ab8fa23 ]

In the EEVDF framework with Run-to-Parity protection, `se-&gt;vprot` is an
independent variable defining the virtual protection timestamp.

When `reweight_entity()` is called (e.g., via nice/renice), it performs
the following actions to preserve Lag consistency:
 1. Scales `se-&gt;vlag` based on the new weight.
 2. Calls `place_entity()`, which recalculates `se-&gt;vruntime` based on
    the new weight and scaled lag.

However, the current implementation fails to update `se-&gt;vprot`, leading
to mismatches between the task's actual runtime and its expected duration.

Fixes: 63304558ba5d ("sched/eevdf: Curb wakeup-preemption")
Suggested-by: Zhang Qiao &lt;zhangqiao22@huawei.com&gt;
Signed-off-by: Wang Tao &lt;wangtao554@huawei.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Tested-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Tested-by: Shubhang Kaushik &lt;shubhang@os.amperecomputing.com&gt;
Link: https://patch.msgid.link/20260120123113.3518950-1-wangtao554@huawei.com
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Only set slice protection at pick time</title>
<updated>2026-03-12T11:09:12Z</updated>
<author>
<name>Peter Zijlstra</name>
</author>
<published>2026-01-23T15:49:09Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=ee54b5ba72d421a7d51c9d580b7b015b45ff8846'/>
<id>urn:sha1:ee54b5ba72d421a7d51c9d580b7b015b45ff8846</id>
<content type='text'>
[ Upstream commit bcd74b2ffdd0a2233adbf26b65c62fc69a809c8e ]

We should not (re)set slice protection in the sched_change pattern
which calls put_prev_task() / set_next_task().

Fixes: 63304558ba5d ("sched/eevdf: Curb wakeup-preemption")
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Tested-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Tested-by: Shubhang Kaushik &lt;shubhang@os.amperecomputing.com&gt;
Link: https://patch.msgid.link/20260219080624.561421378%40infradead.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Fix zero_vruntime tracking</title>
<updated>2026-03-12T11:09:12Z</updated>
<author>
<name>Peter Zijlstra</name>
</author>
<published>2026-02-09T14:28:16Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=99673934a89febe664e704550216638dcb2336a8'/>
<id>urn:sha1:99673934a89febe664e704550216638dcb2336a8</id>
<content type='text'>
[ Upstream commit b3d99f43c72b56cf7a104a364e7fb34b0702828b ]

It turns out that zero_vruntime tracking is broken when there is but a single
task running. Current update paths are through __{en,de}queue_entity(), and
when there is but a single task, pick_next_task() will always return that one
task, and put_prev_set_next_task() will end up in neither function.

This can cause entity_key() to grow indefinitely large and cause overflows,
leading to much pain and suffering.

Furtermore, doing update_zero_vruntime() from __{de,en}queue_entity(), which
are called from {set_next,put_prev}_entity() has problems because:

 - set_next_entity() calls __dequeue_entity() before it does cfs_rq-&gt;curr = se.
   This means the avg_vruntime() will see the removal but not current, missing
   the entity for accounting.

 - put_prev_entity() calls __enqueue_entity() before it does cfs_rq-&gt;curr =
   NULL. This means the avg_vruntime() will see the addition *and* current,
   leading to double accounting.

Both cases are incorrect/inconsistent.

Noting that avg_vruntime is already called on each {en,de}queue, remove the
explicit avg_vruntime() calls (which removes an extra 64bit division for each
{en,de}queue) and have avg_vruntime() update zero_vruntime itself.

Additionally, have the tick call avg_vruntime() -- discarding the result, but
for the side-effect of updating zero_vruntime.

While there, optimize avg_vruntime() by noting that the average of one value is
rather trivial to compute.

Test case:
  # taskset -c -p 1 $$
  # taskset -c 2 bash -c 'while :; do :; done&amp;'
  # cat /sys/kernel/debug/sched/debug | awk '/^cpu#/ {P=0} /^cpu#2,/ {P=1} {if (P) print $0}' | grep -e zero_vruntime -e "^&gt;"

PRE:
    .zero_vruntime                 : 31316.407903
  &gt;R            bash   487     50787.345112   E       50789.145972           2.800000     50780.298364        16     120         0.000000         0.000000         0.000000        /
    .zero_vruntime                 : 382548.253179
  &gt;R            bash   487    427275.204288   E      427276.003584           2.800000    427268.157540        23     120         0.000000         0.000000         0.000000        /

POST:
    .zero_vruntime                 : 17259.709467
  &gt;R            bash   526     17259.709467   E       17262.509467           2.800000     16915.031624         9     120         0.000000         0.000000         0.000000        /
    .zero_vruntime                 : 18702.723356
  &gt;R            bash   526     18702.723356   E       18705.523356           2.800000     18358.045513         9     120         0.000000         0.000000         0.000000        /

Fixes: 79f3f9bedd14 ("sched/eevdf: Fix min_vruntime vs avg_vruntime")
Reported-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Tested-by: Shubhang Kaushik &lt;shubhang@os.amperecomputing.com&gt;
Link: https://patch.msgid.link/20260219080624.438854780%40infradead.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Introduce and use the vruntime_cmp() and vruntime_op() wrappers for wrapped-signed aritmetics</title>
<updated>2026-03-12T11:09:12Z</updated>
<author>
<name>Ingo Molnar</name>
</author>
<published>2025-12-02T15:10:32Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=423b750d87d2e162139c53be709a60a59829f928'/>
<id>urn:sha1:423b750d87d2e162139c53be709a60a59829f928</id>
<content type='text'>
[ Upstream commit 5758e48eefaf111d7764d8f1c8b666140fe5fa27 ]

We have to be careful with vruntime comparisons and subtraction,
due to the possibility of wrapping, so we have macros like:

   #define vruntime_gt(field, lse, rse) ({ (s64)((lse)-&gt;field - (rse)-&gt;field) &gt; 0; })

Which is used like this:

		if (vruntime_gt(min_vruntime, se, rse))
			se-&gt;min_vruntime = rse-&gt;min_vruntime;

Replace this with an easier to read pattern that uses the regular
arithmetics operators:

		if (vruntime_cmp(se-&gt;min_vruntime, "&gt;", rse-&gt;min_vruntime))
			se-&gt;min_vruntime = rse-&gt;min_vruntime;

Also replace vruntime subtractions with vruntime_op():

	-       delta = (s64)(sea-&gt;vruntime - seb-&gt;vruntime) +
	-               (s64)(cfs_rqb-&gt;zero_vruntime_fi - cfs_rqa-&gt;zero_vruntime_fi);
	+       delta = vruntime_op(sea-&gt;vruntime, "-", seb-&gt;vruntime) +
	+               vruntime_op(cfs_rqb-&gt;zero_vruntime_fi, "-", cfs_rqa-&gt;zero_vruntime_fi);

In the vruntime_cmp() and vruntime_op() macros use Use __builtin_strcmp(),
because of __HAVE_ARCH_STRCMP might turn off the compiler optimizations
we rely on here to catch usage bugs.

No change in functionality.

Signed-off-by: Ingo Molnar &lt;mingo@kernel.org&gt;
Stable-dep-of: b3d99f43c72b ("sched/fair: Fix zero_vruntime tracking")
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Rename cfs_rq::avg_vruntime to ::sum_w_vruntime, and helper functions</title>
<updated>2026-03-12T11:09:12Z</updated>
<author>
<name>Ingo Molnar</name>
</author>
<published>2025-12-02T15:09:23Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=028084eca16915b66946671219810f538d269309'/>
<id>urn:sha1:028084eca16915b66946671219810f538d269309</id>
<content type='text'>
[ Upstream commit dcbc9d3f0e594223275a18f7016001889ad35eff ]

The ::avg_vruntime field is a  misnomer: it says it's an
'average vruntime', but in reality it's the momentary sum
of the weighted vruntimes of all queued tasks, which is
at least a division away from being an average.

This is clear from comments about the math of fair scheduling:

    * \Sum (v_i - v0) * w_i := cfs_rq-&gt;avg_vruntime

This confusion is increased by the cfs_avg_vruntime() function,
which does perform the division and returns a true average.

The sum of all weighted vruntimes should be named thusly,
so rename the field to ::sum_w_vruntime. (As arguably
::sum_weighted_vruntime would be a bit of a mouthful.)

Understanding the scheduler is hard enough already, without
extra layers of obfuscated naming. ;-)

Also rename related helper functions:

  sum_vruntime_add()    =&gt; sum_w_vruntime_add()
  sum_vruntime_sub()    =&gt; sum_w_vruntime_sub()
  sum_vruntime_update() =&gt; sum_w_vruntime_update()

With the notable exception of cfs_avg_vruntime(), which
was named accurately.

Signed-off-by: Ingo Molnar &lt;mingo@kernel.org&gt;
Link: https://patch.msgid.link/20251201064647.1851919-7-mingo@kernel.org
Stable-dep-of: b3d99f43c72b ("sched/fair: Fix zero_vruntime tracking")
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Rename cfs_rq::avg_load to cfs_rq::sum_weight</title>
<updated>2026-03-12T11:09:11Z</updated>
<author>
<name>Ingo Molnar</name>
</author>
<published>2025-11-26T11:09:16Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=50890a9a01f34037ac3ec2541ddae5ded6ba320f'/>
<id>urn:sha1:50890a9a01f34037ac3ec2541ddae5ded6ba320f</id>
<content type='text'>
[ Upstream commit 4ff674fa986c27ec8a0542479258c92d361a2566 ]

The ::avg_load field is a long-standing misnomer: it says it's an
'average load', but in reality it's the momentary sum of the load
of all currently runnable tasks. We'd have to also perform a
division by nr_running (or use time-decay) to arrive at any sort
of average value.

This is clear from comments about the math of fair scheduling:

    *              \Sum w_i := cfs_rq-&gt;avg_load

The sum of all weights is ... the sum of all weights, not
the average of all weights.

To make it doubly confusing, there's also an ::avg_load
in the load-balancing struct sg_lb_stats, which *is* a
true average.

The second part of the field's name is a minor misnomer
as well: it says 'load', and it is indeed a load_weight
structure as it shares code with the load-balancer - but
it's only in an SMP load-balancing context where
load = weight, in the fair scheduling context the primary
purpose is the weighting of different nice levels.

So rename the field to ::sum_weight instead, which makes
the terminology of the EEVDF math match up with our
implementation of it:

    *              \Sum w_i := cfs_rq-&gt;sum_weight

Signed-off-by: Ingo Molnar &lt;mingo@kernel.org&gt;
Link: https://patch.msgid.link/20251201064647.1851919-6-mingo@kernel.org
Stable-dep-of: b3d99f43c72b ("sched/fair: Fix zero_vruntime tracking")
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Have SD_SERIALIZE affect newidle balancing</title>
<updated>2026-02-11T12:41:45Z</updated>
<author>
<name>Peter Zijlstra</name>
</author>
<published>2025-11-17T16:13:09Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=13de38aa3ea7aacad3c0bc1312d46582c866dc72'/>
<id>urn:sha1:13de38aa3ea7aacad3c0bc1312d46582c866dc72</id>
<content type='text'>
commit 522fb20fbdbe48ed98f587d628637ff38ececd2d upstream.

Also serialize the possiblty much more frequent newidle balancing for
the 'expensive' domains that have SD_BALANCE set.

Initial benchmarking by K Prateek and Tim showed no negative effect.

Split out from the larger patch moving sched_balance_running around
for ease of bisect and such.

Suggested-by: Shrikanth Hegde &lt;sshegde@linux.ibm.com&gt;
Seconded-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Link: https://lkml.kernel.org/r/df068896-82f9-458d-8fff-5a2f654e8ffd@amd.com
Link: https://patch.msgid.link/6fed119b723c71552943bfe5798c93851b30a361.1762800251.git.tim.c.chen@linux.intel.com
Signed-off-by: Tim Chen &lt;tim.c.chen@linux.intel.com&gt;
Signed-off-by: Greg Kroah-Hartman &lt;gregkh@linuxfoundation.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Skip sched_balance_running cmpxchg when balance is not due</title>
<updated>2026-02-11T12:41:45Z</updated>
<author>
<name>Tim Chen</name>
</author>
<published>2025-11-10T18:47:35Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=de7cb4282dafc1af7b3ef02338df54a6d7e3e5a8'/>
<id>urn:sha1:de7cb4282dafc1af7b3ef02338df54a6d7e3e5a8</id>
<content type='text'>
commit 3324b2180c17b21c31c16966cc85ca41a7c93703 upstream.

The NUMA sched domain sets the SD_SERIALIZE flag by default, allowing
only one NUMA load balancing operation to run system-wide at a time.

Currently, each sched group leader directly under NUMA domain attempts
to acquire the global sched_balance_running flag via cmpxchg() before
checking whether load balancing is due or whether it is the designated
load balancer for that NUMA domain. On systems with a large number
of cores, this causes significant cache contention on the shared
sched_balance_running flag.

This patch reduces unnecessary cmpxchg() operations by first checking
that the balancer is the designated leader for a NUMA domain from
should_we_balance(), and the balance interval has expired before
trying to acquire sched_balance_running to load balance a NUMA
domain.

On a 2-socket Granite Rapids system with sub-NUMA clustering enabled,
running an OLTP workload, 7.8% of total CPU cycles were previously spent
in sched_balance_domain() contending on sched_balance_running before
this change.

         : 104              static __always_inline int arch_atomic_cmpxchg(atomic_t *v, int old, int new)
         : 105              {
         : 106              return arch_cmpxchg(&amp;v-&gt;counter, old, new);
    0.00 :   ffffffff81326e6c:       xor    %eax,%eax
    0.00 :   ffffffff81326e6e:       mov    $0x1,%ecx
    0.00 :   ffffffff81326e73:       lock cmpxchg %ecx,0x2394195(%rip)        # ffffffff836bb010 &lt;sched_balance_running&gt;
         : 110              sched_balance_domains():
         : 12234            if (atomic_cmpxchg_acquire(&amp;sched_balance_running, 0, 1))
   99.39 :   ffffffff81326e7b:       test   %eax,%eax
    0.00 :   ffffffff81326e7d:       jne    ffffffff81326e99 &lt;sched_balance_domains+0x209&gt;
         : 12238            if (time_after_eq(jiffies, sd-&gt;last_balance + interval)) {
    0.00 :   ffffffff81326e7f:       mov    0x14e2b3a(%rip),%rax        # ffffffff828099c0 &lt;jiffies_64&gt;
    0.00 :   ffffffff81326e86:       sub    0x48(%r14),%rax
    0.00 :   ffffffff81326e8a:       cmp    %rdx,%rax

After applying this fix, sched_balance_domain() is gone from the profile
and there is a 5% throughput improvement.

[peterz: made it so that redo retains the 'lock' and split out the
         CPU_NEWLY_IDLE change to a separate patch]
Signed-off-by: Tim Chen &lt;tim.c.chen@linux.intel.com&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Reviewed-by: Chen Yu &lt;yu.c.chen@intel.com&gt;
Reviewed-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Reviewed-by: Shrikanth Hegde &lt;sshegde@linux.ibm.com&gt;
Reviewed-by: K Prateek Nayak &lt;kprateek.nayak@amd.com&gt;
Reviewed-by: Srikar Dronamraju &lt;srikar@linux.ibm.com&gt;
Tested-by: Mohini Narkhede &lt;mohini.narkhede@intel.com&gt;
Tested-by: Shrikanth Hegde &lt;sshegde@linux.ibm.com&gt;
Link: https://patch.msgid.link/6fed119b723c71552943bfe5798c93851b30a361.1762800251.git.tim.c.chen@linux.intel.com
Signed-off-by: Greg Kroah-Hartman &lt;gregkh@linuxfoundation.org&gt;
</content>
</entry>
<entry>
<title>sched/fair: Fix pelt clock sync when entering idle</title>
<updated>2026-01-30T09:32:20Z</updated>
<author>
<name>Vincent Guittot</name>
</author>
<published>2026-01-21T16:33:17Z</published>
<link rel='alternate' type='text/html' href='https://git.0xinfinity.dev/distro/kernel.git/commit/?id=79a074be9b57e921f307bdd48c453d46e372c22e'/>
<id>urn:sha1:79a074be9b57e921f307bdd48c453d46e372c22e</id>
<content type='text'>
[ Upstream commit 98c88dc8a1ace642d9021b103b28cba7b51e3abc ]

Samuel and Alex reported regressions of the util_avg of RT rq with
commit 17e3e88ed0b6 ("sched/fair: Fix pelt lost idle time detection").
It happens that fair is updating and syncing the pelt clock with task one
when pick_next_task_fair() fails to pick a task but before the prev
scheduling class got a chance to update its pelt signals.

Move update_idle_rq_clock_pelt() in set_next_task_idle() which is called
after prev class has been called.

Fixes: 17e3e88ed0b6 ("sched/fair: Fix pelt lost idle time detection")
Closes: https://lore.kernel.org/all/CAG2KctpO6VKS6GN4QWDji0t92_gNBJ7HjjXrE+6H+RwRXt=iLg@mail.gmail.com/
Closes: https://lore.kernel.org/all/8cf19bf0e0054dcfed70e9935029201694f1bb5a.camel@mediatek.com/
Reported-by: Samuel Wu &lt;wusamuel@google.com&gt;
Reported-by: Alex Hoh &lt;Alex.Hoh@mediatek.com&gt;
Signed-off-by: Vincent Guittot &lt;vincent.guittot@linaro.org&gt;
Signed-off-by: Peter Zijlstra (Intel) &lt;peterz@infradead.org&gt;
Tested-by: Samuel Wu &lt;wusamuel@google.com&gt;
Tested-by: Alex Hoh &lt;Alex.Hoh@mediatek.com&gt;
Link: https://patch.msgid.link/20260121163317.505635-1-vincent.guittot@linaro.org
Signed-off-by: Sasha Levin &lt;sashal@kernel.org&gt;
</content>
</entry>
</feed>
