1. Oct 12, 2023
  2. Oct 11, 2023
    • Colin Ian King's avatar
      sched/headers: Remove comment referring to rq::cpu_load, since this has been removed · b19fdb16
      Colin Ian King authored
      There is a comment that refers to cpu_load, however, this cpu_load was
      removed with:
      
        55627e3c
      
       ("sched/core: Remove rq->cpu_load[]")
      
      ... back in 2019. The comment does not make sense with respect to this
      removed array, so remove the comment.
      
      Signed-off-by: default avatarColin Ian King <colin.i.king@gmail.com>
      Signed-off-by: default avatarIngo Molnar <mingo@kernel.org>
      Link: https://lore.kernel.org/r/20231010155744.1381065-1-colin.i.king@gmail.com
      b19fdb16
    • Mel Gorman's avatar
      sched/numa: Complete scanning of inactive VMAs when there is no alternative · f169c62f
      Mel Gorman authored
      
      
      VMAs are skipped if there is no recent fault activity but this represents
      a chicken-and-egg problem as there may be no fault activity if the PTEs
      are never updated to trap NUMA hints. There is an indirect reliance on
      scanning to be forced early in the lifetime of a task but this may fail
      to detect changes in phase behaviour. Force inactive VMAs to be scanned
      when all other eligible VMAs have been updated within the same scan
      sequence.
      
      Test results in general look good with some changes in performance, both
      negative and positive, depending on whether the additional scanning and
      faulting was beneficial or not to the workload. The autonuma benchmark
      workload NUMA01_THREADLOCAL was picked for closer examination. The workload
      creates two processes with numerous threads and thread-local storage that
      is zero-filled in a loop. It exercises the corner case where unrelated
      threads may skip VMAs that are thread-local to another thread and still
      has some VMAs that inactive while the workload executes.
      
      The VMA skipping activity frequency with and without the patch:
      
      	6.6.0-rc2-sched-numabtrace-v1
      	=============================
      	    649 reason=scan_delay
      	  9,094 reason=unsuitable
      	 48,915 reason=shared_ro
      	143,919 reason=inaccessible
      	193,050 reason=pid_inactive
      
      	6.6.0-rc2-sched-numabselective-v1
      	=============================
      	    146 reason=seq_completed
      	    622 reason=ignore_pid_inactive
      
      	    624 reason=scan_delay
      	  6,570 reason=unsuitable
      	 16,101 reason=shared_ro
      	 27,608 reason=inaccessible
      	 41,939 reason=pid_inactive
      
      Note that with the patch applied, the PID activity is ignored
      (ignore_pid_inactive) to ensure a VMA with some activity is completely
      scanned. In addition, a small number of VMAs are scanned when no other
      eligible VMA is available during a single scan window (seq_completed).
      The number of times a VMA is skipped due to no PID activity from the
      scanning task (pid_inactive) drops dramatically. It is expected that
      this will increase the number of PTEs updated for NUMA hinting faults
      as well as hinting faults but these represent PTEs that would otherwise
      have been missed. The tradeoff is scan+fault overhead versus improving
      locality due to migration.
      
      On a 2-socket Cascade Lake test machine, the time to complete the
      workload is as follows;
      
                                                       6.6.0-rc2              6.6.0-rc2
                                             sched-numabtrace-v1 sched-numabselective-v1
        Min       elsp-NUMA01_THREADLOCAL      174.22 (   0.00%)      117.64 (  32.48%)
        Amean     elsp-NUMA01_THREADLOCAL      175.68 (   0.00%)      123.34 *  29.79%*
        Stddev    elsp-NUMA01_THREADLOCAL        1.20 (   0.00%)        4.06 (-238.20%)
        CoeffVar  elsp-NUMA01_THREADLOCAL        0.68 (   0.00%)        3.29 (-381.70%)
        Max       elsp-NUMA01_THREADLOCAL      177.18 (   0.00%)      128.03 (  27.74%)
      
      The time to complete the workload is reduced by almost 30%:
      
                           6.6.0-rc2   6.6.0-rc2
                        sched-numabtrace-v1 sched-numabselective-v1 /
        Duration User       91201.80    63506.64
        Duration System      2015.53     1819.78
        Duration Elapsed     1234.77      868.37
      
      In this specific case, system CPU time was not increased but it's not
      universally true.
      
      From vmstat, the NUMA scanning and fault activity is as follows;
      
                                              6.6.0-rc2      6.6.0-rc2
                                    sched-numabtrace-v1 sched-numabselective-v1
        Ops NUMA base-page range updates       64272.00    26374386.00
        Ops NUMA PTE updates                   36624.00       55538.00
        Ops NUMA PMD updates                      54.00       51404.00
        Ops NUMA hint faults                   15504.00       75786.00
        Ops NUMA hint local faults %           14860.00       56763.00
        Ops NUMA hint local percent               95.85          74.90
        Ops NUMA pages migrated                 1629.00     6469222.00
      
      Both the number of PTE updates and hint faults is dramatically
      increased. While this is superficially unfortunate, it represents
      ranges that were simply skipped without the patch. As a result
      of the scanning and hinting faults, many more pages were also
      migrated but as the time to completion is reduced, the overhead
      is offset by the gain.
      
      Signed-off-by: default avatarMel Gorman <mgorman@techsingularity.net>
      Signed-off-by: default avatarIngo Molnar <mingo@kernel.org>
      Tested-by: default avatarRaghavendra K T <raghavendra.kt@amd.com>
      Link: https://lore.kernel.org/r/20231010083143.19593-7-mgorman@techsingularity.net
      f169c62f
    • Mel Gorman's avatar
      sched/numa: Complete scanning of partial VMAs regardless of PID activity · b7a5b537
      Mel Gorman authored
      
      
      NUMA Balancing skips VMAs when the current task has not trapped a NUMA
      fault within the VMA. If the VMA is skipped then mm->numa_scan_offset
      advances and a task that is trapping faults within the VMA may never
      fully update PTEs within the VMA.
      
      Force tasks to update PTEs for partially scanned PTEs. The VMA will
      be tagged for NUMA hints by some task but this removes some of the
      benefit of tracking PID activity within a VMA. A follow-on patch
      will mitigate this problem.
      
      The test cases and machines evaluated did not trigger the corner case so
      the performance results are neutral with only small changes within the
      noise from normal test-to-test variance. However, the next patch makes
      the corner case easier to trigger.
      
      Signed-off-by: default avatarMel Gorman <mgorman@techsingularity.net>
      Signed-off-by: default avatarIngo Molnar <mingo@kernel.org>
      Tested-by: default avatarRaghavendra K T <raghavendra.kt@amd.com>
      Link: https://lore.kernel.org/r/20231010083143.19593-6-mgorman@techsingularity.net
      b7a5b537
  3. Oct 10, 2023
  4. Oct 09, 2023
  5. Oct 07, 2023
  6. Oct 06, 2023
    • Xuewen Yan's avatar
      cpufreq: schedutil: Update next_freq when cpufreq_limits change · 9e0bc36a
      Xuewen Yan authored
      
      
      When cpufreq's policy is 'single', there is a scenario that will
      cause sg_policy's next_freq to be unable to update.
      
      When the CPU's util is always max, the cpufreq will be max,
      and then if we change the policy's scaling_max_freq to be a
      lower freq, indeed, the sg_policy's next_freq need change to
      be the lower freq, however, because the cpu_is_busy, the next_freq
      would keep the max_freq.
      
      For example:
      
      The cpu7 is a single CPU:
      
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # while true;do done& [1] 4737
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # taskset -p 80 4737
        pid 4737's current affinity mask: ff
        pid 4737's new affinity mask: 80
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # cat scaling_max_freq
        2301000
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # cat scaling_cur_freq
        2301000
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # echo 2171000 > scaling_max_freq
        unisoc:/sys/devices/system/cpu/cpufreq/policy7 # cat scaling_max_freq
        2171000
      
      At this time, the sg_policy's next_freq would stay at 2301000, which
      is wrong.
      
      To fix this, add a check for the ->need_freq_update flag.
      
      [ mingo: Clarified the changelog. ]
      
      Co-developed-by: default avatarGuohua Yan <guohua.yan@unisoc.com>
      Signed-off-by: default avatarXuewen Yan <xuewen.yan@unisoc.com>
      Signed-off-by: default avatarGuohua Yan <guohua.yan@unisoc.com>
      Signed-off-by: default avatarIngo Molnar <mingo@kernel.org>
      Acked-by: default avatar"Rafael J. Wysocki" <rafael@kernel.org>
      Link: https://lore.kernel.org/r/20230719130527.8074-1-xuewen.yan@unisoc.com
      9e0bc36a
  7. Oct 04, 2023
  8. Oct 03, 2023
    • Peter Zijlstra's avatar
      sched/eevdf: Fix avg_vruntime() · 650cad56
      Peter Zijlstra authored
      The expectation is that placing a task at avg_vruntime() makes it
      eligible. Turns out there is a corner case where this is not the case.
      
      Specifically, avg_vruntime() relies on the fact that integer division
      is a flooring function (eg. it discards the remainder). By this
      property the value returned is slightly left of the true average.
      
      However! when the average is a negative (relative to min_vruntime) the
      effect is flipped and it becomes a ceil, with the result that the
      returned value is just right of the average and thus not eligible.
      
      Fixes: af4cf404
      
       ("sched/fair: Add cfs_rq::avg_vruntime")
      Signed-off-by: default avatarPeter Zijlstra (Intel) <peterz@infradead.org>
      650cad56
    • Peter Zijlstra's avatar
      sched/eevdf: Also update slice on placement · 2f2fc17b
      Peter Zijlstra authored
      Tasks that never consume their full slice would not update their slice value.
      This means that tasks that are spawned before the sysctl scaling keep their
      original (UP) slice length.
      
      Fixes: 147f3efa
      
       ("sched/fair: Implement an EEVDF-like scheduling policy")
      Signed-off-by: default avatarPeter Zijlstra (Intel) <peterz@infradead.org>
      Link: https://lkml.kernel.org/r/20230915124822.847197830@noisy.programming.kicks-ass.net
      2f2fc17b
    • Kir Kolyshkin's avatar
      sched/headers: Move 'struct sched_param' out of uapi, to work around glibc/musl breakage · d844fe65
      Kir Kolyshkin authored
      Both glibc and musl define 'struct sched_param' in sched.h, while kernel
      has it in uapi/linux/sched/types.h, making it cumbersome to use
      sched_getattr(2) or sched_setattr(2) from userspace.
      
      For example, something like this:
      
      	#include <sched.h>
      	#include <linux/sched/types.h>
      
      	struct sched_attr sa;
      
      will result in "error: redefinition of ‘struct sched_param’" (note the
      code doesn't need sched_param at all -- it needs struct sched_attr
      plus some stuff from sched.h).
      
      The situation is, glibc is not going to provide a wrapper for
      sched_{get,set}attr, thus the need to include linux/sched_types.h
      directly, which leads to the above problem.
      
      Thus, the userspace is left with a few sub-par choices when it wants to
      use e.g. sched_setattr(2), such as maintaining a copy of struct
      sched_attr definition, or using some other ugly tricks.
      
      OTOH, 'struct sched_param' is well known, defined in POSIX, and it won't
      be ever changed (as that would break backward compatibility).
      
      So, while 'struct sched_param' is indeed part of the kernel uapi,
      exposing it the way it's done now creates an issue, and hiding it
      (like this patch does) fixes that issue, hopefully without creating
      another one: common userspace software rely on libc headers, and as
      for "special" software (like libc), it looks like glibc and musl
      do not rely on kernel headers for 'struct sched_param' definition
      (but let's Cc their mailing lists in case it's otherwise).
      
      The alternative to this patch would be to move struct sched_attr to,
      say, linux/sched.h, or linux/sched/attr.h (the new file).
      
      Oh, and here is the previous attempt to fix the issue:
      
        https://lore.kernel.org/all/20200528135552.GA87103@google.com/
      
      
      
      While I support Linus arguments, the issue is still here
      and needs to be fixed.
      
      [ mingo: Linus is right, this shouldn't be needed - but on the other
               hand I agree that this header is not really helpful to
      	 user-space as-is. So let's pretend that
      	 <uapi/linux/sched/types.h> is only about sched_attr, and
      	 call this commit a workaround for user-space breakage
      	 that it in reality is ... Also, remove the Fixes tag. ]
      
      Signed-off-by: default avatarKir Kolyshkin <kolyshkin@gmail.com>
      Signed-off-by: default avatarIngo Molnar <mingo@kernel.org>
      Cc: Linus Torvalds <torvalds@linux-foundation.org>
      Cc: Peter Zijlstra <peterz@infradead.org>
      Link: https://lore.kernel.org/r/20230808030357.1213829-1-kolyshkin@gmail.com
      d844fe65
  9. Oct 02, 2023