1. Mar 25, 2022
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: fix comments · 59d7cb27
      Jiajian Ye authored
      Two adjustments are made:
      
      1. Correct a grammatical error: replace the "what" in "Do the job what
         you want to debug" with "that".
      
      2. Replace "has not been" with "has been" in the description of the -f
         option: According to Commit b1c9ba071e7d ("tools/vm/page_owner_sort.c:
         fix the instructions for use"), the description of the "-f" option is
         "Filter out the information of blocks whose memory has been released."
      
      Link: https://lkml.kernel.org/r/20220301151438.166118-1-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      59d7cb27
    • Yixuan Cao's avatar
      tools/vm/page_owner_sort.c: fix the instructions for use · 49e495a0
      Yixuan Cao authored
      I noticed a discrepancy between the usage method and the code logic.
      
      If we enable the -f option, it should be "Filter out the information of
      blocks whose memory has been released".
      
      Link: https://lkml.kernel.org/r/20220219143106.2805-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Muchun Song <songmuchun@bytedance.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      49e495a0
    • Yixuan Cao's avatar
      mm/page_owner.c: record tgid · bf215eab
      Yixuan Cao authored
      In a single-threaded process, the pid in kernel task_struct is the same
      as the tgid, which can mark the process of page allocation.  But in a
      multithreaded process, only the task_struct of the thread leader has the
      same pid as tgid, and the pids of other threads are different from tgid.
      Therefore, tgid is recorded to provide effective information for
      debugging and data statistics of multithreaded programs.
      
      This can also be achieved by observing the task name (executable file
      name) for a specific process.  However, when the same program is started
      multiple times, the task name is the same and the tgid is different.
      Therefore, in the debugging of multi-threaded programs, combined with
      the task name and tgid, more accurate runtime information of a certain
      run of the program can be obtained.
      
      Link: https://lkml.kernel.org/r/20220219180450.2399-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Waiman Long <longman@redhat.com>
      Cc: Rafael Aquini <aquini@redhat.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      bf215eab
    • Waiman Long's avatar
      mm/page_owner: record task command name · 865ed6a3
      Waiman Long authored
      The page_owner information currently includes the pid of the calling
      task.  That is useful as long as the task is still running.  Otherwise,
      the number is meaningless.  To have more information about the
      allocating tasks that had exited by the time the page_owner information
      is retrieved, we need to store the command name of the task.
      
      Add a new comm field into page_owner structure to store the command name
      and display it when the page_owner information is retrieved.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-5-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: David Rientjes <rientjes@google.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Mike Rapoport <rppt@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Sergey Senozhatsky <senozhatsky@chromium.org>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      865ed6a3
    • Waiman Long's avatar
      mm/page_owner: print memcg information · fcf89358
      Waiman Long authored
      It was found that a number of offline memcgs were not freed because they
      were pinned by some charged pages that were present.  Even "echo 1 >
      /proc/sys/vm/drop_caches" wasn't able to free those pages.  These
      offline but not freed memcgs tend to increase in number over time with
      the side effect that percpu memory consumption as shown in /proc/meminfo
      also increases over time.
      
      In order to find out more information about those pages that pin offline
      memcgs, the page_owner feature is extended to print memory cgroup
      information especially whether the cgroup is offline or not.  RCU read
      lock is taken when memcg is being accessed to make sure that it won't be
      freed.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-4-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Acked-by: default avatarRoman Gushchin <guro@fb.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Sergey Senozhatsky <senozhatsky@chromium.org>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      fcf89358
    • Waiman Long's avatar
      mm/page_owner: use scnprintf() to avoid excessive buffer overrun check · 3ebc4397
      Waiman Long authored
      The snprintf() function can return a length greater than the given input
      size.  That will require a check for buffer overrun after each
      invocation of snprintf().  scnprintf(), on the other hand, will never
      return a greater length.
      
      By using scnprintf() in selected places, we can avoid some buffer
      overrun checks except after stack_depot_snprint() and after the last
      snprintf().
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-3-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Reviewed-by: default avatarSergey Senozhatsky <senozhatsky@chromium.org>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      3ebc4397
    • Waiman Long's avatar
      lib/vsprintf: avoid redundant work with 0 size · ef62c8ff
      Waiman Long authored
      Patch series "mm/page_owner: Extend page_owner to show memcg information", v4.
      
      While debugging the constant increase in percpu memory consumption on a
      system that spawned large number of containers, it was found that a lot
      of offline mem_cgroup structures remained in place without being freed.
      Further investigation indicated that those mem_cgroup structures were
      pinned by some pages.
      
      In order to find out what those pages are, the existing page_owner
      debugging tool is extended to show memory cgroup information and whether
      those memcgs are offline or not.  With the enhanced page_owner tool, the
      following is a typical page that pinned the mem_cgroup structure in my
      test case:
      
        Page allocated via order 0, mask 0x1100cca(GFP_HIGHUSER_MOVABLE), pid 162970 (podman), ts 1097761405537 ns, free_ts 1097760838089 ns
        PFN 1925700 type Movable Block 3761 type Movable Flags 0x17ffffc00c001c(uptodate|dirty|lru|reclaim|swapbacked|node=0|zone=2|lastcpupid=0x1fffff)
          prep_new_page+0xac/0xe0
          get_page_from_freelist+0x1327/0x14d0
          __alloc_pages+0x191/0x340
          alloc_pages_vma+0x84/0x250
          shmem_alloc_page+0x3f/0x90
          shmem_alloc_and_acct_page+0x76/0x1c0
          shmem_getpage_gfp+0x281/0x940
          shmem_write_begin+0x36/0xe0
          generic_perform_write+0xed/0x1d0
          __generic_file_write_iter+0xdc/0x1b0
          generic_file_write_iter+0x5d/0xb0
          new_sync_write+0x11f/0x1b0
          vfs_write+0x1ba/0x2a0
          ksys_write+0x59/0xd0
          do_syscall_64+0x37/0x80
          entry_SYSCALL_64_after_hwframe+0x44/0xae
        Charged to offline memcg libpod-conmon-15e4f9c758422306b73b2dd99f9d50a5ea53cbb16b4a13a2c2308a4253cc0ec8.
      
      So the page was not freed because it was part of a shmem segment.  That
      is useful information that can help users to diagnose similar problems.
      
      With cgroup v1, /proc/cgroups can be read to find out the total number
      of memory cgroups (online + offline).  With cgroup v2, the cgroup.stat
      of the root cgroup can be read to find the number of dying cgroups (most
      likely pinned by dying memcgs).
      
      The page_owner feature is not supposed to be enabled for production
      system due to its memory overhead.  However, if it is suspected that
      dying memcgs are increasing over time, a test environment with
      page_owner enabled can then be set up with appropriate workload for
      further analysis on what may be causing the increasing number of dying
      memcgs.
      
      This patch (of 4):
      
      For *scnprintf(), vsnprintf() is always called even if the input size is
      0.  That is a waste of time, so just return 0 in this case.
      
      Note that vsnprintf() will never return -1 to indicate an error.  So
      skipping the call to vsnprintf() when size is 0 will have no functional
      impact at all.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-1-longman@redhat.com
      Link: https://lkml.kernel.org/r/20220202203036.744010-2-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Reviewed-by: default avatarSergey Senozhatsky <senozhatsky@chromium.org>
      Acked-by: default avatarRoman Gushchin <guro@fb.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: David Rientjes <rientjes@google.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      ef62c8ff
    • Shuah Khan's avatar
      Documentation/vm/page_owner.rst: fix unexpected indentation warns · 2e944985
      Shuah Khan authored
      Fix Unexpected indentation warns in page_owner:
      
        Documentation/vm/page_owner.rst:92: WARNING: Unexpected indentation.
        Documentation/vm/page_owner.rst:96: WARNING: Unexpected indentation.
        Documentation/vm/page_owner.rst:107: WARNING: Unexpected indentation.
      
      Link: https://lkml.kernel.org/r/20211215001929.47866-1-skhan@linuxfoundation.org
      
      
      Signed-off-by: default avatarShuah Khan <skhan@linuxfoundation.org>
      Cc: Jonathan Corbet <corbet@lwn.net>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      2e944985
    • Shenghong Han's avatar
      Documentation/vm/page_owner.rst: update the documentation · 57f2b54a
      Shenghong Han authored
      Update the documentation of ``page_owner``.
      
      [akpm@linux-foundation.org: small grammatical tweaks]
      
      Link: https://lkml.kernel.org/r/20211214134736.2569-1-hanshenghong2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarShenghong Han <hanshenghong2019@email.szu.edu.cn>
      Cc: Jonathan Corbet <corbet@lwn.net>
      Cc: Vlastimil Babka <vbabka@suse.cz>
      Cc: Georgi Djakov <georgi.djakov@linaro.org>
      Cc: Liam Mark <lmark@codeaurora.org>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Xiaoming Ni <nixiaoming@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      57f2b54a
    • Yixuan Cao's avatar
      tools/vm/page_owner_sort.c: delete invalid duplicate code · 41ed6434
      Yixuan Cao authored
      I noticed that there is two invalid lines of duplicate code.  It's better
      to delete it.
      
      Link: https://lkml.kernel.org/r/20211213095743.3630-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Mark Brown <broonie@kernel.org>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      41ed6434
    • Shenghong Han's avatar
      tools/vm/page_owner_sort.c: two trivial fixes · e7a3f677
      Shenghong Han authored
      1) There is an unused variable. It's better to delete it.
      2) One case is missing in the usage().
      
      Link: https://lkml.kernel.org/r/20211213164518.2461-1-hanshenghong2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarShenghong Han <hanshenghong2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      e7a3f677
    • Chongxi Zhao's avatar
      tools/vm/page_owner_sort.c: support sorting pid and time · 8f9c447e
      Chongxi Zhao authored
      When viewing the page owner information, we expect that the information
      can be sorted by PID, so that we can quickly combine PID with the program
      to check the information together.
      
      We also expect that the information can be sorted by time.  Time sorting
      helps to view the running status of the program according to the time
      interval when the program hangs up.
      
      Finally, we hope to pass the page_ owner_ Sort.  C can reduce part of the
      output and only output the plate information whose memory has not been
      released, which can make us locate the problem of the program faster.
      Therefore, the following adjustments have been made:
      
      1. Add the static functions search_pattern and check_regcomp to
         improve the cleanliness.
      
      2. Add member attributes and their corresponding sorting methods.  In
         terms of comparison time, int will overflow because the data of ull is
         too large, so the ternary operator is used
      
      3. Add the -f parameter to filter out the information of blocks whose
         memory has not been released
      
      Link: https://lkml.kernel.org/r/20211206165653.5093-1-zhaochongxi2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarChongxi Zhao <zhaochongxi2019@email.szu.edu.cn>
      Reviewed-by: default avatarSean Anderson <seanga2@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      8f9c447e
    • Yinan Zhang's avatar
      tools/vm/page_owner_sort.c: add switch between culling by stacktrace and txt · cd75ea0e
      Yinan Zhang authored
      Culling by comparing stacktrace would casue loss of some information.  For
      example, if there exists 2 blocks which have the same stacktrace and the
      different head info
      
        Page allocated via order 0, mask 0x108c48(...), pid 73696,
          ts 1578829190639010 ns, free_ts 1576583851324450 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
        Page allocated via order 0, mask 0x108c48(...), pid 61806,
          ts 1354113726046100 ns, free_ts 1354104926841400 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
      After culling, it would be like this
      
        2 times, 2 pages:
        Page allocated via order 0, mask 0x108c48(...), pid 73696,
          ts 1578829190639010 ns, free_ts 1576583851324450 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
      The info of second block missed.  So, add -c to turn on culling by
      stacktrace.  By default, it will cull by txt.
      
      Link: https://lkml.kernel.org/r/20211129145658.2491-1-zhangyinan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      cd75ea0e
    • Sean Anderson's avatar
      tools/vm/page_owner_sort.c: support sorting by stack trace · 82f5ebc2
      Sean Anderson authored
      This adds the ability to sort by stacktraces.  This is helpful when
      comparing multiple dumps of page_owner taken at different times, since
      blocks will not be reordered if they were allocated/free'd.
      
      Link: https://lkml.kernel.org/r/20211124193709.1805776-2-seanga2@gmail.com
      
      
      Signed-off-by: default avatarSean Anderson <seanga2@gmail.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      82f5ebc2
    • Sean Anderson's avatar
      tools/vm/page_owner_sort.c: sort by stacktrace before culling · ba5a396b
      Sean Anderson authored
      The contents of page_owner have changed to include more information than
      the stack trace.  On a modern kernel, the blocks look like
      
        Page allocated via order 0, mask 0x0(), pid 1, ts 165564237 ns, free_ts 0 ns
          register_early_stack+0x4b/0x90
          init_page_owner+0x39/0x250
          kernel_init_freeable+0x11e/0x242
          kernel_init+0x16/0x130
      
      Sorting by the contents of .txt will result in almost no repeated pages,
      as the pid, ts, and free_ts will almost never be the same.  Instead,
      sort by the contents of the stack trace, which we assume to be whatever
      is after the first line.
      
      [seanga2@gmail.com: fix NULL-pointer dereference when comparing stack traces]
        Link: https://lkml.kernel.org/r/20211125162653.1855958-1-seanga2@gmail.com
      
      Link: https://lkml.kernel.org/r/20211124193709.1805776-1-seanga2@gmail.com
      
      
      Signed-off-by: default avatarSean Anderson <seanga2@gmail.com>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      ba5a396b
    • Linus Torvalds's avatar
      Merge branch 'akpm' (patches from Andrew) · 52deda95
      Linus Torvalds authored
      Merge more updates from Andrew Morton:
       "Various misc subsystems, before getting into the post-linux-next
        material.
      
        41 patches.
      
        Subsystems affected by this patch series: procfs, misc, core-kernel,
        lib, checkpatch, init, pipe, minix, fat, cgroups, kexec, kdump,
        taskstats, panic, kcov, resource, and ubsan"
      
      * emailed patches from Andrew Morton <akpm@linux-foundation.org>: (41 commits)
        Revert "ubsan, kcsan: Don't combine sanitizer with kcov on clang"
        kernel/resource: fix kfree() of bootmem memory again
        kcov: properly handle subsequent mmap calls
        kcov: split ioctl handling into locked and unlocked parts
        panic: move panic_print before kmsg dumpers
        panic: add option to dump all CPUs backtraces in panic_print
        docs: sysctl/kernel: add missing bit to panic_print
        taskstats: remove unneeded dead assignment
        kasan: no need to unset panic_on_warn in end_report()
        ubsan: no need to unset panic_on_warn in ubsan_epilogue()
        panic: unset panic_on_war...
      52deda95
    • Linus Torvalds's avatar
      Merge tag 'net-next-5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next · 169e7776
      Linus Torvalds authored
      Pull networking updates from Jakub Kicinski:
       "The sprinkling of SPI drivers is because we added a new one and Mark
        sent us a SPI driver interface conversion pull request.
      
        Core
        ----
      
         - Introduce XDP multi-buffer support, allowing the use of XDP with
           jumbo frame MTUs and combination with Rx coalescing offloads (LRO).
      
         - Speed up netns dismantling (5x) and lower the memory cost a little.
           Remove unnecessary per-netns sockets. Scope some lists to a netns.
           Cut down RCU syncing. Use batch methods. Allow netdev registration
           to complete out of order.
      
         - Support distinguishing timestamp types (ingress vs egress) and
           maintaining them across packet scrubbing points (e.g. redirect).
      
         - Continue the work of annotating packet drop reasons throughout the
           stack.
      
         - Switch netdev error counters from an atomic to dynamically
           allocated per-CPU counters.
      
         - Rework a few preemp...
      169e7776
    • Linus Torvalds's avatar
      Merge tag 'vfio-v5.18-rc1' of https://github.com/awilliam/linux-vfio · 7403e6d8
      Linus Torvalds authored
      Pull VFIO updates from Alex Williamson:
      
       - Introduce new device migration uAPI and implement device specific
         mlx5 vfio-pci variant driver supporting new protocol (Jason
         Gunthorpe, Yishai Hadas, Leon Romanovsky)
      
       - New HiSilicon acc vfio-pci variant driver, also supporting migration
         interface (Shameer Kolothum, Longfang Liu)
      
       - D3hot fixes for vfio-pci-core (Abhishek Sahu)
      
       - Document new vfio-pci variant driver acceptance criteria
         (Alex Williamson)
      
       - Fix UML build unresolved ioport_{un}map() functions
         (Alex Williamson)
      
       - Fix MAINTAINERS due to header movement (Lukas Bulwahn)
      
      * tag 'vfio-v5.18-rc1' of https://github.com/awilliam/linux-vfio: (31 commits)
        vfio-pci: Provide reviewers and acceptance criteria for variant drivers
        MAINTAINERS: adjust entry for header movement in hisilicon qm driver
        hisi_acc_vfio_pci: Use its own PCI reset_done error handler
        hisi_acc_vfio_pci: Add support for VFIO live migration
        crypto: hisilicon/qm: Set the VF QM state register
        hisi_acc_vfio_pci: Add helper to retrieve the struct pci_driver
        hisi_acc_vfio_pci: Restrict access to VF dev BAR2 migration region
        hisi_acc_vfio_pci: add new vfio_pci driver for HiSilicon ACC devices
        hisi_acc_qm: Move VF PCI device IDs to common header
        crypto: hisilicon/qm: Move few definitions to common header
        crypto: hisilicon/qm: Move the QM header to include/linux
        vfio/mlx5: Fix to not use 0 as NULL pointer
        PCI/IOV: Fix wrong kernel-doc identifier
        vfio/mlx5: Use its own PCI reset_done error handler
        vfio/pci: Expose vfio_pci_core_aer_err_detected()
        vfio/mlx5: Implement vfio_pci driver for mlx5 devices
        vfio/mlx5: Expose migration commands over mlx5 device
        vfio: Remove migration protocol v1 documentation
        vfio: Extend the device migration protocol with RUNNING_P2P
        vfio: Define device migration protocol v2
        ...
      7403e6d8
    • Linus Torvalds's avatar
      Merge tag 'hyperv-next-signed-20220322' of... · 66711cfe
      Linus Torvalds authored
      Merge tag 'hyperv-next-signed-20220322' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux
      
      Pull hyperv updates from Wei Liu:
       "Minor patches from various people"
      
      * tag 'hyperv-next-signed-20220322' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux:
        x86/hyperv: Output host build info as normal Windows version number
        hv_balloon: rate-limit "Unhandled message" warning
        drivers: hv: log when enabling crash_kexec_post_notifiers
        hv_utils: Add comment about max VMbus packet size in VSS driver
        Drivers: hv: Compare cpumasks and not their weights in init_vp_index()
        Drivers: hv: Rename 'alloced' to 'allocated'
        Drivers: hv: vmbus: Use struct_size() helper in kmalloc()
      66711cfe
    • Linus Torvalds's avatar
      Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm · 1ebdbeb0
      Linus Torvalds authored
      Pull kvm updates from Paolo Bonzini:
       "ARM:
         - Proper emulation of the OSLock feature of the debug architecture
      
         - Scalibility improvements for the MMU lock when dirty logging is on
      
         - New VMID allocator, which will eventually help with SVA in VMs
      
         - Better support for PMUs in heterogenous systems
      
         - PSCI 1.1 support, enabling support for SYSTEM_RESET2
      
         - Implement CONFIG_DEBUG_LIST at EL2
      
         - Make CONFIG_ARM64_ERRATUM_2077057 default y
      
         - Reduce the overhead of VM exit when no interrupt is pending
      
         - Remove traces of 32bit ARM host support from the documentation
      
         - Updated vgic selftests
      
         - Various cleanups, doc updates and spelling fixes
      
        RISC-V:
         - Prevent KVM_COMPAT from being selected
      
         - Optimize __kvm_riscv_switch_to() implementation
      
         - RISC-V SBI v0.3 support
      
        s390:
         - memop selftest
      
         - fix SCK locking
      
         - adapter interruptions virtualization for secure guests
      
         - add Claudio Imbrenda as maintainer
      
         - first step to do proper storage key checking
      
        x86:
         - Continue switching kvm_x86_ops to static_call(); introduce
           static_call_cond() and __static_call_ret0 when applicable.
      
         - Cleanup unused arguments in several functions
      
         - Synthesize AMD 0x80000021 leaf
      
         - Fixes and optimization for Hyper-V sparse-bank hypercalls
      
         - Implement Hyper-V's enlightened MSR bitmap for nested SVM
      
         - Remove MMU auditing
      
         - Eager splitting of page tables (new aka "TDP" MMU only) when dirty
           page tracking is enabled
      
         - Cleanup the implementation of the guest PGD cache
      
         - Preparation for the implementation of Intel IPI virtualization
      
         - Fix some segment descriptor checks in the emulator
      
         - Allow AMD AVIC support on systems with physical APIC ID above 255
      
         - Better API to disable virtualization quirks
      
         - Fixes and optimizations for the zapping of page tables:
      
            - Zap roots in two passes, avoiding RCU read-side critical
              sections that last too long for very large guests backed by 4
              KiB SPTEs.
      
            - Zap invalid and defunct roots asynchronously via
              concurrency-managed work queue.
      
            - Allowing yielding when zapping TDP MMU roots in response to the
              root's last reference being put.
      
            - Batch more TLB flushes with an RCU trick. Whoever frees the
              paging structure now holds RCU as a proxy for all vCPUs running
              in the guest, i.e. to prolongs the grace period on their behalf.
              It then kicks the the vCPUs out of guest mode before doing
              rcu_read_unlock().
      
        Generic:
         - Introduce __vcalloc and use it for very large allocations that need
           memcg accounting"
      
      * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (246 commits)
        KVM: use kvcalloc for array allocations
        KVM: x86: Introduce KVM_CAP_DISABLE_QUIRKS2
        kvm: x86: Require const tsc for RT
        KVM: x86: synthesize CPUID leaf 0x80000021h if useful
        KVM: x86: add support for CPUID leaf 0x80000021
        KVM: x86: do not use KVM_X86_OP_OPTIONAL_RET0 for get_mt_mask
        Revert "KVM: x86/mmu: Zap only TDP MMU leafs in kvm_zap_gfn_range()"
        kvm: x86/mmu: Flush TLB before zap_gfn_range releases RCU
        KVM: arm64: fix typos in comments
        KVM: arm64: Generalise VM features into a set of flags
        KVM: s390: selftests: Add error memop tests
        KVM: s390: selftests: Add more copy memop tests
        KVM: s390: selftests: Add named stages for memop test
        KVM: s390: selftests: Add macro as abstraction for MEM_OP
        KVM: s390: selftests: Split memop tests
        KVM: s390x: fix SCK locking
        RISC-V: KVM: Implement SBI HSM suspend call
        RISC-V: KVM: Add common kvm_riscv_vcpu_wfi() function
        RISC-V: Add SBI HSM suspend related defines
        RISC-V: KVM: Implement SBI v0.3 SRST extension
        ...
      1ebdbeb0
    • Linus Torvalds's avatar
      Merge tag 'tomoyo-pr-20220322' of git://git.osdn.net/gitroot/tomoyo/tomoyo-test1 · efee6c79
      Linus Torvalds authored
      Pull tomoyo update from Tetsuo Handa:
       "Avoid unnecessarily leaking kernel command line arguments"
      
      * tag 'tomoyo-pr-20220322' of git://git.osdn.net/gitroot/tomoyo/tomoyo-test1:
        TOMOYO: fix __setup handlers return values
      efee6c79
    • Linus Torvalds's avatar
      Merge tag 'flexible-array-transformations-5.18-rc1' of... · 3ce62cf4
      Linus Torvalds authored
      Merge tag 'flexible-array-transformations-5.18-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gustavoars/linux
      
      Pull flexible-array transformations from Gustavo Silva:
       "Treewide patch that replaces zero-length arrays with flexible-array
        members.
      
        This has been baking in linux-next for a whole development cycle"
      
      * tag 'flexible-array-transformations-5.18-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gustavoars/linux:
        treewide: Replace zero-length arrays with flexible-array members
      3ce62cf4
    • Linus Torvalds's avatar
      Merge tag 'prlimit-tasklist_lock-for-v5.18' of... · cd4699c5
      Linus Torvalds authored
      Merge tag 'prlimit-tasklist_lock-for-v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiederm/user-namespace
      
      Pull tasklist_lock optimizations from Eric Biederman:
       "prlimit and getpriority tasklist_lock optimizations
      
        The tasklist_lock popped up as a scalability bottleneck on some
        testing workloads. The readlocks in do_prlimit and set/getpriority are
        not necessary in all cases.
      
        Based on a cycles profile, it looked like ~87% of the time was spent
        in the kernel, ~42% of which was just trying to get *some* spinlock
        (queued_spin_lock_slowpath, not necessarily the tasklist_lock).
      
        The big offenders (with rough percentages in cycles of the overall
        trace):
         - do_wait 11%
         - setpriority 8% (done previously in commit 7f8ca0ed)
         - kill 8%
         - do_exit 5%
         - clone 3%
         - prlimit64 2%   (this patchset)
         - getrlimit 1%   (this patchset)
      
        I can't easily test this patchset on the original workload for various
        reasons. Instead, I used the microbenchmark below to at least verify
        there was some improvement. This patchset had a 28% speedup (12% from
        baseline to set/getprio, then another 14% for prlimit).
      
        This series used to do the setpriority case, but an almost identical
        change was merged as commit 7f8ca0ed ("kernel/sys.c: only take
        tasklist_lock for get/setpriority(PRIO_PGRP)") so that has been
        dropped from here.
      
        One interesting thing is that my libc's getrlimit() was calling
        prlimit64, so hoisting the read_lock(tasklist_lock) into sys_prlimit64
        had no effect - it essentially optimized the older syscalls only. I
        didn't do that in this patchset, but figured I'd mention it since it
        was an option from the previous patch's discussion"
      
      micobenchmark.c:
      ---------------
      	int main(int argc, char **argv)
      	{
      		pid_t child;
      		struct rlimit rlim[1];
      
      		fork(); fork(); fork(); fork(); fork(); fork();
      
      		for (int i = 0; i < 5000; i++) {
      			child = fork();
      			if (child < 0)
      				exit(1);
      			if (child > 0) {
      				usleep(1000);
      				kill(child, SIGTERM);
      				waitpid(child, NULL, 0);
      			} else {
      				for (;;) {
      					setpriority(PRIO_PROCESS, 0,
      						    getpriority(PRIO_PROCESS, 0));
      					getrlimit(RLIMIT_CPU, rlim);
      				}
      			}
      		}
      
      		return 0;
      	}
      
      Link: https://lore.kernel.org/lkml/20211213220401.1039578-1-brho@google.com/ [v1]
      Link: https://lore.kernel.org/lkml/20220105212828.197013-1-brho@google.com/ [v2]
      Link: https://lore.kernel.org/lkml/20220106172041.522167-1-brho@google.com/ [v3]
      
      * tag 'prlimit-tasklist_lock-for-v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiederm/user-namespace:
        prlimit: do not grab the tasklist_lock
        prlimit: make do_prlimit() static
      cd4699c5
    • Linus Torvalds's avatar
      Merge tag 'fs.rt.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux · 2e2d4650
      Linus Torvalds authored
      Pull mount attributes PREEMPT_RT update from Christian Brauner:
       "This contains Sebastian's fix to make changing mount
        attributes/getting write access compatible with CONFIG_PREEMPT_RT.
      
        The change only applies when users explicitly opt-in to real-time via
        CONFIG_PREEMPT_RT otherwise things are exactly as before. We've waited
        quite a long time with this to make sure folks could take a good look"
      
      * tag 'fs.rt.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux:
        fs/namespace: Boost the mount_lock.lock owner instead of spinning on PREEMPT_RT.
      2e2d4650
    • Linus Torvalds's avatar
      Merge tag 'fs.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux · 15f2e3d6
      Linus Torvalds authored
      Pull mount_setattr updates from Christian Brauner:
       "This contains a few more patches to massage the mount_setattr()
        codepaths and one minor fix to reuse a helper we added some time back.
      
        The final two patches do similar cleanups in different ways. One patch
        is mine and the other is Al's who was nice enough to give me a branch
        for it.
      
        Since his came in later and my branch had been sitting in -next for
        quite some time we just put his on top instead of swap them"
      
      * tag 'fs.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux:
        mount_setattr(): clean the control flow and calling conventions
        fs: clean up mount_setattr control flow
        fs: don't open-code mnt_hold_writers()
        fs: simplify check in mount_setattr_commit()
        fs: add mnt_allow_writers() and simplify mount_setattr_prepare()
      15f2e3d6
  2. Mar 24, 2022