1. Mar 25, 2022
    • Andrey Konovalov's avatar
      kasan, page_alloc: deduplicate should_skip_kasan_poison · 94ae8b83
      Andrey Konovalov authored
      Patch series "kasan, vmalloc, arm64: add vmalloc tagging support for SW/HW_TAGS", v6.
      
      This patchset adds vmalloc tagging support for SW_TAGS and HW_TAGS
      KASAN modes.
      
      About half of patches are cleanups I went for along the way.  None of them
      seem to be important enough to go through stable, so I decided not to
      split them out into separate patches/series.
      
      The patchset is partially based on an early version of the HW_TAGS
      patchset by Vincenzo that had vmalloc support.  Thus, I added a
      Co-developed-by tag into a few patches.
      
      SW_TAGS vmalloc tagging support is straightforward.  It reuses all of the
      generic KASAN machinery, but uses shadow memory to store tags instead of
      magic values.  Naturally, vmalloc tagging requires adding a few
      kasan_reset_tag() annotations to the vmalloc code.
      
      HW_TAGS vmalloc tagging support stands out.  HW_TAGS KASAN is based on Arm
      MTE, which can only assigns tags to physical memory.  As a result, HW_TAGS
      KASAN only tags vmalloc() allocations, which are backed by page_alloc
      memory.  It ignores vmap() and others.
      
      This patch (of 39):
      
      Currently, should_skip_kasan_poison() has two definitions: one for when
      CONFIG_DEFERRED_STRUCT_PAGE_INIT is enabled, one for when it's not.
      
      Instead of duplicating the checks, add a deferred_pages_enabled() helper
      and use it in a single should_skip_kasan_poison() definition.
      
      Also move should_skip_kasan_poison() closer to its caller and clarify all
      conditions in the comment.
      
      Link: https://lkml.kernel.org/r/cover.1643047180.git.andreyknvl@google.com
      Link: https://lkml.kernel.org/r/658b79f5fb305edaf7dc16bc52ea870d3220d4a8.1643047180.git.andreyknvl@google.com
      
      
      Signed-off-by: default avatarAndrey Konovalov <andreyknvl@google.com>
      Acked-by: default avatarMarco Elver <elver@google.com>
      Cc: Alexander Potapenko <glider@google.com>
      Cc: Dmitry Vyukov <dvyukov@google.com>
      Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
      Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
      Cc: Catalin Marinas <catalin.marinas@arm.com>
      Cc: Will Deacon <will@kernel.org>
      Cc: Mark Rutland <mark.rutland@arm.com>
      Cc: Peter Collingbourne <pcc@google.com>
      Cc: Evgenii Stepanov <eugenis@google.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      94ae8b83
    • Anshuman Khandual's avatar
      mm/migration: add trace events for base page and HugeTLB migrations · 4cc79b33
      Anshuman Khandual authored
      This adds two trace events for base page and HugeTLB page migrations.
      These events, closely follow the implementation details like setting and
      removing of PTE migration entries, which are essential operations for
      migration.  The new CREATE_TRACE_POINTS in <mm/rmap.c> covers both
      <events/migration.h> and <events/tlb.h> based trace events.  Hence drop
      redundant CREATE_TRACE_POINTS from other places which could have otherwise
      conflicted during build.
      
      Link: https://lkml.kernel.org/r/1643368182-9588-3-git-send-email-anshuman.khandual@arm.com
      
      
      Signed-off-by: default avatarAnshuman Khandual <anshuman.khandual@arm.com>
      Reported-by: default avatarkernel test robot <lkp@intel.com>
      Cc: Steven Rostedt <rostedt@goodmis.org>
      Cc: Ingo Molnar <mingo@redhat.com>
      Cc: Zi Yan <ziy@nvidia.com>
      Cc: Naoya Horiguchi <naoya.horiguchi@nec.com>
      Cc: John Hubbard <jhubbard@nvidia.com>
      Cc: Matthew Wilcox <willy@infradead.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      4cc79b33
    • Anshuman Khandual's avatar
      mm/migration: add trace events for THP migrations · 283fd6fe
      Anshuman Khandual authored
      Patch series "mm/migration: Add trace events", v3.
      
      This adds trace events for all migration scenarios including base page,
      THP and HugeTLB.
      
      This patch (of 3):
      
      This adds two trace events for PMD based THP migration without split.
      These events closely follow the implementation details like setting and
      removing of PMD migration entries, which are essential operations for THP
      migration.  This moves CREATE_TRACE_POINTS into generic THP from powerpc
      for these new trace events to be available on other platforms as well.
      
      Link: https://lkml.kernel.org/r/1643368182-9588-1-git-send-email-anshuman.khandual@arm.com
      Link: https://lkml.kernel.org/r/1643368182-9588-2-git-send-email-anshuman.khandual@arm.com
      
      
      Signed-off-by: default avatarAnshuman Khandual <anshuman.khandual@arm.com>
      Cc: Steven Rostedt <rostedt@goodmis.org>
      Cc: Ingo Molnar <mingo@redhat.com>
      Cc: Zi Yan <ziy@nvidia.com>
      Cc: Naoya Horiguchi <naoya.horiguchi@nec.com>
      Cc: John Hubbard <jhubbard@nvidia.com>
      Cc: Matthew Wilcox <willy@infradead.org>
      Cc: Michael Ellerman <mpe@ellerman.id.au>
      Cc: Paul Mackerras <paulus@samba.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      283fd6fe
    • Hugh Dickins's avatar
      mm/thp: fix NR_FILE_MAPPED accounting in page_*_file_rmap() · 5d543f13
      Hugh Dickins authored
      NR_FILE_MAPPED accounting in mm/rmap.c (for /proc/meminfo "Mapped" and
      /proc/vmstat "nr_mapped" and the memcg's memory.stat "mapped_file") is
      slightly flawed for file or shmem huge pages.
      
      It is well thought out, and looks convincing, but there's a racy case when
      the careful counting in page_remove_file_rmap() (without page lock) gets
      discarded.  So that in a workload like two "make -j20" kernel builds under
      memory pressure, with cc1 on hugepage text, "Mapped" can easily grow by a
      spurious 5MB or more on each iteration, ending up implausibly bigger than
      most other numbers in /proc/meminfo.  And, hypothetically, might grow to
      the point of seriously interfering in mm/vmscan.c's heuristics, which do
      take NR_FILE_MAPPED into some consideration.
      
      Fixed by moving the __mod_lruvec_page_state() down to where it will not be
      missed before return (and I've grown a bit tired of that oft-repeated
      but-not-everywhere comment on the __ness: it gets lost in the move here).
      
      Does page_add_file_rmap() need the same change?  I suspect not, because
      page lock is held in all relevant cases, and its skipping case looks safe;
      but it's much easier to be sure, if we do make the same change.
      
      Link: https://lkml.kernel.org/r/e02e52a1-8550-a57c-ed29-f51191ea2375@google.com
      Fixes: dd78fedd
      
       ("rmap: support file thp")
      Signed-off-by: default avatarHugh Dickins <hughd@google.com>
      Reviewed-by: default avatarYang Shi <shy828301@gmail.com>
      Cc: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      5d543f13
    • Hugh Dickins's avatar
      mm: filemap_unaccount_folio() large skip mapcount fixup · 85207ad8
      Hugh Dickins authored
      The page_mapcount_reset() when folio_mapped() while mapping_exiting() was
      devised long before there were huge or compound pages in the cache.  It is
      still valid for small pages, but not at all clear what's right to check
      and reset on large pages.  Just don't try when folio_test_large().
      
      Link: https://lkml.kernel.org/r/879c4426-4122-da9c-1a86-697f2c9a083@google.com
      
      
      Signed-off-by: default avatarHugh Dickins <hughd@google.com>
      Cc: Matthew Wilcox <willy@infradead.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      85207ad8
    • Hugh Dickins's avatar
      mm: delete __ClearPageWaiters() · bb43b14b
      Hugh Dickins authored
      The PG_waiters bit is not included in PAGE_FLAGS_CHECK_AT_FREE, and
      vmscan.c's free_unref_page_list() callers rely on that not to generate
      bad_page() alerts.  So __page_cache_release(), put_pages_list() and
      release_pages() (and presumably copy-and-pasted free_zone_device_page())
      are redundant and misleading to make a special point of clearing it (as
      the "__" implies, it could only safely be used on the freeing path).
      
      Delete __ClearPageWaiters().  Remark on this in one of the "possible"
      comments in folio_wake_bit(), and delete the superfluous comments.
      
      Link: https://lkml.kernel.org/r/3eafa969-5b1a-accf-88fe-318784c791a@google.com
      
      
      Signed-off-by: default avatarHugh Dickins <hughd@google.com>
      Tested-by: default avatarYu Zhao <yuzhao@google.com>
      Reviewed-by: default avatarYang Shi <shy828301@gmail.com>
      Reviewed-by: default avatarDavid Hildenbrand <david@redhat.com>
      Cc: Matthew Wilcox <willy@infradead.org>
      Cc: Nicholas Piggin <npiggin@gmail.com>
      Cc: Yu Zhao <yuzhao@google.com>
      Cc: Michal Hocko <mhocko@suse.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      bb43b14b
    • Mike Rapoport's avatar
      selftest/vm: add helpers to detect PAGE_SIZE and PAGE_SHIFT · 6f6a841f
      Mike Rapoport authored
      PAGE_SIZE is not 4096 in many configurations, particularly ppc64 uses 64K
      pages in majority of cases.
      
      Add helpers to detect PAGE_SIZE and PAGE_SHIFT dynamically.
      
      Without this tests are broken w.r.t reading /proc/self/pagemap
      
          if (pread(pagemap_fd, ent, sizeof(ent),
                    (uintptr_t)ptr >> (PAGE_SHIFT - 3)) != sizeof(ent))
                    err(2, "read pagemap");
      
      Link: https://lkml.kernel.org/r/20220307054355.149820-2-aneesh.kumar@linux.ibm.com
      
      
      Signed-off-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Signed-off-by: default avatarAneesh Kumar K.V <aneesh.kumar@linux.ibm.com>
      Cc: Shuah Khan <shuah@kernel.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      6f6a841f
    • Aneesh Kumar K.V's avatar
      selftest/vm: add util.h and and move helper functions there · 90647d9d
      Aneesh Kumar K.V authored
      Avoid code duplication by adding util.h.  No functional change in this
      patch.
      
      Link: https://lkml.kernel.org/r/20220307054355.149820-1-aneesh.kumar@linux.ibm.com
      
      
      Signed-off-by: default avatarAneesh Kumar K.V <aneesh.kumar@linux.ibm.com>
      Cc: Shuah Khan <shuah@kernel.org>
      Cc: Mike Rapoport <rppt@linux.ibm.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      90647d9d
    • Christoph Hellwig's avatar
      mm: unexport page_init_poison · 1a9762b2
      Christoph Hellwig authored
      page_init_poison is only used in core MM code, so unexport it.
      
      Link: https://lkml.kernel.org/r/20220207063446.1833404-1-hch@lst.de
      
      
      Signed-off-by: default avatarChristoph Hellwig <hch@lst.de>
      Reviewed-by: default avatarDavid Hildenbrand <david@redhat.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      1a9762b2
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: support for user-defined culling rules · 9c8a0a8e
      Jiajian Ye authored
      When viewing page owner information, we may want to cull blocks of
      information with our own rules.  So it is important to enhance culling
      function to provide the support for customizing culling rules.
      Therefore, following adjustments are made:
      
      1. Add --cull option to support the culling of blocks of information
         with user-defined culling rules.
      
      	./page_owner_sort <input> <output> --cull=<rules>
      	./page_owner_sort <input> <output> --cull <rules>
      
        <rules> is a single argument in the form of a comma-separated list to
        specify individual culling rules, by the sequence of keys k1,k2, ....
        Mixed use of abbreviated and complete-form of keys is allowed.
      
        For reference, please see the document(Documentation/vm/page_owner.rst).
      
      Now, assuming two blocks in the input file are as follows:
      
      	Page allocated via order 0, mask xxxx, pid 1, tgid 1 (task_name_demo)
      	PFN xxxx
      	 prep_new_page+0xd0/0xf8
      	 get_page_from_freelist+0x4a0/0x1290
      	 __alloc_pages+0x168/0x340
      	 alloc_pages+0xb0/0x158
      
      	Page allocated via order 0, mask xxxx, pid 32, tgid 32 (task_name_demo)
      	PFN xxxx
      	 prep_new_page+0xd0/0xf8
      	 get_page_from_freelist+0x4a0/0x1290
      	 __alloc_pages+0x168/0x340
      	 alloc_pages+0xb0/0x158
      
      If we want to cull the blocks by stacktrace and task command name, we can
      use this command:
      
      	./page_owner_sort <input> <output> --cull=stacktrace,name
      
      The output would be like:
      
      	2 times, 2 pages, task_comm_name: task_name_demo
      	 prep_new_page+0xd0/0xf8
      	 get_page_from_freelist+0x4a0/0x1290
      	 __alloc_pages+0x168/0x340
      	 alloc_pages+0xb0/0x158
      
      As we can see, these two blocks are culled successfully, for they share
      the same pid and task command name.
      
      However, if we want to cull the blocks by pid, stacktrace and task command
      name, we can this command:
      
      	./page_owner_sort <input> <output> --cull=stacktrace,name,pid
      
      The output would be like:
      
      	1 times, 1 pages, PID 1, task_comm_name: task_name_demo
      	 prep_new_page+0xd0/0xf8
      	 get_page_from_freelist+0x4a0/0x1290
      	 __alloc_pages+0x168/0x340
      	 alloc_pages+0xb0/0x158
      
      	1 times, 1 pages, PID 32, task_comm_name: task_name_demo
      	 prep_new_page+0xd0/0xf8
      	 get_page_from_freelist+0x4a0/0x1290
      	 __alloc_pages+0x168/0x340
      	 alloc_pages+0xb0/0x158
      
      As we can see, these two blocks are failed to cull, for their PIDs are
      different.
      
      2. Add explanations of --cull options to the document.
      
      This work is coauthored by
      	Yixuan Cao
      	Shenghong Han
      	Yinan Zhang
      	Chongxi Zhao
      	Yuhong Feng
      
      Link: https://lkml.kernel.org/r/20220312145834.624-1-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Shenghong Han <hanshenghong2019@email.szu.edu.cn>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Chongxi Zhao <zhaochongxi2019@email.szu.edu.cn>
      Cc: Yuhong Feng <yuhongf@szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Sean Anderson <seanga2@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      9c8a0a8e
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: support for selecting by PID, TGID or task command name · 8ea8613a
      Jiajian Ye authored
      When viewing page owner information, we may also need to select the blocks
      by PID, TGID or task command name, which helps to get more accurate page
      allocation information as needed.
      
      Therefore, following adjustments are made:
      
      1. Add three new options, including --pid, --tgid and --name, to support
         the selection of information blocks by a specific pid, tgid and task
         command name. In addtion, multiple options are allowed to be used at
         the same time.
      
      	./page_owner_sort [input] [output] --pid <PID>
      	./page_owner_sort [input] [output] --tgid <TGID>
      	./page_owner_sort [input] [output] --name <TASK_COMMAND_NAME>
      
         Assuming a scenario when a multi-threaded program, ./demo (PID =
         5280), is running, and ./demo creates a child process (PID = 5281).
      
      	$ps
      	PID   TTY        TIME   CMD
      	5215  pts/0    00:00:00  bash
      	5280  pts/0    00:00:00  ./demo
      	5281  pts/0    00:00:00  ./demo
      	5282  pts/0    00:00:00  ps
      
         It would be better to filter out the records with tgid=5280 and the
         task name "demo" when debugging the parent process, and the specific
         usage is
      
      	./page_owner_sort [input] [output] --tgid 5280 --name demo
      
      2. Add explanations of three new options, including --pid, --tgid and
         --name, to the document.
      
      This work is coauthored by
      	Shenghong Han <hanshenghong2019@email.szu.edu.cn>,
      	Yixuan Cao <caoyixuan2019@email.szu.edu.cn>,
      	Yinan Zhang <zhangyinan2019@email.szu.edu.cn>,
      	Chongxi Zhao <zhaochongxi2019@email.szu.edu.cn>,
      	Yuhong Feng <yuhongf@szu.edu.cn>.
      
      Link: https://lkml.kernel.org/r/1646835223-7584-1-git-send-email-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      8ea8613a
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort: support for sorting by task command name · 194d52d7
      Jiajian Ye authored
      When viewing page owner information, we may also need to the block to be
      sorted by task command name.  Therefore, the following adjustments are
      made:
      
      1. Add a member variable to record task command name of block.
      
      2. Add a new -n option to sort the information of blocks by task command
         name.
      
      3. Add -n option explanation in the document.
      
      Link: https://lkml.kernel.org/r/20220306030640.43054-2-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: <zhaochongxi2019@email.szu.edu.cn>
      Cc: <hanshenghong2019@email.szu.edu.cn>
      Cc: <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      194d52d7
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort: fix three trivival places · 578d8f27
      Jiajian Ye authored
      The following adjustments are made:
      
      1. Instead of using another array to cull the blocks after sorting,
         reuse the old array.  So there is no need to malloc a new array.
      
      2. When enabling '-f' option to filter out the blocks which have been
         released, only add those have not been released in the list, rather
         than add all of blocks in the list and then do the filtering when
         printing the result.
      
      3. When enabling '-c' option to cull the blocks by comparing
         stacktrace, print the stacetrace rather than the total block.
      
      Link: https://lkml.kernel.org/r/20220306030640.43054-1-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: <hanshenghong2019@email.szu.edu.cn>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: <zhangyinan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      578d8f27
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: support sorting by tgid and update documentation · cf3c2c86
      Jiajian Ye authored
      When the "page owner" information is read, the information sorted
      by TGID is expected.
      
      As a result, the following adjustments have been made:
      
      1. Add a new -P option to sort the information of blocks by TGID in
         ascending order.
      
      2. Adjust the order of member variables in block_list strust to avoid
         one 4 byte hole.
      
      3. Add -P option explanation in the document.
      
      Link: https://lkml.kernel.org/r/20220301151438.166118-3-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      cf3c2c86
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: add a security check · 56465a38
      Jiajian Ye authored
      Add a security check after using malloc() to allocate memory.
      
      Link: https://lkml.kernel.org/r/20220301151438.166118-2-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      56465a38
    • Jiajian Ye's avatar
      tools/vm/page_owner_sort.c: fix comments · 59d7cb27
      Jiajian Ye authored
      Two adjustments are made:
      
      1. Correct a grammatical error: replace the "what" in "Do the job what
         you want to debug" with "that".
      
      2. Replace "has not been" with "has been" in the description of the -f
         option: According to Commit b1c9ba071e7d ("tools/vm/page_owner_sort.c:
         fix the instructions for use"), the description of the "-f" option is
         "Filter out the information of blocks whose memory has been released."
      
      Link: https://lkml.kernel.org/r/20220301151438.166118-1-yejiajian2018@email.szu.edu.cn
      
      
      Signed-off-by: default avatarJiajian Ye <yejiajian2018@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Yixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      59d7cb27
    • Yixuan Cao's avatar
      tools/vm/page_owner_sort.c: fix the instructions for use · 49e495a0
      Yixuan Cao authored
      I noticed a discrepancy between the usage method and the code logic.
      
      If we enable the -f option, it should be "Filter out the information of
      blocks whose memory has been released".
      
      Link: https://lkml.kernel.org/r/20220219143106.2805-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Muchun Song <songmuchun@bytedance.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      49e495a0
    • Yixuan Cao's avatar
      mm/page_owner.c: record tgid · bf215eab
      Yixuan Cao authored
      In a single-threaded process, the pid in kernel task_struct is the same
      as the tgid, which can mark the process of page allocation.  But in a
      multithreaded process, only the task_struct of the thread leader has the
      same pid as tgid, and the pids of other threads are different from tgid.
      Therefore, tgid is recorded to provide effective information for
      debugging and data statistics of multithreaded programs.
      
      This can also be achieved by observing the task name (executable file
      name) for a specific process.  However, when the same program is started
      multiple times, the task name is the same and the tgid is different.
      Therefore, in the debugging of multi-threaded programs, combined with
      the task name and tgid, more accurate runtime information of a certain
      run of the program can be obtained.
      
      Link: https://lkml.kernel.org/r/20220219180450.2399-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Waiman Long <longman@redhat.com>
      Cc: Rafael Aquini <aquini@redhat.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      bf215eab
    • Waiman Long's avatar
      mm/page_owner: record task command name · 865ed6a3
      Waiman Long authored
      The page_owner information currently includes the pid of the calling
      task.  That is useful as long as the task is still running.  Otherwise,
      the number is meaningless.  To have more information about the
      allocating tasks that had exited by the time the page_owner information
      is retrieved, we need to store the command name of the task.
      
      Add a new comm field into page_owner structure to store the command name
      and display it when the page_owner information is retrieved.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-5-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: David Rientjes <rientjes@google.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Mike Rapoport <rppt@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Sergey Senozhatsky <senozhatsky@chromium.org>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      865ed6a3
    • Waiman Long's avatar
      mm/page_owner: print memcg information · fcf89358
      Waiman Long authored
      It was found that a number of offline memcgs were not freed because they
      were pinned by some charged pages that were present.  Even "echo 1 >
      /proc/sys/vm/drop_caches" wasn't able to free those pages.  These
      offline but not freed memcgs tend to increase in number over time with
      the side effect that percpu memory consumption as shown in /proc/meminfo
      also increases over time.
      
      In order to find out more information about those pages that pin offline
      memcgs, the page_owner feature is extended to print memory cgroup
      information especially whether the cgroup is offline or not.  RCU read
      lock is taken when memcg is being accessed to make sure that it won't be
      freed.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-4-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Acked-by: default avatarRoman Gushchin <guro@fb.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Sergey Senozhatsky <senozhatsky@chromium.org>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      fcf89358
    • Waiman Long's avatar
      mm/page_owner: use scnprintf() to avoid excessive buffer overrun check · 3ebc4397
      Waiman Long authored
      The snprintf() function can return a length greater than the given input
      size.  That will require a check for buffer overrun after each
      invocation of snprintf().  scnprintf(), on the other hand, will never
      return a greater length.
      
      By using scnprintf() in selected places, we can avoid some buffer
      overrun checks except after stack_depot_snprint() and after the last
      snprintf().
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-3-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Reviewed-by: default avatarSergey Senozhatsky <senozhatsky@chromium.org>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      3ebc4397
    • Waiman Long's avatar
      lib/vsprintf: avoid redundant work with 0 size · ef62c8ff
      Waiman Long authored
      Patch series "mm/page_owner: Extend page_owner to show memcg information", v4.
      
      While debugging the constant increase in percpu memory consumption on a
      system that spawned large number of containers, it was found that a lot
      of offline mem_cgroup structures remained in place without being freed.
      Further investigation indicated that those mem_cgroup structures were
      pinned by some pages.
      
      In order to find out what those pages are, the existing page_owner
      debugging tool is extended to show memory cgroup information and whether
      those memcgs are offline or not.  With the enhanced page_owner tool, the
      following is a typical page that pinned the mem_cgroup structure in my
      test case:
      
        Page allocated via order 0, mask 0x1100cca(GFP_HIGHUSER_MOVABLE), pid 162970 (podman), ts 1097761405537 ns, free_ts 1097760838089 ns
        PFN 1925700 type Movable Block 3761 type Movable Flags 0x17ffffc00c001c(uptodate|dirty|lru|reclaim|swapbacked|node=0|zone=2|lastcpupid=0x1fffff)
          prep_new_page+0xac/0xe0
          get_page_from_freelist+0x1327/0x14d0
          __alloc_pages+0x191/0x340
          alloc_pages_vma+0x84/0x250
          shmem_alloc_page+0x3f/0x90
          shmem_alloc_and_acct_page+0x76/0x1c0
          shmem_getpage_gfp+0x281/0x940
          shmem_write_begin+0x36/0xe0
          generic_perform_write+0xed/0x1d0
          __generic_file_write_iter+0xdc/0x1b0
          generic_file_write_iter+0x5d/0xb0
          new_sync_write+0x11f/0x1b0
          vfs_write+0x1ba/0x2a0
          ksys_write+0x59/0xd0
          do_syscall_64+0x37/0x80
          entry_SYSCALL_64_after_hwframe+0x44/0xae
        Charged to offline memcg libpod-conmon-15e4f9c758422306b73b2dd99f9d50a5ea53cbb16b4a13a2c2308a4253cc0ec8.
      
      So the page was not freed because it was part of a shmem segment.  That
      is useful information that can help users to diagnose similar problems.
      
      With cgroup v1, /proc/cgroups can be read to find out the total number
      of memory cgroups (online + offline).  With cgroup v2, the cgroup.stat
      of the root cgroup can be read to find the number of dying cgroups (most
      likely pinned by dying memcgs).
      
      The page_owner feature is not supposed to be enabled for production
      system due to its memory overhead.  However, if it is suspected that
      dying memcgs are increasing over time, a test environment with
      page_owner enabled can then be set up with appropriate workload for
      further analysis on what may be causing the increasing number of dying
      memcgs.
      
      This patch (of 4):
      
      For *scnprintf(), vsnprintf() is always called even if the input size is
      0.  That is a waste of time, so just return 0 in this case.
      
      Note that vsnprintf() will never return -1 to indicate an error.  So
      skipping the call to vsnprintf() when size is 0 will have no functional
      impact at all.
      
      Link: https://lkml.kernel.org/r/20220202203036.744010-1-longman@redhat.com
      Link: https://lkml.kernel.org/r/20220202203036.744010-2-longman@redhat.com
      
      
      Signed-off-by: default avatarWaiman Long <longman@redhat.com>
      Acked-by: default avatarDavid Rientjes <rientjes@google.com>
      Reviewed-by: default avatarSergey Senozhatsky <senozhatsky@chromium.org>
      Acked-by: default avatarRoman Gushchin <guro@fb.com>
      Acked-by: default avatarRafael Aquini <aquini@redhat.com>
      Acked-by: default avatarMike Rapoport <rppt@linux.ibm.com>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Vladimir Davydov <vdavydov.dev@gmail.com>
      Cc: Petr Mladek <pmladek@suse.com>
      Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
      Cc: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      Cc: Rasmus Villemoes <linux@rasmusvillemoes.dk>
      Cc: Ira Weiny <ira.weiny@intel.com>
      Cc: David Rientjes <rientjes@google.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      ef62c8ff
    • Shuah Khan's avatar
      Documentation/vm/page_owner.rst: fix unexpected indentation warns · 2e944985
      Shuah Khan authored
      Fix Unexpected indentation warns in page_owner:
      
        Documentation/vm/page_owner.rst:92: WARNING: Unexpected indentation.
        Documentation/vm/page_owner.rst:96: WARNING: Unexpected indentation.
        Documentation/vm/page_owner.rst:107: WARNING: Unexpected indentation.
      
      Link: https://lkml.kernel.org/r/20211215001929.47866-1-skhan@linuxfoundation.org
      
      
      Signed-off-by: default avatarShuah Khan <skhan@linuxfoundation.org>
      Cc: Jonathan Corbet <corbet@lwn.net>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      2e944985
    • Shenghong Han's avatar
      Documentation/vm/page_owner.rst: update the documentation · 57f2b54a
      Shenghong Han authored
      Update the documentation of ``page_owner``.
      
      [akpm@linux-foundation.org: small grammatical tweaks]
      
      Link: https://lkml.kernel.org/r/20211214134736.2569-1-hanshenghong2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarShenghong Han <hanshenghong2019@email.szu.edu.cn>
      Cc: Jonathan Corbet <corbet@lwn.net>
      Cc: Vlastimil Babka <vbabka@suse.cz>
      Cc: Georgi Djakov <georgi.djakov@linaro.org>
      Cc: Liam Mark <lmark@codeaurora.org>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Xiaoming Ni <nixiaoming@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      57f2b54a
    • Yixuan Cao's avatar
      tools/vm/page_owner_sort.c: delete invalid duplicate code · 41ed6434
      Yixuan Cao authored
      I noticed that there is two invalid lines of duplicate code.  It's better
      to delete it.
      
      Link: https://lkml.kernel.org/r/20211213095743.3630-1-caoyixuan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYixuan Cao <caoyixuan2019@email.szu.edu.cn>
      Cc: Mark Brown <broonie@kernel.org>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      41ed6434
    • Shenghong Han's avatar
      tools/vm/page_owner_sort.c: two trivial fixes · e7a3f677
      Shenghong Han authored
      1) There is an unused variable. It's better to delete it.
      2) One case is missing in the usage().
      
      Link: https://lkml.kernel.org/r/20211213164518.2461-1-hanshenghong2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarShenghong Han <hanshenghong2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      e7a3f677
    • Chongxi Zhao's avatar
      tools/vm/page_owner_sort.c: support sorting pid and time · 8f9c447e
      Chongxi Zhao authored
      When viewing the page owner information, we expect that the information
      can be sorted by PID, so that we can quickly combine PID with the program
      to check the information together.
      
      We also expect that the information can be sorted by time.  Time sorting
      helps to view the running status of the program according to the time
      interval when the program hangs up.
      
      Finally, we hope to pass the page_ owner_ Sort.  C can reduce part of the
      output and only output the plate information whose memory has not been
      released, which can make us locate the problem of the program faster.
      Therefore, the following adjustments have been made:
      
      1. Add the static functions search_pattern and check_regcomp to
         improve the cleanliness.
      
      2. Add member attributes and their corresponding sorting methods.  In
         terms of comparison time, int will overflow because the data of ull is
         too large, so the ternary operator is used
      
      3. Add the -f parameter to filter out the information of blocks whose
         memory has not been released
      
      Link: https://lkml.kernel.org/r/20211206165653.5093-1-zhaochongxi2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarChongxi Zhao <zhaochongxi2019@email.szu.edu.cn>
      Reviewed-by: default avatarSean Anderson <seanga2@gmail.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      8f9c447e
    • Yinan Zhang's avatar
      tools/vm/page_owner_sort.c: add switch between culling by stacktrace and txt · cd75ea0e
      Yinan Zhang authored
      Culling by comparing stacktrace would casue loss of some information.  For
      example, if there exists 2 blocks which have the same stacktrace and the
      different head info
      
        Page allocated via order 0, mask 0x108c48(...), pid 73696,
          ts 1578829190639010 ns, free_ts 1576583851324450 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
        Page allocated via order 0, mask 0x108c48(...), pid 61806,
          ts 1354113726046100 ns, free_ts 1354104926841400 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
      After culling, it would be like this
      
        2 times, 2 pages:
        Page allocated via order 0, mask 0x108c48(...), pid 73696,
          ts 1578829190639010 ns, free_ts 1576583851324450 ns
          prep_new_page+0x80/0xb8
          get_page_from_freelist+0x924/0xee8
          __alloc_pages+0x138/0xc18
          alloc_pages+0x80/0xf0
          __page_cache_alloc+0x90/0xc8
      
      The info of second block missed.  So, add -c to turn on culling by
      stacktrace.  By default, it will cull by txt.
      
      Link: https://lkml.kernel.org/r/20211129145658.2491-1-zhangyinan2019@email.szu.edu.cn
      
      
      Signed-off-by: default avatarYinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Sean Anderson <seanga2@gmail.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      cd75ea0e
    • Sean Anderson's avatar
      tools/vm/page_owner_sort.c: support sorting by stack trace · 82f5ebc2
      Sean Anderson authored
      This adds the ability to sort by stacktraces.  This is helpful when
      comparing multiple dumps of page_owner taken at different times, since
      blocks will not be reordered if they were allocated/free'd.
      
      Link: https://lkml.kernel.org/r/20211124193709.1805776-2-seanga2@gmail.com
      
      
      Signed-off-by: default avatarSean Anderson <seanga2@gmail.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      82f5ebc2
    • Sean Anderson's avatar
      tools/vm/page_owner_sort.c: sort by stacktrace before culling · ba5a396b
      Sean Anderson authored
      The contents of page_owner have changed to include more information than
      the stack trace.  On a modern kernel, the blocks look like
      
        Page allocated via order 0, mask 0x0(), pid 1, ts 165564237 ns, free_ts 0 ns
          register_early_stack+0x4b/0x90
          init_page_owner+0x39/0x250
          kernel_init_freeable+0x11e/0x242
          kernel_init+0x16/0x130
      
      Sorting by the contents of .txt will result in almost no repeated pages,
      as the pid, ts, and free_ts will almost never be the same.  Instead,
      sort by the contents of the stack trace, which we assume to be whatever
      is after the first line.
      
      [seanga2@gmail.com: fix NULL-pointer dereference when comparing stack traces]
        Link: https://lkml.kernel.org/r/20211125162653.1855958-1-seanga2@gmail.com
      
      Link: https://lkml.kernel.org/r/20211124193709.1805776-1-seanga2@gmail.com
      
      
      Signed-off-by: default avatarSean Anderson <seanga2@gmail.com>
      Cc: Changhee Han <ch0.han@lge.com>
      Cc: Tang Bin <tangbin@cmss.chinamobile.com>
      Cc: Zhang Shengju <zhangshengju@cmss.chinamobile.com>
      Cc: Zhenliang Wei <weizhenliang@huawei.com>
      Cc: Stephen Rothwell <sfr@canb.auug.org.au>
      Cc: Yinan Zhang <zhangyinan2019@email.szu.edu.cn>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      ba5a396b
    • Linus Torvalds's avatar
      Merge branch 'akpm' (patches from Andrew) · 52deda95
      Linus Torvalds authored
      Merge more updates from Andrew Morton:
       "Various misc subsystems, before getting into the post-linux-next
        material.
      
        41 patches.
      
        Subsystems affected by this patch series: procfs, misc, core-kernel,
        lib, checkpatch, init, pipe, minix, fat, cgroups, kexec, kdump,
        taskstats, panic, kcov, resource, and ubsan"
      
      * emailed patches from Andrew Morton <akpm@linux-foundation.org>: (41 commits)
        Revert "ubsan, kcsan: Don't combine sanitizer with kcov on clang"
        kernel/resource: fix kfree() of bootmem memory again
        kcov: properly handle subsequent mmap calls
        kcov: split ioctl handling into locked and unlocked parts
        panic: move panic_print before kmsg dumpers
        panic: add option to dump all CPUs backtraces in panic_print
        docs: sysctl/kernel: add missing bit to panic_print
        taskstats: remove unneeded dead assignment
        kasan: no need to unset panic_on_warn in end_report()
        ubsan: no need to unset panic_on_warn in ubsan_epilogue()
        panic: unset panic_on_war...
      52deda95
    • Linus Torvalds's avatar
      Merge tag 'net-next-5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next · 169e7776
      Linus Torvalds authored
      Pull networking updates from Jakub Kicinski:
       "The sprinkling of SPI drivers is because we added a new one and Mark
        sent us a SPI driver interface conversion pull request.
      
        Core
        ----
      
         - Introduce XDP multi-buffer support, allowing the use of XDP with
           jumbo frame MTUs and combination with Rx coalescing offloads (LRO).
      
         - Speed up netns dismantling (5x) and lower the memory cost a little.
           Remove unnecessary per-netns sockets. Scope some lists to a netns.
           Cut down RCU syncing. Use batch methods. Allow netdev registration
           to complete out of order.
      
         - Support distinguishing timestamp types (ingress vs egress) and
           maintaining them across packet scrubbing points (e.g. redirect).
      
         - Continue the work of annotating packet drop reasons throughout the
           stack.
      
         - Switch netdev error counters from an atomic to dynamically
           allocated per-CPU counters.
      
         - Rework a few preempt_disable(), local_irq_save() and busy waiting
           sections problematic on PREEMPT_RT.
      
         - Extend the ref_tracker to allow catching use-after-free bugs.
      
        BPF
        ---
      
         - Introduce "packing allocator" for BPF JIT images. JITed code is
           marked read only, and used to be allocated at page granularity.
           Custom allocator allows for more efficient memory use, lower iTLB
           pressure and prevents identity mapping huge pages from getting
           split.
      
         - Make use of BTF type annotations (e.g. __user, __percpu) to enforce
           the correct probe read access method, add appropriate helpers.
      
         - Convert the BPF preload to use light skeleton and drop the
           user-mode-driver dependency.
      
         - Allow XDP BPF_PROG_RUN test infra to send real packets, enabling
           its use as a packet generator.
      
         - Allow local storage memory to be allocated with GFP_KERNEL if
           called from a hook allowed to sleep.
      
         - Introduce fprobe (multi kprobe) to speed up mass attachment (arch
           bits to come later).
      
         - Add unstable conntrack lookup helpers for BPF by using the BPF
           kfunc infra.
      
         - Allow cgroup BPF progs to return custom errors to user space.
      
         - Add support for AF_UNIX iterator batching.
      
         - Allow iterator programs to use sleepable helpers.
      
         - Support JIT of add, and, or, xor and xchg atomic ops on arm64.
      
         - Add BTFGen support to bpftool which allows to use CO-RE in kernels
           without BTF info.
      
         - Large number of libbpf API improvements, cleanups and deprecations.
      
        Protocols
        ---------
      
         - Micro-optimize UDPv6 Tx, gaining up to 5% in test on dummy netdev.
      
         - Adjust TSO packet sizes based on min_rtt, allowing very low latency
           links (data centers) to always send full-sized TSO super-frames.
      
         - Make IPv6 flow label changes (AKA hash rethink) more configurable,
           via sysctl and setsockopt. Distinguish between server and client
           behavior.
      
         - VxLAN support to "collect metadata" devices to terminate only
           configured VNIs. This is similar to VLAN filtering in the bridge.
      
         - Support inserting IPv6 IOAM information to a fraction of frames.
      
         - Add protocol attribute to IP addresses to allow identifying where
           given address comes from (kernel-generated, DHCP etc.)
      
         - Support setting socket and IPv6 options via cmsg on ping6 sockets.
      
         - Reject mis-use of ECN bits in IP headers as part of DSCP/TOS.
           Define dscp_t and stop taking ECN bits into account in fib-rules.
      
         - Add support for locked bridge ports (for 802.1X).
      
         - tun: support NAPI for packets received from batched XDP buffs,
           doubling the performance in some scenarios.
      
         - IPv6 extension header handling in Open vSwitch.
      
         - Support IPv6 control message load balancing in bonding, prevent
           neighbor solicitation and advertisement from using the wrong port.
           Support NS/NA monitor selection similar to existing ARP monitor.
      
         - SMC
            - improve performance with TCP_CORK and sendfile()
            - support auto-corking
            - support TCP_NODELAY
      
         - MCTP (Management Component Transport Protocol)
            - add user space tag control interface
            - I2C binding driver (as specified by DMTF DSP0237)
      
         - Multi-BSSID beacon handling in AP mode for WiFi.
      
         - Bluetooth:
            - handle MSFT Monitor Device Event
            - add MGMT Adv Monitor Device Found/Lost events
      
         - Multi-Path TCP:
            - add support for the SO_SNDTIMEO socket option
            - lots of selftest cleanups and improvements
      
         - Increase the max PDU size in CAN ISOTP to 64 kB.
      
        Driver API
        ----------
      
         - Add HW counters for SW netdevs, a mechanism for devices which
           offload packet forwarding to report packet statistics back to
           software interfaces such as tunnels.
      
         - Select the default NIC queue count as a fraction of number of
           physical CPU cores, instead of hard-coding to 8.
      
         - Expose devlink instance locks to drivers. Allow device layer of
           drivers to use that lock directly instead of creating their own
           which always runs into ordering issues in devlink callbacks.
      
         - Add header/data split indication to guide user space enabling of
           TCP zero-copy Rx.
      
         - Allow configuring completion queue event size.
      
         - Refactor page_pool to enable fragmenting after allocation.
      
         - Add allocation and page reuse statistics to page_pool.
      
         - Improve Multiple Spanning Trees support in the bridge to allow
           reuse of topologies across VLANs, saving HW resources in switches.
      
         - DSA (Distributed Switch Architecture):
            - replay and offload of host VLAN entries
            - offload of static and local FDB entries on LAG interfaces
            - FDB isolation and unicast filtering
      
        New hardware / drivers
        ----------------------
      
         - Ethernet:
            - LAN937x T1 PHYs
            - Davicom DM9051 SPI NIC driver
            - Realtek RTL8367S, RTL8367RB-VB switch and MDIO
            - Microchip ksz8563 switches
            - Netronome NFP3800 SmartNICs
            - Fungible SmartNICs
            - MediaTek MT8195 switches
      
         - WiFi:
            - mt76: MediaTek mt7916
            - mt76: MediaTek mt7921u USB adapters
            - brcmfmac: Broadcom BCM43454/6
      
         - Mobile:
            - iosm: Intel M.2 7360 WWAN card
      
        Drivers
        -------
      
         - Convert many drivers to the new phylink API built for split PCS
           designs but also simplifying other cases.
      
         - Intel Ethernet NICs:
            - add TTY for GNSS module for E810T device
            - improve AF_XDP performance
            - GTP-C and GTP-U filter offload
            - QinQ VLAN support
      
         - Mellanox Ethernet NICs (mlx5):
            - support xdp->data_meta
            - multi-buffer XDP
            - offload tc push_eth and pop_eth actions
      
         - Netronome Ethernet NICs (nfp):
            - flow-independent tc action hardware offload (police / meter)
            - AF_XDP
      
         - Other Ethernet NICs:
            - at803x: fiber and SFP support
            - xgmac: mdio: preamble suppression and custom MDC frequencies
            - r8169: enable ASPM L1.2 if system vendor flags it as safe
            - macb/gem: ZynqMP SGMII
            - hns3: add TX push mode
            - dpaa2-eth: software TSO
            - lan743x: multi-queue, mdio, SGMII, PTP
            - axienet: NAPI and GRO support
      
         - Mellanox Ethernet switches (mlxsw):
            - source and dest IP address rewrites
            - RJ45 ports
      
         - Marvell Ethernet switches (prestera):
            - basic routing offload
            - multi-chain TC ACL offload
      
         - NXP embedded Ethernet switches (ocelot & felix):
            - PTP over UDP with the ocelot-8021q DSA tagging protocol
            - basic QoS classification on Felix DSA switch using dcbnl
            - port mirroring for ocelot switches
      
         - Microchip high-speed industrial Ethernet (sparx5):
            - offloading of bridge port flooding flags
            - PTP Hardware Clock
      
         - Other embedded switches:
            - lan966x: PTP Hardward Clock
            - qca8k: mdio read/write operations via crafted Ethernet packets
      
         - Qualcomm 802.11ax WiFi (ath11k):
            - add LDPC FEC type and 802.11ax High Efficiency data in radiotap
            - enable RX PPDU stats in monitor co-exist mode
      
         - Intel WiFi (iwlwifi):
            - UHB TAS enablement via BIOS
            - band disablement via BIOS
            - channel switch offload
            - 32 Rx AMPDU sessions in newer devices
      
         - MediaTek WiFi (mt76):
            - background radar detection
            - thermal management improvements on mt7915
            - SAR support for more mt76 platforms
            - MBSSID and 6 GHz band on mt7915
      
         - RealTek WiFi:
            - rtw89: AP mode
            - rtw89: 160 MHz channels and 6 GHz band
            - rtw89: hardware scan
      
         - Bluetooth:
            - mt7921s: wake on Bluetooth, SCO over I2S, wide-band-speed (WBS)
      
         - Microchip CAN (mcp251xfd):
            - multiple RX-FIFOs and runtime configurable RX/TX rings
            - internal PLL, runtime PM handling simplification
            - improve chip detection and error handling after wakeup"
      
      * tag 'net-next-5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next: (2521 commits)
        llc: fix netdevice reference leaks in llc_ui_bind()
        drivers: ethernet: cpsw: fix panic when interrupt coaleceing is set via ethtool
        ice: don't allow to run ice_send_event_to_aux() in atomic ctx
        ice: fix 'scheduling while atomic' on aux critical err interrupt
        net/sched: fix incorrect vlan_push_eth dest field
        net: bridge: mst: Restrict info size queries to bridge ports
        net: marvell: prestera: add missing destroy_workqueue() in prestera_module_init()
        drivers: net: xgene: Fix regression in CRC stripping
        net: geneve: add missing netlink policy and size for IFLA_GENEVE_INNER_PROTO_INHERIT
        net: dsa: fix missing host-filtered multicast addresses
        net/mlx5e: Fix build warning, detected write beyond size of field
        iwlwifi: mvm: Don't fail if PPAG isn't supported
        selftests/bpf: Fix kprobe_multi test.
        Revert "rethook: x86: Add rethook x86 implementation"
        Revert "arm64: rethook: Add arm64 rethook implementation"
        Revert "powerpc: Add rethook support"
        Revert "ARM: rethook: Add rethook arm implementation"
        netdevice: add missing dm_private kdoc
        net: bridge: mst: prevent NULL deref in br_mst_info_size()
        selftests: forwarding: Use same VRF for port and VLAN upper
        ...
      169e7776
    • Linus Torvalds's avatar
      Merge tag 'vfio-v5.18-rc1' of https://github.com/awilliam/linux-vfio · 7403e6d8
      Linus Torvalds authored
      Pull VFIO updates from Alex Williamson:
      
       - Introduce new device migration uAPI and implement device specific
         mlx5 vfio-pci variant driver supporting new protocol (Jason
         Gunthorpe, Yishai Hadas, Leon Romanovsky)
      
       - New HiSilicon acc vfio-pci variant driver, also supporting migration
         interface (Shameer Kolothum, Longfang Liu)
      
       - D3hot fixes for vfio-pci-core (Abhishek Sahu)
      
       - Document new vfio-pci variant driver acceptance criteria
         (Alex Williamson)
      
       - Fix UML build unresolved ioport_{un}map() functions
         (Alex Williamson)
      
       - Fix MAINTAINERS due to header movement (Lukas Bulwahn)
      
      * tag 'vfio-v5.18-rc1' of https://github.com/awilliam/linux-vfio: (31 commits)
        vfio-pci: Provide reviewers and acceptance criteria for variant drivers
        MAINTAINERS: adjust entry for header movement in hisilicon qm driver
        hisi_acc_vfio_pci: Use its own PCI reset_done error handler
        hisi_acc_vfio_pci: Add support for VFIO live migration
        crypto: hisilicon/qm: Set the VF QM state register
        hisi_acc_vfio_pci: Add helper to retrieve the struct pci_driver
        hisi_acc_vfio_pci: Restrict access to VF dev BAR2 migration region
        hisi_acc_vfio_pci: add new vfio_pci driver for HiSilicon ACC devices
        hisi_acc_qm: Move VF PCI device IDs to common header
        crypto: hisilicon/qm: Move few definitions to common header
        crypto: hisilicon/qm: Move the QM header to include/linux
        vfio/mlx5: Fix to not use 0 as NULL pointer
        PCI/IOV: Fix wrong kernel-doc identifier
        vfio/mlx5: Use its own PCI reset_done error handler
        vfio/pci: Expose vfio_pci_core_aer_err_detected()
        vfio/mlx5: Implement vfio_pci driver for mlx5 devices
        vfio/mlx5: Expose migration commands over mlx5 device
        vfio: Remove migration protocol v1 documentation
        vfio: Extend the device migration protocol with RUNNING_P2P
        vfio: Define device migration protocol v2
        ...
      7403e6d8
    • Linus Torvalds's avatar
      Merge tag 'hyperv-next-signed-20220322' of... · 66711cfe
      Linus Torvalds authored
      Merge tag 'hyperv-next-signed-20220322' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux
      
      Pull hyperv updates from Wei Liu:
       "Minor patches from various people"
      
      * tag 'hyperv-next-signed-20220322' of git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux:
        x86/hyperv: Output host build info as normal Windows version number
        hv_balloon: rate-limit "Unhandled message" warning
        drivers: hv: log when enabling crash_kexec_post_notifiers
        hv_utils: Add comment about max VMbus packet size in VSS driver
        Drivers: hv: Compare cpumasks and not their weights in init_vp_index()
        Drivers: hv: Rename 'alloced' to 'allocated'
        Drivers: hv: vmbus: Use struct_size() helper in kmalloc()
      66711cfe
    • Linus Torvalds's avatar
      Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm · 1ebdbeb0
      Linus Torvalds authored
      Pull kvm updates from Paolo Bonzini:
       "ARM:
         - Proper emulation of the OSLock feature of the debug architecture
      
         - Scalibility improvements for the MMU lock when dirty logging is on
      
         - New VMID allocator, which will eventually help with SVA in VMs
      
         - Better support for PMUs in heterogenous systems
      
         - PSCI 1.1 support, enabling support for SYSTEM_RESET2
      
         - Implement CONFIG_DEBUG_LIST at EL2
      
         - Make CONFIG_ARM64_ERRATUM_2077057 default y
      
         - Reduce the overhead of VM exit when no interrupt is pending
      
         - Remove traces of 32bit ARM host support from the documentation
      
         - Updated vgic selftests
      
         - Various cleanups, doc updates and spelling fixes
      
        RISC-V:
         - Prevent KVM_COMPAT from being selected
      
         - Optimize __kvm_riscv_switch_to() implementation
      
         - RISC-V SBI v0.3 support
      
        s390:
         - memop selftest
      
         - fix SCK locking
      
         - adapter interruptions virtualization for secure guests
      
         - add Claudio Imbrenda as maintainer
      
         - first step to do proper storage key checking
      
        x86:
         - Continue switching kvm_x86_ops to static_call(); introduce
           static_call_cond() and __static_call_ret0 when applicable.
      
         - Cleanup unused arguments in several functions
      
         - Synthesize AMD 0x80000021 leaf
      
         - Fixes and optimization for Hyper-V sparse-bank hypercalls
      
         - Implement Hyper-V's enlightened MSR bitmap for nested SVM
      
         - Remove MMU auditing
      
         - Eager splitting of page tables (new aka "TDP" MMU only) when dirty
           page tracking is enabled
      
         - Cleanup the implementation of the guest PGD cache
      
         - Preparation for the implementation of Intel IPI virtualization
      
         - Fix some segment descriptor checks in the emulator
      
         - Allow AMD AVIC support on systems with physical APIC ID above 255
      
         - Better API to disable virtualization quirks
      
         - Fixes and optimizations for the zapping of page tables:
      
            - Zap roots in two passes, avoiding RCU read-side critical
              sections that last too long for very large guests backed by 4
              KiB SPTEs.
      
            - Zap invalid and defunct roots asynchronously via
              concurrency-managed work queue.
      
            - Allowing yielding when zapping TDP MMU roots in response to the
              root's last reference being put.
      
            - Batch more TLB flushes with an RCU trick. Whoever frees the
              paging structure now holds RCU as a proxy for all vCPUs running
              in the guest, i.e. to prolongs the grace period on their behalf.
              It then kicks the the vCPUs out of guest mode before doing
              rcu_read_unlock().
      
        Generic:
         - Introduce __vcalloc and use it for very large allocations that need
           memcg accounting"
      
      * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (246 commits)
        KVM: use kvcalloc for array allocations
        KVM: x86: Introduce KVM_CAP_DISABLE_QUIRKS2
        kvm: x86: Require const tsc for RT
        KVM: x86: synthesize CPUID leaf 0x80000021h if useful
        KVM: x86: add support for CPUID leaf 0x80000021
        KVM: x86: do not use KVM_X86_OP_OPTIONAL_RET0 for get_mt_mask
        Revert "KVM: x86/mmu: Zap only TDP MMU leafs in kvm_zap_gfn_range()"
        kvm: x86/mmu: Flush TLB before zap_gfn_range releases RCU
        KVM: arm64: fix typos in comments
        KVM: arm64: Generalise VM features into a set of flags
        KVM: s390: selftests: Add error memop tests
        KVM: s390: selftests: Add more copy memop tests
        KVM: s390: selftests: Add named stages for memop test
        KVM: s390: selftests: Add macro as abstraction for MEM_OP
        KVM: s390: selftests: Split memop tests
        KVM: s390x: fix SCK locking
        RISC-V: KVM: Implement SBI HSM suspend call
        RISC-V: KVM: Add common kvm_riscv_vcpu_wfi() function
        RISC-V: Add SBI HSM suspend related defines
        RISC-V: KVM: Implement SBI v0.3 SRST extension
        ...
      1ebdbeb0
    • Linus Torvalds's avatar
      Merge tag 'tomoyo-pr-20220322' of git://git.osdn.net/gitroot/tomoyo/tomoyo-test1 · efee6c79
      Linus Torvalds authored
      Pull tomoyo update from Tetsuo Handa:
       "Avoid unnecessarily leaking kernel command line arguments"
      
      * tag 'tomoyo-pr-20220322' of git://git.osdn.net/gitroot/tomoyo/tomoyo-test1:
        TOMOYO: fix __setup handlers return values
      efee6c79
    • Linus Torvalds's avatar
      Merge tag 'flexible-array-transformations-5.18-rc1' of... · 3ce62cf4
      Linus Torvalds authored
      Merge tag 'flexible-array-transformations-5.18-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gustavoars/linux
      
      Pull flexible-array transformations from Gustavo Silva:
       "Treewide patch that replaces zero-length arrays with flexible-array
        members.
      
        This has been baking in linux-next for a whole development cycle"
      
      * tag 'flexible-array-transformations-5.18-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gustavoars/linux:
        treewide: Replace zero-length arrays with flexible-array members
      3ce62cf4
    • Linus Torvalds's avatar
      Merge tag 'prlimit-tasklist_lock-for-v5.18' of... · cd4699c5
      Linus Torvalds authored
      Merge tag 'prlimit-tasklist_lock-for-v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiederm/user-namespace
      
      Pull tasklist_lock optimizations from Eric Biederman:
       "prlimit and getpriority tasklist_lock optimizations
      
        The tasklist_lock popped up as a scalability bottleneck on some
        testing workloads. The readlocks in do_prlimit and set/getpriority are
        not necessary in all cases.
      
        Based on a cycles profile, it looked like ~87% of the time was spent
        in the kernel, ~42% of which was just trying to get *some* spinlock
        (queued_spin_lock_slowpath, not necessarily the tasklist_lock).
      
        The big offenders (with rough percentages in cycles of the overall
        trace):
         - do_wait 11%
         - setpriority 8% (done previously in commit 7f8ca0ed)
         - kill 8%
         - do_exit 5%
         - clone 3%
         - prlimit64 2%   (this patchset)
         - getrlimit 1%   (this patchset)
      
        I can't easily test this patchset on the original workload for various
        reasons. Instead, I used the microbenchmark below to at least verify
        there was some improvement. This patchset had a 28% speedup (12% from
        baseline to set/getprio, then another 14% for prlimit).
      
        This series used to do the setpriority case, but an almost identical
        change was merged as commit 7f8ca0ed ("kernel/sys.c: only take
        tasklist_lock for get/setpriority(PRIO_PGRP)") so that has been
        dropped from here.
      
        One interesting thing is that my libc's getrlimit() was calling
        prlimit64, so hoisting the read_lock(tasklist_lock) into sys_prlimit64
        had no effect - it essentially optimized the older syscalls only. I
        didn't do that in this patchset, but figured I'd mention it since it
        was an option from the previous patch's discussion"
      
      micobenchmark.c:
      ---------------
      	int main(int argc, char **argv)
      	{
      		pid_t child;
      		struct rlimit rlim[1];
      
      		fork(); fork(); fork(); fork(); fork(); fork();
      
      		for (int i = 0; i < 5000; i++) {
      			child = fork();
      			if (child < 0)
      				exit(1);
      			if (child > 0) {
      				usleep(1000);
      				kill(child, SIGTERM);
      				waitpid(child, NULL, 0);
      			} else {
      				for (;;) {
      					setpriority(PRIO_PROCESS, 0,
      						    getpriority(PRIO_PROCESS, 0));
      					getrlimit(RLIMIT_CPU, rlim);
      				}
      			}
      		}
      
      		return 0;
      	}
      
      Link: https://lore.kernel.org/lkml/20211213220401.1039578-1-brho@google.com/ [v1]
      Link: https://lore.kernel.org/lkml/20220105212828.197013-1-brho@google.com/ [v2]
      Link: https://lore.kernel.org/lkml/20220106172041.522167-1-brho@google.com/ [v3]
      
      * tag 'prlimit-tasklist_lock-for-v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiederm/user-namespace:
        prlimit: do not grab the tasklist_lock
        prlimit: make do_prlimit() static
      cd4699c5
    • Linus Torvalds's avatar
      Merge tag 'fs.rt.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux · 2e2d4650
      Linus Torvalds authored
      Pull mount attributes PREEMPT_RT update from Christian Brauner:
       "This contains Sebastian's fix to make changing mount
        attributes/getting write access compatible with CONFIG_PREEMPT_RT.
      
        The change only applies when users explicitly opt-in to real-time via
        CONFIG_PREEMPT_RT otherwise things are exactly as before. We've waited
        quite a long time with this to make sure folks could take a good look"
      
      * tag 'fs.rt.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux:
        fs/namespace: Boost the mount_lock.lock owner instead of spinning on PREEMPT_RT.
      2e2d4650
    • Linus Torvalds's avatar
      Merge tag 'fs.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux · 15f2e3d6
      Linus Torvalds authored
      Pull mount_setattr updates from Christian Brauner:
       "This contains a few more patches to massage the mount_setattr()
        codepaths and one minor fix to reuse a helper we added some time back.
      
        The final two patches do similar cleanups in different ways. One patch
        is mine and the other is Al's who was nice enough to give me a branch
        for it.
      
        Since his came in later and my branch had been sitting in -next for
        quite some time we just put his on top instead of swap them"
      
      * tag 'fs.v5.18' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux:
        mount_setattr(): clean the control flow and calling conventions
        fs: clean up mount_setattr control flow
        fs: don't open-code mnt_hold_writers()
        fs: simplify check in mount_setattr_commit()
        fs: add mnt_allow_writers() and simplify mount_setattr_prepare()
      15f2e3d6