1. Feb 23, 2024
    • Shakeel Butt's avatar
      mm: writeback: ratelimit stat flush from mem_cgroup_wb_stats · d9b3ce87
      Shakeel Butt authored
      One of our workloads (Postgres 14) has regressed when migrated from 5.10
      to 6.1 upstream kernel.  The regression can be reproduced by sysbench's
      oltp_write_only benchmark.  It seems like the always on rstat flush in
      mem_cgroup_wb_stats() is causing the regression.  So, rate limit that
      specific rstat flush.  One potential consequence would be the dirty
      throttling might be decided on stale memcg stats.  However from our
      benchmarks and production traffic we have not observed any change in the
      dirty throttling behavior of the application.
      
      Link: https://lkml.kernel.org/r/20240118184235.618164-1-shakeelb@google.com
      Fixes: 2d146aa3
      
       ("mm: memcontrol: switch to rstat")
      Signed-off-by: default avatarShakeel Butt <shakeelb@google.com>
      Acked-by: default avatarJohannes Weiner <hannes@cmpxchg.org>
      Acked-by: default avatarRoman Gushchin <roman.gushchin@linux.dev>
      Cc: Jan Kara <jack@suse.cz>
      Cc: Jens Axboe <axboe@kernel.dk>
      Cc: Michal Hocko <mhocko@kernel.org>
      Cc: Muchun Song <muchun.song@linux.dev>
      Cc: Tejun Heo <tj@kernel.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      d9b3ce87
    • Kefeng Wang's avatar
      mm: memory: move mem_cgroup_charge() into alloc_anon_folio() · 085ff35e
      Kefeng Wang authored
      The GFP flags from vma_thp_gfp_mask() according to user configuration only
      used for large folio allocation but not for memory cgroup charge, and
      GFP_KERNEL is used for both order-0 and large order folio when memory
      cgroup charge at present.  However, mem_cgroup_charge() uses the GFP flags
      in a fairly sophisticated way.  In addition to checking
      gfpflags_allow_blocking(), it pays attention to __GFP_NORETRY and
      __GFP_RETRY_MAYFAIL to ensure that processes within this memcg do not
      exceed their quotas.
      
      So we'd better to move mem_cgroup_charge() into alloc_anon_folio(),
      
      1) it will make us to allocate as much as possible large order folio,
         because we could try the next order if mem_cgroup_charge() fails,
         although the memcg's memory usage is close to its limits.
      
      2) using same GFP flags for allocation and charge is to be consistent
         with PMD THP firstly, in addition, according to GFP flag returned from
         vma_thp_gfp_mask(), GFP_TRANSHUGE_LIGHT could make us skip direct
         reclaim, _GFP_NORETRY will make us skip mem_cgroup_oom() and won't
         trigger memory cgroup oom from large order(order <= COSTLY_ORDER) folio
         charging.
      
      Link: https://lkml.kernel.org/r/20240122011612.501029-1-wangkefeng.wang@huawei.com
      Link: https://lkml.kernel.org/r/20240117103954.2756050-1-wangkefeng.wang@huawei.com
      
      
      Signed-off-by: default avatarKefeng Wang <wangkefeng.wang@huawei.com>
      Reviewed-by: default avatarRyan Roberts <ryan.roberts@arm.com>
      Cc: David Hildenbrand <david@redhat.com>
      Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
      Cc: Michal Hocko <mhocko@suse.com>
      Cc: Roman Gushchin <roman.gushchin@linux.dev>
      Cc: Johannes Weiner <hannes@cmpxchg.org>
      Cc: Shakeel Butt <shakeelb@google.com>
      Cc: Muchun Song <songmuchun@bytedance.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      085ff35e
    • Ryan Roberts's avatar
      tools/mm: add thpmaps script to dump THP usage info · 2444172c
      Ryan Roberts authored
      With the proliferation of large folios for file-backed memory, and more
      recently the introduction of multi-size THP for anonymous memory, it is
      becoming useful to be able to see exactly how large folios are mapped into
      processes.  For some architectures (e.g.  arm64), if most memory is mapped
      using contpte-sized and -aligned blocks, TLB usage can be optimized so
      it's useful to see where these requirements are and are not being met.
      
      thpmaps is a Python utility that reads /proc/<pid>/smaps,
      /proc/<pid>/pagemap and /proc/kpageflags to print information about how
      transparent huge pages (both file and anon) are mapped to a specified
      process or cgroup.  It aims to help users debug and optimize their
      workloads.  In future we may wish to introduce stats directly into the
      kernel (e.g.  smaps or similar), but for now this provides a short term
      solution without the need to introduce any new ABI.
      
      Run with help option for a full listing of the arguments:
      
          # ./t...
      2444172c
    • Ronald Monthero's avatar
      mm/zswap: improve with alloc_workqueue() call · 8409a385
      Ronald Monthero authored
      The core-api create_workqueue is deprecated, this patch replaces the
      create_workqueue with alloc_workqueue.  The previous implementation
      workqueue of zswap was a bounded workqueue, this patch uses
      alloc_workqueue() to create an unbounded workqueue.  The WQ_UNBOUND
      attribute is desirable making the workqueue not localized to a specific
      cpu so that the scheduler is free to exercise improvisations in any
      demanding scenarios for offloading cpu time slices for workqueues.  For
      example if any other workqueues of the same primary cpu had to be served
      which are WQ_HIGHPRI and WQ_CPU_INTENSIVE.  Also Unbound workqueue happens
      to be more efficient in a system during memory pressure scenarios in
      comparison to a bounded workqueue.
      
      shrink_wq = alloc_workqueue("zswap-shrink",
                           WQ_UNBOUND|WQ_MEM_RECLAIM, 1);
      
      Overall the change suggested in this patch should be seamless and does not
      alter the existing behavior, other than the improvisation to be an
      unbounded workqueue.
      
      Link: https://lkml.kernel.org/r/20240116133145.12454-1-debug.penguin32@gmail.com
      
      
      Signed-off-by: default avatarRonald Monthero <debug.penguin32@gmail.com>
      Acked-by: default avatarNhat Pham <nphamcs@gmail.com>
      Acked-by: default avatarJohannes Weiner <hannes@cmpxchg.org>
      Cc: Chris Li <chrisl@kernel.org>
      Cc: Dan Streetman <ddstreet@ieee.org>
      Cc: Seth Jennings <sjenning@redhat.com>
      Cc: Vitaly Wool <vitaly.wool@konsulko.com>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      8409a385
    • Pankaj Raghav's avatar
      readahead: use ilog2 instead of a while loop in page_cache_ra_order() · e03c16fb
      Pankaj Raghav authored
      A while loop is used to adjust the new_order to be lower than the
      ra->size.  ilog2 could be used to do the same instead of using a loop.
      
      ilog2 typically resolves to a bit scan reverse instruction.  This is
      particularly useful when ra->size is smaller than the 2^new_order as it
      resolves in one instruction instead of looping to find the new_order.
      
      No functional changes.
      
      Link: https://lkml.kernel.org/r/20240115102523.2336742-1-kernel@pankajraghav.com
      
      
      Signed-off-by: default avatarPankaj Raghav <p.raghav@samsung.com>
      Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
      Signed-off-by: default avatarAndrew Morton <akpm@linux-foundation.org>
      e03c16fb
  2. Feb 22, 2024
  3. Feb 21, 2024