1. Sep 05, 2022
    • Yishai Hadas's avatar
      IB/core: Fix a nested dead lock as part of ODP flow · 85eaeb50
      Yishai Hadas authored
      Fix a nested dead lock as part of ODP flow by using mmput_async().
      
      From the below call trace [1] can see that calling mmput() once we have
      the umem_odp->umem_mutex locked as required by
      ib_umem_odp_map_dma_and_lock() might trigger in the same task the
      exit_mmap()->__mmu_notifier_release()->mlx5_ib_invalidate_range() which
      may dead lock when trying to lock the same mutex.
      
      Moving to use mmput_async() will solve the problem as the above
      exit_mmap() flow will be called in other task and will be executed once
      the lock will be available.
      
      [1]
      [64843.077665] task:kworker/u133:2  state:D stack:    0 pid:80906 ppid:
      2 flags:0x00004000
      [64843.077672] Workqueue: mlx5_ib_page_fault mlx5_ib_eqe_pf_action [mlx5_ib]
      [64843.077719] Call Trace:
      [64843.077722]  <TASK>
      [64843.077724]  __schedule+0x23d/0x590
      [64843.077729]  schedule+0x4e/0xb0
      [64843.077735]  schedule_preempt_disabled+0xe/0x10
      [64843.077740]  __mutex_lock.constprop.0+0x263/0x490
      [64843.077747]  __mutex_lock_slowpath+0x13/0x20
      [64843.077752]  mutex_lock+0x34/0x40
      [64843.077758]  mlx5_ib_invalidate_range+0x48/0x270 [mlx5_ib]
      [64843.077808]  __mmu_notifier_release+0x1a4/0x200
      [64843.077816]  exit_mmap+0x1bc/0x200
      [64843.077822]  ? walk_page_range+0x9c/0x120
      [64843.077828]  ? __cond_resched+0x1a/0x50
      [64843.077833]  ? mutex_lock+0x13/0x40
      [64843.077839]  ? uprobe_clear_state+0xac/0x120
      [64843.077860]  mmput+0x5f/0x140
      [64843.077867]  ib_umem_odp_map_dma_and_lock+0x21b/0x580 [ib_core]
      [64843.077931]  pagefault_real_mr+0x9a/0x140 [mlx5_ib]
      [64843.077962]  pagefault_mr+0xb4/0x550 [mlx5_ib]
      [64843.077992]  pagefault_single_data_segment.constprop.0+0x2ac/0x560
      [mlx5_ib]
      [64843.078022]  mlx5_ib_eqe_pf_action+0x528/0x780 [mlx5_ib]
      [64843.078051]  process_one_work+0x22b/0x3d0
      [64843.078059]  worker_thread+0x53/0x410
      [64843.078065]  ? process_one_work+0x3d0/0x3d0
      [64843.078073]  kthread+0x12a/0x150
      [64843.078079]  ? set_kthread_struct+0x50/0x50
      [64843.078085]  ret_from_fork+0x22/0x30
      [64843.078093]  </TASK>
      
      Fixes: 36f30e48
      
       ("IB/core: Improve ODP to use hmm_range_fault()")
      Reviewed-by: default avatarMaor Gottlieb <maorg@nvidia.com>
      Signed-off-by: default avatarYishai Hadas <yishaih@nvidia.com>
      Link: https://lore.kernel.org/r/74d93541ea533ef7daec6f126deb1072500aeb16.1661251841.git.leonro@nvidia.com
      
      
      Signed-off-by: default avatarLeon Romanovsky <leon@kernel.org>
      85eaeb50
  2. Sep 04, 2022
    • Linus Walleij's avatar
      RDMA/siw: Pass a pointer to virt_to_page() · 0d1b756a
      Linus Walleij authored
      Functions that work on a pointer to virtual memory such as
      virt_to_pfn() and users of that function such as
      virt_to_page() are supposed to pass a pointer to virtual
      memory, ideally a (void *) or other pointer. However since
      many architectures implement virt_to_pfn() as a macro,
      this function becomes polymorphic and accepts both a
      (unsigned long) and a (void *).
      
      If we instead implement a proper virt_to_pfn(void *addr)
      function the following happens (occurred on arch/arm):
      
      drivers/infiniband/sw/siw/siw_qp_tx.c:32:23: warning: incompatible
        integer to pointer conversion passing 'dma_addr_t' (aka 'unsigned int')
        to parameter of type 'const void *' [-Wint-conversion]
      drivers/infiniband/sw/siw/siw_qp_tx.c:32:37: warning: passing argument
        1 of 'virt_to_pfn' makes pointer from integer without a cast
        [-Wint-conversion]
      drivers/infiniband/sw/siw/siw_qp_tx.c:538:36: warning: incompatible
        integer to pointer conversion passing 'unsigned long long'
        to parameter of type 'const void *' [-Wint-conversion]
      
      Fix this with an explicit cast. In one case where the SIW
      SGE uses an unaligned u64 we need a double cast modifying the
      virtual address (va) to a platform-specific uintptr_t before
      casting to a (void *).
      
      Fixes: b9be6f18
      
       ("rdma/siw: transmit path")
      Cc: linux-rdma@vger.kernel.org
      Signed-off-by: default avatarLinus Walleij <linus.walleij@linaro.org>
      Link: https://lore.kernel.org/r/20220902215918.603761-1-linus.walleij@linaro.org
      
      
      Signed-off-by: default avatarLeon Romanovsky <leon@kernel.org>
      0d1b756a
  3. Sep 01, 2022
    • yangx.jy@fujitsu.com's avatar
      RDMA/srp: Set scmnd->result only when scmnd is not NULL · 12f35199
      yangx.jy@fujitsu.com authored
      This change fixes the following kernel NULL pointer dereference
      which is reproduced by blktests srp/007 occasionally.
      
      BUG: kernel NULL pointer dereference, address: 0000000000000170
      PGD 0 P4D 0
      Oops: 0002 [#1] PREEMPT SMP NOPTI
      CPU: 0 PID: 9 Comm: kworker/0:1H Kdump: loaded Not tainted 6.0.0-rc1+ #37
      Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.15.0-29-g6a62e0cb0dfe-prebuilt.qemu.org 04/01/2014
      Workqueue:  0x0 (kblockd)
      RIP: 0010:srp_recv_done+0x176/0x500 [ib_srp]
      Code: 00 4d 85 ff 0f 84 52 02 00 00 48 c7 82 80 02 00 00 00 00 00 00 4c 89 df 4c 89 14 24 e8 53 d3 4a f6 4c 8b 14 24 41 0f b6 42 13 <41> 89 87 70 01 00 00 41 0f b6 52 12 f6 c2 02 74 44 41 8b 42 1c b9
      RSP: 0018:ffffaef7c0003e28 EFLAGS: 00000282
      RAX: 0000000000000000 RBX: ffff9bc9486dea60 RCX: 0000000000000000
      RDX: 0000000000000102 RSI: ffffffffb76bbd0e RDI: 00000000ffffffff
      RBP: ffff9bc980099a00 R08: 0000000000000001 R09: 0000000000000001
      R10: ffff9bca53ef0000 R11: ffff9bc980099a10 R12: ffff9bc956e14000
      R13: ffff9bc9836b9cb0 R14: ffff9bc9557b4480 R15: 0000000000000000
      FS:  0000000000000000(0000) GS:ffff9bc97ec00000(0000) knlGS:0000000000000000
      CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
      CR2: 0000000000000170 CR3: 0000000007e04000 CR4: 00000000000006f0
      Call Trace:
       <IRQ>
       __ib_process_cq+0xb7/0x280 [ib_core]
       ib_poll_handler+0x2b/0x130 [ib_core]
       irq_poll_softirq+0x93/0x150
       __do_softirq+0xee/0x4b8
       irq_exit_rcu+0xf7/0x130
       sysvec_apic_timer_interrupt+0x8e/0xc0
       </IRQ>
      
      Fixes: ad215aae ("RDMA/srp: Make struct scsi_cmnd and struct srp_request adjacent")
      Link: https://lore.kernel.org/r/20220831081626.18712-1-yangx.jy@fujitsu.com
      
      
      Signed-off-by: default avatarXiao Yang <yangx.jy@fujitsu.com>
      Acked-by: default avatarBart Van Assche <bvanassche@acm.org>
      Signed-off-by: default avatarLeon Romanovsky <leon@kernel.org>
      12f35199
  4. Aug 30, 2022
  5. Aug 29, 2022
  6. Aug 28, 2022
    • Shiraz Saleem's avatar
      RDMA/irdma: Fix drain SQ hang with no completion · ead54ced
      Shiraz Saleem authored
      SW generated completions for outstanding WRs posted on SQ
      after QP is in error target the wrong CQ. This causes the
      ib_drain_sq to hang with no completion.
      
      Fix this to generate completions on the right CQ.
      
      [  863.969340] INFO: task kworker/u52:2:671 blocked for more than 122 seconds.
      [  863.979224]       Not tainted 5.14.0-130.el9.x86_64 #1
      [  863.986588] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
      [  863.996997] task:kworker/u52:2   state:D stack:    0 pid:  671 ppid:     2 flags:0x00004000
      [  864.007272] Workqueue: xprtiod xprt_autoclose [sunrpc]
      [  864.014056] Call Trace:
      [  864.017575]  __schedule+0x206/0x580
      [  864.022296]  schedule+0x43/0xa0
      [  864.026736]  schedule_timeout+0x115/0x150
      [  864.032185]  __wait_for_common+0x93/0x1d0
      [  864.037717]  ? usleep_range_state+0x90/0x90
      [  864.043368]  __ib_drain_sq+0xf6/0x170 [ib_core]
      [  864.049371]  ? __rdma_block_iter_next+0x80/0x80 [ib_core]
      [  864.056240]  ib_drain_sq+0x66/0x70 [ib_core]
      [  864.062003]  rpcrdma_xprt_disconnect+0x82/0x3b0 [rpcrdma]
      [  864.069365]  ? xprt_prepare_transmit+0x5d/0xc0 [sunrpc]
      [  864.076386]  xprt_rdma_close+0xe/0x30 [rpcrdma]
      [  864.082593]  xprt_autoclose+0x52/0x100 [sunrpc]
      [  864.088718]  process_one_work+0x1e8/0x3c0
      [  864.094170]  worker_thread+0x50/0x3b0
      [  864.099109]  ? rescuer_thread+0x370/0x370
      [  864.104473]  kthread+0x149/0x170
      [  864.109022]  ? set_kthread_struct+0x40/0x40
      [  864.114713]  ret_from_fork+0x22/0x30
      
      Fixes: 81091d76 ("RDMA/irdma: Add SW mechanism to generate completions on error")
      Link: https://lore.kernel.org/r/20220824154358.117-1-shiraz.saleem@intel.com
      
      
      Reported-by: default avatarKamal Heib <kamalheib1@gmail.com>
      Signed-off-by: default avatarShiraz Saleem <shiraz.saleem@intel.com>
      Signed-off-by: default avatarLeon Romanovsky <leon@kernel.org>
      ead54ced
  7. Aug 21, 2022
  8. Aug 16, 2022
  9. Aug 15, 2022
    • Linus Torvalds's avatar
      Linux 6.0-rc1 · 568035b0
      Linus Torvalds authored
      568035b0
    • Yury Norov's avatar
      radix-tree: replace gfp.h inclusion with gfp_types.h · 9f162193
      Yury Norov authored
      
      
      Radix tree header includes gfp.h for __GFP_BITS_SHIFT only. Now we
      have gfp_types.h for this.
      
      Fixes powerpc allmodconfig build:
      
         In file included from include/linux/nodemask.h:97,
                          from include/linux/mmzone.h:17,
                          from include/linux/gfp.h:7,
                          from include/linux/radix-tree.h:12,
                          from include/linux/idr.h:15,
                          from include/linux/kernfs.h:12,
                          from include/linux/sysfs.h:16,
                          from include/linux/kobject.h:20,
                          from include/linux/pci.h:35,
                          from arch/powerpc/kernel/prom_init.c:24:
         include/linux/random.h: In function 'add_latent_entropy':
      >> include/linux/random.h:25:46: error: 'latent_entropy' undeclared (first use in this function); did you mean 'add_latent_entropy'?
            25 |         add_device_randomness((const void *)&latent_entropy, sizeof(latent_entropy));
               |                                              ^~~~~~~~~~~~~~
               |                                              add_latent_entropy
         include/linux/random.h:25:46: note: each undeclared identifier is reported only once for each function it appears in
      
      Reported-by: default avatarkernel test robot <lkp@intel.com>
      CC: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
      CC: Andrew Morton <akpm@linux-foundation.org>
      CC: Jason A. Donenfeld <Jason@zx2c4.com>
      Signed-off-by: default avatarYury Norov <yury.norov@gmail.com>
      Signed-off-by: default avatarLinus Torvalds <torvalds@linux-foundation.org>
      9f162193
    • Linus Torvalds's avatar
      Merge tag 'pull-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs · 74cbb480
      Linus Torvalds authored
      Pull vfs lseek fix from Al Viro:
       "Fix proc_reg_llseek() breakage. Always had been possible if somebody
        left NULL ->proc_lseek, became a practical issue now"
      
      * tag 'pull-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs:
        take care to handle NULL ->proc_lseek()
      74cbb480
    • Al Viro's avatar
      take care to handle NULL ->proc_lseek() · 3f61631d
      Al Viro authored
      Easily done now, just by clearing FMODE_LSEEK in ->f_mode
      during proc_reg_open() for such entries.
      
      Fixes: 868941b1
      
       "fs: remove no_llseek"
      Signed-off-by: default avatarAl Viro <viro@zeniv.linux.org.uk>
      3f61631d
    • Linus Torvalds's avatar
      Merge tag 'for-linus-6.0-rc1b-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip · 5d6a0f4d
      Linus Torvalds authored
      Pull more xen updates from Juergen Gross:
      
       - fix the handling of the "persistent grants" feature negotiation
         between Xen blkfront and Xen blkback drivers
      
       - a cleanup of xen.config and adding xen.config to Xen section in
         MAINTAINERS
      
       - support HVMOP_set_evtchn_upcall_vector, which is more compliant to
         "normal" interrupt handling than the global callback used up to now
      
       - further small cleanups
      
      * tag 'for-linus-6.0-rc1b-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/xen/tip:
        MAINTAINERS: add xen config fragments to XEN HYPERVISOR sections
        xen: remove XEN_SCRUB_PAGES in xen.config
        xen/pciback: Fix comment typo
        xen/xenbus: fix return type in xenbus_file_read()
        xen-blkfront: Apply 'feature_persistent' parameter when connect
        xen-blkback: Apply 'feature_persistent' parameter when connect
        xen-blkback: fix persistent grants negotiation
        x86/xen: Add support for HVMOP_set_evtchn_upcall_vector
      5d6a0f4d
    • Linus Torvalds's avatar
      Merge tag 'perf-tools-fixes-for-v6.0-2022-08-13' of... · 96f86ff0
      Linus Torvalds authored
      Merge tag 'perf-tools-fixes-for-v6.0-2022-08-13' of git://git.kernel.org/pub/scm/linux/kernel/git/acme/linux
      
      Pull more perf tool updates from Arnaldo Carvalho de Melo:
      
       - 'perf c2c' now supports ARM64, adjust its output to cope with
         differences with what is in x86_64. Now go find false sharing on
         ARM64 (at least Neoverse) as well!
      
       - Refactor the JSON processing, making the output more compact and thus
         reducing the size of the resulting perf binary
      
       - Improvements for 'perf offcpu' profiling, including tracking child
         processes
      
       - Update Intel JSON metrics and events files for broadwellde,
         broadwellx, cascadelakex, haswellx, icelakex, ivytown, jaketown,
         knightslanding, sapphirerapids, skylakex and snowridgex
      
       - Add 'perf stat' JSON output and a 'perf test' entry for it
      
       - Ignore memfd and anonymous mmap events if jitdump present
      
       - Refactor 'perf test' shell tests allowing subdirs
      
       - Fix an error handling path in 'parse_perf_probe_command()'
      
       - Fixes for the guest Intel PT tracing patchkit in the 1st batch of
         this merge window
      
       - Print debuginfod queries if -v option is used, to explain delays in
         processing when debuginfo servers are enabled to fetch DSOs with
         richer symbol tables
      
       - Improve error message for 'perf record -p not_existing_pid'
      
       - Fix openssl and libbpf feature detection
      
       - Add PMU pai_crypto event description for IBM z16 on 'perf list'
      
       - Fix typos and duplicated words on comments in various places
      
      * tag 'perf-tools-fixes-for-v6.0-2022-08-13' of git://git.kernel.org/pub/scm/linux/kernel/git/acme/linux: (81 commits)
        perf test: Refactor shell tests allowing subdirs
        perf vendor events: Update events for snowridgex
        perf vendor events: Update events and metrics for skylakex
        perf vendor events: Update metrics for sapphirerapids
        perf vendor events: Update events for knightslanding
        perf vendor events: Update metrics for jaketown
        perf vendor events: Update metrics for ivytown
        perf vendor events: Update events and metrics for icelakex
        perf vendor events: Update events and metrics for haswellx
        perf vendor events: Update events and metrics for cascadelakex
        perf vendor events: Update events and metrics for broadwellx
        perf vendor events: Update metrics for broadwellde
        perf jevents: Fold strings optimization
        perf jevents: Compress the pmu_events_table
        perf metrics: Copy entire pmu_event in find metric
        perf pmu-events: Hide the pmu_events
        perf pmu-events: Don't assume pmu_event is an array
        perf pmu-events: Move test events/metrics to JSON
        perf test: Use full metric resolution
        perf pmu-events: Hide pmu_events_map
        ...
      96f86ff0
  10. Aug 14, 2022