1. Dec 15, 2023
    • Linus Torvalds's avatar
      Merge tag '6.7-rc5-smb3-client-fixes' of git://git.samba.org/sfrench/cifs-2.6 · 3f716859
      Linus Torvalds authored
      Pull smb client fixes from Steve French:
       "Address OOBs and NULL dereference found by Dr. Morris's recent
        analysis and fuzzing.
      
        All marked for stable as well"
      
      * tag '6.7-rc5-smb3-client-fixes' of git://git.samba.org/sfrench/cifs-2.6:
        smb: client: fix OOB in smb2_query_reparse_point()
        smb: client: fix NULL deref in asn1_ber_decoder()
        smb: client: fix potential OOBs in smb2_parse_contexts()
        smb: client: fix OOB in receive_encrypted_standard()
      3f716859
    • Linus Torvalds's avatar
      Merge tag 'platform-drivers-x86-v6.7-4' of... · 976600c6
      Linus Torvalds authored
      Merge tag 'platform-drivers-x86-v6.7-4' of git://git.kernel.org/pub/scm/linux/kernel/git/pdx86/platform-drivers-x86
      
      Pull x86 platform driver fixes from Ilpo Järvinen:
      
       - tablet-mode-switch events fix
      
       - kernel-doc warning fixes
      
      * tag 'platform-drivers-x86-v6.7-4' of git://git.kernel.org/pub/scm/linux/kernel/git/pdx86/platform-drivers-x86:
        platform/x86: intel_ips: fix kernel-doc formatting
        platform/x86: thinkpad_acpi: fix kernel-doc warnings
        platform/x86: intel-vbtn: Fix missing tablet-mode-switch events
      976600c6
    • Linus Torvalds's avatar
      Merge tag 'net-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net · c7402612
      Linus Torvalds authored
      Pull networking fixes from Paolo Abeni:
      "Current release - regressions:
      
         - tcp: fix tcp_disordered_ack() vs usec TS resolution
      
        Current release - new code bugs:
      
         - dpll: sanitize possible null pointer dereference in
           dpll_pin_parent_pin_set()
      
         - eth: octeon_ep: initialise control mbox tasks before using APIs
      
        Previous releases - regressions:
      
         - io_uring/af_unix: disable sending io_uring over sockets
      
         - eth: mlx5e:
             - TC, don't offload post action rule if not supported
             - fix possible deadlock on mlx5e_tx_timeout_work
      
         - eth: iavf: fix iavf_shutdown to call iavf_remove instead iavf_close
      
         - eth: bnxt_en: fix skb recycling logic in bnxt_deliver_skb()
      
         - eth: ena: fix DMA syncing in XDP path when SWIOTLB is on
      
         - eth: team: fix use-after-free when an option instance allocation
           fails
      
        Previous releases - always broken:
      
         - neighbour: don't let neigh_forced_gc() disable preemption for long
      
         - net: prevent mss overflow in skb_segment()
      
         - ipv6: support reporting otherwise unknown prefix flags in
           RTM_NEWPREFIX
      
         - tcp: remove acked SYN flag from packet in the transmit queue
           correctly
      
         - eth: octeontx2-af:
             - fix a use-after-free in rvu_nix_register_reporters
             - fix promisc mcam entry action
      
         - eth: dwmac-loongson: make sure MDIO is initialized before use
      
         - eth: atlantic: fix double free in ring reinit logic"
      
      * tag 'net-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (62 commits)
        net: atlantic: fix double free in ring reinit logic
        appletalk: Fix Use-After-Free in atalk_ioctl
        net: stmmac: Handle disabled MDIO busses from devicetree
        net: stmmac: dwmac-qcom-ethqos: Fix drops in 10M SGMII RX
        dpaa2-switch: do not ask for MDB, VLAN and FDB replay
        dpaa2-switch: fix size of the dma_unmap
        net: prevent mss overflow in skb_segment()
        vsock/virtio: Fix unsigned integer wrap around in virtio_transport_has_space()
        Revert "tcp: disable tcp_autocorking for socket when TCP_NODELAY flag is set"
        MIPS: dts: loongson: drop incorrect dwmac fallback compatible
        stmmac: dwmac-loongson: drop useless check for compatible fallback
        stmmac: dwmac-loongson: Make sure MDIO is initialized before use
        tcp: disable tcp_autocorking for socket when TCP_NODELAY flag is set
        dpll: sanitize possible null pointer dereference in dpll_pin_parent_pin_set()
        net: ena: Fix XDP redirection error
        net: ena: Fix DMA syncing in XDP path when SWIOTLB is on
        net: ena: Fix xdp drops handling due to multibuf packets
        net: ena: Destroy correct number of xdp queues upon failure
        net: Remove acked SYN flag from packet in the transmit queue correctly
        qed: Fix a potential use-after-free in qed_cxt_tables_alloc
        ...
      c7402612
    • Linus Torvalds's avatar
      Merge tag 'for-6.7-rc5-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux · bdb2701f
      Linus Torvalds authored
      Pull btrfs fixes from David Sterba:
        "Some fixes to quota accounting code, mostly around error handling and
         correctness:
      
         - free reserves on various error paths, after IO errors or
           transaction abort
      
         - don't clear reserved range at the folio release time, it'll be
           properly cleared after final write
      
         - fix integer overflow due to int used when passing around size of
           freed reservations
      
         - fix a regression in squota accounting that missed some cases with
           delayed refs"
      
      * tag 'for-6.7-rc5-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
        btrfs: ensure releasing squota reserve on head refs
        btrfs: don't clear qgroup reserved bit in release_folio
        btrfs: free qgroup pertrans reserve on transaction abort
        btrfs: fix qgroup_free_reserved_data int overflow
        btrfs: free qgroup reserve when ORDERED_IOERR is set
      bdb2701f
  2. Dec 14, 2023
  3. Dec 13, 2023
    • David S. Miller's avatar
      Merge branch 'stmmac-bug-fixes' · 2513974c
      David S. Miller authored
      Yanteng Si says:
      
      ====================
      stmmac: Some bug fixes
      
      * Put Krzysztof's patch into my thread, pick Conor's Reviewed-by
        tag and Jiaxun's Acked-by tag.(prev version is RFC patch)
      
      * I fixed an Oops related to mdio, mainly to ensure that
        mdio is initialized before use, because it will be used
        in a series of patches I am working on.
      
      see <https://lore.kernel.org/loongarch/cover.1699533745.git.siyanteng@loongson.cn/T/#t
      
      >
      ====================
      
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      2513974c
    • Krzysztof Kozlowski's avatar
      MIPS: dts: loongson: drop incorrect dwmac fallback compatible · 4907a3f5
      Krzysztof Kozlowski authored
      
      
      Device binds to proper PCI ID (LOONGSON, 0x7a03), already listed in DTS,
      so checking for some other compatible does not make sense.  It cannot be
      bound to unsupported platform.
      
      Drop useless, incorrect (space in between) and undocumented compatible.
      
      Signed-off-by: default avatarKrzysztof Kozlowski <krzysztof.kozlowski@linaro.org>
      Signed-off-by: default avatarYanteng Si <siyanteng@loongson.cn>
      Reviewed-by: default avatarConor Dooley <conor.dooley@microchip.com>
      Acked-by: default avatarJiaxun Yang <jiaxun.yang@flygoat.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      4907a3f5
    • Krzysztof Kozlowski's avatar
      stmmac: dwmac-loongson: drop useless check for compatible fallback · 31fea092
      Krzysztof Kozlowski authored
      
      
      Device binds to proper PCI ID (LOONGSON, 0x7a03), already listed in DTS,
      so checking for some other compatible does not make sense.  It cannot be
      bound to unsupported platform.
      
      Drop useless, incorrect (space in between) and undocumented compatible.
      
      Signed-off-by: default avatarKrzysztof Kozlowski <krzysztof.kozlowski@linaro.org>
      Signed-off-by: default avatarYanteng Si <siyanteng@loongson.cn>
      Reviewed-by: default avatarConor Dooley <conor.dooley@microchip.com>
      Acked-by: default avatarJiaxun Yang <jiaxun.yang@flygoat.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      31fea092
    • Yanteng Si's avatar
      stmmac: dwmac-loongson: Make sure MDIO is initialized before use · e87d3a13
      Yanteng Si authored
      Generic code will use mdio. If it is not initialized before use,
      the kernel will Oops.
      
      Fixes: 30bba69d
      
       ("stmmac: pci: Add dwmac support for Loongson")
      Signed-off-by: default avatarYanteng Si <siyanteng@loongson.cn>
      Signed-off-by: default avatarFeiyang Chen <chenfeiyang@loongson.cn>
      Reviewed-by: default avatarAndrew Lunn <andrew@lunn.ch>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      e87d3a13
    • Salvatore Dipietro's avatar
      tcp: disable tcp_autocorking for socket when TCP_NODELAY flag is set · f3f32a35
      Salvatore Dipietro authored
      Based on the tcp man page, if TCP_NODELAY is set, it disables Nagle's algorithm
      and packets are sent as soon as possible. However in the `tcp_push` function
      where autocorking is evaluated the `nonagle` value set by TCP_NODELAY is not
      considered which can trigger unexpected corking of packets and induce delays.
      
      For example, if two packets are generated as part of a server's reply, if the
      first one is not transmitted on the wire quickly enough, the second packet can
      trigger the autocorking in `tcp_push` and be delayed instead of sent as soon as
      possible. It will either wait for additional packets to be coalesced or an ACK
      from the client before transmitting the corked packet. This can interact badly
      if the receiver has tcp delayed acks enabled, introducing 40ms extra delay in
      completion times. It is not always possible to control who has delayed acks
      set, but it is possible to adjust when and how autocorking is triggered.
      Patch prevents autocorking if the TCP_NODELAY flag is set on the socket.
      
      Patch has been tested using an AWS c7g.2xlarge instance with Ubuntu 22.04 and
      Apache Tomcat 9.0.83 running the basic servlet below:
      
      import java.io.IOException;
      import java.io.OutputStreamWriter;
      import java.io.PrintWriter;
      import javax.servlet.ServletException;
      import javax.servlet.http.HttpServlet;
      import javax.servlet.http.HttpServletRequest;
      import javax.servlet.http.HttpServletResponse;
      
      public class HelloWorldServlet extends HttpServlet {
          @Override
          protected void doGet(HttpServletRequest request, HttpServletResponse response)
            throws ServletException, IOException {
              response.setContentType("text/html;charset=utf-8");
              OutputStreamWriter osw = new OutputStreamWriter(response.getOutputStream(),"UTF-8");
              String s = "a".repeat(3096);
              osw.write(s,0,s.length());
              osw.flush();
          }
      }
      
      Load was applied using  wrk2 (https://github.com/kinvolk/wrk2) from an AWS
      c6i.8xlarge instance.  With the current auto-corking behavior and TCP_NODELAY
      set an additional 40ms latency from P99.99+ values are observed.  With the
      patch applied we see no occurrences of 40ms latencies. The patch has also been
      tested with iperf and uperf benchmarks and no regression was observed.
      
      # No patch with tcp_autocorking=1 and TCP_NODELAY set on all sockets
      ./wrk -t32 -c128 -d40s --latency -R10000  http://172.31.49.177:8080/hello/hello'
        ...
       50.000%    0.91ms
       75.000%    1.12ms
       90.000%    1.46ms
       99.000%    1.73ms
       99.900%    1.96ms
       99.990%   43.62ms   <<< 40+ ms extra latency
       99.999%   48.32ms
      100.000%   49.34ms
      
      # With patch
      ./wrk -t32 -c128 -d40s --latency -R10000  http://172.31.49.177:8080/hello/hello'
        ...
       50.000%    0.89ms
       75.000%    1.13ms
       90.000%    1.44ms
       99.000%    1.67ms
       99.900%    1.78ms
       99.990%    2.27ms   <<< no 40+ ms extra latency
       99.999%    3.71ms
      100.000%    4.57ms
      
      Fixes: f54b3111
      
       ("tcp: auto corking")
      Signed-off-by: default avatarSalvatore Dipietro <dipiets@amazon.com>
      Reviewed-by: default avatarEric Dumazet <edumazet@google.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      f3f32a35
    • Linus Torvalds's avatar
      Merge tag 'hid-for-linus-2023121201' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid · 88035e56
      Linus Torvalds authored
      Pull HID fixes from Jiri Kosina:
      
       - Lenovo ThinkPad TrackPoint Keyboard II firmware-specific regression
         fix (Mikhail Khvainitski)
      
       - device-specific fixes (various authors)
      
      * tag 'hid-for-linus-2023121201' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid:
        HID: apple: Add "hfd.cn" and "WKB603" to the list of non-apple keyboards
        HID: lenovo: Restrict detection of patched firmware only to USB cptkbd
        HID: Add quirk for Labtec/ODDOR/aikeec handbrake
        HID: i2c-hid: Add IDEA5002 to i2c_hid_acpi_blacklist[]
        mailmap: add address mapping for Jiri Kosina
      88035e56
    • Jiri Pirko's avatar
      dpll: sanitize possible null pointer dereference in dpll_pin_parent_pin_set() · 65c95f78
      Jiri Pirko authored
      
      
      User may not pass DPLL_A_PIN_STATE attribute in the pin set operation
      message. Sanitize that by checking if the attr pointer is not null
      and process the passed state attribute value only in that case.
      
      Reported-by: default avatarXingyuan Mo <hdthky0@gmail.com>
      Fixes: 9d71b54b
      
       ("dpll: netlink: Add DPLL framework base functions")
      Signed-off-by: default avatarJiri Pirko <jiri@nvidia.com>
      Acked-by: default avatarVadim Fedorenko <vadim.fedorenko@linux.dev>
      Link: https://lore.kernel.org/r/20231211083758.1082853-1-jiri@resnulli.us
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      65c95f78
    • Jakub Kicinski's avatar
      Merge branch 'ena-driver-xdp-bug-fixes' · 154bb2fa
      Jakub Kicinski authored
      David Arinzon says:
      
      ====================
      ENA driver XDP bug fixes
      
      This patchset contains multiple XDP-related bug fixes
      in the ENA driver.
      ====================
      
      Link: https://lore.kernel.org/r/20231211062801.27891-1-darinzon@amazon.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      154bb2fa
    • David Arinzon's avatar
      net: ena: Fix XDP redirection error · 4ab138ca
      David Arinzon authored
      When sending TX packets, the meta descriptor can be all zeroes
      as no meta information is required (as in XDP).
      
      This patch removes the validity check, as when
      `disable_meta_caching` is enabled, such TX packets will be
      dropped otherwise.
      
      Fixes: 0e3a3f6d
      
       ("net: ena: support new LLQ acceleration mode")
      Signed-off-by: default avatarShay Agroskin <shayagr@amazon.com>
      Signed-off-by: default avatarDavid Arinzon <darinzon@amazon.com>
      Link: https://lore.kernel.org/r/20231211062801.27891-5-darinzon@amazon.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      4ab138ca
    • David Arinzon's avatar
      net: ena: Fix DMA syncing in XDP path when SWIOTLB is on · d7601170
      David Arinzon authored
      This patch fixes two issues:
      
      Issue 1
      -------
      Description
      ```````````
      Current code does not call dma_sync_single_for_cpu() to sync data from
      the device side memory to the CPU side memory before the XDP code path
      uses the CPU side data.
      This causes the XDP code path to read the unset garbage data in the CPU
      side memory, resulting in incorrect handling of the packet by XDP.
      
      Solution
      ````````
      1. Add a call to dma_sync_single_for_cpu() before the XDP code starts to
         use the data in the CPU side memory.
      2. The XDP code verdict can be XDP_PASS, in which case there is a
         fallback to the non-XDP code, which also calls
         dma_sync_single_for_cpu().
         To avoid calling dma_sync_single_for_cpu() twice:
      2.1. Put the dma_sync_single_for_cpu() in the code in such a place where
           it happens before XDP and non-XDP code.
      2.2. Remove the calls to dma_sync_single_for_cpu() in the non-XDP code
           for the first buffer only (rx_copybreak and non-rx_copybreak
           cases), since the new call that was added covers these cases.
           The call to dma_sync_single_for_cpu() for the second buffer and on
           stays because only the first buffer is handled by the newly added
           dma_sync_single_for_cpu(). And there is no need for special
           handling of the second buffer and on for the XDP path since
           currently the driver supports only single buffer packets.
      
      Issue 2
      -------
      Description
      ```````````
      In case the XDP code forwarded the packet (ENA_XDP_FORWARDED),
      ena_unmap_rx_buff_attrs() is called with attrs set to 0.
      This means that before unmapping the buffer, the internal function
      dma_unmap_page_attrs() will also call dma_sync_single_for_cpu() on
      the whole buffer (not only on the data part of it).
      This sync is both wasteful (since a sync was already explicitly
      called before) and also causes a bug, which will be explained
      using the below diagram.
      
      The following diagram shows the flow of events causing the bug.
      The order of events is (1)-(4) as shown in the diagram.
      
      CPU side memory area
      
           (3)convert_to_xdp_frame() initializes the
              headroom with xdpf metadata
                            ||
                            \/
                ___________________________________
               |                                   |
       0       |                                   V                       4K
       ---------------------------------------------------------------------
       | xdpf->data      | other xdpf       |   < data >   | tailroom ||...|
       |                 | fields           |              | GARBAGE  ||   |
       ---------------------------------------------------------------------
      
                         /\                        /\
                         ||                        ||
         (4)ena_unmap_rx_buff_attrs() calls     (2)dma_sync_single_for_cpu()
            dma_sync_single_for_cpu() on the       copies data from device
            whole buffer page, overwriting         side to CPU side memory
            the xdpf->data with GARBAGE.           ||
       0                                                                   4K
       ---------------------------------------------------------------------
       | headroom                           |   < data >   | tailroom ||...|
       | GARBAGE                            |              | GARBAGE  ||   |
       ---------------------------------------------------------------------
      
      Device side memory area                      /\
                                                   ||
                                     (1) device writes RX packet data
      
      After the call to ena_unmap_rx_buff_attrs() in (4), the xdpf->data
      becomes corrupted, and so when it is later accessed in
      ena_clean_xdp_irq()->xdp_return_frame(), it causes a page fault,
      crashing the kernel.
      
      Solution
      ````````
      Explicitly tell ena_unmap_rx_buff_attrs() not to call
      dma_sync_single_for_cpu() by passing it the ENA_DMA_ATTR_SKIP_CPU_SYNC
      flag.
      
      Fixes: f7d625ad
      
       ("net: ena: Add dynamic recycling mechanism for rx buffers")
      Signed-off-by: default avatarArthur Kiyanovski <akiyano@amazon.com>
      Signed-off-by: default avatarDavid Arinzon <darinzon@amazon.com>
      Link: https://lore.kernel.org/r/20231211062801.27891-4-darinzon@amazon.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      d7601170
    • David Arinzon's avatar
      net: ena: Fix xdp drops handling due to multibuf packets · 505b1a88
      David Arinzon authored
      Current xdp code drops packets larger than ENA_XDP_MAX_MTU.
      This is an incorrect condition since the problem is not the
      size of the packet, rather the number of buffers it contains.
      
      This commit:
      
      1. Identifies and drops XDP multi-buffer packets at the
         beginning of the function.
      2. Increases the xdp drop statistic when this drop occurs.
      3. Adds a one-time print that such drops are happening to
         give better indication to the user.
      
      Fixes: 838c93dc
      
       ("net: ena: implement XDP drop support")
      Signed-off-by: default avatarArthur Kiyanovski <akiyano@amazon.com>
      Signed-off-by: default avatarDavid Arinzon <darinzon@amazon.com>
      Link: https://lore.kernel.org/r/20231211062801.27891-3-darinzon@amazon.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      505b1a88
    • David Arinzon's avatar
      net: ena: Destroy correct number of xdp queues upon failure · 41db6f99
      David Arinzon authored
      The ena_setup_and_create_all_xdp_queues() function freed all the
      resources upon failure, after creating only xdp_num_queues queues,
      instead of freeing just the created ones.
      
      In this patch, the only resources that are freed, are the ones
      allocated right before the failure occurs.
      
      Fixes: 548c4940
      
       ("net: ena: Implement XDP_TX action")
      Signed-off-by: default avatarShahar Itzko <itzko@amazon.com>
      Signed-off-by: default avatarDavid Arinzon <darinzon@amazon.com>
      Link: https://lore.kernel.org/r/20231211062801.27891-2-darinzon@amazon.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      41db6f99
    • Dong Chenchen's avatar
      net: Remove acked SYN flag from packet in the transmit queue correctly · f99cd562
      Dong Chenchen authored
      syzkaller report:
      
       kernel BUG at net/core/skbuff.c:3452!
       invalid opcode: 0000 [#1] PREEMPT SMP KASAN PTI
       CPU: 0 PID: 0 Comm: swapper/0 Not tainted 6.7.0-rc4-00009-gbee0e776-dirty #135
       RIP: 0010:skb_copy_and_csum_bits (net/core/skbuff.c:3452)
       Call Trace:
       icmp_glue_bits (net/ipv4/icmp.c:357)
       __ip_append_data.isra.0 (net/ipv4/ip_output.c:1165)
       ip_append_data (net/ipv4/ip_output.c:1362 net/ipv4/ip_output.c:1341)
       icmp_push_reply (net/ipv4/icmp.c:370)
       __icmp_send (./include/net/route.h:252 net/ipv4/icmp.c:772)
       ip_fragment.constprop.0 (./include/linux/skbuff.h:1234 net/ipv4/ip_output.c:592 net/ipv4/ip_output.c:577)
       __ip_finish_output (net/ipv4/ip_output.c:311 net/ipv4/ip_output.c:295)
       ip_output (net/ipv4/ip_output.c:427)
       __ip_queue_xmit (net/ipv4/ip_output.c:535)
       __tcp_transmit_skb (net/ipv4/tcp_output.c:1462)
       __tcp_retransmit_skb (net/ipv4/tcp_output.c:3387)
       tcp_retransmit_skb (net/ipv4/tcp_output.c:3404)
       tcp_retransmit_timer (net/ipv4/tcp_timer.c:604)
       tcp_write_timer (./include/linux/spinlock.h:391 net/ipv4/tcp_timer.c:716)
      
      The panic issue was trigered by tcp simultaneous initiation.
      The initiation process is as follows:
      
            TCP A                                            TCP B
      
        1.  CLOSED                                           CLOSED
      
        2.  SYN-SENT     --> <SEQ=100><CTL=SYN>              ...
      
        3.  SYN-RECEIVED <-- <SEQ=300><CTL=SYN>              <-- SYN-SENT
      
        4.               ... <SEQ=100><CTL=SYN>              --> SYN-RECEIVED
      
        5.  SYN-RECEIVED --> <SEQ=100><ACK=301><CTL=SYN,ACK> ...
      
        // TCP B: not send challenge ack for ack limit or packet loss
        // TCP A: close
      	tcp_close
      	   tcp_send_fin
                    if (!tskb && tcp_under_memory_pressure(sk))
                        tskb = skb_rb_last(&sk->tcp_rtx_queue); //pick SYN_ACK packet
                 TCP_SKB_CB(tskb)->tcp_flags |= TCPHDR_FIN;  // set FIN flag
      
        6.  FIN_WAIT_1  --> <SEQ=100><ACK=301><END_SEQ=102><CTL=SYN,FIN,ACK> ...
      
        // TCP B: send challenge ack to SYN_FIN_ACK
      
        7.               ... <SEQ=301><ACK=101><CTL=ACK>   <-- SYN-RECEIVED //challenge ack
      
        // TCP A:  <SND.UNA=101>
      
        8.  FIN_WAIT_1 --> <SEQ=101><ACK=301><END_SEQ=102><CTL=SYN,FIN,ACK> ... // retransmit panic
      
      	__tcp_retransmit_skb  //skb->len=0
      	    tcp_trim_head
      		len = tp->snd_una - TCP_SKB_CB(skb)->seq // len=101-100
      		    __pskb_trim_head
      			skb->data_len -= len // skb->len=-1, wrap around
      	    ... ...
      	    ip_fragment
      		icmp_glue_bits //BUG_ON
      
      If we use tcp_trim_head() to remove acked SYN from packet that contains data
      or other flags, skb->len will be incorrectly decremented. We can remove SYN
      flag that has been acked from rtx_queue earlier than tcp_trim_head(), which
      can fix the problem mentioned above.
      
      Fixes: 1da177e4
      
       ("Linux-2.6.12-rc2")
      Co-developed-by: default avatarEric Dumazet <edumazet@google.com>
      Signed-off-by: default avatarEric Dumazet <edumazet@google.com>
      Signed-off-by: default avatarDong Chenchen <dongchenchen2@huawei.com>
      Link: https://lore.kernel.org/r/20231210020200.1539875-1-dongchenchen2@huawei.com
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      f99cd562
    • Dinghao Liu's avatar
      qed: Fix a potential use-after-free in qed_cxt_tables_alloc · b65d52ac
      Dinghao Liu authored
      qed_ilt_shadow_alloc() will call qed_ilt_shadow_free() to
      free p_hwfn->p_cxt_mngr->ilt_shadow on error. However,
      qed_cxt_tables_alloc() accesses the freed pointer on failure
      of qed_ilt_shadow_alloc() through calling qed_cxt_mngr_free(),
      which may lead to use-after-free. Fix this issue by setting
      p_mngr->ilt_shadow to NULL in qed_ilt_shadow_free().
      
      Fixes: fe56b9e6
      
       ("qed: Add module with basic common support")
      Reviewed-by: default avatarPrzemek Kitszel <przemyslaw.kitszel@intel.com>
      Signed-off-by: default avatarDinghao Liu <dinghao.liu@zju.edu.cn>
      Link: https://lore.kernel.org/r/20231210045255.21383-1-dinghao.liu@zju.edu.cn
      
      
      Signed-off-by: default avatarJakub Kicinski <kuba@kernel.org>
      b65d52ac
    • Linus Torvalds's avatar
      Merge tag 'ext4_for_linus-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4 · cf52eed7
      Linus Torvalds authored
      Pull ext4 fixes from Ted Ts'o:
       "Fix various bugs / regressions for ext4, including a soft lockup, a
        WARN_ON, and a BUG"
      
      * tag 'ext4_for_linus-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4:
        jbd2: fix soft lockup in journal_finish_inode_data_buffers()
        ext4: fix warning in ext4_dio_write_end_io()
        jbd2: increase the journal IO's priority
        jbd2: correct the printing of write_flags in jbd2_write_superblock()
        ext4: prevent the normalized size from exceeding EXT_MAX_BLOCKS
      cf52eed7
    • Slawomir Laba's avatar
      iavf: Fix iavf_shutdown to call iavf_remove instead iavf_close · 7ae42ef3
      Slawomir Laba authored
      Make the flow for pci shutdown be the same to the pci remove.
      
      iavf_shutdown was implementing an incomplete version
      of iavf_remove. It misses several calls to the kernel like
      iavf_free_misc_irq, iavf_reset_interrupt_capability, iounmap
      that might break the system on reboot or hibernation.
      
      Implement the call of iavf_remove directly in iavf_shutdown to
      close this gap.
      
      Fixes below error messages (dmesg) during shutdown stress tests -
      [685814.900917] ice 0000:88:00.0: MAC 02:d0:5f:82:43:5d does not exist for
       VF 0
      [685814.900928] ice 0000:88:00.0: MAC 33:33:00:00:00:01 does not exist for
      VF 0
      
      Reproduction:
      
      1. Create one VF interface:
      echo 1 > /sys/class/net/<interface_name>/device/sriov_numvfs
      
      2. Run live dmesg on the host:
      dmesg -wH
      
      3. On SUT, script below steps into vf_namespace_assignment.sh
      
      <#!/bin/sh> // Remove <>. Git removes # line
      if=<VF name> (edit this per VF name)
      loop=0
      
      while true; do
      
      echo test round $loop
      let loop++
      
      ip netns add ns$loop
      ip link set dev $if up
      ip link set dev $if netns ns$loop
      ip netns exec ns$loop ip link set dev $if up
      ip netns exec ns$loop ip link set dev $if netns 1
      ip netns delete ns$loop
      
      done
      
      4. Run the script for at least 1000 iterations on SUT:
      ./vf_namespace_assignment.sh
      
      Expected result:
      No errors in dmesg.
      
      Fixes: 129cf89e
      
       ("iavf: rename functions and structs to new name")
      Signed-off-by: default avatarSlawomir Laba <slawomirx.laba@intel.com>
      Reviewed-by: default avatarMichal Swiatkowski <michal.swiatkowski@linux.intel.com>
      Reviewed-by: default avatarAhmed Zaki <ahmed.zaki@intel.com>
      Reviewed-by: default avatarJesse Brandeburg <jesse.brandeburg@intel.com>
      Co-developed-by: default avatarRanganatha Rao <ranganatha.rao@intel.com>
      Signed-off-by: default avatarRanganatha Rao <ranganatha.rao@intel.com>
      Tested-by: default avatarRafal Romanowski <rafal.romanowski@intel.com>
      Signed-off-by: default avatarTony Nguyen <anthony.l.nguyen@intel.com>
      7ae42ef3
    • Piotr Gardocki's avatar
      iavf: Handle ntuple on/off based on new state machines for flow director · 09d23b89
      Piotr Gardocki authored
      ntuple-filter feature on/off:
      Default is on. If turned off, the filters will be removed from both
      PF and iavf list. The removal is irrespective of current filter state.
      
      Steps to reproduce:
      -------------------
      
      1. Ensure ntuple is on.
      
      ethtool -K enp8s0 ntuple-filters on
      
      2. Create a filter to receive the traffic into non-default rx-queue like 15
      and ensure traffic is flowing into queue into 15.
      Now, turn off ntuple. Traffic should not flow to configured queue 15.
      It should flow to default RX queue.
      
      Fixes: 0dbfbabb
      
       ("iavf: Add framework to enable ethtool ntuple filters")
      Signed-off-by: default avatarPiotr Gardocki <piotrx.gardocki@intel.com>
      Reviewed-by: default avatarLarysa Zaremba <larysa.zaremba@intel.com>
      Signed-off-by: default avatarRanganatha Rao <ranganatha.rao@intel.com>
      Tested-by: default avatarRafal Romanowski <rafal.romanowski@intel.com>
      Signed-off-by: default avatarTony Nguyen <anthony.l.nguyen@intel.com>
      09d23b89
    • Piotr Gardocki's avatar
      iavf: Introduce new state machines for flow director · 3a0b5a29
      Piotr Gardocki authored
      New states introduced:
      
       IAVF_FDIR_FLTR_DIS_REQUEST
       IAVF_FDIR_FLTR_DIS_PENDING
       IAVF_FDIR_FLTR_INACTIVE
      
      Current FDIR state machines (SM) are not adequate to handle a few
      scenarios in the link DOWN/UP event, reset event and ntuple-feature.
      
      For example, when VF link goes DOWN and comes back UP administratively,
      the expectation is that previously installed filters should also be
      restored. But with current SM, filters are not restored.
      So with new SM, during link DOWN filters are marked as INACTIVE in
      the iavf list but removed from PF. After link UP, SM will transition
      from INACTIVE to ADD_REQUEST to restore the filter.
      
      Similarly, with VF reset, filters will be removed from the PF, but
      marked as INACTIVE in the iavf list. Filters will be restored after
      reset completion.
      
      Steps to reproduce:
      -------------------
      
      1. Create a VF. Here VF is enp8s0.
      
      2. Assign IP addresses to VF and link partner and ping continuously
      from remote. Here remote IP is 1.1.1.1.
      
      3. Check default RX Queue of traffic.
      
      ethtool -S enp8s0 | grep -E "rx-[[:digit:]]+\.packets"
      
      4. Add filter - change default RX Queue (to 15 here)
      
      ethtool -U ens8s0 flow-type ip4 src-ip 1.1.1.1 action 15 loc 5
      
      5. Ensure filter gets added and traffic is received on RX queue 15 now.
      
      Link event testing:
      -------------------
      6. Bring VF link down and up. If traffic flows to configured queue 15,
      test is success, otherwise it is a failure.
      
      Reset event testing:
      --------------------
      7. Reset the VF. If traffic flows to configured queue 15, test is success,
      otherwise it is a failure.
      
      Fixes: 0dbfbabb
      
       ("iavf: Add framework to enable ethtool ntuple filters")
      Signed-off-by: default avatarPiotr Gardocki <piotrx.gardocki@intel.com>
      Reviewed-by: default avatarLarysa Zaremba <larysa.zaremba@intel.com>
      Signed-off-by: default avatarRanganatha Rao <ranganatha.rao@intel.com>
      Tested-by: default avatarRafal Romanowski <rafal.romanowski@intel.com>
      Signed-off-by: default avatarTony Nguyen <anthony.l.nguyen@intel.com>
      3a0b5a29
    • Linus Torvalds's avatar
      Merge tag 'fuse-fixes-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse · eaadbbaa
      Linus Torvalds authored
      Pull fuse fixes from Miklos Szeredi:
      
       - Fix a couple of potential crashes, one introduced in 6.6 and one
         in 5.10
      
       - Fix misbehavior of virtiofs submounts on memory pressure
      
       - Clarify naming in the uAPI for a recent feature
      
      * tag 'fuse-fixes-6.7-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse:
        fuse: disable FOPEN_PARALLEL_DIRECT_WRITES with FUSE_DIRECT_IO_ALLOW_MMAP
        fuse: dax: set fc->dax to NULL in fuse_dax_conn_free()
        fuse: share lookup state between submount and its parent
        docs/fuse-io: Document the usage of DIRECT_IO_ALLOW_MMAP
        fuse: Rename DIRECT_IO_RELAX to DIRECT_IO_ALLOW_MMAP
      eaadbbaa
    • Linus Torvalds's avatar
      Merge tag '6.7-rc5-ksmbd-server-fixes' of git://git.samba.org/ksmbd · 8b8cd4be
      Linus Torvalds authored
      Pull smb server fixes from Steve French:
      
       - Memory leak fix (in lock error path)
      
       - Two fixes for create with allocation size
      
       - FIx for potential UAF in lease break error path
      
       - Five directory lease (caching) fixes found during additional recent
         testing
      
      * tag '6.7-rc5-ksmbd-server-fixes' of git://git.samba.org/ksmbd:
        ksmbd: fix wrong name of SMB2_CREATE_ALLOCATION_SIZE
        ksmbd: fix wrong allocation size update in smb2_open()
        ksmbd: avoid duplicate opinfo_put() call on error of smb21_lease_break_ack()
        ksmbd: lazy v2 lease break on smb2_write()
        ksmbd: send v2 lease break notification for directory
        ksmbd: downgrade RWH lease caching state to RH for directory
        ksmbd: set v2 lease capability
        ksmbd: set epoch in create context v2 lease
        ksmbd: fix memory leak in smb2_lock()
      8b8cd4be
  4. Dec 12, 2023
    • Ye Bin's avatar
      jbd2: fix soft lockup in journal_finish_inode_data_buffers() · 6c02757c
      Ye Bin authored
      
      
      There's issue when do io test:
      WARN: soft lockup - CPU#45 stuck for 11s! [jbd2/dm-2-8:4170]
      CPU: 45 PID: 4170 Comm: jbd2/dm-2-8 Kdump: loaded Tainted: G  OE
      Call trace:
       dump_backtrace+0x0/0x1a0
       show_stack+0x24/0x30
       dump_stack+0xb0/0x100
       watchdog_timer_fn+0x254/0x3f8
       __hrtimer_run_queues+0x11c/0x380
       hrtimer_interrupt+0xfc/0x2f8
       arch_timer_handler_phys+0x38/0x58
       handle_percpu_devid_irq+0x90/0x248
       generic_handle_irq+0x3c/0x58
       __handle_domain_irq+0x68/0xc0
       gic_handle_irq+0x90/0x320
       el1_irq+0xcc/0x180
       queued_spin_lock_slowpath+0x1d8/0x320
       jbd2_journal_commit_transaction+0x10f4/0x1c78 [jbd2]
       kjournald2+0xec/0x2f0 [jbd2]
       kthread+0x134/0x138
       ret_from_fork+0x10/0x18
      
      Analyzed informations from vmcore as follows:
      (1) There are about 5k+ jbd2_inode in 'commit_transaction->t_inode_list';
      (2) Now is processing the 855th jbd2_inode;
      (3) JBD2 task has TIF_NEED_RESCHED flag;
      (4) There's no pags in address_space around the 855th jbd2_inode;
      (5) There are some process is doing drop caches;
      (6) Mounted with 'nodioread_nolock' option;
      (7) 128 CPUs;
      
      According to informations from vmcore we know 'journal->j_list_lock' spin lock
      competition is fierce. So journal_finish_inode_data_buffers() maybe process
      slowly. Theoretically, there is scheduling point in the filemap_fdatawait_range_keep_errors().
      However, if inode's address_space has no pages which taged with PAGECACHE_TAG_WRITEBACK,
      will not call cond_resched(). So may lead to soft lockup.
      journal_finish_inode_data_buffers
        filemap_fdatawait_range_keep_errors
          __filemap_fdatawait_range
            while (index <= end)
              nr_pages = pagevec_lookup_range_tag(&pvec, mapping, &index, end, PAGECACHE_TAG_WRITEBACK);
              if (!nr_pages)
                 break;    --> If 'nr_pages' is equal zero will break, then will not call cond_resched()
              for (i = 0; i < nr_pages; i++)
                wait_on_page_writeback(page);
              cond_resched();
      
      To solve above issue, add scheduling point in the journal_finish_inode_data_buffers();
      
      Signed-off-by: default avatarYe Bin <yebin10@huawei.com>
      Reviewed-by: default avatarJan Kara <jack@suse.cz>
      Link: https://lore.kernel.org/r/20231211112544.3879780-1-yebin10@huawei.com
      
      
      Signed-off-by: default avatarTheodore Ts'o <tytso@mit.edu>
      6c02757c