1. Jan 07, 2022
  2. Jan 06, 2022
    • Daniel Borkmann's avatar
      veth: Do not record rx queue hint in veth_xmit · 710ad98c
      Daniel Borkmann authored
      Laurent reported that they have seen a significant amount of TCP retransmissions
      at high throughput from applications residing in network namespaces talking to
      the outside world via veths. The drops were seen on the qdisc layer (fq_codel,
      as per systemd default) of the phys device such as ena or virtio_net due to all
      traffic hitting a _single_ TX queue _despite_ multi-queue device. (Note that the
      setup was _not_ using XDP on veths as the issue is generic.)
      
      More specifically, after edbea922 ("veth: Store queue_mapping independently
      of XDP prog presence") which made it all the way back to v4.19.184+,
      skb_record_rx_queue() would set skb->queue_mapping to 1 (given 1 RX and 1 TX
      queue by default for veths) instead of leaving at 0.
      
      This is eventually retained and callbacks like ena_select_queue() will also pick
      single queue via netdev_core_pick_tx()'s ndo_select_queue() once all the traffic
      is forwarded to that device via upper stack or other means. Similarly, for others
      not implementing ndo_select_queue() if XPS is disabled, netdev_pick_tx() might
      call into the skb_tx_hash() and check for prior skb_rx_queue_recorded() as well.
      
      In general, it is a _bad_ idea for virtual devices like veth to mess around with
      queue selection [by default]. Given dev->real_num_tx_queues is by default 1,
      the skb->queue_mapping was left untouched, and so prior to edbea922 the
      netdev_core_pick_tx() could do its job upon __dev_queue_xmit() on the phys device.
      
      Unbreak this and restore prior behavior by removing the skb_record_rx_queue()
      from veth_xmit() altogether.
      
      If the veth peer has an XDP program attached, then it would return the first RX
      queue index in xdp_md->rx_queue_index (unless configured in non-default manner).
      However, this is still better than breaking the generic case.
      
      Fixes: edbea922 ("veth: Store queue_mapping independently of XDP prog presence")
      Fixes: 638264dc
      
       ("veth: Support per queue XDP ring")
      Reported-by: default avatarLaurent Bernaille <laurent.bernaille@datadoghq.com>
      Signed-off-by: default avatarDaniel Borkmann <daniel@iogearbox.net>
      Cc: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
      Cc: Toshiaki Makita <toshiaki.makita1@gmail.com>
      Cc: Eric Dumazet <eric.dumazet@gmail.com>
      Cc: Paolo Abeni <pabeni@redhat.com>
      Cc: John Fastabend <john.fastabend@gmail.com>
      Cc: Willem de Bruijn <willemb@google.com>
      Acked-by: default avatarJohn Fastabend <john.fastabend@gmail.com>
      Reviewed-by: default avatarEric Dumazet <edumazet@google.com>
      Acked-by: default avatarToshiaki Makita <toshiaki.makita1@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      710ad98c
    • Greg Kroah-Hartman's avatar
      ethernet: ibmveth: use default_groups in kobj_type · c288bc0d
      Greg Kroah-Hartman authored
      There are currently 2 ways to create a set of sysfs files for a
      kobj_type, through the default_attrs field, and the default_groups
      field.  Move the ibmveth sysfs code to use default_groups
      field which has been the preferred way since aa30f47c
      
       ("kobject: Add
      support for default attribute groups to kobj_type") so that we can soon
      get rid of the obsolete default_attrs field.
      
      Cc: Michael Ellerman <mpe@ellerman.id.au>
      Cc: Benjamin Herrenschmidt <benh@kernel.crashing.org>
      Cc: Paul Mackerras <paulus@samba.org>
      Cc: Cristobal Forno <cforno12@linux.ibm.com>
      Cc: "David S. Miller" <davem@davemloft.net>
      Cc: Jakub Kicinski <kuba@kernel.org>
      Cc: linuxppc-dev@lists.ozlabs.org
      Cc: netdev@vger.kernel.org
      Signed-off-by: default avatarGreg Kroah-Hartman <gregkh@linuxfoundation.org>
      Reviewed-by: default avatarTyrel Datwyler <tyreld@linux.ibm.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      c288bc0d
    • Jiapeng Chong's avatar
      sfc: Use swap() instead of open coding it · 0cf765fb
      Jiapeng Chong authored
      
      
      Clean the following coccicheck warning:
      
      ./drivers/net/ethernet/sfc/efx_channels.c:870:36-37: WARNING opportunity
      for swap().
      
      ./drivers/net/ethernet/sfc/efx_channels.c:824:36-37: WARNING opportunity
      for swap().
      
      Reported-by: default avatarAbaci Robot <abaci@linux.alibaba.com>
      Signed-off-by: default avatarJiapeng Chong <jiapeng.chong@linux.alibaba.com>
      Acked-by: default avatarMartin Habets <habetsm.xilinx@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      0cf765fb
    • Tom Rix's avatar
      ethtool: use phydev variable · ccd21ec5
      Tom Rix authored
      
      
      In ethtool_get_phy_stats(), the phydev varaible is set to
      dev->phydev but dev->phydev is still used.  Replace
      dev->phydev uses with phydev.
      
      Signed-off-by: default avatarTom Rix <trix@redhat.com>
      Reviewed-by: default avatarAndrew Lunn <andrew@lunn.ch>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      ccd21ec5
    • Russell King (Oracle)'s avatar
      net: macb: use .mac_select_pcs() interface · 8876769b
      Russell King (Oracle) authored
      
      
      Convert the PCS selection to use mac_select_pcs, which allows the PCS
      to perform any validation it needs.
      
      We must use separate phylink_pcs instances for the USX and SGMII PCS,
      rather than just changing the "ops" pointer before re-setting it to
      phylink as this interface queries the PCS, rather than requesting it
      to be changed.
      
      Acked-by: default avatarNicolas Ferre <nicolas.ferre@microchip.com>
      Signed-off-by: default avatarRussell King (Oracle) <rmk+kernel@armlinux.org.uk>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      8876769b
    • Coco Li's avatar
      gro: add ability to control gro max packet size · eac1b93c
      Coco Li authored
      
      
      Eric Dumazet suggested to allow users to modify max GRO packet size.
      
      We have seen GRO being disabled by users of appliances (such as
      wifi access points) because of claimed bufferbloat issues,
      or some work arounds in sch_cake, to split GRO/GSO packets.
      
      Instead of disabling GRO completely, one can chose to limit
      the maximum packet size of GRO packets, depending on their
      latency constraints.
      
      This patch adds a per device gro_max_size attribute
      that can be changed with ip link command.
      
      ip link set dev eth0 gro_max_size 16000
      
      Suggested-by: default avatarEric Dumazet <edumazet@google.com>
      Signed-off-by: default avatarCoco Li <lixiaoyan@google.com>
      Signed-off-by: default avatarEric Dumazet <edumazet@google.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      eac1b93c
    • Miroslav Lichvar's avatar
      net: fix SOF_TIMESTAMPING_BIND_PHC to work with multiple sockets · 007747a9
      Miroslav Lichvar authored
      When multiple sockets using the SOF_TIMESTAMPING_BIND_PHC flag received
      a packet with a hardware timestamp (e.g. multiple PTP instances in
      different PTP domains using the UDPv4/v6 multicast or L2 transport),
      the timestamps received on some sockets were corrupted due to repeated
      conversion of the same timestamp (by the same or different vclocks).
      
      Fix ptp_convert_timestamp() to not modify the shared skb timestamp
      and return the converted timestamp as a ktime_t instead. If the
      conversion fails, return 0 to not confuse the application with
      timestamps corresponding to an unexpected PHC.
      
      Fixes: d7c08826
      
       ("net: socket: support hardware timestamp conversion to PHC bound")
      Signed-off-by: default avatarMiroslav Lichvar <mlichvar@redhat.com>
      Cc: Yangbo Lu <yangbo.lu@nxp.com>
      Cc: Richard Cochran <richardcochran@gmail.com>
      Acked-by: default avatarRichard Cochran <richardcochran@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      007747a9
    • Vladimir Oltean's avatar
      net: dsa: warn about dsa_port and dsa_switch bit fields being non atomic · 1b26d364
      Vladimir Oltean authored
      As discussed during review here:
      https://patchwork.kernel.org/project/netdevbpf/patch/20220105132141.2648876-3-vladimir.oltean@nxp.com/
      
      
      
      we should inform developers about pitfalls of concurrent access to the
      boolean properties of dsa_switch and dsa_port, now that they've been
      converted to bit fields. No other measure than a comment needs to be
      taken, since the code paths that update these bit fields are not
      concurrent with each other.
      
      Suggested-by: default avatarFlorian Fainelli <f.fainelli@gmail.com>
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      1b26d364
    • Vladimir Oltean's avatar
      net: dsa: don't enumerate dsa_switch and dsa_port bit fields using commas · 63cfc657
      Vladimir Oltean authored
      This is a cosmetic incremental fixup to commits
      7787ff77 ("net: dsa: merge all bools of struct dsa_switch into a single u32")
      bde82f38 ("net: dsa: merge all bools of struct dsa_port into a single u8")
      
      The desire to make this change was enunciated after posting these
      patches here:
      https://patchwork.kernel.org/project/netdevbpf/cover/20220105132141.2648876-1-vladimir.oltean@nxp.com/
      
      
      
      but due to a slight timing overlap (message posted at 2:28 p.m. UTC,
      merge commit is at 2:46 p.m. UTC), that comment was missed and the
      changes were applied as-is.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      63cfc657
    • David S. Miller's avatar
      Merge branch 'dsa-init-cleanups' · af8c6db1
      David S. Miller authored
      Vladimir Oltean says:
      
      ====================
      DSA initialization cleanups
      
      These patches contain miscellaneous work that makes the DSA init code
      path symmetric with the teardown path, and some additional patches
      carried by Ansuel Smith for his register access over Ethernet work, but
      those patches can be applied as-is too.
      https://patchwork.kernel.org/project/netdevbpf/patch/20211214224409.5770-3-ansuelsmth@gmail.com/
      
      
      ====================
      
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      af8c6db1
    • Vladimir Oltean's avatar
      net: dsa: setup master before ports · 11fd667d
      Vladimir Oltean authored
      
      
      It is said that as soon as a network interface is registered, all its
      resources should have already been prepared, so that it is available for
      sending and receiving traffic. One of the resources needed by a DSA
      slave interface is the master.
      
      dsa_tree_setup
      -> dsa_tree_setup_ports
         -> dsa_port_setup
            -> dsa_slave_create
               -> register_netdevice
      -> dsa_tree_setup_master
         -> dsa_master_setup
            -> sets up master->dsa_ptr, which enables reception
      
      Therefore, there is a short period of time after register_netdevice()
      during which the master isn't prepared to pass traffic to the DSA layer
      (master->dsa_ptr is checked by eth_type_trans). Same thing during
      unregistration, there is a time frame in which packets might be missed.
      
      Note that this change opens us to another race: dsa_master_find_slave()
      will get invoked potentially earlier than the slave creation, and later
      than the slave deletion. Since dp->slave starts off as a NULL pointer,
      the earlier calls aren't a problem, but the later calls are. To avoid
      use-after-free, we should zeroize dp->slave before calling
      dsa_slave_destroy().
      
      In practice I cannot really test real life improvements brought by this
      change, since in my systems, netdevice creation races with PHY autoneg
      which takes a few seconds to complete, and that masks quite a few races.
      Effects might be noticeable in a setup with fixed links all the way to
      an external system.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      11fd667d
    • Vladimir Oltean's avatar
      net: dsa: first set up shared ports, then non-shared ports · 1e3f407f
      Vladimir Oltean authored
      After commit a57d8c21
      
       ("net: dsa: flush switchdev workqueue before
      tearing down CPU/DSA ports"), the port setup and teardown procedure
      became asymmetric.
      
      The fact of the matter is that user ports need the shared ports to be up
      before they can be used for CPU-initiated termination. And since we
      register net devices for the user ports, those won't be functional until
      we also call the setup for the shared (CPU, DSA) ports. But we may do
      that later, depending on the port numbering scheme of the hardware we
      are dealing with.
      
      It just makes sense that all shared ports are brought up before any user
      port is. I can't pinpoint any issue due to the current behavior, but
      let's change it nonetheless, for consistency's sake.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      1e3f407f
    • Vladimir Oltean's avatar
      net: dsa: hold rtnl_mutex when calling dsa_master_{setup,teardown} · c146f9bc
      Vladimir Oltean authored
      
      
      DSA needs to simulate master tracking events when a binding is first
      with a DSA master established and torn down, in order to give drivers
      the simplifying guarantee that ->master_state_change calls are made
      only when the master's readiness state to pass traffic changes.
      master_state_change() provide a operational bool that DSA driver can use
      to understand if DSA master is operational or not.
      To avoid races, we need to block the reception of
      NETDEV_UP/NETDEV_CHANGE/NETDEV_GOING_DOWN events in the netdev notifier
      chain while we are changing the master's dev->dsa_ptr (this changes what
      netdev_uses_dsa(dev) reports).
      
      The dsa_master_setup() and dsa_master_teardown() functions optionally
      require the rtnl_mutex to be held, if the tagger needs the master to be
      promiscuous, these functions call dev_set_promiscuity(). Move the
      rtnl_lock() from that function and make it top-level.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Reviewed-by: default avatarFlorian Fainelli <f.fainelli@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      c146f9bc
    • Vladimir Oltean's avatar
      net: dsa: stop updating master MTU from master.c · a1ff94c2
      Vladimir Oltean authored
      
      
      At present there are two paths for changing the MTU of the DSA master.
      
      The first is:
      
      dsa_tree_setup
      -> dsa_tree_setup_ports
         -> dsa_port_setup
            -> dsa_slave_create
               -> dsa_slave_change_mtu
                  -> dev_set_mtu(master)
      
      The second is:
      
      dsa_tree_setup
      -> dsa_tree_setup_master
         -> dsa_master_setup
            -> dev_set_mtu(dev)
      
      So the dev_set_mtu() call from dsa_master_setup() has been effectively
      superseded by the dsa_slave_change_mtu(slave_dev, ETH_DATA_LEN) that is
      done from dsa_slave_create() for each user port. The later function also
      updates the master MTU according to the largest user port MTU from the
      tree. Therefore, updating the master MTU through a separate code path
      isn't needed.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Reviewed-by: default avatarFlorian Fainelli <f.fainelli@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      a1ff94c2
    • Vladimir Oltean's avatar
      net: dsa: merge rtnl_lock sections in dsa_slave_create · e31dbd3b
      Vladimir Oltean authored
      
      
      Currently dsa_slave_create() has two sequences of rtnl_lock/rtnl_unlock
      in a row. Remove the rtnl_unlock() and rtnl_lock() in between, such that
      the operation can execute slighly faster.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Reviewed-by: default avatarFlorian Fainelli <f.fainelli@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      e31dbd3b
    • Vladimir Oltean's avatar
      net: dsa: reorder PHY initialization with MTU setup in slave.c · 904e112a
      Vladimir Oltean authored
      
      
      In dsa_slave_create() there are 2 sections that take rtnl_lock():
      MTU change and netdev registration. They are separated by PHY
      initialization.
      
      There isn't any strict ordering requirement except for the fact that
      netdev registration should be last. Therefore, we can perform the MTU
      change a bit later, after the PHY setup. A future change will then be
      able to merge the two rtnl_lock sections into one.
      
      Signed-off-by: default avatarVladimir Oltean <vladimir.oltean@nxp.com>
      Reviewed-by: default avatarFlorian Fainelli <f.fainelli@gmail.com>
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      904e112a
    • David S. Miller's avatar
      Merge branch 'master' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec-next · d093d17c
      David S. Miller authored
      
      
      Steffen Klassert says:
      
      ====================
      pull request (net-next): ipsec-next 2022-01-06
      
      1) Fix some clang_analyzer warnings about never read variables.
         From luo penghao.
      
      2) Check for pols[0] only once in xfrm_expand_policies().
         From Jean Sacren.
      
      3) The SA curlft.use_time was updated only on SA cration time.
         Update whenever the SA is used. From Antony Antony
      
      4) Add support for SM3 secure hash.
         From Xu Jia.
      
      5) Add support for SM4 symmetric cipher algorithm.
         From Xu Jia.
      
      6) Add a rate limit for SA mapping change messages.
         From Antony Antony.
      ====================
      
      Signed-off-by: default avatarDavid S. Miller <davem@davemloft.net>
      d093d17c
    • Toke Høiland-Jørgensen's avatar
      xdp: Add xdp_do_redirect_frame() for pre-computed xdp_frames · 1372d34c
      Toke Høiland-Jørgensen authored
      
      
      Add an xdp_do_redirect_frame() variant which supports pre-computed
      xdp_frame structures. This will be used in bpf_prog_run() to avoid having
      to write to the xdp_frame structure when the XDP program doesn't modify the
      frame boundaries.
      
      Signed-off-by: default avatarToke Høiland-Jørgensen <toke@redhat.com>
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      Link: https://lore.kernel.org/bpf/20220103150812.87914-6-toke@redhat.com
      1372d34c
    • Toke Høiland-Jørgensen's avatar
      xdp: Move conversion to xdp_frame out of map functions · d53ad5d8
      Toke Høiland-Jørgensen authored
      
      
      All map redirect functions except XSK maps convert xdp_buff to xdp_frame
      before enqueueing it. So move this conversion of out the map functions
      and into xdp_do_redirect(). This removes a bit of duplicated code, but more
      importantly it makes it possible to support caller-allocated xdp_frame
      structures, which will be added in a subsequent commit.
      
      Signed-off-by: default avatarToke Høiland-Jørgensen <toke@redhat.com>
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      Link: https://lore.kernel.org/bpf/20220103150812.87914-5-toke@redhat.com
      d53ad5d8
    • Toke Høiland-Jørgensen's avatar
      page_pool: Store the XDP mem id · 64693ec7
      Toke Høiland-Jørgensen authored
      
      
      Store the XDP mem ID inside the page_pool struct so it can be retrieved
      later for use in bpf_prog_run().
      
      Signed-off-by: default avatarToke Høiland-Jørgensen <toke@redhat.com>
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      Acked-by: default avatarJesper Dangaard Brouer <brouer@redhat.com>
      Link: https://lore.kernel.org/bpf/20220103150812.87914-4-toke@redhat.com
      64693ec7
    • Toke Høiland-Jørgensen's avatar
      page_pool: Add callback to init pages when they are allocated · 35b2e549
      Toke Høiland-Jørgensen authored
      
      
      Add a new callback function to page_pool that, if set, will be called every
      time a new page is allocated. This will be used from bpf_test_run() to
      initialise the page data with the data provided by userspace when running
      XDP programs with redirect turned on.
      
      Signed-off-by: default avatarToke Høiland-Jørgensen <toke@redhat.com>
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      Acked-by: default avatarJohn Fastabend <john.fastabend@gmail.com>
      Acked-by: default avatarJesper Dangaard Brouer <brouer@redhat.com>
      Link: https://lore.kernel.org/bpf/20220103150812.87914-3-toke@redhat.com
      35b2e549
    • Toke Høiland-Jørgensen's avatar
      xdp: Allow registering memory model without rxq reference · 4a48ef70
      Toke Høiland-Jørgensen authored
      
      
      The functions that register an XDP memory model take a struct xdp_rxq as
      parameter, but the RXQ is not actually used for anything other than pulling
      out the struct xdp_mem_info that it embeds. So refactor the register
      functions and export variants that just take a pointer to the xdp_mem_info.
      
      This is in preparation for enabling XDP_REDIRECT in bpf_prog_run(), using a
      page_pool instance that is not connected to any network device.
      
      Signed-off-by: default avatarToke Høiland-Jørgensen <toke@redhat.com>
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      Link: https://lore.kernel.org/bpf/20220103150812.87914-2-toke@redhat.com
      4a48ef70
    • Alexei Starovoitov's avatar
      Merge branch 'samples/bpf: xdpsock app enhancements' · 640a171c
      Alexei Starovoitov authored
      Ong Boon says:
      
      ====================
      
      First of all, sorry for taking more time to get back to this series and
      thanks to all valuble feedback in series-1 at [1] from Jesper and Song
      Liu.
      
      Since then I have looked into what Jesper suggested in [2] and worked on
      revising the patch series into several patches for ease of review:
      
      v1->v2:
      1/7: [No change]. Add VLAN tag (ID & Priority) to the generated Tx-Only
           frames.
      
      2/7: [No change]. Add DMAC and SMAC setting to the generated Tx-Only
           frames. If parameters are not set, previous DMAC and SMAC are used.
      
      3/7: [New]. Add support for selecting different CLOCK for clock_gettime()
           used in get_nsecs.
      
      4/7: [New]. This is a total rework from series-1 3/4-patch [3]. It uses
           clock_nanosleep() suggested by Jesper. In addition, added statistic
           for Tx schedule variance under application stat (-a|--app-stats).
           Make the cyclic Tx operation and --poll mode to be mutually-
           exclusive. Still, the ability to specify TX cycle time and used
           together with batch size and packet count remain the same.
      
      5/7: [New]. Add the support for TX process schedule policy and priority
           setting. By default, SCHED_OTHER policy is used. This too is matching
           the schedule policy setting in [2].
      
      6/7: [Change]. This is update from series-1 4/4-patch [4]. Added TX clean
           process time-out in 1s granularity with configurable retries count
           (-O|--retries).
      
      7/7: [New]. Added timestamp for TX packet following pktgen_hdr format
           matching the implementation in [2]. However, the sequence ID remains
           the same as it is instead of process schedule diff in [2].
      
      To summarize on what program options have been added with v2 series
      using an example below:-
      
       DMAC (-G)                 = fa:8d:f1:e2:0b:e8
       SMAC (-H)                 = ce:17:07:17:3e:3a
      
       VLAN tagged (-V)
       VLAN ID (-J)              = 12
       VLAN Pri (-K)             = 3
      
       Tx Queue (-q)             = 3
       Cycle Time in us (-T)     = 1000
       Batch (-b)                = 2
       Packet Count              = 6
       Tx schedule policy (-W)   = FIFO
       Tx schedule priority (-U) = 50
       Clock selection (-w)      = REALTIME
      
       Tx timeout retries(-O)    = 5
       Tx timestamp (-y)
       Cyclic Tx schedule stat (-a)
      
      Note: xdpsock sets UDP dest-port and src-port to 0x1000 as default.
      
       Sending Board
       =============
       $ xdpsock -i eth0 -t -N -z -H ce:17:07:17:3e:3a -G fa:8d:f1:e2:0b:e8 \
         -V -J 12 -K 3 -q 3 \
         -T 1000 -b 2 -C 6 -W FIFO -U 50 -w REALTIME \
         -O 5 -y -a
      
        sock0@eth0:3 txonly xdp-drv
                          pps            pkts           0.00
       rx                 0              0
       tx                 0              6
      
                          calls/s        count
       rx empty polls     0              0
       fill fail polls    0              0
       copy tx sendtos    0              0
       tx wakeup sendtos  0              5
       opt polls          0              0
      
                          period     min        ave        max        cycle
       Cyclic TX          1000000    31033      32009      33397      3
      
       Receiving Board
       ===============
       $ tcpdump -nei eth0 udp port 0x1000 -vv -Q in -X \
          --time-stamp-precision nano
      tcpdump: listening on eth0, link-type EN10MB (Ethernet), capture size 262144 bytes
      03:46:40.520111580 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e997 be9b e955  ...............U
              0x0020:  0000 0000 61cd 2ba1 0006 987c            ....a.+....|
      03:46:40.520112163 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e996 be9b e955  ...............U
              0x0020:  0000 0001 61cd 2ba1 0006 987c            ....a.+....|
      03:46:40.521066860 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e5af be9b e955  ...............U
              0x0020:  0000 0002 61cd 2ba1 0006 9c62            ....a.+....b
      03:46:40.521067012 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e5ae be9b e955  ...............U
              0x0020:  0000 0003 61cd 2ba1 0006 9c62            ....a.+....b
      03:46:40.522061935 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e1c5 be9b e955  ...............U
              0x0020:  0000 0004 61cd 2ba1 0006 a04a            ....a.+....J
      03:46:40.522062173 ce:17:07:17:3e:3a > fa:8d:f1:e2:0b:e8, ethertype 802.1Q (0x8100), length 62: vlan 12, p 3, ethertype IPv4, (tos 0x0, ttl 64, id 0, offset 0, flags [none], proto UDP (17), length 44)
          10.10.10.16.4096 > 10.10.10.32.4096: [udp sum ok] UDP, length 16
              0x0000:  4500 002c 0000 0000 4011 527e 0a0a 0a10  E..,....@.R~....
              0x0010:  0a0a 0a20 1000 1000 0018 e1c4 be9b e955  ...............U
              0x0020:  0000 0005 61cd 2ba1 0006 a04a            ....a.+....J
      
      I have tested the above with both tagged and untagged packet format and
      based on the timestamp in tcpdump found that the timing of the batch
      cyclic transmission is correct.
      
      Appreciate if community can give the patch series v2 a try and point out
      any gap.
      
      Thanks
      Boon Leong
      
      [1] https://patchwork.kernel.org/project/netdevbpf/cover/20211124091821.3916046-1-boon.leong.ong@intel.com/
      [2] https://github.com/netoptimizer/network-testing/blob/master/src/udp_pacer.c
      [3] https://patchwork.kernel.org/project/netdevbpf/patch/20211124091821.3916046-4-boon.leong.ong@intel.com/
      [4] https://patchwork.kernel.org/project/netdevbpf/patch/20211124091821.3916046-5-boon.leong.ong@intel.com/
      
      
      ====================
      
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      640a171c