1. Jul 12, 2023
  2. Jul 11, 2023
    • Andrii Nakryiko's avatar
      libbpf: Fix realloc API handling in zero-sized edge cases · 8a0260db
      Andrii Nakryiko authored
      realloc() and reallocarray() can either return NULL or a special
      non-NULL pointer, if their size argument is zero. This requires a bit
      more care to handle NULL-as-valid-result situation differently from
      NULL-as-error case. This has caused real issues before ([0]), and just
      recently bit again in production when performing bpf_program__attach_usdt().
      
      This patch fixes 4 places that do or potentially could suffer from this
      mishandling of NULL, including the reported USDT-related one.
      
      There are many other places where realloc()/reallocarray() is used and
      NULL is always treated as an error value, but all those have guarantees
      that their size is always non-zero, so those spot don't need any extra
      handling.
      
        [0] d08ab82f ("libbpf: Fix double-free when linker processes empty sections")
      
      Fixes: 999783c8 ("libbpf: Wire up spec management and other arch-independent USDT logic")
      Fixes: b63b3c49 ("libbpf: Add bpf_program__set_insns function")
      Fixes: 697f104d ("libbpf: Support custom SEC() handlers")
      Fixes: b1268826
      
       ("libbpf: Change the order of data and text relocations.")
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Signed-off-by: default avatarDaniel Borkmann <daniel@iogearbox.net>
      Link: https://lore.kernel.org/bpf/20230711024150.1566433-1-andrii@kernel.org
      8a0260db
    • David Vernet's avatar
      bpf,docs: Create new standardization subdirectory · 4d496be9
      David Vernet authored
      The BPF standardization effort is actively underway with the IETF. As
      described in the BPF Working Group (WG) charter in [0], there are a
      number of proposed documents, some informational and some proposed
      standards, that will be drafted as part of the standardization effort.
      
      [0]: https://datatracker.ietf.org/wg/bpf/about/
      
      
      
      Though the specific documents that will formally be standardized will
      exist as Internet Drafts (I-D) and WG documents in the BPF WG
      datatracker page, the source of truth from where those documents will be
      generated will reside in the kernel documentation tree (originating in
      the bpf-next tree).
      
      Because these documents will be used to generate the I-D and WG
      documents which will be standardized with the IETF, they are a bit
      special as far as kernel-tree documentation goes:
      
      - They will be dual licensed with LGPL-2.1 OR BSD-2-Clause
      - IETF I-D and WG documents (the documents which will actually be
        standardized) will be auto-generated from these documents.
      
      In order to keep things clearly organized in the BPF documentation tree,
      and to make it abundantly clear where standards-related documentation
      needs to go, we should move standards-relevant documents into a separate
      standardization/ subdirectory.
      
      Signed-off-by: default avatarDavid Vernet <void@manifault.com>
      Link: https://lore.kernel.org/r/20230710183027.15132-1-void@manifault.com
      
      
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      4d496be9
    • Andrii Nakryiko's avatar
      Merge branch 'bpftool: Fix skeletons compilation for older kernels' · 19f4b532
      Andrii Nakryiko authored
      Quentin Monnet says:
      
      ====================
      At runtime, bpftool may run its own BPF programs to get the pids of
      processes referencing BPF programs, or to profile programs. The skeletons
      for these programs rely on a vmlinux.h header and may fail to compile when
      building bpftool on hosts running older kernels, where some structs or
      enums are not defined. In this set, we address this issue by using local
      definitions for struct perf_event, struct bpf_perf_link,
      BPF_LINK_TYPE_PERF_EVENT (pids.bpf.c) and struct bpf_perf_event_value
      (profiler.bpf.c).
      
      This set contains patches 1 to 3 from Alexander Lobakin's series, "bpf:
      random unpopular userspace fixes (32 bit et al)" (v2) [0], from April 2022.
      An additional patch defines a local version of BPF_LINK_TYPE_PERF_EVENT in
      bpftool's pids.bpf.c.
      
      [0] https://lore.kernel.org/bpf/20220421003152.339542-1-alobakin@pm.me/
      
      
      
      v2: Fixed description (CO-RE for container_of()) in patch 2.
      
      Cc: Alexander Lobakin <aleksander.lobakin@intel.com>
      Cc: Michal Suchánek <msuchanek@suse.de>
      
      Alexander Lobakin (3):
        bpftool: use a local copy of perf_event to fix accessing ::bpf_cookie
        bpftool: define a local bpf_perf_link to fix accessing its fields
        bpftool: use a local bpf_perf_event_value to fix accessing its fields
      ====================
      
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      19f4b532
    • Alexander Lobakin's avatar
      bpftool: Use a local bpf_perf_event_value to fix accessing its fields · 658ac068
      Alexander Lobakin authored
      Fix the following error when building bpftool:
      
        CLANG   profiler.bpf.o
        CLANG   pid_iter.bpf.o
      skeleton/profiler.bpf.c:18:21: error: invalid application of 'sizeof' to an incomplete type 'struct bpf_perf_event_value'
              __uint(value_size, sizeof(struct bpf_perf_event_value));
                                 ^     ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
      tools/bpf/bpftool/bootstrap/libbpf/include/bpf/bpf_helpers.h:13:39: note: expanded from macro '__uint'
      tools/bpf/bpftool/bootstrap/libbpf/include/bpf/bpf_helper_defs.h:7:8: note: forward declaration of 'struct bpf_perf_event_value'
      struct bpf_perf_event_value;
             ^
      
      struct bpf_perf_event_value is being used in the kernel only when
      CONFIG_BPF_EVENTS is enabled, so it misses a BTF entry then.
      Define struct bpf_perf_event_value___local with the
      `preserve_access_index` attribute inside the pid_iter BPF prog to
      allow compiling on any configs. It is a full mirror of a UAPI
      structure, so is compatible both with and w/o CO-RE.
      bpf_perf_event_read_value() requires a pointer of the original type,
      so a cast is needed.
      
      Fixes: 47c09d6a
      
       ("bpftool: Introduce "prog profile" command")
      Suggested-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Signed-off-by: default avatarAlexander Lobakin <alobakin@pm.me>
      Signed-off-by: default avatarQuentin Monnet <quentin@isovalent.com>
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Link: https://lore.kernel.org/bpf/20230707095425.168126-5-quentin@isovalent.com
      658ac068
    • Quentin Monnet's avatar
      bpftool: Use a local copy of BPF_LINK_TYPE_PERF_EVENT in pid_iter.bpf.c · 44ba7b30
      Quentin Monnet authored
      In order to allow the BPF program in bpftool's pid_iter.bpf.c to compile
      correctly on hosts where vmlinux.h does not define
      BPF_LINK_TYPE_PERF_EVENT (running kernel versions lower than 5.15, for
      example), define and use a local copy of the enum value. This requires
      LLVM 12 or newer to build the BPF program.
      
      Fixes: cbdaf71f
      
       ("bpftool: Add bpf_cookie to link output")
      Signed-off-by: default avatarQuentin Monnet <quentin@isovalent.com>
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Link: https://lore.kernel.org/bpf/20230707095425.168126-4-quentin@isovalent.com
      44ba7b30
    • Alexander Lobakin's avatar
      bpftool: Define a local bpf_perf_link to fix accessing its fields · 67a43462
      Alexander Lobakin authored
      When building bpftool with !CONFIG_PERF_EVENTS:
      
      skeleton/pid_iter.bpf.c:47:14: error: incomplete definition of type 'struct bpf_perf_link'
              perf_link = container_of(link, struct bpf_perf_link, link);
                          ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
      tools/bpf/bpftool/bootstrap/libbpf/include/bpf/bpf_helpers.h:74:22: note: expanded from macro 'container_of'
                      ((type *)(__mptr - offsetof(type, member)));    \
                                         ^~~~~~~~~~~~~~~~~~~~~~
      tools/bpf/bpftool/bootstrap/libbpf/include/bpf/bpf_helpers.h:68:60: note: expanded from macro 'offsetof'
       #define offsetof(TYPE, MEMBER)  ((unsigned long)&((TYPE *)0)->MEMBER)
                                                        ~~~~~~~~~~~^
      skeleton/pid_iter.bpf.c:44:9: note: forward declaration of 'struct bpf_perf_link'
              struct bpf_perf_link *perf_link;
                     ^
      
      &bpf_perf_link is being defined and used only under the ifdef.
      Define struct bpf_perf_link___local with the `preserve_access_index`
      attribute inside the pid_iter BPF prog to allow compiling on any
      configs. CO-RE will substitute it with the real struct bpf_perf_link
      accesses later on.
      container_of() uses offsetof(), which does the necessary CO-RE
      relocation if the field is specified with `preserve_access_index` - as
      is the case for struct bpf_perf_link___local.
      
      Fixes: cbdaf71f
      
       ("bpftool: Add bpf_cookie to link output")
      Suggested-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Signed-off-by: default avatarAlexander Lobakin <alobakin@pm.me>
      Signed-off-by: default avatarQuentin Monnet <quentin@isovalent.com>
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Link: https://lore.kernel.org/bpf/20230707095425.168126-3-quentin@isovalent.com
      67a43462
    • Alexander Lobakin's avatar
      bpftool: use a local copy of perf_event to fix accessing :: Bpf_cookie · 4cbeeb0d
      Alexander Lobakin authored
      When CONFIG_PERF_EVENTS is not set, struct perf_event remains empty.
      However, the structure is being used by bpftool indirectly via BTF.
      This leads to:
      
      skeleton/pid_iter.bpf.c:49:30: error: no member named 'bpf_cookie' in 'struct perf_event'
              return BPF_CORE_READ(event, bpf_cookie);
                     ~~~~~~~~~~~~~~~~~~~~~^~~~~~~~~~~
      
      ...
      
      skeleton/pid_iter.bpf.c:49:9: error: returning 'void' from a function with incompatible result type '__u64' (aka 'unsigned long long')
              return BPF_CORE_READ(event, bpf_cookie);
                     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
      
      Tools and samples can't use any CONFIG_ definitions, so the fields
      used there should always be present.
      Define struct perf_event___local with the `preserve_access_index`
      attribute inside the pid_iter BPF prog to allow compiling on any
      configs. CO-RE will substitute it with the real struct perf_event
      accesses later on.
      
      Fixes: cbdaf71f
      
       ("bpftool: Add bpf_cookie to link output")
      Suggested-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Signed-off-by: default avatarAlexander Lobakin <alobakin@pm.me>
      Signed-off-by: default avatarQuentin Monnet <quentin@isovalent.com>
      Signed-off-by: default avatarAndrii Nakryiko <andrii@kernel.org>
      Link: https://lore.kernel.org/bpf/20230707095425.168126-2-quentin@isovalent.com
      4cbeeb0d
  3. Jul 09, 2023
  4. Jul 08, 2023
  5. Jul 07, 2023
  6. Jul 06, 2023
    • Hou Tao's avatar
      selftests/bpf: Add benchmark for bpf memory allocator · fd283ab1
      Hou Tao authored
      
      
      The benchmark could be used to compare the performance of hash map
      operations and the memory usage between different flavors of bpf memory
      allocator (e.g., no bpf ma vs bpf ma vs reuse-after-gp bpf ma). It also
      could be used to check the performance improvement or the memory saving
      provided by optimization.
      
      The benchmark creates a non-preallocated hash map which uses bpf memory
      allocator and shows the operation performance and the memory usage of
      the hash map under different use cases:
      (1) overwrite
      Each CPU overwrites nonoverlapping part of hash map. When each CPU
      completes overwriting of 64 elements in hash map, it increases the
      op_count.
      (2) batch_add_batch_del
      Each CPU adds then deletes nonoverlapping part of hash map in batch.
      When each CPU adds and deletes 64 elements in hash map, it increases
      the op_count twice.
      (3) add_del_on_diff_cpu
      Each two-CPUs pair adds and deletes nonoverlapping part of map
      cooperatively. When each CPU adds or deletes 64 elements in hash map,
      it will increase the op_count.
      
      The following is the benchmark results when comparing between different
      flavors of bpf memory allocator. These tests are conducted on a KVM guest
      with 8 CPUs and 16 GB memory. The command line below is used to do all
      the following benchmarks:
      
        ./bench htab-mem --use-case $name ${OPTS} -w3 -d10 -a -p8
      
      These results show that preallocated hash map has both better performance
      and smaller memory footprint.
      
      (1) non-preallocated + no bpf memory allocator (v6.0.19)
      use kmalloc() + call_rcu
      
      overwrite            per-prod-op: 11.24 ± 0.07k/s, avg mem: 82.64 ± 26.32MiB, peak mem: 119.18MiB
      batch_add_batch_del  per-prod-op: 18.45 ± 0.10k/s, avg mem: 50.47 ± 14.51MiB, peak mem: 94.96MiB
      add_del_on_diff_cpu  per-prod-op: 14.50 ± 0.03k/s, avg mem: 4.64 ± 0.73MiB, peak mem: 7.20MiB
      
      (2) preallocated
      OPTS=--preallocated
      
      overwrite            per-prod-op: 191.42 ± 0.09k/s, avg mem: 1.24 ± 0.00MiB, peak mem: 1.49MiB
      batch_add_batch_del  per-prod-op: 221.83 ± 0.17k/s, avg mem: 1.23 ± 0.00MiB, peak mem: 1.49MiB
      add_del_on_diff_cpu  per-prod-op: 39.66 ± 0.31k/s, avg mem: 1.47 ± 0.13MiB, peak mem: 1.75MiB
      
      (3) normal bpf memory allocator
      
      overwrite            per-prod-op: 126.59 ± 0.02k/s, avg mem: 2.26 ± 0.00MiB, peak mem: 2.74MiB
      batch_add_batch_del  per-prod-op: 83.37 ± 0.20k/s, avg mem: 2.14 ± 0.17MiB, peak mem: 2.74MiB
      add_del_on_diff_cpu  per-prod-op: 21.25 ± 0.24k/s, avg mem: 17.50 ± 3.32MiB, peak mem: 28.87MiB
      
      Acked-by: default avatarJohn Fastabend <john.fastabend@gmail.com>
      Signed-off-by: default avatarHou Tao <houtao1@huawei.com>
      Link: https://lore.kernel.org/r/20230704025039.938914-1-houtao@huaweicloud.com
      
      
      Signed-off-by: default avatarAlexei Starovoitov <ast@kernel.org>
      fd283ab1
  7. Jul 05, 2023
  8. Jul 01, 2023
  9. Jun 30, 2023
    • Kui-Feng Lee's avatar
      selftests/bpf: Verify that the cgroup_skb filters receive expected packets. · 539c7e67
      Kui-Feng Lee authored
      
      
      This test case includes four scenarios:
      
      1. Connect to the server from outside the cgroup and close the connection
         from outside the cgroup.
      2. Connect to the server from outside the cgroup and close the connection
         from inside the cgroup.
      3. Connect to the server from inside the cgroup and close the connection
         from outside the cgroup.
      4. Connect to the server from inside the cgroup and close the connection
         from inside the cgroup.
      
      The test case is to verify that cgroup_skb/{egress, ingress} filters
      receive expected packets including SYN, SYN/ACK, ACK, FIN, and FIN/ACK.
      
      Signed-off-by: default avatarKui-Feng Lee <kuifeng@meta.com>
      Signed-off-by: default avatarDaniel Borkmann <daniel@iogearbox.net>
      Link: https://lore.kernel.org/bpf/20230624014600.576756-3-kuifeng@meta.com
      539c7e67
    • Kui-Feng Lee's avatar
      bpf, net: Check skb ownership against full socket. · 223f5f79
      Kui-Feng Lee authored
      Check skb ownership of an skb against full sockets instead of request_sock.
      
      The filters were called only if an skb is owned by the sock that the skb is
      sent out through. In another words, skb->sk should point to the sock that
      it is sending through its egress. However, the filters would miss SYN/ACK
      skbs that they are owned by a request_sock but sent through the listener
      sock, that is the socket listening incoming connections.
      
      However, the listener socket is also the full socket of the request socket.
      We should use the full socket as the owner socket of an skb instead.
      
      What is the ownership check for?
      ================================
      
      BPF_CGROUP_RUN_PROG_INET_EGRESS() checked sk == skb->sk to ensure the
      ownership of an skb. Alexei referred to a mailing list conversation [0]
      that took place a few years ago. In that conversation, Daniel Borkmann
      stated that:
      
          Wouldn't that mean however, when you go through stacked devices that
          you'd run the same eBPF cgroup program for skb->sk multiple times?
      
      According to what Daniel said, the ownership check mentioned earlier
      presumably prevents multiple calls of egress filters caused by an skb.
      
      A test that reproduce this scenario shows that the BPF cgroup egress
      programs can be called multiple times for one skb if this ownership
      check is not there. So, we can not just remove this check.
      
      Test Stacked Devices
      ====================
      
      We use L2TP to build an environment of stacked devices. L2TP (Layer 2
      Tunneling Protocol) is a tunneling protocol used to support virtual private
      networks (VPNs). It relays encapsulated packets; for example in UDP, to its
      peer by using a socket.
      
      Using L2TP, packets are first sent through the IP stack and should then
      arrive at an L2TP device. The device will expand its skb header to
      encapsulate the packet. The skb will be sent back to the IP stack using
      the socket that was made for the L2TP session. After that, the routing
      process will occur once more, but this time for a new destination.
      
      We changed tools/testing/selftests/net/l2tp.sh to set up a test environment
      using L2TP. The run_ping() function in l2tp.sh is where the main change
      occurred.
      
          run_ping()
          {
              local desc="$1"
      
              sleep 10
              run_cmd host-1 ${ping6} -s 227 -c 4 -i 10 -I fc00:101::1
              fc00:101::2
              log_test $? 0 "IPv6 route through L2TP tunnel ${desc}"
              sleep 10
          }
      
      The test will use L2TP devices to send PING messages. These messages will
      have a message size of 227 bytes as a special label to distinguish them.
      This is not an ideal solution, but works.
      
      During the execution of the test script, bpftrace was attached to
      ip6_finish_output() and l2tp_xmit_skb():
      
          bpftrace -e '
            kfunc:ip6_finish_output {
              time("%H:%M:%S: ");
              printf("ip6_finish_output skb=%p skb->len=%d cgroup=%p sk=%p
                      skb->sk=%p\n", args->skb, args->skb->len,
                     args->sk->sk_cgrp_data.cgroup, args->sk, args->skb->sk); }
            kfunc:l2tp_xmit_skb {
              time("%H:%M:%S: ");
              printf("l2tp_xmit_skb skb=%p sk=%p\n", args->skb,
      	       args->session->tunnel->sock); }'
      
      The following is part of the output messages printed by bpftrace:
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e600 skb->len=275
                    cgroup=0xffff88810741f800 sk=0xffff888105f3b900
                    skb->sk=0xffff888105f3b900
      
          16:35:20: l2tp_xmit_skb skb=0xffff888103d8e600 sk=0xffff888103dd6300
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e600 skb->len=337
                    cgroup=0xffff88810741f800 sk=0xffff888103dd6300
                    skb->sk=0xffff888105f3b900
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e600 skb->len=337
                    cgroup=(nil) sk=(nil) skb->sk=(nil)
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e000 skb->len=275
                    cgroup=0xffffffff837741d0 sk=0xffff888101fe0000
                    skb->sk=0xffff888101fe0000
      
          16:35:20: l2tp_xmit_skb skb=0xffff888103d8e000 sk=0xffff888103483180
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e000 skb->len=337
                    cgroup=0xffff88810741f800 sk=0xffff888103483180
                    skb->sk=0xffff888101fe0000
      
          16:35:20: ip6_finish_output skb=0xffff888103d8e000 skb->len=337
                    cgroup=(nil) sk=(nil) skb->sk=(nil)
      
      The first four entries describe a PING message that was sent using the ping
      command, whereas the following four entries describe the response received.
      Multiple sockets are used to send one skb, including the socket used by the
      L2TP session. This can be observed.
      
      Based on this information, it seems that the ownership check is designed to
      avoid multiple calls of egress filters caused by a single skb.
      
        [0] https://lore.kernel.org/all/58193E9D.7040201@iogearbox.net/
      
      
      
      Signed-off-by: default avatarKui-Feng Lee <kuifeng@meta.com>
      Signed-off-by: default avatarDaniel Borkmann <daniel@iogearbox.net>
      Link: https://lore.kernel.org/bpf/20230624014600.576756-2-kuifeng@meta.com
      223f5f79
    • Stanislav Fomichev's avatar
      selftests/bpf: Add test to exercise typedef walking · 2597a25c
      Stanislav Fomichev authored
      
      
      Add new bpf_fentry_test_sinfo with skb_shared_info argument and try to
      access frags.
      
      Signed-off-by: default avatarStanislav Fomichev <sdf@google.com>
      Signed-off-by: default avatarDaniel Borkmann <daniel@iogearbox.net>
      Acked-by: default avatarYonghong Song <yhs@fb.com>
      Link: https://lore.kernel.org/bpf/20230626212522.2414485-2-sdf@google.com
      2597a25c