1. Apr 14, 2023
    • Nico Weber's avatar
      [gn] Port fcc5f9ef (GenVT) · c2e27be8
      Nico Weber authored
      c2e27be8
    • Nikita Popov's avatar
      [FunctionAttrs] Fix nounwind inference for landingpads · 9fe78db4
      Nikita Popov authored
      Currently, FunctionAttrs treats landingpads as non-throwing, and
      will infer nounwind for functions with landingpads (assuming they
      can't unwind in some other way, e.g. via resum). There are two
      problems with this:
      
      * Non-cleanup landingpads with catch/filter clauses do not
        necessarily catch all exceptions. Unless there are catch ptr null
        or filter [0 x ptr] zeroinitializer clauses, we should assume
        that we may unwind past this landingpad. This seems like an
        outright bug.
      * Cleanup landingpads are skipped during phase one unwinding, so
        we effectively need to support unwinding past them. Marking these
        nounwind is technically correct, but not compatible with how
        unwinding works in reality.
      
      Fixes https://github.com/llvm/llvm-project/issues/61945.
      
      Differential Revision: https://reviews.llvm.org/D147694
      9fe78db4
    • Florian Hahn's avatar
      [Matrix] Add dot product tests with builtin loads with variable strides · 677b0d33
      Florian Hahn authored
      Extra tests for D147330.
      677b0d33
    • Kai Luo's avatar
      [PowerPC] Update `incr` after resetting the register in MI · eee024bf
      Kai Luo authored
      After performing signed extension, we update the register in MI. We should also update `incr` register which is tracking the register in `MI`.
      
      Fixes https://github.com/llvm/llvm-project/issues/61882.
      
      Reviewed By: #powerpc, shchenz
      
      Differential Revision: https://reviews.llvm.org/D147594
      eee024bf
    • Nikita Popov's avatar
      [InstCombine] Support multiple comparisons in foldAllocaCmp() · a7597451
      Nikita Popov authored
      foldAllocaCmp() needs to fold all comparisons of an alloca at the
      same time, to ensure that there is a consistent view of the alloca
      address. Currently, it folds "all" comparisons by limiting to the
      case where there is only one. This patch switches the algorithm to
      instead actually collect and fold all comparisons.
      
      Something we need to be careful about here is that there may be
      comparisons where both sides of the icmp are based on the alloca.
      Such comparisons are comparing offsets of the alloca, and as such
      can be ignored here, but shouldn't be folded to false.
      
      Differential Revision: https://reviews.llvm.org/D144492
      a7597451
    • Nikita Popov's avatar
    • Christian Sigg's avatar
      Fix bazel build after e5f50bd2 · d45888ec
      Christian Sigg authored
      d45888ec
    • Peter Smith's avatar
      [LLD][ARM] Handle .ARM.exidx sections at non-zero output sec offset · 7a2000ac
      Peter Smith authored
      Embedded systems that do not use an ELF loader locate the
      .ARM.exidx exception table via linker defined __exidx_start and
      __exidx_end rather than use the PT_ARM_EXIDX program header. This
      means that some linker scripts such as the picolibc C library's
      linker script, do not have the .ARM.exidx sections at offset 0 in
      the OutputSection. For example:
      
      .except_unordered : {
          . = ALIGN(8);
          PROVIDE(__exidx_start = .);
          *(.ARM.exidx*)
          PROVIDE(__exidx_end = .);
      } >flash AT>flash :text
      
      This is within the specification of Arm exception tables, and is
      handled correctly by ld.bfd.
      
      This patch has 2 parts. The first updates the writing of the data
      of the .ARM.exidx SyntheticSection to account for a non-zero
      OutputSection offset. The second part makes the PT_ARM_EXIDX program
      header generation a special case so that it covers only the
      SyntheticSection and not the parent OutputSection. While not strictly
      necessary for programs locating the exception tables via the symbols
      it may cause ELF utilities that locate the exception tables via
      the PT_ARM_EXIDX program header to fail. This does not seem to be the
      case for GNU and LLVM readelf which seems to look for the
      SHT_ARM_EXIDX section.
      
      Differential Revision: https://reviews.llvm.org/D148033
      7a2000ac
    • Kristof Beyls's avatar
      6b19efe2
    • David Stuttard's avatar
      [AMDGPU] Add backend support for new PAL ELF Metadata 3.0 · fc83f1de
      David Stuttard authored
      PAL Metadata 3.0 introduces an explicit structure in metadata for the
      programmable registers written out by the compiler backend.
      Rather than using opaque registers which can change between different
      architectures and requires encoding the bitfield information in the backend,
      which may change between versions.
      
      This is the initial minimal implementation that enables the use of PAL Metadata
      3.0.
      
      The change itself should be NFC for non-PAL, although the way RSRC2 register is
      handled has been changed slightly.
      
      The test is fairly minimal, but checks that the metadata format looks as
      expected and verifies a couple of special cases such as tgid_[xyz]_en handling
      and PsInputAddr/Ena which also change to explicit fields.
      
      Differential Revision: https://reviews.llvm.org/D147143
      fc83f1de
    • Nikita Popov's avatar
      [LangRef][Local] dereferenceable metadata violation is UB · e4251fc6
      Nikita Popov authored
      I believe !dereferencable violation is immediate undefined behavior,
      but this was not explicitly spelled out in LangRef. We already
      assume that !dereferenceable is implicitly !noundef and cannot
      return poison in isGuaranteedNotToBeUndefOrPoison().
      
      The reason why we made dereferenceable implicitly noundef is that
      the purpose of this metadata is to allow speculation, and that
      would not be legal on a potential poison pointer.
      
      Differential Revision: https://reviews.llvm.org/D148202
      e4251fc6
    • Nikita Popov's avatar
      c508e933
    • Diana Picus's avatar
      [AMDGPU] Don't S_MOV_B32 into $scc · b9ba0536
      Diana Picus authored
      The peephole optimizer tries to replace
      ```
      %n:sgpr_32 = S_MOV_B32 x
      $scc = COPY %n
      ```
      with a `S_MOV_B32` directly into `$scc`.
      
      This crashes because `S_MOV_B32` cannot take `$scc` as input.
      
      We currently generate code like this from GlobalISel when lowering a
      G_BRCOND with a constant condition. We should probably look into
      removing this kind of branch altogether, but until then we should at
      least not crash.
      
      This patch fixes the issue by making sure we don't apply the peephole
      optimization when trying to move into a physical register that
      doesn't belong to the correct register class.
      
      Differential Revision: https://reviews.llvm.org/D148117
      b9ba0536
    • Nikita Popov's avatar
      [Coroutines] Directly remove unnecessary lifetime intrinsics · 243e62b9
      Nikita Popov authored
      The insertSpills() code will currently skip lifetime intrinsic users
      when replacing the alloca with a frame reference. Rather than
      leaving behind the dead lifetime intrinsics working on the old
      alloca, directly remove them. This makes sure the alloca can be
      dropped as well.
      
      I noticed this as a regression when converting tests to opaque
      pointers. Without opaque pointers, this code didn't really do
      anything, because there would usually be a bitcast in between.
      The lifetimes would get rewritten to the frame pointer. With
      opaque pointers, this code now triggers and leaves behind users
      of the old allocas.
      
      Differential Revision: https://reviews.llvm.org/D148240
      243e62b9
    • Krasimir Georgiev's avatar
      [sanitizer] adapt for 75f1f158 · de4c038c
      Krasimir Georgiev authored
       No functional changes intended.
      de4c038c
    • Tobias Gysi's avatar
      [mlir][llvm] Move the LLVM dialect definition (NFC). · e5f50bd2
      Tobias Gysi authored
      The revision separates out the LLVM dialect definition in a separate
      tablegen file and ensures the LLVMOpBase.td can include the attributes
      defined by LLVMAttrDefs.td. The change allows us to use LLVM dialect
      attributes in the definition of the intrinsic and memory operation
      base classes, e.g. to represent alias analysis metadata using
      attributes.
      
      Reviewed By: Dinistro
      
      Differential Revision: https://reviews.llvm.org/D148007
      e5f50bd2
    • Tobias Hieta's avatar
      Revert "[clang-format] Handle object instansiation in if-statements" · 104cd749
      Tobias Hieta authored
      This reverts commit 70de684d.
      
      This causes a regression as described in #61785
      104cd749
    • Jie Fu's avatar
      [llvm-exegesis] Fix -Wc++98-compat-extra-semi in BenchmarkRunner.cpp (NFC) · 62a0049a
      Jie Fu authored
      /data/llvm-project/llvm/tools/llvm-exegesis/lib/BenchmarkRunner.cpp:66:2: error: extra ';' outside of a function is incompatible with C++98 [-Werror,-W
      c++98-compat-extra-semi]
      };
       ^
      1 error generated.
      62a0049a
    • Aiden Grossman's avatar
      [llvm-exegesis] Refactor common parts out of FunctionExecutorImpl · d2280594
      Aiden Grossman authored
      This patch refactors some code out of FunctionExecutorImpl into the base
      class that should be common across all implementations of
      FunctionExecutor. Particularly, this patch factors out
      accumulateCounterValues, and also factors out runAndSample, moving
      implementation specific code into a new runWithCounter function. This
      makes adding new implementations of FunctinExecutor easier.
      
      Reviewed By: gchatelet
      
      Differential Revision: https://reviews.llvm.org/D148079
      d2280594
    • Nathan Ridge's avatar
      28575f41
    • Aiden Grossman's avatar
      [llvm-exegesis][NFC] remove runAndMeasure · 999a8b8c
      Aiden Grossman authored
      This completes the FIXME listed in FunctionExecutor in regards to
      deprecating this function. It simply makes the appropriate call into
      runAndSample and grabs the first counter value. This patch completely
      removes the function, moving that logic into the callers (currently only
      uopsBenchmarkRunner). This makes creating new FunctionExecutors easier
      as an implementation no longer needs to worry about this detail.
      
      Reviewed By: gchatelet
      
      Differential Revision: https://reviews.llvm.org/D147878
      999a8b8c
    • Vlad Serebrennikov's avatar
      [clang] Add test for CWG1894 and CWG2199 · 576c7524
      Vlad Serebrennikov authored
      [[https://wg21.link/p1787 | P1787]]: CWG1894 and its duplicate CWG2199 are resolved per Richard’s proposal for [[ https://listarchives.isocpp.org/cgi-bin/wg21/message?wg=core&msg=28415 | “dr407 still leaves open questions about typedef / tag hiding” ]], using generic conflicting-declaration rules even for typedef, and discarding a redundant typedef-name when looking up an elaborated-type-specifier.
      Wording: See changes to [dcl.typedef], [basic.lookup.elab], and [basic.lookup]/4.
      
      Generic conflicting-declaration rules are specified in changes to [basic.scope.scope]. [[ https://cplusplus.github.io/CWG/issues/407.html | CWG407]], [[ https://cplusplus.github.io/CWG/issues/1894.html | CWG1894 ]], and [[ https://cplusplus.github.io/CWG/issues/2199.html | CWG2199 ]] discuss how elaborated type specifiers interact with typedefs, using directives, and using declarations. Since existing test for CWG407 covers examples provided in CWG1894 and CWG2199, and does it in accordance with P1787, I reused parts of it.
      
      Reviewed By: #clang-language-wg, cor3ntin
      
      Differential Revision: https://reviews.llvm.org/D148136
      576c7524
    • Nathan Ridge's avatar
      [clangd] Inactive regions support via dedicated protocol · 3f6a904b
      Nathan Ridge authored
      This implements the server side of the approach discussed at
      https://github.com/clangd/vscode-clangd/pull/193#issuecomment-1044315732
      
      Differential Revision: https://reviews.llvm.org/D143974
      3f6a904b
    • Job Noorman's avatar
      [JITLink][RISCV] Handle R_RISCV_CALL_PLT fixups · 4752787c
      Job Noorman authored
      In the default link configuration, PLT stubs are created automatically
      for R_RISCV_CALL_PLT relocations and the relocation itself is
      transformed to R_RISCV_CALL (PerGraphGOTAndPLTStubsBuilder_ELF_riscv).
      Only the latter is later handled when applying fixups and the former is
      simply ignored.
      
      This patch proposes to handle R_RISCV_CALL_PLT anyway when applying
      fixups to support custom configurations that do not need automatic PLT
      creation. An example of this is BOLT where PLT entries from the input
      binary are reused (D147544).
      
      Reviewed By: StephenFan
      
      Differential Revision: https://reviews.llvm.org/D148238
      4752787c
    • Aiden Grossman's avatar
      [MLGO] Change MBB Profile Dump from using MBB numbers to MBB IDs · 35714e3a
      Aiden Grossman authored
      Currenty, setting the -mbb-profile-dump dumps a CSV file with blocks
      inside an individual function identified by their MBB numbers. This
      patch changes the MBBs to be identified by their ID which is set at MBB
      creation and not changed afterwards, making it inherently stable
      throughout the backend. This alleviates concerns with the MBB IDs
      changing between the profile dump and what ends up in the final object
      file. The MBBs inside the SHT_LLVM_BB_ADDR_MAP sections are also
      identified using their MBB ID rather than number, so if we want to match
      them up we need to identify the MBBs here by number.
      
      Reviewed By: mtrofin, rahmanl
      
      Differential Revision: https://reviews.llvm.org/D147366
      35714e3a
    • Fangrui Song's avatar
      200f82c2
    • Jean Perier's avatar
      [flang] fix fir.array_coor of fir.box with component references · 3ce7e4b2
      Jean Perier authored
      When dealing with "derived_array(j)%component" where derived_array
      is not a contiguous array, but for which we know the extent, lowering
      generates a fir.array_coor op on a !fir.box<!fir.array<cst x T>> with
      a fir.slice containing "j" in the component path.
      
      Codegen first computes "derived_array(j)" address using the byte
      strides inside the descriptor, and then computes the offset of "j"
      from that address with a second GEP.
      The type of the address in that second GEP matters since "j" is passed
      in the GEP via an index indicating its component position in the type.
      
      The code was using the LLVM type of "derived_array" instead of
      "derived_array(j)".
      In general, with fir.box, the extent ("cst" above) is unknown and those
      types match. But if the extent of "derived_array" is a compile time
      constant, its LLVM type will be [cst x T] instead of T*, and the produced
      GEP will compute the address of the nth T instead of the nth component
      inside T leading to undefined behaviors.
      
      Fix this by computing the element type for the second GEP.
      
      Differential Revision: https://reviews.llvm.org/D148226
      3ce7e4b2
    • Jean Perier's avatar
      [flang] Change TYPE(*) arrays passing convention · 25ce9867
      Jean Perier authored
      - Fix the BIND(C) assumed-shape case: TYPE(*) assumed shape are passed
        via CFI_cdesc_t according to Fortran 2018 standard 18.3.6 point 2 (5).
      - Align the none BIND(C) case with the BIND(C) case. There is little
        point passing TYPE(*) assumed size via descriptor, use a simple
        address. C710 ensures there is no way the knowledge of the actual
        type will be required when manipulating the dummy.
      
      Differential Revision: https://reviews.llvm.org/D148130
      25ce9867
    • Jean Perier's avatar
      [flang] Fold IS_CONTIGUOUS for TYPE(*) when possible · 6f5df419
      Jean Perier authored
      TYPE(*) arguments fell through in IS_CONTIGUOUS folding
      because they are not Expr<SomeType>. Expose entry point for
      symbols in IsContiguous and use that.
      
      The added test revealed that IS_CONTIGUOUS was folded to
      false for assumed rank arguments. Fix this: the contiguity of
      assumed rank without the CONTIGUOUS argument can only be
      verified at runtime.
      
      Differential Revision: https://reviews.llvm.org/D148128
      6f5df419
    • Serge Pavlov's avatar
      [symbolizer] Change error message if module not found · 75f1f158
      Serge Pavlov authored
      If llvm-symbolize did not find module, the error looked like:
      
          LLVMSymbolizer: error reading file: No such file or directory
      
      This message does not follow common practice: LLVMSymbolizer is not an
      utility name. Also the message did not not contain the name of missed file.
      
      With this change the error message looks differently:
      
          llvm-symbolizer: error: 'abc': No such file or directory
      
      This format is closer to messages produced by other utilities and allow
      proper coloring.
      
      Differential Revision: https://reviews.llvm.org/D148032
      75f1f158
    • Max Kazantsev's avatar
      [IRCE][NFC] Refactor parseRangeCheckICmp to compute SCEVs instead of Values · a39b807d
      Max Kazantsev authored
      The motivation is to make an opportunity to compute and return
      expressions after parsing ICmp into a range check (e.g. Length + 1).
      
      Patch by Aleksandr Popov!
      
      Differential Revision: https://reviews.llvm.org/D148205
      a39b807d
    • Karl-Johan Karlsson's avatar
      [compiler-rt] Fix signed shift overflows in absvdi2.c, absvsi2.c, negvdi2.c and negvsi2.c · 854686f0
      Karl-Johan Karlsson authored
      When compiling compiler-rt with -fsanitize=undefined and running testcases you
      end up with the following warnings:
      
      UBSan: absvdi2.c:21:23: left shift of 1 by 63 places cannot be represented in type 'di_int' (aka 'long long')
      UBSan: absvsi2.c:21:23: left shift of 1 by 31 places cannot be represented in type 'si_int' (aka 'long')
      UBSan: negvdi2.c:20:32: left shift of 1 by 63 places cannot be represented in type 'di_int' (aka 'long long')
      UBSan: negvsi2.c:20:32: left shift of 1 by 31 places cannot be represented in type 'si_int' (aka 'long')
      
      This can be avoided by doing the shift in a matching unsigned variant of the
      type.
      
      This was found in an out of tree target.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D146932
      854686f0
    • Thurston Dang's avatar
      Re-land 'ASan: move allocator base to avoid conflict with high-entropy ASLR for x86-64 Linux' · fb77ca05
      Thurston Dang authored
      D147984 was reverted because it broke lit tests on Mac. This revision is based on D147984
      but maintains the old behavior for Apple.
      
      Note that, per the follow-up discussion with MaskRay in D147984, this patch excludes Apple
      but includes other platforms (e.g., aarch64, MIPS64) and OSes (e.g., FreeBSD, S390X), not just
      x86-64 Linux.
      
      Original commit message from D147984:
      
      Users have discovered [*] that when CONFIG_ARCH_MMAP_RND_BITS == 32,
      it will frequently conflict with ASan's allocator on x86-64 Linux, because the
      PIE program segment base address of 0x555555555554 plus an ASLR shift of up to
      ((2**32) * 4K == 0x100000000000) will sometimes exceed ASan's hardcoded
      base address of 0x600000000000. We fix this by simply moving the allocator base
      to 0x500000000000, which is below the PIE program segment base address. This is
      cleaner than trying to move it to another location that is sandwiched between
      the PIE program and library segments, because if either of those grow too large,
      it will collide with the allocator region.
      
      Note that we will never need to change this base address again (unless we want to increase
      the size of the allocator), because ASLR cannot be set above 32-bits for x86-64 Linux (the
      PIE program segment and library segments would collide with each other; see also
      ARCH_MMAP_RND_BITS_MAX in https://github.com/torvalds/linux/blob/master/arch/x86/Kconfig).
      
      [*] see https://b.corp.google.com/issues/276925478
      and https://groups.google.com/a/google.com/g/chrome-os-gardeners/c/BbfzCP3dEeo/m/h3C_vVUxCQAJ
      
      Differential Revision: https://reviews.llvm.org/D148280
      fb77ca05
    • Sam Clegg's avatar
      [lld][WebAssembly] stub objects: Fix handling of LTO libcall dependencies · d43f0889
      Sam Clegg authored
      This actually simplifies the code by performs a pre-pass of the stub
      objects prior to LTO.
      
      This should be the final change needed before we can make the switch
      on the emscripten side: https://github.com/emscripten-core/emscripten/pull/18905
      
      Differential Revision: https://reviews.llvm.org/D148287
      d43f0889
    • wangpc's avatar
      [TableGen] Allow references to class template arguments in defvar · fd5d0a88
      wangpc authored
      We can't refer to template arguments for defvar statements in class
      definitions, or it will report some errors like:
      
      ```
      error: Variable not defined: 'xxx'.
      ```
      
      The key point here is we used to pass nullptr to `ParseValue` in
      `ParseDefvar`. As a result, we can't refer to template arguments
      since `CurRec` is nullptr in `ParseIDValue`.
      
      So we add an argument `CurRec` to `ParseDefvar` and provide it
      when parsing defvar statements in class definitions.
      
      Reviewed By: tra, simon_tatham
      
      Differential Revision: https://reviews.llvm.org/D148197
      fd5d0a88
    • varconst's avatar
      [libc++] Fix `generate_ignore_format.sh` and regenerate the file. · 99e52b68
      varconst authored
      The script incorrectly produced double slashes in paths, e.g.
      `libcxx/src//thread.cpp`.
      99e52b68
    • Jie Fu's avatar
      [lldb][test] Fix -Wsign-compare in RegisterFlagsTest.cpp (NFC) · 0ae342f4
      Jie Fu authored
      /data/llvm-project/third-party/unittest/googletest/include/gtest/gtest.h:1526:11: error: comparison of integers of different signs: 'const unsigned long long' and 'const int' [-Werror,-Wsign-compare]
        if (lhs == rhs) {
            ~~~ ^  ~~~
      /data/llvm-project/third-party/unittest/googletest/include/gtest/gtest.h:1553:12: note: in instantiation of function template specialization 'testing::internal::CmpHelperEQ<unsigned long long, int>' requested here
          return CmpHelperEQ(lhs_expression, rhs_expression, lhs, rhs);
                 ^
      /data/llvm-project/lldb/unittests/Target/RegisterFlagsTest.cpp:128:3: note: in instantiation of function template specialization 'testing::internal::EqHelper::Compare<unsigned long long, int, nullptr>' requested here
        ASSERT_EQ(0x12345678ULL, rf.ReverseFieldOrder(0x12345678));
        ^
      /data/llvm-project/third-party/unittest/googletest/include/gtest/gtest.h:2056:32: note: expanded from macro 'ASSERT_EQ'
                              ...
      0ae342f4
    • V Donaldson's avatar
      Revert "[flang] REAL(KIND=3) and COMPLEX(KIND=3) descriptors" · 4add0e3d
      V Donaldson authored
      This reverts commit 17a4fcec.
      4add0e3d
    • Yeting Kuo's avatar
      [RISCV] Support vector strict_fsetcc/fsetccs. · d83620d1
      Yeting Kuo authored
      The patch supports vector strict_fsetcc/fsetccs. Instead of revserving fflags,
      the method to implement scalar quiet compares, the patch implement quiet
      compares by masking the signaling compares when either input is NaN [0].
      
      [0]: https://github.com/riscv/riscv-v-spec/blob/master/v-spec.adoc#vector-floating-point-compare-instructions
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D147998
      d83620d1
    • V Donaldson's avatar
      [flang] REAL(KIND=3) and COMPLEX(KIND=3) descriptors · 17a4fcec
      V Donaldson authored
      Update descriptor generation to correctly set the `type` field for
      REAL(3) and COMPLEX(3) objects.
      17a4fcec