1. Nov 15, 2023
  2. Nov 14, 2023
    • Alexey Bataev's avatar
      [SLP]Emit actual bitwidth for analyzed MinBitwidth nodes, NFCI. · f6ae50f7
      Alexey Bataev authored
      SLP includes analysis for the minimum bitwidth, the actual integer
      operations can be emitted. It allows to reduce register pressure and
      improve perf. Currently, it includes only cost model and the next
      transformation relies on InstructionCombiner. Better to do it directly
      in SLP, it allows to reduce compile time and fix cost model issues.
      f6ae50f7
    • Qiongsi Wu's avatar
      [SelectionDAG] Handling Oversized Alloca Types under 32 bit Mode to Avoid Code... · c8b11091
      Qiongsi Wu authored
      [SelectionDAG] Handling Oversized Alloca Types under 32 bit Mode to Avoid Code Generator Crash (#71472)
      
      Situations may arise leading to negative `NumElements` argument of an
      `alloca` instruction. In this case the `NumElements` is treated as a
      large unsigned value. Such large arrays may cause the size constant to
      overflow during code generation under 32 bit mode, leading to a crash.
      This PR limits the constant's bit width to the width of the pointer on
      the target. With this fix,
      ```
      alloca i32, i32 -1
      ```
      and
      ```
      alloca [4294967295 x i32], i32 1
      ```
      generates the exact same PowerPC assembly code under 32 bit mode.
      c8b11091
    • Aaron Ballman's avatar
      Fix the NATVIS visualizer for FileEntry · d554355d
      Aaron Ballman authored
      d554355d
    • Timm Baeder's avatar
      [clang][Interp] Fix stack peek offset for This ptr (#70663) · 216dfd5f
      Timm Baeder authored
      `Function::getArgSize()` include both the instance and the RVO pointer,
      so we need to subtract here.
      216dfd5f
    • Utkarsh Saxena's avatar
      Revert "[STLExtras] Remove incorrect hack to make indexed_accessor_range... · 94d6699b
      Utkarsh Saxena authored
      Revert "[STLExtras] Remove incorrect hack to make indexed_accessor_range operator== compatible with C++20" (#72265)
      
      Reverts llvm/llvm-project#72220
      
      This breaks C++20 build bot. Need to see if upgrading to clang-17 in the
      build bot would solve the issue.
      94d6699b
    • Momchil Velikov's avatar
      [CFIFixup] Allow function prologues to span more than one basic block (#68984) · 33374c44
      Momchil Velikov authored
      The CFIFixup pass assumes a function prologue is contained in a single
      basic block. This assumption is broken with upcoming support for stack
      probing (`-fstack-clash-protection`) in AArch64 - the emitted probing
      sequence in a prologue may contain loops, i.e. more than one basic
      block. The generated CFG is not arbitrary though:
       * CFI instructions are outside of any loops
      * for any two CFI instructions of the function prologue one dominates
      and is post-dominated by the other
      
      Thus, for the prologue CFI instructions, if one is executed then all are
      executed, there is a total order of executions, and the last instruction
      in that order can be considered the end of the prologoue for the purpose
      of inserting the initial `.cfi_remember_state` directive.
      
      That last instruction is found by finding the first block in the
      post-order traversal which contains prologue CFI instructions.
      33374c44
    • Alexey Bataev's avatar
      [SLP][NFCI]Improve compile time by using SmallBitVector and filtering · d4cec1ce
      Alexey Bataev authored
      trees with phis/buildvectors only.
      d4cec1ce
    • David Truby's avatar
      [flang] Add dependency to all runtime types to main target on Windows · 1256d1d1
      David Truby authored
      This patch fixes a small bug where the new flang runtime types for
      Windows (static, static_dbg, etc) are not built when the FortranRuntime
      is requested by adding the missing dependency.
      1256d1d1
    • David Spickett's avatar
      [GitHub] Add --fail to curl commands (#72238) · a39a28d2
      David Spickett authored
      This means that if we try to download a missing file, we do not get a
      document with the same file name, but containing only the http response
      code.
      
      ```
      $ curl -O -L --fail https://raw.githubusercontent.com/llvm/llvm-project/main/.github/workflows/not-a-file.py
        % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                       Dload  Upload   Total   Spent    Left  Speed
        0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0
      curl: (22) The requested URL returned error: 404
      $ $?
      22: command not found
      ```
      
      Which will be less confusing than python complaining about the file
      contents.
      a39a28d2
    • elhewaty's avatar
      [InstCombine] Fold xored one-complemented operand comparisons (#69882) · daddf402
      elhewaty authored
      - [InstCombine] Add test coverage for comparisons of operands including
      one-complemented oparands(NFC).
      - [InstCombine] Fold xored one-complemented operand comparisons.
      Alive2: https://alive2.llvm.org/ce/z/PZMJeB
      Fixes #69803.
      daddf402
    • Jeremy Morse's avatar
      [DebugInfo][RemoveDIs] Add flag to use "new" debug-info in opt (#71937) · da843aa0
      Jeremy Morse authored
      Our option to turn on the non-intrinsic form of debug-info
      (`--experimental-debuginfo-iterators`) currently requires that LLVM is
      built with the `LLVM_EXPERIMENTAL_DEBUGINFO_ITERATORS` cmake flag
      enabled, so that some (slight) performance regressions aren't
      on-by-default during the prototype/testing period. However, we still
      want to be able to _optionally_ run tests, if support is built into
      LLVM.
      
      To allow optionally exercising the non-intrinsic debug-info code, this
      patch adds `--try-experimental-debuginfo-iterators` to opt, which turns
      the `--experimental-debuginfo-iterators` flag on if support is built in,
      or leaves it off. This means we can run tests that:
       * Use normal dbg.value intrinsics if there's no support, or
       * Uses non-instruction DPValues if there is support.
        
      Which means we can start getting test coverage of DPValues/RemoveDIs
      behaviour, from in-tree tests, on our RemoveDIs buildbot. All the code
      to do with automagically converting from one form to the other landed in
      10a9e744.
      da843aa0
    • Utkarsh Saxena's avatar
      [STLExtras] Remove incorrect hack to make indexed_accessor_range operator==... · 2be3fcab
      Utkarsh Saxena authored
      [STLExtras] Remove incorrect hack to make indexed_accessor_range operator== compatible with C++20 (#72220)
      
      This partially reverts c312f025
      
      The motivation behind this is unclear and the change predates the clang
      [implementation](https://github.com/llvm/llvm-project/commit/38b9d313e6945804fffc654f849cfa05ba2c713d)
      of
      [p2468r2](https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2022/p2468r2.html)
      so I am not sure if it was ever intended to work. Rewritten template
      operators were broken since the beginning.
      
      Moreover, moving away from `friend` would be beneficial as these would
      only accepted once clang revises its implementation for fixing
      https://github.com/llvm/llvm-project/issues/70210. It also helps in
      making sure that older compilers still compile LLVM (in C++20).
      2be3fcab
    • Leandro Lupori's avatar
      [flang] Fix flang tests on MacOS (#70811) · 8cccda27
      Leandro Lupori authored
      Adjust some of the tests run by check-flang to make them pass on MacOS,
      either by skipping unsupported tests or by adapting the test for correct
      execution on MacOS.
      
      For now ctofortran.f90 test is marked as unsupported, but it can be
      adapted to run on MacOS once support for -isysroot flag is added to
      flang.
      
      Issues #70805 and #70807 are tracking the failing tests that remain, as
      these reveal real problems with flang.
      8cccda27
    • David Sherwood's avatar
      [CodeGen][AArch64] Set min jump table entries to 13 for AArch64 targets (#71166) · bdc0afc8
      David Sherwood authored
      There are some workloads that are negatively impacted by using jump
      tables when the number of entries is small. The SPEC2017 perlbench
      benchmark is one example of this, where increasing the threshold to
      around 13 gives a ~1.5% improvement on neoverse-v1. I chose the minimum
      threshold based on empirical evidence rather than science, and just
      manually increased the threshold until I got the best performance
      without impacting other workloads. For neoverse-v1 I saw around ~0.2%
      improvement in the SPEC2017 integer geomean, and no overall change for
      neoverse-n1. If we find issues with this threshold later on we can
      always revisit this.
      
      The most significant SPEC2017 score changes on neoverse-v1 were:
      
      500.perlbench_r: +1.6%
      520.omnetpp_r: +0.6%
      
      and the rest saw changes < 0.5%.
      
      I updated CodeGen/AArch64/min-jump-table.ll to reflect the new
      threshold. For most of the affected tests I manually set the min number
      of entries back to 4 on the RUN line because the tests seem to rely upon
      this behaviour.
      bdc0afc8
    • Simon Pilgrim's avatar
      [DAG] foldABSToABD - support abs(*ext(x) - *ext(y)) -> zext(abd*(x, y)) from... · 074e4ae0
      Simon Pilgrim authored
      [DAG] foldABSToABD - support abs(*ext(x) - *ext(y)) -> zext(abd*(x, y)) from different extension source types (#71670)
      
      We currently limit the fold to cases where we're extending from the same source type, but we can safely perform this using the wider of mismatching source types (we're really just interested in having extension bits on both sources), ensuring we don't create additional extensions/truncations.
      074e4ae0
    • Benjamin Maxwell's avatar
      [mlir][ArmSME] Make use of backend function attributes for enabling ZA storage (#71044) · 783ac3b6
      Benjamin Maxwell authored
      Previously, we were inserting za.enable/disable intrinsics for functions
      with the "arm_za" attribute (at the MLIR level), rather than using the
      backend attributes. This was done to avoid a dependency on the SME ABI
      functions from compiler-rt (which have only recently been implemented).
      
      Doing things this way did have correctness issues, for example, calling
      a streaming-mode function from another streaming-mode function (both
      with ZA enabled) would lead to ZA being disabled after returning to the
      caller (where it should still be enabled). Fixing issues like this would
      require re-doing the ABI work already done in the backend within MLIR.
      
      Instead, this patch switches to use the "arm_new_za" (backend) attribute
      for enabling ZA for an MLIR function. For the integration tests, this
      requires some way of linking the SME ABI functions. This is done via the
      `%arm_sme_abi_shlib` lit substitution. By default, this expands to a
      stub implementation of the SME ABI functions, but this can be overridden
      by providing the `ARM_SME_ABI_ROUTINES_SHLIB` CMake cache variable
      (pointing it at an alternative implementation). For now, the ArmSME
      integration tests pass with just stubs, as we don't make use of nested
      ZA-enabled calls.
      
      A future patch may add an option to compiler-rt to build the SME
      builtins into a standalone shared library to allow easily
      building/testing with the actual implementation.
      783ac3b6
    • Simon Pilgrim's avatar
      [X86] Regenerate expand-vp-int-intrinsics.ll · 66845418
      Simon Pilgrim authored
      Add missing X86 checks
      66845418
    • Ivan Kosarev's avatar
      [AMDGPU] Fix subtarget predicates for MUBUF instructions. (#72110) · 9ee68f8f
      Ivan Kosarev authored
      Resolves AsmParser ambiguities, e.g., between BUFFER_WBINVL1_vi and
      BUFFER_WBINVL1_gfx6_gfx7.
      
      Part of <https://github.com/llvm/llvm-project/issues/69256>.
      9ee68f8f
    • Akash Banerjee's avatar
      [MLIR][OpenMP] Changes to function-filtering pass (#71850) · 8701b178
      Akash Banerjee authored
      
      
      Currently, when deleting the device functions in the second stage of filtering during MLIR to LLVM translation we can end up with invalid calls to these functions. This is because of the removal of the EarlyOutliningPass which would have otherwise gotten rid of any such calls.
      
      This patch aims to alter the function filtering pass in the following way:
      	- Any host function is completely removed.
      	- Call to the host function are also removed and their uses replaced with Undef values.
      	- Any host function with target region code is marked to be removed during the the second stage.
      	- Calls to such functions are still removed and their uses replaced with Undef values.
      
      Co-authored-by: default avatarSergio Afonso <sergio.afonsofumero@amd.com>
      8701b178
    • Diana's avatar
      [AMDGPU] Use immediates for stack accesses in chain funcs (#71913) · eb3c02fd
      Diana authored
      Switch to using immediate offsets instead of the SP register to access
      objects on the current stack frame in chain functions. This means we no
      longer need to reserve a SP register just for accesing stack objects and
      it also allows us to set the SP (when one is actually needed) to the
      stack size from the very beginning.
      
      This only works if we use a FixedObject for the ScavengeFI, which is
      what we do for entry functions anyway (and we generally want to keep
      chain functions close to amdgpu_cs behaviour where we don't have a good
      reason to diverge).
      eb3c02fd
    • Anatoly Trosinenko's avatar
      [AArch64][PAC] Refactor aarch64-ptrauth pass (#70446) · 9bc142a0
      Anatoly Trosinenko authored
      Refactor Pointer Authentication pass in preparation for adding more
      PAUTH_* pseudo instructions:
      * dropped early return from runOnMachineFunction() as other PAUTH_*
        instructions need expansion even when pac-ret is disabled
      * refactored runOnMachineFunction() to first collect all the
        instructions of interest without modifying anything and then performing
        changes in the later loops. There are two types of relevant
        instructions: PAUTH_* pseudos that should definitely be replaced by this
        pass and tail call instructions that may require attention if pac-ret is
        enabled
      * made the loop iterating over all of the instructions handle
        instruction bundles by itself: even though this pass still does not
        support bundled TCRETURN* instructions (such as produced by KCFI) it
        does not crash anymore when no support is actually required
      9bc142a0
    • QuietMisdreavus's avatar
      ExtractAPI: use zero-based indices for line/column in symbol graph (#71753) · 63537872
      QuietMisdreavus authored
      Other implementations of the symbol graph format use zero-based indices
      for source locations, which causes problems when combined with clang's
      current one-based indices. This commit sets ExtractAPI's symbol graph
      output to use zero-based indices to align with other implementations.
      
      rdar://107639783
      63537872
    • Matthew Devereau's avatar
      [AArch64][SME2] Add ldr_zt, str_zt builtins and intrinsics (#71795) · cc124498
      Matthew Devereau authored
      Adds the builtins:
      void svldr_zt(uint64_t zt, const void *rn)
      void svstr_zt(uint64_t zt, void *rn)
      
      And the intrinsics:
      call void @llvm.aarch64.sme.ldr.zt(i32, ptr)
      tail call void @llvm.aarch64.sme.str.zt(i32, ptr)
      
      Patch by: Kerry McLaughlin <kerry.mclaughlin@arm.com>
      cc124498
    • S. B. Tam's avatar