1. Jan 12, 2024
    • Chris Bieneman's avatar
      [NFC] Remove trailing whitespace · c2fd5b73
      Chris Bieneman authored
      This seems to be causing problems that I couldn't reproduce locally.
      c2fd5b73
    • Philip Reames's avatar
      [LSR] Require non-zero step when considering wrap around for term folding (#77809) · f5dd70c5
      Philip Reames authored
      The term folding logic needs to prove that the induction variable does
      not cycle through the same set of values so that testing for the value
      of the IV on the exiting iteration is guaranteed to trigger only on that
      iteration. The prior code checked the no-self-wrap property on the IV,
      but this is insufficient as a zero step is trivially no-self-wrap per
      SCEV's definition but does repeat the same series of values.
      
      In the current form, this has the effect of basically disabling lsr's
      term-folding for all non-constant strides. This is still a net
      improvement as we've disabled term-folding entirely, so being able to
      enable it for constant strides is still a net improvement.
      
      As future work, there's two SCEV weakness worth investigating.
      
      First sext (or i32 %a, 1) to i64 does not return true for
      isKnownNonZero. This is because we check only the unsigned range in that
      query. We could either do query pushdown, or check the signed range as
      well. I tried the second locally and it has very broad impact - i.e. we
      have a bunch of missing optimizations here.
      
      Second, zext (or i32 %a, 1) to i64 as the increment to the IV in
      expensive_expand_short_tc causes the addrec to no longer be provably
      no-self-wrap. I didn't investigate this so it might be necessary, but
      the loop structure is such that I find this result surprising.
      f5dd70c5
    • dancing-leaves's avatar
      [lldb] Fix MaxSummaryLength target property type (#72233) · ee457102
      dancing-leaves authored
      There seems to be a regression since
      https://github.com/llvm/llvm-project/commit/6f8b33f6dfd0a0f8d2522b6c832bd6298ae2f3f3.
      `Max String Summary Length` target property is not read properly and the
      default value (1024) is being used instead.
      
      16.0.6:
      ```
      (lldb) settings set target.max-string-summary-length 16
      (lldb) var
      (std::string) longStdString = "0123456789101112131415161718192021222324252627282930313233343536"
      (const char *) longCharPointer = 0x000055555556f310 "0123456789101112131415161718192021222324252627282930313233343536"
      ```
      
      17.0.4:
      ```
      (lldb) settings set target.max-string-summary-length 16
      (lldb) var
      (std::string) longStdString = "0123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180181182183184185186187188189190191192193194195196197198199200201202203204205206207208209210211212213214215216217218219220221222223224225226227228229230231232233234235236237238239240241242243244245246247248249250251252253254255256257258259260261262263264265266267268269270271272273274275276277278279280281282283284285286287288289290291292293294295296297298299300301302303304305306307308309310311312313314315316317318319320321322323324325326327328329330331332333334335336337338339340341342343344345346347348349350351352353354355356357358359360361362363364365366367368369370371372373374375376377"...
      (const char *) longCharPointer = 0x000055555556f310 "*same as line above*"...
      ```
      
      Comparison fails here:
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Interpreter/OptionValue.cpp#L256
      
      Due to the type difference:
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Target/Target.cpp#L4611
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Target/TargetProperties.td#L98
      ee457102
    • Ivan Butygin's avatar
    • erichkeane's avatar
      [OpenACC] Implement 'use_device' clause parsing · dd5ce457
      erichkeane authored
      'use_device' is effectively identical to the 'copy' parsing in that it
      has required parens and no 'special' name, so this is a pretty trivial
      impementation.  There are a number of other similar situation clauses
      I'll do in a followup patch.
      dd5ce457
    • Joseph Huber's avatar
      [Libomptarget] Fix JIT on the NVPTX target by calling ptx manually (#77801) · 3ede817f
      Joseph Huber authored
      Summary:
      Recently a patch added an assertion in the GlobalHandler to indicate
      when an ELF was not used. This began to fire whenever NVPTX JIT was
      used, because the JIT pass output a PTX file instead of an ELF. The
      CUModuleLoad method consumes `.s` internally and compiles it to a cubin,
      however, this is too late as we perform several checks on the ELF
      directly for the presence of certain symbols and to read some necessary
      constants. This results in inconsistent behaviour.
      
      To address this, this patch simply calls `ptxas` manually, similar to
      how `lld` is called for the AMDGPU JIT pass. This is inevitably going to
      be slower than simply passing it to the CUDA routine due to the overhead
      involved in file IO and a fork call, but it's necessary for correctness.
      
      CUDA provides an API for compiling PTX manually. However, this only
      started showing up in CUDA 11.1 and is only provided "officially" in a
      static library. The `libnvidia-ptxjitcompiler.so` next to the CUDA
      driver has the same symbols and can likely be used as a replacement.
      This would be the faster solution. However, given that it's not
      documented it may have some issues.
      3ede817f
    • Luke Lau's avatar
      [RISCV] Add test for strided gather with recursive disjoint or. NFC · 114e6d7b
      Luke Lau authored
      This already gets converted to a strided intrinsic because we currently call
      haveNoCommonBitsSet when checking or instructions, but an upcoming patch will
      change this logic and we want to preserve this case.
      
      Note that this IR is in the form that comes from instcombine. The splats need
      to be inline constexprs, otherwise isSplatValue() will fail. (It can't
      currently handle splats where the shufflevector is an instruction, and the
      insertelement is a constexpr.
      114e6d7b
    • erichkeane's avatar
      [OpenACC] Implement 'copy' Clause · 923f0392
      erichkeane authored
      The copy clause takes a var-list, similar to cache.  This patch
      implements the parsing in terms of how we did cache, and does some
      infrastructure for future clause parsing.
      
      As a part of this, many functions needed to become members of Parser,
      which I anticipated needing to happen in the future anyway.
      923f0392
    • Eleanor Bonnici's avatar
      [lld][ELF] Allow Arm PC-relative relocations in PIC links (#77304) · d21fb06a
      Eleanor Bonnici authored
      The relocations that map to R_ARM_PCA are equivalent to R_PC. They are
      PC-relative and safe to use in shared libraries, but have a different
      relocation code as they are evaluated differently. Now that LLVM may
      generate these relocations in object files, they may occur in
      shared libraries or position-independent executables.
      d21fb06a
    • Florian Hahn's avatar
      3b3da7c7
    • Momchil Velikov's avatar
      [AArch64] Fix missing `pfalse` diagnostic (#77746) · 90eb4e24
      Momchil Velikov authored
      The missing diagnostic causes an ICE when a suffix other than `.B`
      is used in a `pfalse` instruction with a predicate-as-counter operand.
      90eb4e24
    • Mirko Brkušanin's avatar
    • Vladislav Dzhidzhoev's avatar
      [CloneFunction][DebugInfo] Avoid cloning DILocalVariables of inlined functions (#75385) · fc6faa11
      Vladislav Dzhidzhoev authored
      - [DebugMetadata][DwarfDebug] Support function-local types in lexical
      block scopes (4/7)
      - [CloneFunction][DebugInfo] Avoid cloning DILocalVariables of inlined
      functions
      
      This is a follow-up for https://reviews.llvm.org/D144006, fixing a crash
      reported
      in Chromium (https://reviews.llvm.org/D144006#4651955).
      
      The first commit is added for convenience, as it has already been
      accepted.
      
      If DISubpogram was not cloned (e.g. we are cloning a function that has
      other
      functions inlined into it, and subprograms of the inlined functions are
      not supposed to be cloned), it doesn't make sense to clone its
      DILocalVariables as well.
      Otherwise get duplicated DILocalVariables not tracked in their
      subprogram's retainedNodes, that crash LTO with Chromium.
      
      This is meant to be committed along with
      https://reviews.llvm.org/D144006.
      fc6faa11
  2. Jan 11, 2024