1. Jan 12, 2024
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][openacc] Apply mutually exclusive clauses restriction to routine (#77802) · bdfe5d69
      this patch enforce or fix the enforcement of two restrictions from
      section 2.15.1:
      
      > Only the gang, worker, vector, seq and bind clauses may follow a
      device_type clause.
      
      `seq` was not allowed after `device_type` with the current
      implementation.
      
      > Exactly one of the gang, worker, vector, or seq clauses must appear.
      
      This was not properly checked. This patch check correctly for mutually
      exclusion as described in section 2.4. Mutually exclusive clauses may
      appear on the same directive if they apply for different device_type.
      bdfe5d69
    • Krzysztof Parzyszek's avatar
      Revert "[Flang][Parser] Add missing dependencies to CMakeLists.txt (#77483)" · d4473047
      Krzysztof Parzyszek authored
      This reverts commit cc53ec82.
      
      This commit hasn't accomplished anything. The original issue was that
      `DumpTree`, when called from lowering, caused linker errors due to some
      directive-naming functions being absent. Adding FrontendOpenMP to the
      parser library didn't fix that problem, and according to the notes in
      PR #77483, calling `DumpTree` from lowering isn't really supported.
      d4473047
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [mlir][openacc][flang] Simplify gang, vector and worker representation (#77667) · 40f5f905
      The IR representation for gang, vector and worker has grown with the
      support for device_type. This patch simplify the IR representation for
      gang, vector and worker information on the acc.loop operation.
      
      When the only the keyword is present without any values, the information
      is printed at the same place than when there is values. The device_type
      is omitted if there is no values and it is equal to None. Otherwise the
      full information is displayed. First the keyword only device_type
      information and then the values with their device_type.
      40f5f905
    • Joseph Huber's avatar
      [Libomptarget] Fix GPU Dtors referencing possibly deallocated image (#77828) · 37c1a5e3
      Joseph Huber authored
      Summary:
      The constructors and destructors look up a symbol in the ELF quickly to
      determine if they need to be run on the GPU. This allows us to avoid the
      very slow actions required to do the slower lookup using the vendor API.
      
      One problem occurs with how we handle the lifetime of these images.
      Right now there is no invariant to specify the lifetime of the
      underlying binary image that is loaded. In the typical case, this comes
      from the binary itself in the `.llvm.offloading` section, meaning that
      the lifetime of the binary should match the executable itself. This
      would work fine, if it weren't for the fact that the plugin is loaded
      via `dlopen` and can have a teardown order out of sync with the main
      executable.
      
      This was likely what was occuring when this failed on some systems but
      not others. A potential solution would be to simply copy images into
      memory so the runtime does not rely on external references. Another
      would be to manually zero these out after initialization as to prevent
      this mistake from happening accidentally. The former has the benefit of
      making some checks easier, and allowing for constant initialization be
      done on the ELF itself (normally we can't do this because writing to a
      constant section, e.g. .llvm.offloading is a segfault.). The downside
      would be the extra time required to copy the image in bulk (Although we
      are likely doing this in the vendor runtimes as well).
      
      This patch went with a quick solution to simply set a boolean value at
      initialization time if we need to call destructors.
      
      Fixes: https://github.com/llvm/llvm-project/issues/77798
      37c1a5e3
    • Tacet's avatar
      [ASan][libc++] Initialize `__r_` variable with lambda (#77394) · 75efddba
      Tacet authored
      This commit is a refactor (increases readability) and optimization fix.
      
      This is a fixed commit of
      https://github.com/llvm/llvm-project/pull/76200 First reverthed here:
      https://github.com/llvm/llvm-project/commit/1ea7a56057492d9da1124787a9855cc2edca7df9
      
      Please, check original PR for details.
      
      The difference is a return type of the lambda.
      
      Original description:
      
      This commit addresses optimization and instrumentation challenges
      encountered within comma constructors.
      1) _LIBCPP_STRING_INTERNAL_MEMORY_ACCESS does not work in comma
      constructors.
      2) Code inside comma constructors is not always correctly optimized.
      Problematic code examples:
      - `: __r_(((__str.__is_long() ? 0 : (__str.__annotate_delete(), 0)),
      std::move(__str.__r_))) {`
      - `: __r_(__r_([&](){ if(!__s.__is_long()) __s.__annotate_delete();
      return std::move(__s.__r_);}())) {`
      
      However, lambda with argument seems to be correctly optimized. This
      patch uses that fact.
      
      Use of lambda based on idea from @ldionne.
      75efddba
    • Zequan Wu's avatar
      Revert "[asan] Enable StackSafetyAnalysis by default" · e7f79487
      Zequan Wu authored
      This reverts commit 51fbab13.
      This causes the compiler to crash. Will file a issue to track the status.
      e7f79487
    • Kazu Hirata's avatar
      [Dialect] Fix a warning · 3e82663b
      Kazu Hirata authored
      This patch fixes:
      
        mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp:3154:8: error: unused
        variable 'rank' [-Werror,-Wunused-variable]
      3e82663b
    • Amir Ayupov's avatar
      [BOLT] Encode BAT using ULEB128 (#76899) · 565f40d6
      Amir Ayupov authored
      Reduces BAT section size, bytes:
      - large binary: 38676872 -> 23262524 (0.60x),
      - medium binary (trunk clang): 5938004 -> 3213504 (0.54x),
      - small binary (X86/bolt-address-translation.test): 1436 -> 680 (0.47x).
      
      Test Plan: Updated bolt/test/X86/bolt-address-translation.test
      565f40d6
    • Amir Ayupov's avatar
      [BOLT] Add BOLT Address Translation documentation (#76899) · a7cf0a1f
      Amir Ayupov authored
      Test Plan: Open the page in browser
      a7cf0a1f
    • Kazu Hirata's avatar
      [flang] Fix a warning · fb094471
      Kazu Hirata authored
      This patch fixes:
      
        flang/unittests/Runtime/CommandTest.cpp:702:14: error: variable
        length arrays are a C99 feature [-Werror,-Wvla-extension]
      fb094471
    • Kazu Hirata's avatar
      [Format] Fix a warning · cf3421de
      Kazu Hirata authored
      This patch fixes:
      
        clang/unittests/Format/TokenAnnotatorTest.cpp:2181:29: error: lambda
        capture 'Style' is not used [-Werror,-Wunused-lambda-capture]
      cf3421de
    • Kazu Hirata's avatar
      [flang] Fix a warning · 18734f60
      Kazu Hirata authored
      This patch fixes:
      
        flang/runtime/extensions.cpp:111:12: error: variable length arrays
        are a C99 feature [-Werror,-Wvla-extension]
      18734f60
    • Hirofumi Nakamura's avatar
      [clang-format] TableGen keywords support. (#77477) · 0cc31579
      Hirofumi Nakamura authored
      Add TableGen keywords to the additional keyword list of the formatter.
      
      This pull request is the splited part from
      https://github.com/llvm/llvm-project/pull/76059 .
      0cc31579
    • Amir Ayupov's avatar
      [BOLT][NFC] Print BAT section size (#76897) · 2bb511e2
      Amir Ayupov authored
      Test Plan: Updated bolt/test/X86/bolt-address-translation.test
      2bb511e2
    • Usman Nadeem's avatar
      [AArch64][SVE2] Generate XAR (#77160) · c3e3aa9c
      Usman Nadeem authored
      Bitwise exclusive OR and rotate right by immediate
      
      Select xar (x, y, imm) for the following pattern:
          or (shl (xor x, y), nBits-imm), (shr (xor x, y), imm)
      
      This is essentially:
          rotr (xor(x, y), imm)
      c3e3aa9c
    • Felix Schneider's avatar
      [mlir][memref] Transpose: allow affine map layouts in result, extend folder (#76294) · 4619e21c
      Felix Schneider authored
      Currently, the `memref.transpose` verifier forces the result type of the
      Op to have an explicit `StridedLayoutAttr` via the method
      `inferTransposeResultType`. This means that the example Op
      given in the documentation is actually invalid because it uses an `AffineMap`
      to specify the layout.
      It also means that we can't "un-transpose" a transposed memref back to
      the implicit layout form, because the verifier will always enforce the
      explicit strided layout.
      
      This patch makes the following changes:
      
      1. The verifier checks whether the canonicalized strided layout of the
      result Type is identitcal to the canonicalized infered result type
      layout. This way, it's only important that the two Types have the same
      strided layout, not necessarily the same representation of it.
      2. The folder is extended to support folding away the trivial case of
      identity permutation and to fold one transposition into another by
      composing the permutation maps.
      4619e21c
    • Felix Schneider's avatar
      [mlir][affine] Add dependency on `UBDialect` for `PoisonAttr` (#77691) · 061b777c
      Felix Schneider authored
      The folder for `AffineApplyOp` will try creating a `PoisonAttr`
      under certain circumstances. However, this will result in a crash if the
      `UBDialect` isn't loaded.
      
      This patch adds a dependency of `AffineDialect` on `UBDialect`.
      061b777c
    • Mats Petersson's avatar
      Add more ZA modes (#77361) · 21e1bf2d
      Mats Petersson authored
      Add more ZA modes
          
       Adds the arm_shared_za and arm_preserves_za attributes to the existing
       arm_new_za attribute. The functionality already exists in LLVM, so just
       "linking the pieces together".
          
      For more details see:
      https://arm-software.github.io/acle/main/acle.html#sme-attributes-relating-to-za
      21e1bf2d
    • Chris Bieneman's avatar
      [NFC] Remove trailing whitespace · c2fd5b73
      Chris Bieneman authored
      This seems to be causing problems that I couldn't reproduce locally.
      c2fd5b73
    • Philip Reames's avatar
      [LSR] Require non-zero step when considering wrap around for term folding (#77809) · f5dd70c5
      Philip Reames authored
      The term folding logic needs to prove that the induction variable does
      not cycle through the same set of values so that testing for the value
      of the IV on the exiting iteration is guaranteed to trigger only on that
      iteration. The prior code checked the no-self-wrap property on the IV,
      but this is insufficient as a zero step is trivially no-self-wrap per
      SCEV's definition but does repeat the same series of values.
      
      In the current form, this has the effect of basically disabling lsr's
      term-folding for all non-constant strides. This is still a net
      improvement as we've disabled term-folding entirely, so being able to
      enable it for constant strides is still a net improvement.
      
      As future work, there's two SCEV weakness worth investigating.
      
      First sext (or i32 %a, 1) to i64 does not return true for
      isKnownNonZero. This is because we check only the unsigned range in that
      query. We could either do query pushdown, or check the signed range as
      well. I tried the second locally and it has very broad impact - i.e. we
      have a bunch of missing optimizations here.
      
      Second, zext (or i32 %a, 1) to i64 as the increment to the IV in
      expensive_expand_short_tc causes the addrec to no longer be provably
      no-self-wrap. I didn't investigate this so it might be necessary, but
      the loop structure is such that I find this result surprising.
      f5dd70c5
    • dancing-leaves's avatar
      [lldb] Fix MaxSummaryLength target property type (#72233) · ee457102
      dancing-leaves authored
      There seems to be a regression since
      https://github.com/llvm/llvm-project/commit/6f8b33f6dfd0a0f8d2522b6c832bd6298ae2f3f3.
      `Max String Summary Length` target property is not read properly and the
      default value (1024) is being used instead.
      
      16.0.6:
      ```
      (lldb) settings set target.max-string-summary-length 16
      (lldb) var
      (std::string) longStdString = "0123456789101112131415161718192021222324252627282930313233343536"
      (const char *) longCharPointer = 0x000055555556f310 "0123456789101112131415161718192021222324252627282930313233343536"
      ```
      
      17.0.4:
      ```
      (lldb) settings set target.max-string-summary-length 16
      (lldb) var
      (std::string) longStdString = "0123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180181182183184185186187188189190191192193194195196197198199200201202203204205206207208209210211212213214215216217218219220221222223224225226227228229230231232233234235236237238239240241242243244245246247248249250251252253254255256257258259260261262263264265266267268269270271272273274275276277278279280281282283284285286287288289290291292293294295296297298299300301302303304305306307308309310311312313314315316317318319320321322323324325326327328329330331332333334335336337338339340341342343344345346347348349350351352353354355356357358359360361362363364365366367368369370371372373374375376377"...
      (const char *) longCharPointer = 0x000055555556f310 "*same as line above*"...
      ```
      
      Comparison fails here:
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Interpreter/OptionValue.cpp#L256
      
      Due to the type difference:
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Target/Target.cpp#L4611
      
      https://github.com/llvm/llvm-project/blob/9cb1673fa5d267148ac81ee31b37f1d2f7c0f2b8/lldb/source/Target/TargetProperties.td#L98
      ee457102
    • Ivan Butygin's avatar
    • erichkeane's avatar
      [OpenACC] Implement 'use_device' clause parsing · dd5ce457
      erichkeane authored
      'use_device' is effectively identical to the 'copy' parsing in that it
      has required parens and no 'special' name, so this is a pretty trivial
      impementation.  There are a number of other similar situation clauses
      I'll do in a followup patch.
      dd5ce457
    • Joseph Huber's avatar
      [Libomptarget] Fix JIT on the NVPTX target by calling ptx manually (#77801) · 3ede817f
      Joseph Huber authored
      Summary:
      Recently a patch added an assertion in the GlobalHandler to indicate
      when an ELF was not used. This began to fire whenever NVPTX JIT was
      used, because the JIT pass output a PTX file instead of an ELF. The
      CUModuleLoad method consumes `.s` internally and compiles it to a cubin,
      however, this is too late as we perform several checks on the ELF
      directly for the presence of certain symbols and to read some necessary
      constants. This results in inconsistent behaviour.
      
      To address this, this patch simply calls `ptxas` manually, similar to
      how `lld` is called for the AMDGPU JIT pass. This is inevitably going to
      be slower than simply passing it to the CUDA routine due to the overhead
      involved in file IO and a fork call, but it's necessary for correctness.
      
      CUDA provides an API for compiling PTX manually. However, this only
      started showing up in CUDA 11.1 and is only provided "officially" in a
      static library. The `libnvidia-ptxjitcompiler.so` next to the CUDA
      driver has the same symbols and can likely be used as a replacement.
      This would be the faster solution. However, given that it's not
      documented it may have some issues.
      3ede817f
    • Luke Lau's avatar
      [RISCV] Add test for strided gather with recursive disjoint or. NFC · 114e6d7b
      Luke Lau authored
      This already gets converted to a strided intrinsic because we currently call
      haveNoCommonBitsSet when checking or instructions, but an upcoming patch will
      change this logic and we want to preserve this case.
      
      Note that this IR is in the form that comes from instcombine. The splats need
      to be inline constexprs, otherwise isSplatValue() will fail. (It can't
      currently handle splats where the shufflevector is an instruction, and the
      insertelement is a constexpr.
      114e6d7b
    • erichkeane's avatar
      [OpenACC] Implement 'copy' Clause · 923f0392
      erichkeane authored
      The copy clause takes a var-list, similar to cache.  This patch
      implements the parsing in terms of how we did cache, and does some
      infrastructure for future clause parsing.
      
      As a part of this, many functions needed to become members of Parser,
      which I anticipated needing to happen in the future anyway.
      923f0392
    • Eleanor Bonnici's avatar
      [lld][ELF] Allow Arm PC-relative relocations in PIC links (#77304) · d21fb06a
      Eleanor Bonnici authored
      The relocations that map to R_ARM_PCA are equivalent to R_PC. They are
      PC-relative and safe to use in shared libraries, but have a different
      relocation code as they are evaluated differently. Now that LLVM may
      generate these relocations in object files, they may occur in
      shared libraries or position-independent executables.
      d21fb06a
    • Florian Hahn's avatar
      3b3da7c7
    • Momchil Velikov's avatar
      [AArch64] Fix missing `pfalse` diagnostic (#77746) · 90eb4e24
      Momchil Velikov authored
      The missing diagnostic causes an ICE when a suffix other than `.B`
      is used in a `pfalse` instruction with a predicate-as-counter operand.
      90eb4e24
    • Mirko Brkušanin's avatar
    • Vladislav Dzhidzhoev's avatar
      [CloneFunction][DebugInfo] Avoid cloning DILocalVariables of inlined functions (#75385) · fc6faa11
      Vladislav Dzhidzhoev authored
      - [DebugMetadata][DwarfDebug] Support function-local types in lexical
      block scopes (4/7)
      - [CloneFunction][DebugInfo] Avoid cloning DILocalVariables of inlined
      functions
      
      This is a follow-up for https://reviews.llvm.org/D144006, fixing a crash
      reported
      in Chromium (https://reviews.llvm.org/D144006#4651955).
      
      The first commit is added for convenience, as it has already been
      accepted.
      
      If DISubpogram was not cloned (e.g. we are cloning a function that has
      other
      functions inlined into it, and subprograms of the inlined functions are
      not supposed to be cloned), it doesn't make sense to clone its
      DILocalVariables as well.
      Otherwise get duplicated DILocalVariables not tracked in their
      subprogram's retainedNodes, that crash LTO with Chromium.
      
      This is meant to be committed along with
      https://reviews.llvm.org/D144006.
      fc6faa11
  2. Jan 11, 2024
    • Leandro Lupori's avatar
      31ce0f1d
    • Teresa Johnson's avatar
    • HaohaiWen's avatar
      b6fc463d
    • Guillaume Chatelet's avatar
      [libc][NFC] Use 16-byte indices for _mmXXX_shuffle_epi8 (#77781) · 57948542
      Guillaume Chatelet authored
      This is less confusing since the implementation only cares about the 4
      lower bits.
      57948542
    • Louis Dionne's avatar
      [runtimes] Use LLVM libunwind from libc++abi by default (#77687) · 8f90e693
      Louis Dionne authored
      I recently came across LIBCXXABI_USE_LLVM_UNWINDER and was surprised to
      notice it was disabled by default. Since we build libunwind by default
      and ship it in the LLVM toolchain, it would seem to make sense that
      libc++ and libc++abi rely on libunwind for unwinding instead of using
      the system-provided unwinding library (if any).
      
      Most importantly, using the system unwinder implies that libc++abi is
      ABI compatible with that system unwinder, which is not necessarily the
      case. Hence, it makes a lot more sense to instead default to using the
      known-to-be-compatible LLVM unwinder, and let vendors manually select a
      different unwinder if desired.
      
      As a follow-up change, we should probably apply the same default to
      compiler-rt.
      
      Differential Revision: https://reviews.llvm.org/D150897
      Fixes #77662
      rdar://120801778
      8f90e693
    • Luke Lau's avatar
      3b3ee1f5
    • Rainer Orth's avatar
      [flang] Handle missing LOGIN_NAME_MAX definition in runtime (#77775) · 731b2956
      Rainer Orth authored
      18af032c broke the Solaris build:
      ```
      /vol/llvm/src/llvm-project/dist/flang/runtime/extensions.cpp:60:24: error: use of undeclared identifier 'LOGIN_NAME_MAX'
         60 |   const int nameMaxLen{LOGIN_NAME_MAX + 1};
            |                        ^
      /vol/llvm/src/llvm-project/dist/flang/runtime/extensions.cpp:61:12: warning: variable length arrays in C++ are a Clang extension [-Wvla-cxx-extension]
         61 |   char str[nameMaxLen];
            |            ^~~~~~~~~~
      /vol/llvm/src/llvm-project/dist/flang/runtime/extensions.cpp:61:12: note: initializer of 'nameMaxLen' is unknown
      /vol/llvm/src/llvm-project/dist/flang/runtime/extensions.cpp:60:13: note: declared here
         60 |   const int nameMaxLen{LOGIN_NAME_MAX + 1};
            |             ^
      ```
      `flang/unittests/Runtime/CommandTest.cpp` has the same issue.
      
      As documented in Solaris 11.4 `limits.h(3HEAD)`, `LOGIN_NAME_MAX` can be
      undefined. To determine the value, `sysconf(3C)` needs to be used
      instead.
      
      Beside that portable method, Solaris also provides a non-standard
      `LOGNAME_MAX` which could be used, but I've preferred the standard route
      instead which would support other targets with the same issue.
      
      Tested on `amd64-pc-solaris2.11` and `x86_64-pc-linux-gnu`.
      731b2956
    • Alexey Bataev's avatar
      [SLP]Do not require external uses for roots and single use for other... · 18473eb1
      Alexey Bataev authored
      [SLP]Do not require external uses for roots and single use for other instructions in computeMinimumValueSizes. (#72679)
      
      After changes, that does not require support from InstCombine, we can
      drop some extra requirements for values-to-be-demoted. No need to check
      for external uses for roots/other instructions, just check that the
        no non-vectorized insertelement instruction, which may require
        widening.
      
      Review: https://github.com/llvm/llvm-project/pull/72679
      18473eb1
    • Teresa Johnson's avatar
      [MemProf] Handle missing tail call frames (#75823) · 26a8664e
      Teresa Johnson authored
      If tail call optimization was not disabled for the profiled binary, the
      call contexts will be missing frames for tail calls. Handle this by
      performing a limited search through tail call edges for the profiled
      callee when a discontinuity is detected. The search depth is adjustable
      but defaults to 5.
      
      If we are able to identify a short sequence of tail calls, update the
      graph for those calls. In the case of ThinLTO, synthesize the necessary
      CallsiteInfos for carrying the cloning information to the backends.
      26a8664e