1. May 06, 2021
    • Jay Foad's avatar
    • Nemanja Ivanovic's avatar
      [PowerPC] Provide some P8-specific altivec overloads for P7 · ed87f512
      Nemanja Ivanovic authored
      This adds additional support for XL compatibility. There are a number
      of functions in altivec.h that produce a single instruction (or a
      very short sequence) for Power8 but can be done on Power7 without
      scalarization. XL provides these implementations.
      This patch adds the following overloads for doubleword vectors:
      vec_add
      vec_cmpeq
      vec_cmpgt
      vec_cmpge
      vec_cmplt
      vec_cmple
      vec_sl
      vec_sr
      vec_sra
      ed87f512
    • Paul C. Anagnostopoulos's avatar
    • Anastasia Stulova's avatar
      [OpenCL] Remove subgroups pragma in enqueue kernel and pipe builtins. · c28a6023
      Anastasia Stulova authored
      This patch simplifies the parser and makes the language semantics
      consistent. There is no extension pragma requirement in the spec
      for the subgroup functions in enqueue kernel or pipes and all other
      builtin functions are available without the pragama.
      
      Differential Revision: https://reviews.llvm.org/D100984
      c28a6023
    • Jon Chesterfield's avatar
      [amdgpu-arch] Fix rpath to run from build dir · b24e9f82
      Jon Chesterfield authored
      [amdgpu-arch] Fix rpath to run from build dir
      
      Prior to this, amdgpu-arch has RUNPATH set to $ORIGIN/../lib which works
      for some installs, but not from the build directory where clang executes
      the tool from when running tests.
      
      This cmake option adds the location of the rocr runtime to the RUNPATH
      (note, it amends RUNPATH here, despite the cmake option referring to RPATH)
      to create a binary that runs from build or install location.
      
      Before:
      RUNPATH [$ORIGIN/../lib]
      After:
      RUNPATH [$ORIGIN/../lib:$HOME/llvm-install/lib]
      
      Credit to Greg for knowing this trick and pointing to examples of it in use
      for the aomp build scripts.
      
      Reviewed By: pdhaliwal
      
      Differential Revision: https://reviews.llvm.org/D101926
      b24e9f82
    • Carl Ritson's avatar
      [AMDGPU] Fix WQM failure with single block inactive demote · 67cfefeb
      Carl Ritson authored
      Instruction test for inactive kill/demote needs to be based on
      actual opcode not whether instruction would be lowered to demote.
      
      Reviewed By: piotr
      
      Differential Revision: https://reviews.llvm.org/D101966
      67cfefeb
    • Malhar Jajoo's avatar
      Revert "[ARM] Transforming memcpy to Tail predicated Loop" · fc690777
      Malhar Jajoo authored
      Reverting commit since it causes failure (10462).
      This reverts commit b856f4a2.
      fc690777
    • Benjamin Kramer's avatar
    • David Green's avatar
      [LV] Account for tripcount when calculation vectorization profitability · 4979c904
      David Green authored
      The loop vectorizer will currently assume a large trip count when
      calculating which of several vectorization factors are more profitable.
      That is often not a terrible assumption to make as small trip count
      loops will usually have been fully unrolled. There are cases however
      where we will try to vectorize them, and especially when folding the
      tail by masking can incorrectly choose to vectorize loops that are not
      beneficial, due to the folded tail rounding the iteration count up for
      the vectorized loop.
      
      The motivating example here has a trip count of 5, so either performs 5
      scalar iterations or 2 vector iterations (with VF=4). At a high enough
      trip count the vectorization becomes profitable, but the rounding up to
      2 vector iterations vs only 5 scalar makes it unprofitable.
      
      This adds an alternative cost calculation when we know the max trip
      count and are folding tail by masking, rounding the iteration count up
      to the correct number for the vector width. We still do not account for
      anything like setup cost or the mixture of vector and scalar loops, but
      this is at least an improvement in a few cases that we have had
      reported.
      
      Differential Revision: https://reviews.llvm.org/D101726
      4979c904
    • Ben Dunbobbin's avatar
      [LLD] Improve --strip-all help text · 5dd9f44c
      Ben Dunbobbin authored
      This is a slight improvement to the help text, as I was slightly
      surprised when strip-all did more than remove the symbol table.
      
      Currently, we match gold's help text for strip-all and strip-debug.
      I think that the GNU documentation for these options is not particularly
      clear. However, I have opted to make only a minor change here and keep
      the help text similar to gold's as these are mature options that are
      well understood.
      
      ld.bfd (https://sourceware.org/binutils/docs/ld/Options.html) has a
      similar implication although it defines strip-debug as a subset of
      strip-all. However, felt that noting that strip-all implies strip-debug
      is better; because, with the ld.bfd approach you have to read both the
      --strip-debug and the --strip-all help text to understand the behaviour
      of --strip-all (and the --strip-all help text doesn't indicate that he
      --strip-debug help text is related).
      
      Differential Revision: https://reviews.llvm.org/D101890
      5dd9f44c
    • Christian Sigg's avatar
      [mlir] Add support for ops with regions in 'gpu-async-region' rewriter. · a0d019fc
      Christian Sigg authored
      Reviewed By: herhut
      
      Differential Revision: https://reviews.llvm.org/D101757
      a0d019fc
    • Simon Pilgrim's avatar
      [AMDGPU] Regenerate fp2int tests. NFCI. · 0fdce16e
      Simon Pilgrim authored
      0fdce16e
    • Simon Pilgrim's avatar
      [AMDGPU] Regenerate shift tests. NFCI. · 20e976e2
      Simon Pilgrim authored
      20e976e2
    • Jonas Paulsson's avatar
      [SystemZ] Support builtin_frame_address with packed stack without backchain. · a0da66bc
      Jonas Paulsson authored
      In order to use __builtin_frame_address(0) with packed stack and no
      backchain, the address of where the backchain would have been written is
      returned (like GCC).
      
      This address may either contain a saved register or be unused.
      
      Review: Ulrich Weigand
      
      Differential Revision: https://reviews.llvm.org/D101897
      a0da66bc
    • Kerry McLaughlin's avatar
      [SVE][LoopVectorize] Add support for scalable vectorization of first-order recurrences · 8c9742bd
      Kerry McLaughlin authored
      Adds support for scalable vectorization of loops containing first-order recurrences, e.g:
      ```
      for(int i = 0; i < n; i++)
        b[i] =  a[i] + a[i - 1]
      ```
      This patch changes fixFirstOrderRecurrence for scalable vectors to take vscale into
      account when inserting into and extracting from the last lane of a vector.
      CreateVectorSplice has been added to construct a vector for the recurrence, which
      returns a splice intrinsic for scalable types. For fixed-width the behaviour
      remains unchanged as CreateVectorSplice will return a shufflevector instead.
      
      The tests included here are the same as test/Transform/LoopVectorize/first-order-recurrence.ll
      
      Reviewed By: david-arm, fhahn
      
      Differential Revision: https://reviews.llvm.org/D101076
      8c9742bd
    • Eliza Velasquez's avatar
      [clang-format] Rename common types between C#/JS · cdf33962
      Eliza Velasquez authored
      Reviewed By: curdeius
      
      Differential Revision: https://reviews.llvm.org/D101862
      cdf33962
    • Eliza Velasquez's avatar
      [clang-format] Fix C# nullable-related errors · ec725b30
      Eliza Velasquez authored
      This fixes two errors:
      
      Previously, clang-format was splitting up type identifiers from the
      nullable ?. This changes this behavior so that the type name sticks with
      the operator.
      
      Additionally, nullable operators attached to return types in interface
      functions were not parsed correctly. Digging deeper, it looks like
      interface bodies were being parsed differently than classes and structs,
      causing MustBeDeclaration to be incorrect for interface members. They
      now share the same logic.
      
      One other change is reintroducing the CSharpNullable type independent of
      JsTypeOptionalQuestion. Despite having a similar semantic purpose, their
      actual syntax differs quite a bit.
      
      Reviewed By: MyDeveloperDay, curdeius
      
      Differential Revision: https://reviews.llvm.org/D101860
      ec725b30
    • Eliza Velasquez's avatar
    • Jay Foad's avatar
      [AMDGPU] SIFoldOperands: clean up tryConstantFoldOp · 7c706af0
      Jay Foad authored
      First clean up the strange API of tryConstantFoldOp where it took an
      immediate operand value, but no indication of which operand it was the
      value for.
      
      Second clean up the loop that calls tryConstantFoldOp so that it does
      not have to restart from the beginning every time it folds an
      instruction.
      
      This is NFCI but there are some minor changes caused by the order in
      which things are folded.
      
      Differential Revision: https://reviews.llvm.org/D100031
      7c706af0
    • Andrzej Warzynski's avatar
      [flang] Remove `%f18` from LIT configuration files · 65cd0d6b
      Andrzej Warzynski authored
      `%f18` was originally introduced to represent the old Flang driver,
      `f18`. With the introduction of the new driver, `flang-new`, we have
      been switching to `%flang` (compiler driver) and `%flang_fc1` (frontend
      driver) as more generic alternatives.
      
      As most tests have been portend to use the new LIT variables instead of
      `%f18`, this is good time to remove it from lit.cfg.py. There's only one
      test left that requires the old driver to run. It's updated with:
      ```
      ! REQUIRES: old-flang-driver
      ```
      This way we preserve its semantics while reducing the number of
      variables in LIT configuration.
      
      Differential Revision: https://reviews.llvm.org/D101281
      65cd0d6b
    • Malhar Jajoo's avatar
      [ARM] Transforming memcpy to Tail predicated Loop · b856f4a2
      Malhar Jajoo authored
      This patch converts llvm.memcpy intrinsic into Tail Predicated
      Hardware loops for a target that supports the Arm M-profile
      Vector Extension (MVE).
      
      From an implementation point of view, the patch
      
      - adds an ARM specific SDAG Node (to which the llvm.memcpy intrinsic is lowered to, during first phase of ISel)
      - adds a corresponding TableGen entry to generate a pseudo instruction, with a custom inserter,
        on matching the above node.
      - Adds a custom inserter function that expands the pseudo instruction into MIR suitable
         to be (by later passes) into a WLSTP loop.
      
      Note: A cli option is used to control the conversion of memcpy to TP
      loop and this option is currently disabled by default. It may be enabled
      in the future after further downstream testing.
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D99723
      b856f4a2
    • James Henderson's avatar
      [lit] Report tool path from use_llvm_tool if found via env variable · abe2c906
      James Henderson authored
      Previously, if the search_env argument was specified, and the tool was
      found at that location, the path was not reported, unlike other
      situations when this function was called. Adding the reporting makes the
      function consistent.
      
      Reviewed by: thopre
      
      Differential Revision: https://reviews.llvm.org/D101896
      abe2c906
    • Tim Renouf's avatar
      [llvm-objdump] Use std::make_unique · ab5932ff
      Tim Renouf authored
      Fix up my recent commit rG1128311a to
      use std::make_unique instead of std::unique_ptr(new), as requested by
      David Blaikie.
      
      Differential Revision: https://reviews.llvm.org/D101822
      ab5932ff
    • Guillaume Chatelet's avatar
    • Guillaume Chatelet's avatar
    • Guillaume Chatelet's avatar
    • Guillaume Chatelet's avatar
    • Guillaume Chatelet's avatar
      b4795544
    • Johannes Doerfert's avatar
      [OpenMP] Overhaul `declare target` handling · df729e2b
      Johannes Doerfert authored
      This patch fixes various issues with our prior `declare target` handling
      and extends it to support `omp begin declare target` as well.
      
      This started with PR49649 in mind, trying to provide a way for users to
      avoid the "ref" global use introduced for globals with internal linkage.
      From there it went down the rabbit hole, e.g., all variables, even
      `nohost` ones, were emitted into the device code so it was impossible to
      determine if "ref" was needed late in the game (based on the name only).
      To make it really useful, `begin declare target` was needed as it can
      carry the `device_type`. Not emitting variables eagerly had a ripple
      effect. Finally, the precedence of the (explicit) declare target list
      items needed to be taken into account, that meant we cannot just look
      for any declare target attribute to make a decision. This caused the
      handling of functions to require fixup as well.
      
      I tried to clean up things while I was at it, e.g., we should not "parse
      declarations and defintions" as part of OpenMP parsing, this will always
      break at some point. Instead, we keep track what region we are in and
      act on definitions and declarations instead, this is what we do for
      declare variant and other begin/end directives already.
      
      Highlights:
        - new diagnosis for restrictions specificed in the standard,
        - delayed emission of globals not mentioned in an explicit
          list of a declare target,
        - omission of `nohost` globals on the host and `host` globals on the
          device,
        - no explicit parsing of declarations in-between `omp [begin] declare
          variant` and the corresponding end anymore, regular parsing instead,
        - precedence for explicit mentions in `declare target` lists over
          implicit mentions in the declaration-definition-seq, and
        - `omp allocate` declarations will now replace an earlier emitted
          global, if necessary.
      
      ---
      
      Notes:
      
      The patch is larger than I hoped but it turns out that most changes do
      on their own lead to "inconsistent states", which seem less desirable
      overall.
      
      After working through this I feel the standard should remove the
      explicit declare target forms as the delayed emission is horrible.
      That said, while we delay things anyway, it seems to me we check too
      often for the current status even though that is often not sufficient to
      act upon. There seems to be a lot of duplication that can probably be
      trimmed down. Eagerly emitting some things seems pretty weak as an
      argument to keep so much logic around.
      
      ---
      
      Reviewed By: ABataev
      
      Differential Revision: https://reviews.llvm.org/D101030
      df729e2b
    • Johannes Doerfert's avatar
      [OpenMP] Ensure the DefaultMapperId has a location · 3f145967
      Johannes Doerfert authored
      A user reported an assertion (below) but without a reproducer. I failed to
      create a test myself but from the assertion one can derive the problem.
      I set the DefaultMapperId location now to make sure this doesn't cause
      trouble.
      
      ```
      clang-13: .../DeclTemplate.h:1940:
      void clang::ClassTemplateSpecializationDecl::setPointOfInstantiation(clang::SourceLocation):
      Assertion `Loc.isValid() && "point of instantiation must be valid!"' failed.
      ```
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D100621
      3f145967
    • Johannes Doerfert's avatar
      [OpenMP] Make sure classes work on the device as they do on the host · 5d8d994d
      Johannes Doerfert authored
      We do provide `operator delete(void*)` in `<new>` but it should be
      available by default. This is mostly boilerplate to test it and the
      unconditional include of `<new>` in the header we always in include
      on the device.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D100620
      5d8d994d
    • Navdeep Kumar's avatar
      [MLIR][GPU][NVVM] Add warp synchronous matrix-multiply accumulate ops · 875eb523
      Navdeep Kumar authored
      Add warp synchronous matrix-multiply accumulate ops in GPU and NVVM
      dialect. Add following three ops to GPU dialect :-
        1.) subgroup_mma_load_matrix
        2.) subgroup_mma_store_matrix
        3.) subgroup_mma_compute
      Add following three ops to NVVM dialect :-
        1.) wmma.m16n16k16.load.[a,b,c].[f16,f32].row.stride
        2.) wmma.m16n16k16.store.d.[f16,f32].row.stride
        3.) wmma.m16n16k16.mma.row.row.[f16,f32].[f16,f32]
      
      Reviewed By: bondhugula, ftynse, ThomasRaoux
      
      Differential Revision: https://reviews.llvm.org/D95330
      875eb523
    • Queen Dela Cruz's avatar
      [clangd] Check if macro is already in the IdentifierTable before loading it · 16c78297
      Queen Dela Cruz authored
      Having nested macros in the C code could cause clangd to fail an assert in clang::Preprocessor::setLoadedMacroDirective() and crash.
      
       #1 0x00000000007ace30 PrintStackTraceSignalHandler(void*) /qdelacru/llvm-project/llvm/lib/Support/Unix/Signals.inc:632:1
       #2 0x00000000007aaded llvm::sys::RunSignalHandlers() /qdelacru/llvm-project/llvm/lib/Support/Signals.cpp:76:20
       #3 0x00000000007ac7c1 SignalHandler(int) /qdelacru/llvm-project/llvm/lib/Support/Unix/Signals.inc:407:1
       #4 0x00007f096604db20 __restore_rt (/lib64/libpthread.so.0+0x12b20)
       #5 0x00007f0964b307ff raise (/lib64/libc.so.6+0x377ff)
       #6 0x00007f0964b1ac35 abort (/lib64/libc.so.6+0x21c35)
       #7 0x00007f0964b1ab09 _nl_load_domain.cold.0 (/lib64/libc.so.6+0x21b09)
       #8 0x00007f0964b28de6 (/lib64/libc.so.6+0x2fde6)
       #9 0x0000000001004d1a clang::Preprocessor::setLoadedMacroDirective(clang::IdentifierInfo*, clang::MacroDirective*, clang::MacroDirective*) /qdelacru/llvm-project/clang/lib/Lex/PPMacroExpansion.cpp:116:5
      
      An example of the code that causes the assert failure:
      ```
      ...
      ```
      
      During code completion in clangd, the macros will be loaded in loadMainFilePreambleMacros() by iterating over the macro names and calling PreambleIdentifiers->get(). Since these macro names are store in a StringSet (has StringMap underlying container), the order of the iterator is not guaranteed to be same as the order seen in the source code.
      
      When clangd is trying to resolve nested macros it sometimes attempts to load them out of order which causes a macro to be stored twice. In the example above, ECHO2 macro gets resolved first, but since it uses another macro that has not been resolved it will try to resolve/store that as well. Now there are two MacroDirectives stored in the Preprocessor, ECHO and ECHO2. When clangd tries to load the next macro, ECHO, the preprocessor fails an assert in clang::Preprocessor::setLoadedMacroDirective() because there is already a MacroDirective stored for that macro name.
      
      In this diff, I check if the macro is already inside the IdentifierTable and if it is skip it so that it is not resolved twice.
      
      Reviewed By: kadircet
      
      Differential Revision: https://reviews.llvm.org/D101870
      16c78297
    • Giorgis Georgakoudis's avatar
      [OpenMP][NFC] Refactor Clang OpenMP tests using update_cc_test_checks · 207b08a9
      Giorgis Georgakoudis authored
      This patch refactors a subset of Clang OpenMP tests, generating checklines using the update_cc_test_checks script. This refactoring facilitates updating the Clang OpenMP code generation codebase by automating test generation.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D101849
      207b08a9
    • Jessica Clarke's avatar
      [SelectionDAG][Mips][PowerPC][RISCV][WebAssembly] Teach... · 6c80361b
      Jessica Clarke authored
      [SelectionDAG][Mips][PowerPC][RISCV][WebAssembly] Teach computeKnownBits/ComputeNumSignBits about atomics
      
      Unlike normal loads these don't have an extension field, but we know
      from TargetLowering whether these are sign-extending or zero-extending,
      and so can optimise away unnecessary extensions.
      
      This was noticed on RISC-V, where sign extensions in the calling
      convention would result in unnecessary explicit extension instructions,
      but this also fixes some Mips inefficiencies. PowerPC sees churn in the
      tests as all the zero extensions are only for promoting 32-bit to
      64-bit, but these zero extensions are still not optimised away as they
      should be, likely due to i32 being a legal type.
      
      This also simplifies the WebAssembly code somewhat, which currently
      works around the lack of target-independent combines with some ugly
      patterns that break once they're optimised away.
      
      Re-landed with correct handling in ComputeNumSignBits for Tmp == VTBits,
      where zero-extending atomics were incorrectly returning 0 rather than
      the (slightly confusing) required return value of 1.
      
      Reviewed By: RKSimon, atanasyan
      
      Differential Revision: https://reviews.llvm.org/D101342
      6c80361b
    • Jinsong Ji's avatar
      [BPF][Test] Disable codegen test on AIX · 6bdfcb16
      Jinsong Ji authored
      https://reviews.llvm.org/D101194 changed the default getMultiarchTriple in toolchain.
      So -march=bpf on AIX will get triple of bpf-ibm-aix now,
      this is unexpected and causing test failures.
      
      BPF on AIX is not supported (yet), disable the codegen test on AIX in lit cfg.
      
      Reviewed By: yonghong-song
      
      Differential Revision: https://reviews.llvm.org/D101866
      6bdfcb16
    • Lang Hames's avatar
      abdd14a2
    • Giorgis Georgakoudis's avatar
      [OpenMP] Fix non-determinism in clang copyin codegen · f97b843d
      Giorgis Georgakoudis authored
      Codegen for OpeMP copyin has non-deterministic IR output due to the unspecified evaluation order in a codegen conditional branch, which makes automatic test generation unreliable. This patch refactors codegen code to avoid this non-determinism.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D101952
      f97b843d
    • Lang Hames's avatar
      [ORC] Introduce C API for adding object buffers directly to an object layer. · 7b73cd68
      Lang Hames authored
      This can be useful for clients constructing custom JIT stacks: If the C API
      for your custom stack exposes API to obtain a reference to an object layer
      (e.g. LLVMOrcLLJITGetObjLinkingLayer) then the newly added
      LLVMOrcObjectLayerAddObjectFile and LLVMOrcObjectLayerAddObjectFileWithRT
      functions can be used to add objects directly to that layer.
      7b73cd68
    • Christopher Ferris's avatar
      [scudo] Add initialization for TSDRegistrySharedT · 6fac3425
      Christopher Ferris authored
      Fixes compilation on Android which has a TSDSharedRegistry object in the config.
      
      Reviewed By: cryptoad, vitalybuka
      
      Differential Revision: https://reviews.llvm.org/D101951
      6fac3425