1. Jul 29, 2022
  2. Jul 28, 2022
    • Philip Reames's avatar
      [LV] Don't predicate uniform mem op stores unneccessarily · 82c1b136
      Philip Reames authored
      We already had the reasoning about uniform mem op loads; if the address is accessed at least once, we know the instruction doesn't need predicated to ensure fault safety. For stores, we do need to ensure that the values visible in memory are the same with and without predication. The easiest sub-case to check for is that all the values being stored are the same. Since we know that at least one lane is active, this tells us that the value must be visible.
      
      Warning on confusing terminology: "uniform" vs "uniform mem op" mean two different things here, and this patch is specific to the later. It would *not* be legal to make this same change for merely "uniform" operations.
      
      Differential Revision: https://reviews.llvm.org/D130637
      82c1b136
    • Jon Chesterfield's avatar
    • Prabhdeep Singh Soni's avatar
      [Flang][MLIR][OpenMP] Add support for simdlen clause · f5efa189
      Prabhdeep Singh Soni authored
      This supports lowering from parse-tree to MLIR and translation from
      MLIR to LLVM IR using OMPIRBuilder for OpenMP simdlen clause in SIMD
      construct.
      
      Reviewed By: shraiysh, peixin, arnamoy10
      
      Differential Revision: https://reviews.llvm.org/D130195
      f5efa189
    • Jon Chesterfield's avatar
    • Jon Chesterfield's avatar
      [openmp] Introduce optional plugin init/deinit functions · 1f9d3974
      Jon Chesterfield authored
      Will allow plugins to migrate away from using global variables to
      manage lifetime, which will fix a segfault discovered in relation to D127432
      
      Reviewed By: jhuber6
      
      Differential Revision: https://reviews.llvm.org/D130712
      1f9d3974
    • LLVM GN Syncbot's avatar
      [gn build] Port d52e775b · 59ea2c64
      LLVM GN Syncbot authored
      59ea2c64
    • Liqiang Tao's avatar
      [llvm][ModuleInliner] Add inline cost priority for module inliner · d52e775b
      Liqiang Tao authored
      This patch introduces the inline cost priority into the
      module inliner, which uses the same computation as
      InlineCost.
      
      Reviewed By: kazu
      
      Differential Revision: https://reviews.llvm.org/D130012
      d52e775b
    • LLVM GN Syncbot's avatar
      [gn build] Port c1135943 · cf0196db
      LLVM GN Syncbot authored
      cf0196db
    • Liqiang Tao's avatar
      c1135943
    • Florian Hahn's avatar
      Revert "[X86][DAGISel] Don't widen shuffle element with AVX512" · f912bab1
      Florian Hahn authored
      This reverts commit 5fb41342.
      
      This patch is causing crashes when building llvm-test-suite when
      optimizing for CPUs with AVX512.
      
      Reproducer crashing with llc:
      
          target datalayout = "e-m:o-p270:32:32-p271:32:32-p272:64:64-i64:64-f80:128-n8:16:32:64-S128"
          target triple = "x86_64-apple-macosx"
      
          define i32 @test(<32 x i32> %0) #0 {
          entry:
            %1 = mul <32 x i32> %0, <i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1, i32 1>
            %2 = tail call i32 @llvm.vector.reduce.add.v32i32(<32 x i32> %1)
            ret i32 %2
          }
      
          ; Function Attrs: nocallback nofree nosync nounwind readnone willreturn
          declare i32 @llvm.vector.reduce.add.v32i32(<32 x i32>) #1
      
          attributes #0 = { "min-legal-vector-width"="0" "target-cpu"="skylake-avx512" }
          attributes #1 = { nocallback nofree nosync nounwind readnone willreturn }
      f912bab1
    • Simon Pilgrim's avatar
      [DAG] DAGCombiner::visitTRUNCATE - remove GetDemandedBits call · be488ba7
      Simon Pilgrim authored
      This should now all be handled by SimplifyDemandedBits.
      be488ba7
    • Chris Bieneman's avatar
      [HLSL] Add __builtin_hlsl_create_handle · fe13002b
      Chris Bieneman authored
      This is pretty straightforward, it just adds a builtin to return a
      pointer to a resource handle. This maps to a dx intrinsic.
      
      The shape of this builtin and the underlying intrinsic will likely
      shift a bit as this implementation becomes more feature complete, but
      this is a good basis to get started.
      
      Depends on D128569.
      
      Differential Revision: https://reviews.llvm.org/D130016
      fe13002b
    • Chris Bieneman's avatar
      Start support for HLSL `RWBuffer` · 6e56d0db
      Chris Bieneman authored
      Most of the change here is fleshing out the HLSLExternalSemaSource with
      builder implementations to build the builtin types. Eventually, I may
      move some of this code into tablegen or a more managable declarative
      file but I want to get the AST generation logic ready first.
      
      This code adds two new types into the HLSL AST, `hlsl::Resource` and
      `hlsl::RWBuffer`. The `Resource` type is just a wrapper around a handle
      identifier, and is largely unused in source. It will morph a bit over
      time as I work on getting the source compatability correct, but for now
      it is a reasonable stand-in. The `RWBuffer` type is not ready for use.
      I'm posting this change for review because it adds a lot of
      infrastructure code and is testable.
      
      There is one change to clang code outside the HLSL-specific logic here,
      which addresses a behavior change introduced a long time ago in
      967d4384. That change resulted in unintentionally breaking
      situations where an incomplete template declaration was provided from
      an AST source, and needed to be completed later by the external AST.
      That situation doesn't happen in the normal AST importer flow, but can
      happen when an AST source provides incomplete declarations of
      templates. The solution is to annotate template specializations of
      incomplete types with the HasExternalLexicalSource bit from the base
      template.
      
      Depends on D128012.
      
      Differential Revision: https://reviews.llvm.org/D128569
      6e56d0db
    • Sunho Kim's avatar
      [clang-repl] Disable exception unittest on AIX. · bd08f413
      Sunho Kim authored
      AIX platform was not supported but it was not explicitly checked in exception test as it was excluded by isPPC() check.
      bd08f413
    • Simon Pilgrim's avatar
      [DAG] SelectionDAG::GetDemandedBits - don't simplify opaque constants · ea7f14da
      Simon Pilgrim authored
      I'm actually trying to get rid of GetDemandedBits - but while dismantling it I noticed that we were altering opaque constants. Fixing that causes a FP_TO_INT_SAT regression that should be addressed separately - I'll raise a bug.
      ea7f14da
    • LLVM GN Syncbot's avatar
      [gn build] Port bb7f62bb · e2938024
      LLVM GN Syncbot authored
      e2938024
    • Liqiang Tao's avatar
      [llvm][ModuleInliner] Add inline cost priority for module inliner · bb7f62bb
      Liqiang Tao authored
      This patch introduces the inline cost priority into the
      module inliner, which uses the same computation as
      InlineCost.
      
      Reviewed By: kazu
      
      Differential Revision: https://reviews.llvm.org/D130012
      bb7f62bb
    • David Green's avatar
      [ARM] Remove duplicate fp16 intrinsics · 3b09e532
      David Green authored
      These vdup and vmov float16 intrinsics are being defined in both the
      general section and then again in fp16 under a !aarch64 flag. The
      vdup_lane intrinsics were being defined in both aarch64 and !aarch64
      sections, so have been commoned.  They are defined as macros, so do not
      give duplicate warnings, but removing the duplicates shouldn't alter the
      available intrinsics.
      3b09e532
    • Simon Pilgrim's avatar
      [DAG] Enable ISD::SRL SimplifyMultipleUseDemandedBits handling inside SimplifyDemandedBits · 69d5a038
      Simon Pilgrim authored
      This patch allows SimplifyDemandedBits to call SimplifyMultipleUseDemandedBits in cases where the ISD::SRL source operand has other uses, enabling us to peek through the shifted value if we don't demand all the bits/elts.
      
      This is another step towards removing SelectionDAG::GetDemandedBits and just using TargetLowering::SimplifyMultipleUseDemandedBits.
      
      There a few cases where we end up with extra register moves which I think we can accept in exchange for the increased ILP.
      
      Differential Revision: https://reviews.llvm.org/D77804
      69d5a038
    • Kevin P. Neal's avatar
      Precommit tests for D112256 "[FPEnv][EarlyCSE] Add support for CSE of... · 25a83005
      Kevin P. Neal authored
      Precommit tests for D112256 "[FPEnv][EarlyCSE] Add support for CSE of constrained FP intrinsics, take 2"
      25a83005
    • Amaury Séchet's avatar
      [DAG] Use recursivelyDeleteUnusedNodes in PromoteLoad · 474a8ee0
      Amaury Séchet authored
      It simplifies the code overall and removes the need for manual bookkeeping.
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D130447
      474a8ee0
    • Sebastian Neubauer's avatar
      [CMake][OpenMP] Remove wrong backslash · 50716ba2
      Sebastian Neubauer authored
      outdir is defined in the line above, it will not exist in the install
      command, so it should not be escaped.
      50716ba2
    • Amaury Séchet's avatar
      [DAG] Use recursivelyDeleteUnusedNodes in ReplaceLoadWithPromotedLoad · 7920805b
      Amaury Séchet authored
      It simplifies the code overall and removes the need for manual bookkeeping.
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D130444
      7920805b
    • Alexander Timofeev's avatar
      [AMDGPU] avoid blind converting to VALU REG_SEQUENCE and PHIs · 76d9ae92
      Alexander Timofeev authored
      In the 2e29b013 we introduce a specific solving algorithm
      that analyzes the VGPR to SGPR copies use chains and either lowers
      the copy to v_readfirstlane_b32 or converts the whole chain to VALU forms.
      Same time we still have the code that blindly converts to VALU REG_SEQUENCE and PHIs
      in case they produce SGPR but have VGPRs input operands. In case the REG_SEQUENCE and PHIs
      are in the VGPR to SGPR copy use chain, and this chain was considered long enough to convert
      copy to v_readfistlane_b32, further lowering them to VALU leads to several kinds of issues.
      At first, we have v_readfistlane_b32 which is completely useless because most parts of its use chain
      were moved to VALU forms. Second, we may encounter subtle bugs related to the EXEC-dependent CF
      because of the weird mixing of SALU and VALU instructions.
      This change removes the code that moves REG_SEQUENCE and PHIs to VALU. Instead, we use the fact
      that both REG_SEQUENCE and PHIs have copy semantics. That is, if they define SGPR but have VGPR inputs,
      we insert VGPR to SGPR copies to make them pure SGPR. Then, the new copies are processed by the common
      VGPR to SGPR lowering algorithm.
      This is Part 2 in the series of commits aiming at the massive refactoring of the SIFixSGPRCopies pass.
      
      Reviewed By: rampitec
      
      Differential Revision: https://reviews.llvm.org/D130367
      76d9ae92
    • Sunho Kim's avatar
      [clang-repl] Add host exception support check utility flag. · 3cc3be8f
      Sunho Kim authored
      Add host exception support check utility flag. This is needed to not run tests that require exception support in few buildbots that lacks related symbols for some reason.
      
      Reviewed By: lhames
      
      Differential Revision: https://reviews.llvm.org/D129242
      3cc3be8f
    • Sunho Kim's avatar
      [ORC] Fix weak hidden symbols failure on PPC with runtimedyld · 72ea1a72
      Sunho Kim authored
      Fix "JIT session error: Symbols not found: [ DW.ref.__gxx_personality_v0 ] error" which happens when trying to use exceptions on ppc linux. To do this, it expands AutoClaimSymbols option in RTDyldObjectLinkingLayer to also claim weak symbols before they are tried to be resovled. In ppc linux, DW.ref symbols is emitted as weak hidden symbols in the later stage of MC pipeline. This means when using IRLayer (i.e. LLJIT), IRLayer will not claim responsibility for such symbols and RuntimeDyld will skip defining this symbol even though it couldn't resolve corresponding external symbol.
      
      Reviewed By: sgraenitz
      
      Differential Revision: https://reviews.llvm.org/D129175
      72ea1a72
    • Muhammad Usman Shahid's avatar
      Missing tautological compare warnings due to unary operators · 0cc3c184
      Muhammad Usman Shahid authored
      The patch mainly focuses on the lack of warnings for
      -Wtautological-compare. It works fine for positive numbers but doesn't
      for negative numbers. This is because the warning explicitly checks for
      an IntegerLiteral AST node, but -1 is represented by a UnaryOperator
      with an IntegerLiteral sub-Expr.
      
      For the below code we have warnings:
      
      if (0 == (5 | x)) {}
      
      but not for
      
      if (0 == (-5 | x)) {}
      
      This patch changes the analysis to not look at the AST node directly to
      see if it is an IntegerLiteral, but instead attempts to evaluate the
      expression to see if it is an integer constant expression. This handles
      unary negation signs, but also handles all the other possible operators
      as well.
      
      Fixes #42918
      Differential Revision: https://reviews.llvm.org/D130510
      0cc3c184
    • Dmitry Preobrazhensky's avatar
      [AMDGPU][GFX1030][DOC][NFC] Update assembler syntax description · 955cc56a
      Dmitry Preobrazhensky authored
      Summary of changes:
      - Update FLAT LDS syntax (see https://reviews.llvm.org/D125126)
      955cc56a