1. Aug 22, 2023
    • Lorenzo Chelini's avatar
      [MLIR][Linalg] Respect DPS in `lower_unpack` · d2f2ef84
      Lorenzo Chelini authored
      `tensor.unpack` implements the DPS (Destination Passing Style) interface
      and expects the result to be "stored" in the `outs` operand, but this is
      not the case with the current decomposition as the final operation is a
      `tensor.extract_slice` that does not implement DPS. Add a `linalg.copy`
      to fix the problem.
      
      Reviewed By: springerm
      
      Differential Revision: https://reviews.llvm.org/D158393
      d2f2ef84
    • Nikita Popov's avatar
      [IR] Remove support for and/or constant expressions · 62511340
      Nikita Popov authored
      As part of https://discourse.llvm.org/t/rfc-remove-most-constant-expressions/63179,
      this removes support for and and or constant expressions. Places
      creating such expressions have been migrated in advance, so this
      is mostly API removal and test updates.
      
      Differential Revision: https://reviews.llvm.org/D155924
      62511340
    • Nikita Popov's avatar
      [SCEVExpander] Fix incorrect reuse of more poisonous instructions (PR63763) · 1c6e6432
      Nikita Popov authored
      SCEVExpander tries to reuse existing instruction with the same
      SCEV expression. However, doing this replacement blindly is not
      safe, because the instruction might be more poisonous.
      
      What we were already doing is to drop poison-generating flags on
      the reused instruction. But this is not the only way that more
      poison can be introduced. The poison-generating flag might not
      be directly on the reused instruction, or the poison contribution
      might come from something like 0 * %var, which folds to 0 but can
      still introduce poison.
      
      This patch fixes the issue in a principled way, by determining which
      values can contribute poison to the SCEV expression, and then
      checking whether any additional values can contribute poison to the
      instruction being reused. Poison-generating flags are dropped if
      doing that enables reuse.
      
      This is a pretty big hammer and does cause some regressions in
      tests, but less than I would have expected. I wasn't able to come
      up with a less intrusive fix that still satisfies the correctness
      requirements.
      
      Fixes https://github.com/llvm/llvm-project/issues/63763.
      Fixes https://github.com/llvm/llvm-project/issues/63926.
      Fixes https://github.com/llvm/llvm-project/issues/64333.
      Fixes https://github.com/llvm/llvm-project/issues/63727.
      
      Differential Revision: https://reviews.llvm.org/D158181
      1c6e6432
    • pvanhout's avatar
      [GlobalISel] Rewrite some simple rules using MIR Patterns · 2d87319f
      pvanhout authored
      Rewrites some simple rules that cause little to no codegen regressions as MIR patterns.
      
      I may have missed some easy cases, but some other rules have intentionally been left as-is because bigger
      changes are needed to make them work.
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D157690
      2d87319f
    • Jim Lin's avatar
    • wangpc's avatar
      [docs] Add minutes/docs of RISC-V sync-up call · fa51188f
      wangpc authored
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D158388
      fa51188f
    • Matthias Springer's avatar
      [mlir][bufferization] Improve `bufferizesToElementwiseAccess` · f36e1934
      Matthias Springer authored
      The operands for which elementwise access is relevant can now be specified. All other operands are ignored. This is useful because only two particular operands participate in a RaW conflict. Furthermore, the two tensors no longer must be equivalent to rule out conflicts due to elementwise access. Equivalent tensor sets may be formed after an inplace bufferization decision is made. The two tensors are actually not required to be equivalent. The only important thing is that they have "equivalent" indexing into the same base buffer.
      
      Differential Revision: https://reviews.llvm.org/D158428
      f36e1934
    • Fangrui Song's avatar
      [X86] Clean up GlobalISel headers. NFC · bf6e3936
      Fangrui Song authored
      bf6e3936
    • Congcong Cai's avatar
      [clang-tidy][readability-braces-around-statements] ignore false-positive for... · 9b6859dc
      Congcong Cai authored
      [clang-tidy][readability-braces-around-statements] ignore false-positive for constexpr if statement in lambda expression
      
      Fixed: #64545
      
      When TreeTransform, Stmt in constexpr IfStmt will be transform to NullStmt.
      This NullStmt has the different beginning token.
      This patch add addtional check in checkStmt to handle this case.
      
      Reviewed By: PiotrZSL
      
      Differential Revision: https://reviews.llvm.org/D158480
      9b6859dc
    • Christian Sigg's avatar
      097efdd6
    • Christian Sigg's avatar
    • Christian Ulmann's avatar
      [mlir][LLVM] Fix export call mapping for calls with a result · fdaf2f9c
      Christian Ulmann authored
      This commit adds a missed update of the call mapping in the LLVM export
      for calls with no result. Before, these calls were not inserted and
      thus, the export dropped branch weights on them.
      
      Reviewed By: zero9178
      
      Differential Revision: https://reviews.llvm.org/D158453
      fdaf2f9c
    • Mikael Holmen's avatar
      [test] Add -verify-coalescing to testcase and fix problems · d1e685df
      Mikael Holmen authored
      Apparently the testcase
       coalesce-partial-redundant-reguse-terminator.mir
      was broken in a way that -verify-coalescing detected.
      
      Update the testcase so -verify-coalescing doesn't complain and so
      that it still exposes the problem originally fixed in 6c062b76.
      
      Differential Revision: https://reviews.llvm.org/D158397
      d1e685df
    • huqizhi's avatar
      [clang][ASTImporter]Skip check depth of friend template parameter · 07ab5140
      huqizhi authored
      Depth of the parameter of friend template class declaration in a
      template class is 1, while in the specialization the depth is 0.
      This will cause failure on 'IsStructurallyEquivalent' as a name
      conflict in 'VisitClassTemplateDecl'(see testcase of
      'SkipComparingFriendTemplateDepth'). The patch fix it by ignore
      the depth only in this special case.
      
      Reviewed By: balazske
      
      Differential Revision: https://reviews.llvm.org/D156693
      07ab5140
    • Christian Sigg's avatar
      3eba48b8
    • Kai Luo's avatar
    • Congcong Cai's avatar
      [clang-tidy]mark record initList as non-const param · 1c941244
      Congcong Cai authored
      ```
      struct XY {
        int *x;
        int *y;
      };
      void recordInitList(int *x) {
        XY xy = {x, nullptr};
      }
      ```
      x cannot be const int* becase it in a initialize list which only accept int*
      
      Reviewed By: PiotrZSL
      
      Differential Revision: https://reviews.llvm.org/D158152
      1c941244
    • Douglas Yung's avatar
      a0db7385
    • Fangrui Song's avatar
      [CSKY] Adjust includes in MCTargetDesc to avoid unnecessary CodeGen deps, NFC · 79df15b7
      Fangrui Song authored
      See issue llvm#64166 for more information about layering.
      79df15b7
    • Slava Zakharin's avatar
      [flang] Fixed simplification for FP maxval. · 89b98c13
      Slava Zakharin authored
      On x86, a simplified F128 maxval ends up calling fmaxl that does not
      work properly for F128 arguments. It is probably an LLVM issue, but
      we also should not use arith.maxf if NaN or -0.0 operands are possible.
      The change is to use cmpf and select. Unfortunately, these arith ops
      do not support FastMathFlags currently, so I will have to fix this
      sooner or later (depending on how this affects performance).
      
      Reviewed By: kiranchandramohan
      
      Differential Revision: https://reviews.llvm.org/D158200
      89b98c13
    • YunQiang Su's avatar
      Clang/Gnu: Scan GCC with triple without vendor if vendor is unknown · 6ee3b244
      YunQiang Su authored
      If ScanLibDirForGCCTriple with target triple fails, let's try the triple with
      vendor stripped if vendor is unknown.
      
      Debian always uses triples without a vendor section. In general, triples without
      a vendor section is the most similar aliases than any other aliases.
      
      To archive this, we add a private member TripleNoVendor to
      GCCInstallationDetector.
      
      This modification makes testcases riscv32-toolchain.c and riscv64-toolchain.c
      fail. The reason is that they are wrong: --triple riscv64-unknown-elf tries
      to use riscv64-unknown-linux-gnu first.
      
      This patch accidentally fixes this problem.
      
      We also drop the path delimiter pattern {{/|\\\\}}, as these 2 tests are
      disabled on Windows, and in fact the positions of this pattern are not
      correct.
      
      Reviewed by: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D158183
      6ee3b244
    • Chuanqi Xu's avatar
      [C++20] [Coroutines] Mark await_suspend as noinline if the awaiter is not empty · c4672454
      Chuanqi Xu authored
      Close https://github.com/llvm/llvm-project/issues/56301
      Close https://github.com/llvm/llvm-project/issues/64151
      
      See the summary and the discussion of https://reviews.llvm.org/D157070
      to get the full context.
      
      As @rjmccall pointed out, the key point of the root cause is that
      currently we didn't implement the semantics for '@llvm.coro.save' well
      ("after the await-ready returns false, the coroutine is considered to be
      suspended ") well.
      Since the semantics implies that we (the compiler) shouldn't write the
      spills into the coroutine frame in the await_suspend. But now it is possible
      due to some combinations of the optimizations so the semantics are
      broken. And the inlining is the root optimization of such optimizations.
      So in this patch, we tried to add the `noinline` attribute to the
      await_suspend call.
      
      Also as an optimization, we don't add the `noinline` attribute to the
      await_suspend call if the awaiter is an empty class. Thi...
      c4672454
    • Nicolas Vasilache's avatar
      [mlir] Disentangle dialect and extension registrations. · 7c4e8c6a
      Nicolas Vasilache authored
      This revision avoids the registration of dialect extensions in Pass::getDependentDialects.
      
      Such registration of extensions can be dangerous because `DialectRegistry::isSubsetOf` is
      always guaranteed to return false for extensions (i.e. there is no mechanism to track
      whether a lambda is already in the list of already registered extensions).
      When the context is already in a multi-threaded mode, this is guaranteed to assert.
      
      Arguably a more structured registration mechanism for extensions with a unique ExtensionID
      could be envisioned in the future.
      
      In the process of cleaning this up, multiple usage inconsistencies surfaced around the
      registration of translation extensions that this revision also cleans up.
      
      Reviewed By: springerm
      
      Differential Revision: https://reviews.llvm.org/D157703
      7c4e8c6a
    • Jennifer Yu's avatar
      [SEH] Fix wrong argument passes to the call of OutlinedFinally. · ff08c8e5
      Jennifer Yu authored
      When return out of __try block.  In this test case, currently "false" is
      passed to OutlinedFinally's call.  "true" should be passed to indicate
      abnormal terminations.
      
      The rule: Except _leave and fall-through at the end, all other exits
      in a _try (return/goto/continue/break) are considered as abnormal
      terminations, NormalCleanupDestSlot is used to indicate abnormal
      terminations.
      
      The problem is, during the processing abnormal terminations,
      the ExitSwitch is used.  However, in this case, Existswitch is route out.
      
      One way to fix is to skip route it without a switch. So proper
      abnormal termination's code could be generated.
      
      Differential Revision: https://reviews.llvm.org/D158233
      ff08c8e5
    • Ziqing Luo's avatar
      [-Wunsafe-buffer-usage] Stop generating incorrect fix-its for variable... · b58e5288
      Ziqing Luo authored
      [-Wunsafe-buffer-usage] Stop generating incorrect fix-its for variable declarations with unsupported specifiers
      
      We have to give up on fixing a variable declaration if it has
      specifiers that are not supported yet.  We could support these
      specifiers incrementally using the same approach as how we deal with
      cv-qualifiers. If a fixing variable declaration has a storage
      specifier, instead of trying to find out the source location of the
      specifier or to avoid touching it, we add the keyword to a
      canonicalized place in the fix-it text that replaces the whole
      declaration.
      
      Reviewed by: NoQ (Artem Dergachev), jkorous (Jan Korous)
      
      Differential revision: https://reviews.llvm.org/D156192
      b58e5288
    • Fangrui Song's avatar
      Use InernalAlloc in DemangleCXXABI · 649004ae
      Fangrui Song authored
      This reverts commit 06c74b5e.
      Tested on AddressSanitizer-arm64-darwin that there is no more failure.
      649004ae
    • Gabor Horvath's avatar
      8330116e
    • Christopher Ferris's avatar
      [scudo] Fix definition of SCUDO_SMALL_STACK_DEPOT. · 41a27532
      Christopher Ferris authored
      The SCUDO_FUZZ macro is either defined or not defined. The previous
      code assumed it had a one or zero value, so change the setting of
      SCUDO_SMALL_STACK_DEPOT based on defined(SCUDO_FUZZ).
      
      Reviewed By: Chia-hungDuan
      
      Differential Revision: https://reviews.llvm.org/D158459
      41a27532
    • Ziqing Luo's avatar
      [-Wunsafe-buffer-usage] Refactor to let local variable fix-its and parameter... · 3a67b912
      Ziqing Luo authored
      [-Wunsafe-buffer-usage] Refactor to let local variable fix-its and parameter fix-its share common code
      
      Refactor the code for local variable fix-its so that it reuses the
      code for parameter fix-its, which is in general better. For example,
      cv-qualifiers are supported.
      
      Reviewed by: NoQ (Artem Dergachev), t-rasmud (Rashmi Mudduluru)
      
      Differential revision: https://reviews.llvm.org/D156189
      3a67b912
    • Justin Bogner's avatar
      [DXILBitcodeWriter] Fix handling of an unspecified lower bound in DISubrange · d7f3b238
      Justin Bogner authored
      If the lower bound isn't specified it implies that it's zero.
      
      Differential Revision: https://reviews.llvm.org/D158441
      d7f3b238
    • Craig Topper's avatar
      [RISCV] Check type size for lax conversions between RVV builtin types and... · 33af2f13
      Craig Topper authored
      [RISCV] Check type size for lax conversions between RVV builtin types and VectorType::RVVFixedLengthDataVector.
      
      This code was copied from SVE and modified for RVV. For SVE, there
      is only one size for builtin types so they didn't need to check
      the size. For RVV, due to LMUL there are 7 different sizes of builtin
      types so we do need to check the size.
      
      I'm not sure we should have lax vector conversions at all for RVV.
      That appears to be contributing to https://github.com/llvm/llvm-project/issues/64404
      
      This patch at least fixes the obvious correctness issue.
      This should be backported to LLVM 17.
      
      Reviewed By: jacquesguan
      
      Differential Revision: https://reviews.llvm.org/D157130
      33af2f13
    • Ashley Nelson's avatar
      [WebAssembly] Add multiple memories feature · 86ed8cb8
      Ashley Nelson authored
      Adding to allow users to get this flag into the target features section for
      future use cases.
      
      Reviewed By: tlively, aheejin
      
      Differential Revision: https://reviews.llvm.org/D158409
      86ed8cb8
    • Krzysztof Drewniak's avatar
      [mlir][MemRefToLLVM] Add fmin, fmax to AtomicRMW lowering · 7db18533
      Krzysztof Drewniak authored
      Add cases to the memref.atomicrmw lowering for floating-point min and
      max, since LLVM supports these.
      
      Reviewed By: bondhugula
      
      Differential Revision: https://reviews.llvm.org/D158283
      7db18533
    • Nicole Mazzuca's avatar
      ASan: Add additional wcs* interceptors on Windows · c0c83668
      Nicole Mazzuca authored
      This adds wcs[n]cat, wcs[n]cmp, wcs[n]cpy, and wcschr functions to the
      interception code on Windows; wcs[n]cat was already intercepted, but only on
      POSIX.
      
      Differential Revision: https://reviews.llvm.org/D157038
      c0c83668
    • Eduard Zingerman's avatar
      [BPF] Replace BPFMIPeepholeTruncElim by custom logic in isZExtFree() · 651e6445
      Eduard Zingerman authored
      Replace `BPFMIPeepholeTruncElim` by adding an overload for
      `TargetLowering::isZExtFree()` aware that zero extension is
      free for `ISD::LOAD`.
      
      Short description
      =================
      
      The `BPFMIPeepholeTruncElim` handles two patterns:
      
      Pattern #1:
      
          %1 = LDB %0, ...              %1 = LDB %0, ...
          %2 = AND_ri %1, 0xff      ->  %2 = MOV_ri %1    <-- (!)
      
      Pattern #2:
      
          bb.1:                         bb.1:
            %a = LDB %0, ...              %a = LDB %0, ...
            br %bb3                       br %bb3
          bb.2:                         bb.2:
            %b = LDB %0, ...        ->    %b = LDB %0, ...
            br %bb3                       br %bb3
          bb.3:                         bb.3:
            %1 = PHI %a, %b               %1 = PHI %a, %b
            %2 = AND_ri %1, 0xff          %2 = MOV_ri %1  <-- (!)
      
      Plus variations:
      - AND_ri_32 instead of AND_ri
      - SLL/SLR instead of AND_ri
      - LDH, LDW, LDB32, LDH32, LDW32
      
      Both patterns could be handled by built-in transformations at
      instruction selection phase if suitable `isZExtFree()` implementation
      is provided. The idea is borrowed from `ARMTargetLowering::isZExtFree`.
      
      When evaluating on BPF kernel selftests and remove_truncate_*.ll LLVM
      test cases this revisions performs slightly better than
      BPFMIPeepholeTruncElim, see "Impact" section below for details.
      
      Commit also adds a few test cases to make sure that patterns in
      question are handled.
      
      Long description
      ================
      
      Why this works: Pattern #1
      --------------------------
      
      Consider the following example:
      
          define i1 @foo(ptr %p) {
          entry:
            %a = load i8, ptr %p, align 1
            %cond = icmp eq i8 %a, 0
            ret i1 %cond
          }
      
      Log for `llc -mcpu=v2 -mtriple=bpfel -debug-only=isel` command:
      
          ...
          Type-legalized selection DAG: %bb.0 'foo:entry'
          SelectionDAG has 13 nodes:
            t0: ch,glue = EntryToken
                    t2: i64,ch = CopyFromReg t0, Register:i64 %0
                  t16: i64,ch = load<(load (s8) from %ir.p), anyext from i8> t0, t2, undef:i64
                t19: i64 = and t16, Constant:i64<255>
              t17: i64 = setcc t19, Constant:i64<0>, seteq:ch
            t11: ch,glue = CopyToReg t0, Register:i64 $r0, t17
            t12: ch = BPFISD::RET_GLUE t11, Register:i64 $r0, t11:1
          ...
          Replacing.1 t19: i64 = and t16, Constant:i64<255>
          With: t16: i64,ch = load<(load (s8) from %ir.p), anyext from i8> t0, t2, undef:i64
           and 0 other values
          ...
          Optimized type-legalized selection DAG: %bb.0 'foo:entry'
          SelectionDAG has 11 nodes:
            t0: ch,glue = EntryToken
                  t2: i64,ch = CopyFromReg t0, Register:i64 %0
                t20: i64,ch = load<(load (s8) from %ir.p), zext from i8> t0, t2, undef:i64
              t17: i64 = setcc t20, Constant:i64<0>, seteq:ch
            t11: ch,glue = CopyToReg t0, Register:i64 $r0, t17
            t12: ch = BPFISD::RET_GLUE t11, Register:i64 $r0, t11:1
          ...
      
      Note:
      - Optimized type-legalized selection DAG:
        - `t19 = and t16, 255` had been replaced by `t16` (load).
        - Patterns like `(and (load ... i8), 255)` are replaced by `load`
          in `DAGCombiner::BackwardsPropagateMask` called from
          `DAGCombiner::visitAND`.
        - Similarly patterns like `(shl (srl ..., 56), 56)` are replaced by
          `(and ..., 255)` in `DAGCombiner::visitSRL` (this function is huge,
          look for `TLI.shouldFoldConstantShiftPairToMask()` call).
      
      Why this works: Pattern #2
      --------------------------
      
      Consider the following example:
      
          define i1 @foo(ptr %p) {
          entry:
            %a = load i8, ptr %p, align 1
            br label %next
      
          next:
            %cond = icmp eq i8 %a, 0
            ret i1 %cond
          }
      
      Consider log for `llc -mcpu=v2 -mtriple=bpfel -debug-only=isel` command.
      Log for first basic block:
      
          Initial selection DAG: %bb.0 'foo:entry'
          SelectionDAG has 9 nodes:
            t0: ch,glue = EntryToken
            t3: i64 = Constant<0>
                  t2: i64,ch = CopyFromReg t0, Register:i64 %1
                t5: i8,ch = load<(load (s8) from %ir.p)> t0, t2, undef:i64
              t6: i64 = zero_extend t5
            t8: ch = CopyToReg t0, Register:i64 %0, t6
          ...
          Replacing.1 t6: i64 = zero_extend t5
          With: t9: i64,ch = load<(load (s8) from %ir.p), zext from i8> t0, t2, undef:i64
           and 0 other values
          ...
          Optimized lowered selection DAG: %bb.0 'foo:entry'
          SelectionDAG has 7 nodes:
            t0: ch,glue = EntryToken
                t2: i64,ch = CopyFromReg t0, Register:i64 %1
              t9: i64,ch = load<(load (s8) from %ir.p), zext from i8> t0, t2, undef:i64
            t8: ch = CopyToReg t0, Register:i64 %0, t9
      
      Note:
      - Initial selection DAG:
        - `%a = load ...` is lowered as `t6 = (zero_extend (load ...))`
          w/o special `isZExtFree()` overload added by this commit
          it is instead lowered as `t6 = (any_extend (load ...))`.
        - The decision to generate `zero_extend` or `any_extend` is
          done in `RegsForValue::getCopyToRegs` called from
          `SelectionDAGBuilder::CopyValueToVirtualRegister`:
          - if `isZExtFree()` for load returns true `zero_extend` is used;
          - `any_extend` is used otherwise.
      - Optimized lowered selection DAG:
        - `t6 = (any_extend (load ...))` is replaced by
          `t9 = load ..., zext from i8`
          This is done by `DagCombiner.cpp:tryToFoldExtOfLoad()` called from
          `DAGCombiner::visitZERO_EXTEND`.
      
      Log for second basic block:
      
          Initial selection DAG: %bb.1 'foo:next'
          SelectionDAG has 13 nodes:
            t0: ch,glue = EntryToken
                      t2: i64,ch = CopyFromReg t0, Register:i64 %0
                    t4: i64 = AssertZext t2, ValueType:ch:i8
                  t5: i8 = truncate t4
                t8: i1 = setcc t5, Constant:i8<0>, seteq:ch
              t9: i64 = any_extend t8
            t11: ch,glue = CopyToReg t0, Register:i64 $r0, t9
            t12: ch = BPFISD::RET_GLUE t11, Register:i64 $r0, t11:1
          ...
          Replacing.2 t18: i64 = and t4, Constant:i64<255>
          With: t4: i64 = AssertZext t2, ValueType:ch:i8
          ...
          Type-legalized selection DAG: %bb.1 'foo:next'
          SelectionDAG has 13 nodes:
            t0: ch,glue = EntryToken
                    t2: i64,ch = CopyFromReg t0, Register:i64 %0
                  t4: i64 = AssertZext t2, ValueType:ch:i8
                t18: i64 = and t4, Constant:i64<255>
              t16: i64 = setcc t18, Constant:i64<0>, seteq:ch
            t11: ch,glue = CopyToReg t0, Register:i64 $r0, t16
            t12: ch = BPFISD::RET_GLUE t11, Register:i64 $r0, t11:1
          ...
          Optimized type-legalized selection DAG: %bb.1 'foo:next'
          SelectionDAG has 11 nodes:
            t0: ch,glue = EntryToken
                  t2: i64,ch = CopyFromReg t0, Register:i64 %0
                t4: i64 = AssertZext t2, ValueType:ch:i8
              t16: i64 = setcc t4, Constant:i64<0>, seteq:ch
            t11: ch,glue = CopyToReg t0, Register:i64 $r0, t16
            t12: ch = BPFISD::RET_GLUE t11, Register:i64 $r0, t11:1
          ...
      
      Note:
      - Initial selection DAG:
        - `t0` is an input value for this basic block, it corresponds load
          instruction (`t9`) from the first basic block.
        - It is accessed within basic block via
          `t4` (AssertZext (CopyFromReg t0, ...)).
        - The `AssertZext` is generated by RegsForValue::getCopyFromRegs
          called from SelectionDAGBuilder::getCopyFromRegs, it is generated
          only when `LiveOutInfo` with known number of leading zeros is
          present for `t0`.
        - Known register bits in `LiveOutInfo` are computed by
          `SelectionDAG::computeKnownBits` called from
          `SelectionDAGISel::ComputeLiveOutVRegInfo`.
        - `computeKnownBits()` generates leading zeros information for
          `(load ..., zext from ...)` but *does not* generate leading zeros
          information for `(load ..., anyext from ...)`.
          This is why `isZExtFree()` added in this commit is important.
      - Type-legalized selection DAG:
        - `t5 = truncate t4` is replaced by `t18 = and t4, 255`
      - Optimized type-legalized selection DAG:
        - `t18 = and t4, 255` is replaced by `t4`, this is done by
          `DAGCombiner::SimplifyDemandedBits` called from
          `DAGCombiner::visitAND`, which simplifies patterns like
          `(and (assertzext ...))`
      
      Impact
      ------
      
      This change covers all remove_truncate_*.ll test cases:
      - for -mcpu=v4 there are no changes in the generated code;
      - for -mcpu=v2 code generated for remove_truncate_7 and
        remove_truncate_8 improved slightly, for other tests it is
        unchanged.
      
      For remove_truncate_7:
      
          Before this revision                 After this revision
          --------------------                 -------------------
              r1 <<= 0x20                          r1 <<= 0x20
              r1 >>= 0x20                          r1 >>= 0x20
              if r1 == 0x0 goto +0x2 <LBB0_2>      if r1 == 0x0 goto +0x2 <LBB0_2>
              r1 = *(u32 *)(r2 + 0x0)              r0 = *(u32 *)(r2 + 0x0)
              goto +0x1 <LBB0_3>                   goto +0x1 <LBB0_3>
          <LBB0_2>:                            <LBB0_2>:
              r1 = *(u32 *)(r2 + 0x4)              r0 = *(u32 *)(r2 + 0x4)
          <LBB0_3>:                            <LBB0_3>:
              r0 = r1                              exit
              exit
      
      For remove_truncate_8:
      
          Before this revision                 After this revision
          --------------------                 -------------------
              r2 = *(u32 *)(r1 + 0x0)              r2 = *(u32 *)(r1 + 0x0)
              r3 = r2                              r3 = r2
              r3 <<= 0x20                          r3 <<= 0x20
              r4 = r3                              r3 s>>= 0x20
              r4 s>>= 0x20
              if r4 s> 0x2 goto +0x5 <LBB0_3>      if r3 s> 0x2 goto +0x4 <LBB0_3>
              r4 = *(u32 *)(r1 + 0x4)              r3 = *(u32 *)(r1 + 0x4)
              r3 >>= 0x20
              if r3 >= r4 goto +0x2 <LBB0_3>       if r2 >= r3 goto +0x2 <LBB0_3>
              r2 += 0x2                            r2 += 0x2
              *(u32 *)(r1 + 0x0) = r2              *(u32 *)(r1 + 0x0) = r2
          <LBB0_3>:                            <LBB0_3>:
              r0 = 0x3                             r0 = 0x3
              exit                                 exit
      
      For kernel BPF selftests statistics is as follows: (-mcpu=v4):
      - For -mcpu=v4: 9 out of 655 object files have differences,
        in all cases total number of instructions marginally decreased
        (-27 instructions).
      - For -mcpu=v2: 9 out of 655 object files have differences:
        - For 19 object files number of instruction decreased
          (-129 instruction in total): some redundant `rX &= 0xffff`
          and register to register assignments removed;
        - For 2 object files number of instructions increased +2
          instructions in each file.
      
      Both -mcpu=v2 instruction increases could be reduced to the same
      example:
      
          define void @foo(ptr %p) {
          entry:
            %a = load i32, ptr %p, align 4
            %b = sext i32 %a to i64
            %c = icmp ult i64 1, %b
            br i1 %c, label %next, label %end
      
          next:
            call void inttoptr (i64 62 to ptr)(i32 %a)
            br label %end
      
          end:
            ret void
          }
      
      Note that this example uses value loaded to `%a` both as a sign
      extended (`%b`) and as zero extended (`%a` passed as parameter).
      Here is the difference in final assembly code:
      
          Before this revision          After this revision
          --------------------          -------------------
              r1 = *(u32 *)(r1 + 0)         r1 = *(u32 *)(r1 + 0)
              r1 <<= 32                     r1 <<= 32
              r1 s>>= 32                    r1 s>>= 32
              if r1 < 2 goto <LBB0_2>       if r1 < 2 goto <LBB0_2>
                                            r1 <<= 32
                                            r1 >>= 32
              call 62                       call 62
          <LBB0_2>:                     <LBB0_2>:
              exit                          exit
      
      Before this commit `%a` is passed to call as a sign extended value,
      after this commit `%a` is passed to call as a zero extended value,
      both are correct as 32-bit sub-register is the same.
      
      The difference comes from `DAGCombiner` operation on the initial DAG:
      
      Initial selection DAG before this commit:
      
          t5: i32,ch = load<(load (s32) from %ir.p)> t0, t2, undef:i64
                t6: i64 = any_extend t5         <--------------------- (1)
              t8: ch = CopyToReg t0, Register:i64 %0, t6
                  t9: i64 = sign_extend t5
                t12: i1 = setcc Constant:i64<1>, t9, setult:ch
      
      Initial selection DAG after this commit:
      
          t5: i32,ch = load<(load (s32) from %ir.p)> t0, t2, undef:i64
                t6: i64 = zero_extend t5        <--------------------- (2)
              t8: ch = CopyToReg t0, Register:i64 %0, t6
                  t9: i64 = sign_extend t5
                t12: i1 = setcc Constant:i64<1>, t9, setult:ch
      
      The node `t9` is processed before node `t6` and `load` instruction is
      combined to load with sign extension:
      
          Replacing.1 t9: i64 = sign_extend t5
          With: t30: i64,ch = load<(load (s32) from %ir.p), sext from i32> t0, t2, undef:i64
           and 0 other values
          Replacing.1 t5: i32,ch = load<(load (s32) from %ir.p)> t0, t2, undef:i64
          With: t31: i32 = truncate t30
           and 1 other values
      
      This is done by `DAGCombiner.cpp:tryToFoldExtOfLoad` called from
      `DAGCombiner::visitSIGN_EXTEND`. Note that `t5` is used by `t6` which
      is `any_extend` in (1) and `zero_extend` in (2).
      `tryToFoldExtOfLoad()` rewrites such uses of `t5` differently:
      - `any_extend` is simply removed
      - `zero_extend` is replaced by `and t30, 0xffffffff`, which is later
        converted to a pair of shifts. This pair of shifts survives till the
        end of translation.
      
      Differential Revision: https://reviews.llvm.org/D157870
      651e6445
    • Justin Bogner's avatar
      [DXILBitcodeWriter] Don't create a new abbrev per MDString · 48e0a6f9
      Justin Bogner authored
      We were running out of abbrevs and crashing if there were more than 20
      something strings in metadata, which turned out to be a bug where we
      created an abbrev every time we emitted a string rather than just one
      for the string table.
      
      Differential Revision: https://reviews.llvm.org/D158440
      48e0a6f9
    • Fangrui Song's avatar
      Revert D157750 "[Driver][CodeGen] Properly handle -fsplit-machine-functions... · 77596e6b
      Fangrui Song authored
      Revert D157750 "[Driver][CodeGen] Properly handle -fsplit-machine-functions for fatbinary compilation."
      
      This reverts commit 317a0fe5.
      This reverts commit 30c4b97a.
      
      See post-commit discussions on https://reviews.llvm.org/D157750 that
      we should use a different mechanism to handle the error with --cuda-gpu-arch=
      
      The IR/DiagnosticInfo.cpp, warn_drv_for_elf_only, codegne tests in
      clang/test/Driver, and the following driver behavior (downgrading error
      to warning) changes are undesired.
      ```
      % clang --target=riscv64 -fsplit-machine-functions -c a.c
      warning: -fsplit-machine-functions is not valid for riscv64 [-Wbackend-plugin]
      ```
      77596e6b
    • Med Ismail Bennani's avatar
      [lldb/crashlog] Fix python version requirement issue · 446abb51
      Med Ismail Bennani authored
      In 21a597c3, we fixed a module loading issue by using the new
      `argparse.BooleanOptionalAction`. However, this is only available
      starting python 3.9 and causes test failures on bots that don't fulfill
      this requirement.
      
      To address that, this patch replaces the use of `BooleanOptionalAction`
      by a pair of 2 opposite `store` actions pointing to the same destination
      variable.
      
      Differential Revision: https://reviews.llvm.org/D158452
      
      
      
      Signed-off-by: default avatarMed Ismail Bennani <ismail@bennani.ma>
      446abb51
    • Joseph Huber's avatar
      [libc] Add the 'cpp.new' as a dependency on `atexit` · a69340dd
      Joseph Huber authored
      The `atexit` function depends on the implementations in CPP/new.h but it
      is not listed as a dependency. This causes the GPU build to not include
      it in the `libcgpu.a` file and prevents us from using the startup code
      externally. Simply add it.
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D158447
      a69340dd
    • Aart Bik's avatar
      [mlir][sparse] migrate more to new surface syntax · bb44a6b7
      Aart Bik authored
      Replaced the "NEW_SYNTAX" with the more readable "map"
      (which we may, or may not keep). Minor improvement in
      keyword parsing, migrated a few more examples over.
      
      Reviewed By: Peiming, yinying-lisa-li
      
      Differential Revision: https://reviews.llvm.org/D158325
      bb44a6b7