1. Sep 22, 2023
    • qcolombet's avatar
      [MLIR][linalg] Fix unpack rewriter for dynamic shapes (#67096) · a44b787e
      qcolombet authored
      Prior to this patch, `GeneralizeOuterUnitDimsUnPackOpPattern` would
      assert that we cannot create a `tensor.empty` operation with dynamic
      shapes.
      
      The problem stems from the fact that we were not using the right builder
      for the `tensor.empty` operation. Indeed, each dynamic dim needs to be
      specified by an input variable.
      
      Simply provide the dynamic dimensions to the `tensor.empty` builder to
      fix that.
      a44b787e
    • Florian Hahn's avatar
      [ConstraintElim] Support adding facts from switch terminators. (#67061) · 39d7f700
      Florian Hahn authored
      After 4a5bcbd5, switch instructions can now be handled in a
      straight-forward manner by adding (ICMP_EQ, ConditionVal, CaseVal) for
      te successor blocks per case.
      39d7f700
    • Ivan Kosarev's avatar
      [AMDGPU] Don't suppress printing the .l and .h register suffixes. · c62f208c
      Ivan Kosarev authored
      We don't seem to have a use for the -amdgpu-keep-16-bit-reg-suffixes
      option anymore. Was introduced in <https://reviews.llvm.org/D79435>.
      
      Reviewed By: Joe_Nash, foad
      
      Differential Revision: https://reviews.llvm.org/D156102
      c62f208c
    • Florian Hahn's avatar
      [ConstraintElim] Add additional switch case and use i8 instead of i32. · 0cb35306
      Florian Hahn authored
      Shorten the types used to i8 for cheaper verification and add test case
      where 2 cases have the same destination, as suggested in #67061.
      0cb35306
    • Simon Pilgrim's avatar
      [DAG] getNode() - remove oneuse limit from (zext (trunc (assertzext x))) ->... · b61b2426
      Simon Pilgrim authored
      [DAG] getNode() - remove oneuse limit from (zext (trunc (assertzext x))) -> (assertzext x) fold (REAPPLIED)
      
      Noticed on D159533 and I've finally dealt with the x86 regressions - MatchingStackOffset wasn't peeking through AssertZext nodes while trying to find CopyFromReg/Load sources, it was only removing them if they were part of a (trunc (assertzext x)) pattern.
      
      Reapplied after being reverted at 4389252c - which should be addressed by D159537 / 6d267999
      b61b2426
    • Alcaro's avatar
      docs: Fix misplaced apostrophe (#67103) · 10217b9d
      Alcaro authored
      10217b9d
    • Ivan Kosarev's avatar
      [AMDGPU] Introduce real and keep fake True16 instructions. · 0f864c7b
      Ivan Kosarev authored
      The existing fake True16 instructions using 32-bit VGPRs are supposed to
      co-exist with real ones until all the necessary True16 functionality is
      implemented and relevant tests are updated.
      
      Reviewed By: arsenm, Joe_Nash
      
      Differential Revision: https://reviews.llvm.org/D156101
      0f864c7b
    • Nikita Popov's avatar
      [StackColoring] Handle SEH catch object stack slots conservatively · b3cb4f06
      Nikita Popov authored
      The write to the SEH catch object happens before cleanuppads are
      executed, while the first reference to the object will typically
      be in a catchpad.
      
      If we make use of first-use analysis, we may end up allocating
      an alloca used inside the cleanuppad and the catch object at the
      same stack offset, which would be incorrect.
      
      https://reviews.llvm.org/D86673 was a previous attempt to fix it.
      It used the heuristic "a slot loaded in a WinEH pad and never
      written" to detect catch objects. However, because it checks
      for more than one load (while probably more than zero was
      intended), the fix does not actually work.
      
      The general approach also seems dubious to me, so this patch
      reverts that change entirely, and instead marks all catch object
      slots as conservative (i.e. excluded from first-use analysis)
      based on the WinEHFuncInfo. As far as I can tell we don't need
      any heuristics here, we know exactly which slots are affected.
      
      Fixes https://github.com/llvm/llvm-project/issues/66984.
      b3cb4f06
    • Ivan Kosarev's avatar
      [AMDGPU] Have a subtarget feature to control use of real True16 instructions. · bea56b0b
      Ivan Kosarev authored
      Real True16 instructions are as they are defined in the ISA. Fake True16
      instructions are identical to real ones except that they take 32-bit
      registers as operands and always use their low halves.
      
      Reviewed By: Joe_Nash
      
      Differential Revision: https://reviews.llvm.org/D156100
      bea56b0b
    • Guray Ozen's avatar
      [MLIR][NVGPU] Adding `nvgpu.warpgroup.mma` Op for Hopper GPUs (#65440) · 23882226
      Guray Ozen authored
      This work introduces a new operation called `warpgroup.mma` to the NVGPU
      dialect of MLIR. The purpose of this operation is to facilitate
      warpgroup-level matrix multiply and accumulate (WGMMA) operations on
      Hopper GPUs with sm_90a architecture.
      
      Previously, the `nvvm.wgmma.mma_async` operation was introduced to
      support warpgroup-level matrix operations in NVVM dialect. This op is
      used multiple instances of `nvvm.wgmma.mma_async` to achieve the desired
      shape. The new `nvgpu.warpgroup.mma` operation abstracts this complexity
      and provides a higher-level interface for performing warpgroup-level
      matrix operations.
      
      The `nvgpu.warpgroup.mma` does followings:
      1) Corresponds multiple `wgmma` instructions.
      2) Iterates input matrix descriptors to achieve the desired computation
      shape. 3) Groups and runs `wgmma` instructions asynchronously, and
      eventually waits them. This are done by `wgmma.fence.aligned`,
      `wgmma.commit.group.sync.aligned`, and `wgmma.wait.group.sync.aligned`
      4) Results fragmented matrices
      
      Here's an example usage of the `nvgpu.warpgroup.mma` operation:
      ```
      %wgmmaResult, %wgmmaResult2 = nvgpu.warpgroup.mma %descA, %descB, %acc1, %acc2 {transposeB}: 
            !nvgpu.wgmma.descriptor<tensor = memref<128x64xf16, 3>>, 
            !nvgpu.wgmma.descriptor<tensor = memref<64x128xf16, 3>>, 
            !nvgpu.warpgroup.accumulator< fragmented = vector<64x128xf32>>,
            !nvgpu.warpgroup.accumulator< fragmented = vector<64x128xf32>> 
            -> 
            !nvgpu.warpgroup.accumulator< fragmented = vector<64x128xf32>>, 
            !nvgpu.warpgroup.accumulator< fragmented = vector<64x128xf32>>  
      ```
      
      The op will result following PTX:
      ```
      wgmma.fence.sync.aligned;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f1, %f2,    62 more registers}, %descA,     %descB,     p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f1, %f2,    62 more registers}, %descA+2,   %descB+128, p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f1, %f2,    62 more registers}, %descA+4,   %descB+256, p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f1, %f2,    62 more registers}, %descA+8,   %descB+348, p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f500,%f501, 62 more registers}, %descA+512, %descB,     p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f500,%f501, 62 more registers}, %descA+514, %descB+128, p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f500,%f501, 62 more registers}, %descA+516, %descB+256, p, 1, 1, 0, 1;
      wgmma.mma_async.sync.aligned.m64n128k16.f32.f16.f16 {%f500,%f501, 62 more registers}, %descA+518, %descB+348, p, 1, 1, 0, 1;
      wgmma.commit_group.sync.aligned;
      wgmma.wait_group.sync.aligned 1;
      ```
      
      The Op keeps 
        - first 64 registers (`{%f1, %f2,    62 more registers}`) -> `%acc1` 
      - second 64 registers (`{%f500,%f501, 62 more registers}`) -> `%acc2`.
      23882226
    • Simon Pilgrim's avatar
      [AArch64] Don't rely on (zext (trunc x)) pattern to detect zext_inreg MULL... · 6d267999
      Simon Pilgrim authored
      [AArch64] Don't rely on (zext (trunc x)) pattern to detect zext_inreg MULL patterns - use value tracking directly
      
      As explained on D159533, I'm trying to generalize the "(zext (trunc x)) -> x iff the upper bits are known zero" fold in getNode() and I was seeing assertions in the aarch64 mull matching code as it was assuming these 'zero-extend-inreg' patterns will remain from earlier in LowerMUL.
      
      Instead I've updated selectUmullSmull/skipExtensionForVectorMULL to just use value tracking to detect when the upper bits are known zero, and to insert the truncation nodes later if necessary.
      
      Differential Revision: https://reviews.llvm.org/D159537
      6d267999
    • Simon Pilgrim's avatar
      [AArch64] LowerMUL - use SDValue directly instead of SDNode. NFC. · bba83e20
      Simon Pilgrim authored
      As discussed on D159537, using the SDValue operands directly instead of peeking inside to the SDNode prevents any issues where a non-zero result index has been used.
      bba83e20
    • Matthew Devereau's avatar
      [AArch64][SVE2] Do not emit RSHRNB for large shifts (#66672) · 23ea98f1
      Matthew Devereau authored
      rshrnb's shift amount operand must be between 1-EltSizeInBits. This
      patch stops RSHRNB ISD nodes being emitted in this case
      23ea98f1
    • Sam McCall's avatar
      [dataflow] Parse formulas from text (#66424) · 3f78d6ab
      Sam McCall authored
      My immediate use for this is not in checked-in code, but rather the
      ability to plug printed flow conditions (from analysis logs) back into
      sat solver unittests to reproduce slowness.
      
      It does allow simplifying some of the existing solver tests, though.
      3f78d6ab
    • Timm Bäder's avatar
      [clang][Sema][NFC] _or_null -> _if_present · b2bbf694
      Timm Bäder authored
      b2bbf694
    • Sam McCall's avatar
      [Serialization] Do less redundant work computing affecting module maps (#66933) · 0f050965
      Sam McCall authored
      We're traversing the same chains of module ancestors and include
      locations repeatedly, despite already populating sets that can detect
      it!
      
      This is a problem because translateFile() is expensive. I think we can
      avoid it entirely, but this seems like an improvement either way.
      
      I removed a callback indirection rather than giving it a more
      complicated
      signature, and accordingly renamed the lambdas to be more concrete.
      0f050965
    • jeanPerier's avatar
      [flang][runtime] Finalize polymorphic components using dynamic type (#67040) · efd5cdee
      jeanPerier authored
      Previous code was finalizing polymorphic components according to static
      type (calling the static type final routine, if any).
      
      There is no way (I think) to know from a
      Fortran::runtime::typeInfo::Component if an allocatable component is
      polymorphic or not. So this patch just always uses the dynamic type
      descriptor to check for derived type allocatable component finalization.
      efd5cdee
    • Ivan Kosarev's avatar
      [AMDGPU] Add True16 register classes. · 469b3bfa
      Ivan Kosarev authored
      Reviewed By: rampitec, Joe_Nash
      
      Differential Revision: https://reviews.llvm.org/D156099
      469b3bfa
    • Benjamin Maxwell's avatar
      [mlir][ArmSME] Add tile slice to vector intrinsics (#66910) · cb3a3944
      Benjamin Maxwell authored
      Add support for following vector to tile (MOVA) intrinsics to ArmSME
      dialect:
      ```
      llvm.aarch64.sme.read.vert
      llvm.aarch64.sme.read.horiz
      ```
      This also slightly updates ArmSME_IntrOp to support return values.
      cb3a3944
    • Sam McCall's avatar
      [clangd] Allow --query-driver to match a dot-normalized form of the path (#66757) · 01d3045d
      Sam McCall authored
      (In addition to the un-normalized form, so this is back-compatible)
      01d3045d
    • Tobias Hieta's avatar
      build-docs: Add option to disable doxygen/sphinx docs (#66928) · fe7fe6d3
      Tobias Hieta authored
      Doxygen documentation takes very long to build, when making releases we
      want to get the normal documentation up earlier so that we don't have to
      wait for doxygen documentation.
      
      This PR just adds the flag to disable doxygen builds, I will then later
      make a PR that changes the actions to first build the normal docs and
      another job to build the doxygen docs.
      fe7fe6d3
    • Guray Ozen's avatar
      [mlir] Fix incorrect check in ptx builder test (nfc) (#67097) · d60376bf
      Guray Ozen authored
      The test code had a space between "CHECK" and ":" that prohibits testing. This PR fixes this problem
      d60376bf
    • Tobias Hieta's avatar
      Reland: [Workflow] Add new code format helper. · a1177b0b
      Tobias Hieta authored
      I landed this format helper, but unfortunately, it didn't work because
      of permissions, it could not add comments on a fork's PR. @cor3ntin
      informed me there are fixes for this that you had worked on @tstellar -
      but I didn't have time to read up on it too much. Can you explain what
      changes are needed to get the action to be able to write comments on
      fork's PR?
      a1177b0b
    • Nikita Popov's avatar
      [LSR] Simplify type check for opaque pointers (NFC) · d35e5afc
      Nikita Popov authored
      For pointer types, checking the address space is the same as
      type equality now, so we no longer need the special case.
      d35e5afc
    • David Spickett's avatar
      Revert "[SLP]Use source vector type as the original vector type instead of" · 8f548610
      David Spickett authored
      This reverts commit 9a99944d.
      
      Due to test suite failures on all our SVE buildbots e.g.:
      https://lab.llvm.org/buildbot/#/builders/184/builds/7375
      
      clang: ../llvm/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp:3565:
      InstructionCost llvm::AArch64TTIImpl::getShuffleCost(TTI::ShuffleKind,
      VectorType *, ArrayRef<int>, TTI::TargetCostKind, int, VectorType *,
      ArrayRef<const Value *>): Assertion `Mask.size() == TpNumElts && "Expected Mask and Tp size to match!"' failed.
      8f548610
    • Kazu Hirata's avatar
      [Analysis] Use drop_begin (NFC) · 619c5012
      Kazu Hirata authored
      619c5012
    • Balint Cristian's avatar
      [clang] Enable descriptions for --print-supported-extensions (#66715) · 73779bb2
      Balint Cristian authored
      Enables summary descriptions along with the names of the feature.
      Descriptions here are simply looked up via the available llvm tablegen
      data.
      73779bb2
    • Kazu Hirata's avatar
      [llvm] Use range-based for loops (NFC) · 4c14638b
      Kazu Hirata authored
      4c14638b
    • Nikita Popov's avatar
      [SCEVExpander] Clarify absence of no-op casts (NFC) · 45778602
      Nikita Popov authored
      Remove all the expandCodeFor() uses that specify an explicit type,
      as well as InsertNoopCastOfTo() calls and most uses of
      getEffectiveSCEVType().
      
      The only place where no-op casts can now be inserted are public
      expandCodeFor() uses.
      45778602
    • Heejin Ahn's avatar
      [libunwind][WebAssembly] Support Wasm EH · 058222b2
      Heejin Ahn authored
      This adds Wasm-specific libunwind port to support Wasm exception
      handling (https://github.com/WebAssembly/exception-handling).
      
      Wasm EH requires `__USING_WASM_EXCEPTIONS__` to be defined. This adds
      `Unwind-wasm.c`, which defines libunwind APIs for Wasm. This also adds a
      `thread_local` struct of type `_Unwind_LandingPadContext`, which serves
      as a medium for input/output data between the user code and the
      personality function. How all these work is explained in
      https://github.com/WebAssembly/tool-conventions/blob/main/EHScheme.md.
      (The doc is old and "You Shouldn't Prune Unreachable Resumes" section
      doesn't apply anymore, but otherwise it should be good)
      
      The bulk of these changes was added back in Mar 2020 in
      https://github.com/emscripten-core/emscripten/pull/10577 to emscripten
      repo and has been used ever since. Now we'd like to upstream this so
      that other toolchains that don't use emscripten libraries, e.g., WASI,
      can use this too.
      
      Companion patch: D158918
      
      Reviewed By: dschuff, #libunwind, phosek
      
      Differential Revision: https://reviews.llvm.org/D158919
      058222b2
    • Heejin Ahn's avatar
      [libc++abi][WebAssembly] Support Wasm EH · e6cbba74
      Heejin Ahn authored
      This adds Wasm-specific libc++abi changes to support Wasm exception
      handling (https://github.com/WebAssembly/exception-handling).
      
      Wasm EH requires `__USING_WASM_EXCEPTIONS__` to be defined. Wasm EH's
      LSDA handling mostly shares that of SjLj EH.
      Changes are:
      - In Wasm, a destructor returns its argument.
      - Wasm EH currently only has one phase (search) that does both search
        and cleanup. So added an additional `set_registers` to support that.
      
      The bulk of these changes was added back in Mar 2020 in
      https://github.com/emscripten-core/emscripten/pull/10577 to emscripten
      repo and has been used ever since. Now we'd like to upstream this so
      that other toolchains that don't use emscripten libraries, e.g., WASI,
      can use this too.
      
      Companion patch: D158919
      
      Reviewed By: dschuff, #libc_abi, phosek
      
      Differential Revision: https://reviews.llvm.org/D158918
      e6cbba74
    • Jie Fu's avatar
      [SCEV] Fix -Wunused-variable in ScalarEvolutionExpander.cpp (NFC) · 8ffe73b6
      Jie Fu authored
      llvm-project/llvm/lib/Transforms/Utils/ScalarEvolutionExpander.cpp:1077:15: error: unused variable 'Start' [-Werror,-Wunused-variable]
        const SCEV *Start = Normalized->getStart();
                    ^
      1 error generated.
      8ffe73b6
    • Kazu Hirata's avatar
      [Analysis] Use std::clamp (NFC) · ea91ae5b
      Kazu Hirata authored
      ea91ae5b
    • Fangrui Song's avatar
      Revert "[DAG] getNode() - remove oneuse limit from (zext (trunc (assertzext... · 4389252c
      Fangrui Song authored
      Revert "[DAG] getNode() - remove oneuse limit from (zext (trunc (assertzext x))) -> (assertzext x) fold"
      
      This reverts commit 05926a5a.
      
      Caused AArch64 crash
      
       #12 0x00007f09eec09181 skipExtensionForVectorMULL(llvm::SDNode*, llvm::SelectionDAG&)
       #13 0x00007f09eec08289 llvm::AArch64TargetLowering::LowerMUL(llvm::SDValue, llvm::SelectionDAG&) const
       #14 0x00007f09eec1a3fd llvm::AArch64TargetLowering::LowerOperation(llvm::SDValue, llvm::SelectionDAG&) const
       #15 0x00007f09dc8586a7 (anonymous namespace)::VectorLegalizer::LowerOperationWrapper(llvm::SDNode*, llvm::SmallVectorImpl<llvm::SDValue>&)
      4389252c
    • Nikita Popov's avatar
      [SCEV] Require that addrec operands dominate the loop · 2d8d622c
      Nikita Popov authored
      SCEVExpander currently has special handling for the case where the
      start or the step of an addrec do not dominate the loop header,
      which is not used by any lit test.
      
      Initially I thought that this is entirely dead code, because
      addrec operands are required to be loop invariant. However,
      SCEV currently allows creating an addrec with operands that are
      loop invariant but defined *after* the loop.
      
      This doesn't seem like a useful case to allow, and we don't
      appear to be using this outside a single easy to adjust unit test.
      2d8d622c
    • jeanPerier's avatar
      [flang] Deallocate local allocatable at end of their scopes (#67036) · 0c7d0ad9
      jeanPerier authored
      Implement automatic deallocation of unsaved local alloctables when
      reaching the end of their scope of block as described in Fortran 2018
      9.7.3.2 point 2. and 3.
      
      Uses genDeallocateIfAllocated used for intent(out) deallocation and the
      "function context" already used for finalization at end of scope.
      0c7d0ad9
    • David Green's avatar
      22f423aa
    • Bruno Cardoso Lopes's avatar
      [Clang][LLVM][Coroutines] Prevent __coro_gro from outliving __promise (#66706) · 34415fd6
      Bruno Cardoso Lopes authored
      When dealing with short-circuiting coroutines (e.g. expected), the
      deferred calls that resolve the get_return_object are currently being
      emitted after we delete the coroutine frame.
      
      This was caught by ASAN when using optimizations -O1 and above:
      optimizations after inlining would place the __coro_gro in the heap, and
      subsequent delete of the coroframe followed by the conversion -> BOOM.
      
      This patch forbids the GRO to be placed in the coroutine frame, by
      adding a new metadata node that can be attached to `alloca`
      instructions.
      
      Fix #49843 
      34415fd6
    • Jeff Bailey's avatar
      [libc] Pull more definitions from linux/stat.h (#67071) · c618e131
      Jeff Bailey authored
      For file handling, we need more definitions from
      linux/stat.h, so this pulls them in. It also adjusts other definitions
      to match the kernel's exactly [NFC] so that it's easy to verify that
      there's been no divergence one day when it's time to use linux/stat.h
      directly.
      
      Tested:
      check-libc
      c618e131
    • Balazs Benics's avatar
      [analyzer] Fix taint sink rules for exec-like functions (#66358) · f90e0633
      Balazs Benics authored
      Variadic arguments were not considered as taint sink arguments. I also
      decided to extend the list of exec-like functions.
      
      (Juliet CWE78_OS_Command_Injection__char_connect_socket_execl)
      f90e0633