1. Mar 04, 2021
    • Zequan Wu's avatar
    • River Riddle's avatar
      [mlir][IR] Refactor the internal implementation of Value · 3dfa8614
      River Riddle authored
      The current implementation of Value involves a pointer int pair with several different kinds of owners, i.e. BlockArgumentImpl*, Operation *, TrailingOpResult*. This design arose from the desire to save memory overhead for operations that have a very small number of results (generally 0-2). There are, unfortunately, many problematic aspects of the current implementation that make Values difficult to work with or just inefficient.
      
      Operation result types are stored as a separate array on the Operation. This is very inefficient for many reasons: we use TupleType for multiple results, which can lead to huge amounts of memory usage if multi-result operations change types frequently(they do). It also means that simple methods like Value::getType/Value::setType now require complex logic to get to the desired type.
      
      Value only has one pointer bit free, severely limiting the ability to use it in things like PointerUnion/PointerIntPair. Given that we store the kind of a Value along with the "owner" pointer, we only leave one bit free for users of Value. This creates situations where we end up nesting PointerUnions to be able to use Value in one.
      
      As noted above, most of the methods in Value need to branch on at least 3 different cases which is both inefficient, possibly error prone, and verbose. The current storage of results also creates problems for utilities like ValueRange/TypeRange, which want to efficiently store base pointers to ranges (of which Operation* isn't really useful as one).
      
      This revision greatly simplifies the implementation of Value by the introduction of a new ValueImpl class. This class contains all of the state shared between all of the various derived value classes; i.e. the use list, the type, and the kind. This shared implementation class provides several large benefits:
      
      * Most of the methods on value are now branchless, and often one-liners.
      
      * The "kind" of the value is now stored in ValueImpl instead of Value
      This frees up all of Value's pointer bits, allowing for users to take full advantage of PointerUnion/PointerIntPair/etc. It also allows for storing more operation results as "inline", 6 now instead of 2, freeing up 1 word per new inline result.
      
      * Operation result types are now stored in the result, instead of a side array
      This drops the size of zero-result operations by 1 word. It also removes the memory crushing use of TupleType for operations results (which could lead up to hundreds of megabytes of "dead" TupleTypes in the context). This also allowed restructured ValueRange, making it simpler and one word smaller.
      
      This revision does come with two conceptual downsides:
      * Operation::getResultTypes no longer returns an ArrayRef<Type>
      This conceptually makes some usages slower, as the iterator increment is slightly more complex.
      * OpResult::getOwner is slightly more expensive, as it now requires a little bit of arithmetic
      
      From profiling, neither of the conceptual downsides have resulted in any perceivable hit to performance. Given the advantages of the new design, most compiles are slightly faster.
      
      Differential Revision: https://reviews.llvm.org/D97804
      3dfa8614
    • Jez Ng's avatar
      5d916984
    • Sanjay Patel's avatar
      [Analysis] simplify propagation of FMF in recurrences; NFC · b3f0c265
      Sanjay Patel authored
      This is a mess, but this is hopefully no-functional-change.
      The 'Prev' descriptor is only used for min/max recurrences
      or when starting a match from a phi, so it should not be a
      factor when propagating FMF for fmul/fadd.
      
      The API is confusing (and should be reduced in subsequent steps)
      because the "UnsafeAlgebraInst" appears to actually be a placeholder
      for a recurrence that does NOT have FMF, but we still want to
      treat it as reassociative.
      b3f0c265
    • Stefan Gränitz's avatar
      295ea050
    • David Blaikie's avatar
      Fix use of deprecated API · 4fda0dc1
      David Blaikie authored
      4fda0dc1
    • Louis Dionne's avatar
      [libc++] Temporary hack: disable Apple back-deployment testing · 460953ad
      Louis Dionne authored
      Apple back-deployment testing is currently failing because Green Dragon
      is down. To avoid stalling the whole CI pipeline because of that, I am
      temporarily disabling those jobs until Green Dragon is back, or even
      better we have found a different way to store those small artifacts.
      460953ad
    • MaheshRavishankar's avatar
      [mlir] Remove incorrect folding for SubTensorInsertOp · c118fdcd
      MaheshRavishankar authored
      The SubTensorInsertOp has a requirement that dest type and result
      type match. Just folding the tensor.cast operation violates this and
      creates verification errors during canonicalization. Also fix other
      canonicalization methods that werent inserting casts properly.
      
      Differential Revision: https://reviews.llvm.org/D97800
      c118fdcd
    • Florian Hahn's avatar
      [AArch64] Add implicit uses for operands when expanding BLR_RVMARKER. · 75805dce
      Florian Hahn authored
      Make sure we preserve info about passed arguments as implicit uses, to
      make sure later passes still have access to this information.
      
      This fixes a mis-compile where the machine-combiner would pick an
      incorrect free register.
      75805dce
    • Stefan Gränitz's avatar
      Revert "hack to unbreak check-llvm on win after D97335" in attempt for actual fix · e984c2b0
      Stefan Gränitz authored
      This reverts commit 900f0761 and attempts an actual fix: All failing tests for llvm-jitlink use the `-noexec` flag. The inputs they operate on are not meant for execution on the host system. Looking e.g. at the MachO_test_harness_harnesss.s test, llvm-mc generates input machine code with "x86_64-apple-macosx10.9".
      
      My previous attempt in bbdb4c8c disabled the debug support plugin for Windows targets, but what we would actually want is to disable it on Windows HOSTS.
      
      With the new patch here, I don't do exactly that, but instead follow the approach for the EH frame plugin and include the `-noexec` flag in the condition. It should have the desired effect when it comes to the test suite. It appears a little workaround'ish, but should work reliably for now. I will discuss the issue with Lang and see if we can do better. Thanks @thakis again for the temporary fix.
      e984c2b0
    • Soumi Manna's avatar
      [WebAssembly] Add missing default cases in switch statements · eec7f8f7
      Soumi Manna authored
      unsigned variable 'IntNo' has been declared but not been defined inside function
      EmitWebAssemblyBuiltinExpr().
      
      static code analysis tool complains about uninitialized variable "IntNo" since
      this enters to default branch without setting any intrinsics and calls Function
      *Callee = CGM.getIntrinsic(IntNo).
      
      This patch fixes the problem by adding default cases in switch statements.
      eec7f8f7
    • Jez Ng's avatar
      [lld-macho] Require -arch and -platform_version to always be specified · b63919e1
      Jez Ng authored
      We previously defaulted to x86_64 and an unknown platform, which was fine when
      we only supported one arch and did no platform checks, but that will no longer
      be true going ahead. Therefore, we should require those flags to be specified
      whenever the linker is invoked.
      
      Note that LLD-ELF and ld64 both infer the arch from their input object files,
      but the usefulness of that is questionable since clang will always specify these
      flags, and most of the time `lld` will be invoked via clang.
      
      Reviewed By: #lld-macho, thakis
      
      Differential Revision: https://reviews.llvm.org/D97799
      b63919e1
    • Jez Ng's avatar
      [lld-macho][nfc] Parse more options using getLastArg{Value} · 1168736c
      Jez Ng authored
      The option-iterating loop should be reserved for options whose command-line
      order is important. I think LLD-ELF follows a similar design.
      
      Reviewed By: #lld-macho, smeenai
      
      Differential Revision: https://reviews.llvm.org/D97797
      1168736c
    • Whitney Tsang's avatar
      [LoopUnrollRuntime] Add option to assume the non latch exit block to be · 58d531fd
      Whitney Tsang authored
      predictable.
      
      Reviewed By: Meinersbur, bmahjour
      
      Differential Revision: https://reviews.llvm.org/D97747
      58d531fd
    • Alexey Bataev's avatar
      [Cost]Add tests for boolean and/or reductions, NFC. · 60470ac7
      Alexey Bataev authored
      Tests with the default costs for boolean and/or reductions.
      
      Differential Revision: https://reviews.llvm.org/D97793
      60470ac7
    • Philip Reames's avatar
      Address review comment from D97219 (follow up to 80511565) · 89d331a3
      Philip Reames authored
      Probably should have done this before landing, but I forgot.
      
      Basic idea is to avoid using the SCEV predicate when it doesn't buy us anything.  Also happens to set us up for handling non-add recurrences in the future if desired.
      89d331a3
    • Philip Reames's avatar
      Sink routine for replacing a operand bundle to CallBase [NFC] · 99f54173
      Philip Reames authored
      We had equivalent code for both CallInst and InvokeInst, but never cared about the result type.
      99f54173
    • Philip Reames's avatar
      [LSR] Unify scheduling of existing and inserted addrecs · 80511565
      Philip Reames authored
      LSR goes to some lengths to schedule IV increments such that %iv and %iv.next never need to overlap. This is fairly fundamental to LSRs cost model. LSR assumes that an addrec can be represented with a single register. If %iv and %iv.next have to overlap, then that assumption does not hold.
      
      The bug - which this patch is fixing - is that LSR only does this scheduling for IVs which it inserts, but it's cost model assumes the same for existing IVs that it reuses. It will rewrite existing IV users such that the no-overlap property holds, but will not actually reschedule said IV increment.
      
      As you can see from the relatively lack of test updates, this doesn't actually impact codegen much. The main reason for doing it is to make a follow up patch series which improves post-increment use and scheduling easier to follow.
      
      Differential Revision: https://reviews.llvm.org/D97219
      80511565
    • Jonas Paulsson's avatar
      [SystemZ] Reimplement the i8/i16 compare-and-swap logic. · 7334b3dc
      Jonas Paulsson authored
      Even though the implementation in emitAtomicCmpSwapW() was correct, it made
      Valgrind report an error. Instead of using a RISBG on CmpVal, an LL[CH]R can
      be made on the OldVal, and the problem is avoided.
      
      Review: Ulrich Weigand
      
      Differential Revision: https://reviews.llvm.org/D97604
      7334b3dc
    • Peter Steinfeld's avatar
      [flang] Prohibit MODULE procedures in the global scope · 1c2935a7
      Peter Steinfeld authored
      We were allowing procedures with the MODULE prefix to be declared at the global
      scope.  This is prohibited by C1547 and was causing an internal check of the
      compiler to fail.
      
      I fixed this by adding a check.  I also added a test that would trigger a crash
      without this change.
      
      Differential Revision: https://reviews.llvm.org/D97875
      1c2935a7
    • Hanhan Wang's avatar
      [mlir][linalg] Add depthwise_conv_2d_input_nhwc_filter_hwcf to Linalg TC ops. · 83c56aa4
      Hanhan Wang authored
      Different from the definition in Tensorflow and TOSA, the output is [N,H,W,C,M]. This can make transforms easier in LinAlg because the indexing maps are plain. E.g., to determine if the fill op has dependency between the depthwise conv op, the current pipeline only recognizes the dep if they are all projected affine map.
      
      Reviewed By: asaadaldien
      
      Differential Revision: https://reviews.llvm.org/D97798
      83c56aa4
    • Florian Hahn's avatar
      [AArch64] Move CALL_RVMARKER definition after CALL. · 8c3a70a7
      Florian Hahn authored
      This is a NFC with respect to the generated code. But it fixes a crash
      when using -debug, because of the position in the enum CALL_RVMARKER
      nodes were treated as memops. That caused a crash when printing
      CALL_RVMARKER nodes.
      8c3a70a7
    • Fangrui Song's avatar
      [InstrProfiling] Place __llvm_prf_vnodes and __llvm_prf_names in llvm.used on ELF · a84f4fc0
      Fangrui Song authored
      `__llvm_prf_vnodes` and `__llvm_prf_names` are used by runtime but not
      referenced via relocation in the translation unit.
      
      With `-z start-stop-gc` (LLD 13 (D96914); GNU ld 2.37 https://sourceware.org/bugzilla/show_bug.cgi?id=27451),
      the linker does not let `__start_/__stop_` references retain their sections.
      
      Place `__llvm_prf_vnodes` and `__llvm_prf_names` in `llvm.used` to make
      them retained by the linker.
      
      This patch changes most existing `UsedVars` cases to `CompilerUsedVars`
      to reflect the ideal state - if the binary format properly supports
      section based GC (dead stripping), `llvm.compiler.used` should be sufficient.
      
      `__llvm_prf_vnodes` and `__llvm_prf_names` are switched to `UsedVars`
      since we want them to be unconditionally retained by both compiler and linker.
      
      Behaviors on COFF/Mach-O are not affected.
      
      Reviewed By: davidxl
      
      Differential Revision: https://reviews.llvm.org/D97649
      a84f4fc0
    • Fangrui Song's avatar
      [test] Improve PGO tests · 75df61e9
      Fangrui Song authored
      75df61e9
    • Zequan Wu's avatar
    • Sam McCall's avatar
      [clangd] ObjC fixes for semantic highlighting and xref highlights · 7d2fba8d
      Sam McCall authored
      - highlight references to protocols in class/protocol/extension decls
      - support multi-token selector highlights in semantic + xref highlights
        (method calls and declarations only)
      - In `@interface I(C)`, I now references the interface and C the category
      - highlight uses of interfaces as types
      - added semantic highlightings of protocol names (as "interface") and
        category names (as "namespace").
        These are both standard kinds, maybe "extension" will be standardized...
      - highlight `auto` as "class" when it resolves to an ObjC pointer
      - don't highlight `self` as a variable even though the AST models it as one
      
      Not fixed: uses of protocols in type names (needs some refactoring of
      unrelated code first)
      
      Differential Revision: https://reviews.llvm.org/D97617
      7d2fba8d
    • George Balatsouras's avatar
      [dfsan] Remove hardcoded shadow width in abilist_aggregate.ll · 87e854a5
      George Balatsouras authored
      As a preparation step for fast8 support, we need to update the tests
      to pass in both modes. That requires generalizing the shadow width
      and remove any hard coded references that assume it's always 2 bytes.
      
      Reviewed By: stephan.yichao.zhao
      
      Differential Revision: https://reviews.llvm.org/D97723
      87e854a5
    • Petr Hosek's avatar
      [CMake] Rename RUNTIMES_BUILD to LLVM_RUNTIMES_BUILD · 61a792b3
      Petr Hosek authored
      This avoid potential conflict with other internal variables.
      
      Differential Revision: https://reviews.llvm.org/D97838
      61a792b3
    • Louis Dionne's avatar
      3c62198c
    • Stanislav Mekhanoshin's avatar
      [AMDGPU] Exclude always_inline from max bb threshold · b70c483e
      Stanislav Mekhanoshin authored
      Honor always_inline attribute when processing -amdgpu-inline-max-bb.
      It was lost during the ports of the heuristic. There is no reason
      to honor inline hint, but not always inline.
      
      Differential Revision: https://reviews.llvm.org/D97790
      b70c483e
    • Mehdi Amini's avatar
      Add basic JIT Python Bindings · 13cb4317
      Mehdi Amini authored
      This offers the ability to create a JIT and invoke a function by passing
      ctypes pointers to the argument and the result.
      
      Differential Revision: https://reviews.llvm.org/D97523
      13cb4317
    • Mehdi Amini's avatar
      Add C bindings for mlir::ExecutionEngine · 86c8a785
      Mehdi Amini authored
      This adds minimalistic bindings for the execution engine, allowing to
      invoke the JIT from the C API. This is still quite early and
      experimental and shouldn't be considered stable in any way.
      
      Differential Revision: https://reviews.llvm.org/D96651
      86c8a785
    • Hongtao Yu's avatar
      [CSSPGO][llvm-profgen] Continue disassembling after illegal instruction is seen. · 55356c01
      Hongtao Yu authored
      Previously we errored out when disassembling illegal instructions and there would be no profile generated. In fact illegal instructions are not uncommon and we'd better skip them and print "unknown" instead of erroring out. This matches the behavior of llvm-objdump (see disassembleObject in llvm-objdump.cpp).
      
      Reviewed By: wlei, wenlei
      
      Differential Revision: https://reviews.llvm.org/D97776
      55356c01
    • Choongwoo Han's avatar
      [llvm-cov] Cache file status information · 9d8a3e75
      Choongwoo Han authored
      Currently, getSourceFile accesses file system to check if two paths are
      the same file with a thread lock, which is a huge performance bottleneck
      in some cases. Currently, it's accessing file system size(files) * size(files) times.
      
      Thus, cache file status information, which reduces file system access to size(files) times.
      
      When I tested it with two binaries and 16 cpu cores,
      it saved over 70% of time.
      
      Binary 1: 56 secs -> 3 secs
      Binary 2: 17 hours -> 4 hours
      
      Differential Revision: https://reviews.llvm.org/D97061
      9d8a3e75
    • Elia Geretto's avatar
      [XRay][x86_64] Fix CFI directives in assembly trampolines · 9ee61cf3
      Elia Geretto authored
      This patch modifies the x86_64 XRay trampolines to fix the CFI information
      generated by the assembler. One of the main issues in correcting the CFI
      directives is the `ALIGNED_CALL_RAX` macro, which makes the CFA dependent on
      the alignment of the stack. However, this macro is not really necessary because
      some additional assumptions can be made on the alignment of the stack when the
      trampolines are called. The code has been written as if the stack is guaranteed
      to be 8-bytes aligned; however, it is instead guaranteed to be misaligned by 8
      bytes with respect to a 16-bytes alignment. For this reason, always moving the
      stack pointer by 8 bytes is sufficient to restore the appropriate alignment.
      
      Trampolines that are called from within a function as a result of the builtins
      `__xray_typedevent` and `__xray_customevent` are necessarely called with the
      stack properly aligned so, in this case too, `ALIGNED_CALL_RAX` can be
      eliminated.
      
      Fixes: https://bugs.llvm.org/show_bug.cgi?id=49060
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D96785
      9ee61cf3
    • Louis Dionne's avatar
      [libc++] Use generator expression to simplify the CMake code · 5034d711
      Louis Dionne authored
      A comment was left for when we would require CMake >= 3, which we do now.
      I expect this should be a NFC.
      
      Differential Revision: https://reviews.llvm.org/D97341
      5034d711
    • Louis Dionne's avatar
      [libc++/abi] Replace uses of _NOEXCEPT in src/ by noexcept · 5601305f
      Louis Dionne authored
      We always build the libraries in a Standard mode that supports noexcept,
      so there's no need to use the _NOEXCEPT macro.
      
      Differential Revision: https://reviews.llvm.org/D97700
      5601305f
    • Hanhan Wang's avatar
      [mlir][linalg] Delete unused vars if there are shaped-only operands. · 497b7b8c
      Hanhan Wang authored
      Reviewed By: stella.stamenova
      
      Differential Revision: https://reviews.llvm.org/D97851
      497b7b8c
    • MaheshRavishankar's avatar
      [mlir] Add LinalgInterface method to clone with a given BlockAndValueMapping. · 5d7e0a23
      MaheshRavishankar authored
      Since Linalg operations have regions by default which are not isolated
      from above, add an another method to the interface that will take a
      BlockAndValueMapping to remap the values within the region as well.
      
      Differential Revision: https://reviews.llvm.org/D97709
      5d7e0a23
    • Philip Reames's avatar
      Fix a build warning from ea7d208b · c8cf27e3
      Philip Reames authored
      c8cf27e3