1. Nov 14, 2019
    • Francis Visoiu Mistrih's avatar
    • Craig Topper's avatar
      [TargetLowering] Increase the storage size of NumRegistersForVT to allow the... · 84e83b54
      Craig Topper authored
      [TargetLowering] Increase the storage size of NumRegistersForVT to allow the type break down for v256i1 and other types to be stored correctly
      
      v256i1 on X86 without avx512 breaks down to 256 i8 values when passed between basic blocks. But the NumRegistersForVT was sized at a byte for each VT. This results in 256 being stored as 0.
      
      This patch enlarges the type to 16 bits and adds an assert to ensure that no information is lost when the entry is stored.
      
      Differential Revision: https://reviews.llvm.org/D70138
      84e83b54
    • Simon Atanasyan's avatar
      63bbbcde
    • Simon Atanasyan's avatar
    • Simon Atanasyan's avatar
      3216d284
    • Quentin Colombet's avatar
      [LiveInterval] Allow updating subranges with slightly out-dated IR · de94cda8
      Quentin Colombet authored
      During register coalescing, we update the live-intervals on-the-fly.
      To do that we are in this strange mode where the live-intervals can
      be slightly out-of-sync (more precisely they are forward looking)
      compared to what the IR actually represents.
      This happens because the register coalescer only updates the IR when
      it is done with updating the live-intervals and it has to do it this
      way because updating the IR on-the-fly would actually clobber some
      information on how the live-ranges that are being updated look like.
      
      This is problematic for updates that rely on the IR to accurately
      represents the state of the live-ranges. Right now, we have only
      one of those: stripValuesNotDefiningMask.
      To reconcile this need of out-of-sync IR, this patch introduces a
      new argument to LiveInterval::refineSubRanges that allows the code
      doing the live range updates to reason about how the code should
      look like after the coalescer will have rewritten the registers.
      Essentially this captures how a subregister index with be offseted
      to match its position in a new register class.
      
      E.g., let say we want to merge:
          V1.sub1:<2 x s32> = COPY V2.sub3:<4 x s32>
      
      We do that by choosing a class where sub1:<2 x s32> and sub3:<4 x s32>
      overlap, i.e., by choosing a class where we can find "offset + 1 == 3".
      Put differently we align V2's sub3 with V1's sub1:
          V2: sub0 sub1 sub2 sub3
          V1: <offset>  sub0 sub1
      
      This offset will look like a composed subregidx in the the class:
           V1.(composed sub2 with sub1):<4 x s32> = COPY V2.sub3:<4 x s32>
       =>  V1.(composed sub2 with sub1):<4 x s32> = COPY V2.sub3:<4 x s32>
      
      Now if we didn't rewrite the uses and def of V1, all the checks for V1
      need to account for this offset to match what the live intervals intend
      to capture.
      
      Prior to this patch, we would fail to recognize the uses and def of V1
      and would end up with machine verifier errors: No live segment at def.
      This could lead to miscompile as we would drop some live-ranges and
      thus, miss some interferences.
      
      For this problem to trigger, we need to reach stripValuesNotDefiningMask
      while having a mismatch between the IR and the live-ranges (i.e.,
      we have to apply a subreg offset to the IR.)
      
      This requires the following three conditions:
      1. An update of overlapping subreg lanes: e.g., dsub0 == <ssub0, ssub1>
      2. An update with Tuple registers with a possibility to coalesce the
         subreg index: e.g., v1.dsub_1 == v2.dsub_3
      3. Subreg liveness enabled.
      
      looking at the IR to decide what is alive and what is not, i.e., calling
      stripValuesNotDefiningMask.
      coalescer maintains for the live-ranges information.
      
      None of the targets that currently use subreg liveness (i.e., the targets
      that fulfill #3, Hexagon, AMDGPU, PowerPC, and SystemZ IIRC) expose #1 and
      and #2, so this patch also artificial enables subreg liveness for ARM,
      so that a nice test case can be attached.
      de94cda8
    • Michael Liao's avatar
      [TTI] Fix cast cost on vector types. · 2bf9b9a5
      Michael Liao authored
      - Only split vector types when both src and dst types are splittable.
      2bf9b9a5
    • Francis Visoiu Mistrih's avatar
      [llvm-bcanalyzer] Don't dump the contents if -dump is not passed · 1ca85b3d
      Francis Visoiu Mistrih authored
      With all the previous refactorings this slipped through and now we
      always dump the contents of the bitcode files, even if -dump is not
      passed.
      1ca85b3d
    • Ahmed Bougacha's avatar
      [AArch64][v8.3a] Add missing imp-defs on RETA*. · 7313d7d6
      Ahmed Bougacha authored
      RETA always implicitly uses LR, unlike RET which merely has an
      alias that defaults it to LR.
      Additionally, RETA implicitly uses SP as well, which it uses as
      a discriminator to authenticate LR.
      
      This isn't usually noticeable, because RET_ReallyLR is used in most
      of the backend.  However, the post-RA scheduler, if enabled, will
      cause miscompiles if the imp-uses are missing.
      
      While there, fix a typo in the lone affected testcase.
      7313d7d6
    • Ahmed Bougacha's avatar
      [AArch64][v8.3a] Add LDRA '[xN]!' alias. · 643ac6c0
      Ahmed Bougacha authored
      The instruction definition has been retroactively expanded to
      allow for an alias for '[xN, 0]!' as '[xN]!'.
      That wouldn't make sense on LDR, but does for LDRA.
      643ac6c0
    • Sanjay Patel's avatar
      [SLP] improve test readability; NFC · 142cbe73
      Sanjay Patel authored
      142cbe73
    • Yonghong Song's avatar
      [BPF] fix clang test failure for bpf-attr-preserve-access-index-4.c · 15831580
      Yonghong Song authored
      Depending on different cmake configures, clang may generate different
      IR name for slot variables. Let us use the regex instead of hard
      coding the name. I did the same for other bpf-attr-preserve-access-index
      tests with such an approach, but somehow did not do for this one.
      15831580
    • Edward Jones's avatar
      [RISCV] Use compiler-rt if no GCC installation detected · 3289352e
      Edward Jones authored
      If a GCC installation is not detected, then this attempts to
      use compiler-rt and the compiler-rt crtbegin/crtend
      implementations as a fallback.
      
      Differential Revision: https://reviews.llvm.org/D68407
      3289352e
    • David Stenberg's avatar
      Fix typo in DwarfDebug [NFC] · 7417cc14
      David Stenberg authored
      7417cc14
    • David Tenty's avatar
      Don't set LLVM_NO_DEAD_STRIP on AIX · 8b2b2c08
      David Tenty authored
      Summary:
      when building plugins, as AIX has symbols in it's standard library that
      must be garbage collected or we will see link errors. Export lists will
      handle this instead on AIX.
      
      Reviewers: stevewan, sfertile, jasonliu, xingxue, DiggerLin
      
      Reviewed By: DiggerLin
      
      Subscribers: mgorny, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70130
      8b2b2c08
    • Yonghong Song's avatar
      [BPF] add missing attribute in pragma-attribute-supported-attributes-list.test · f5824799
      Yonghong Song authored
      Add the newly supported BPF specific __attribute__((preserve_access_index)
      in the pragma-attribute-supported-attributes-list.test.
      f5824799
    • Sanjay Patel's avatar
    • Yonghong Song's avatar
      [BPF] Add preserve_access_index attribute for record definition · 4e2ce228
      Yonghong Song authored
      This is a resubmission for the previous reverted commit
      94343604 with the same subject. This commit fixed the
      segfault issue and addressed additional review comments.
      
      This patch introduced a new bpf specific attribute which can
      be added to struct or union definition. For example,
        struct s { ... } __attribute__((preserve_access_index));
        union u { ... } __attribute__((preserve_access_index));
      The goal is to simplify user codes for cases
      where preserve access index happens for certain struct/union,
      so user does not need to use clang __builtin_preserve_access_index
      for every members.
      
      The attribute has no effect if -g is not specified.
      
      When the attribute is specified and -g is specified, any member
      access defined by that structure or union, including array subscript
      access and inner records, will be preserved through
        __builtin_preserve_{array,struct,union}_access_index()
      IR intrinsics, which will enable relocation generation
      in bpf backend.
      
      The following is an example to illustrate the usage:
        -bash-4.4$ cat t.c
        #define __reloc__ __attribute__((preserve_access_index))
        struct s1 {
          int c;
        } __reloc__;
      
        struct s2 {
          union {
            struct s1 b[3];
          };
        } __reloc__;
      
        struct s3 {
          struct s2 a;
        } __reloc__;
      
        int test(struct s3 *arg) {
          return arg->a.b[2].c;
        }
        -bash-4.4$ clang -target bpf -g -S -O2 t.c
      
      A relocation with access string "0:0:0:0:2:0" will be generated
      representing access offset of arg->a.b[2].c.
      
      forward declaration with attribute is also handled properly such
      that the attribute is copied and populated in real record definition.
      
      Differential Revision: https://reviews.llvm.org/D69759
      4e2ce228
    • Matthew Malcomson's avatar
      e5f3760e
  2. Nov 13, 2019
    • Vedant Kumar's avatar
      [profile] Factor out logic for mmap'ing merged profile, NFC · e7aab320
      Vedant Kumar authored
      Split out the logic to get the size of a merged profile and to do a
      compatibility check. This can be shared with both the continuous+merging
      mode implementation, as well as the runtime-allocated counters
      implementation planned for Fuchsia.
      
      Lifted out of D69586.
      
      Differential Revision: https://reviews.llvm.org/D70135
      e7aab320
    • Sanjay Patel's avatar
      [InstCombine] propagate fast-math-flags (FMF) to select when inverting fcmp+select · 3d6b5398
      Sanjay Patel authored
      As noted by the FIXME comment, this is not correct based on our current FMF semantics.
      We should be propagating FMF from the final value in a sequence (in this case the
      'select'). So the behavior even without this patch is wrong, but we did not allow FMF
      on 'select' until recently.
      
      But if we do the correct thing right now in this patch, we'll inevitably introduce
      regressions because we have not wired up FMF propagation for 'phi' and 'select' in
      other passes (like SimplifyCFG) or other places in InstCombine. I'm not seeing a
      better incremental way to make progress.
      
      That said, the potential extra damage over the existing wrong behavior from this
      patch is very limited. AFAIK, the only way to have different FMF on IR in the same
      function is if we have LTO inlined IR from 2 modules that were compiled using
      different fast-math settings.
      
      As seen in the tests, we may actually see some improvements with this patch because
      adding the FMF to the 'select' allows matching to min/max intrinsics that were
      previously missed (in the common case, the 'fcmp' and 'select' should have identical
      FMF to begin with).
      
      Next steps in the transition:
      
          Make similar changes in instcombine as needed.
          Enable phi-to-select FMF propagation in SimplifyCFG.
          Remove dependencies on fcmp with FMF.
          Deprecate FMF on fcmp.
      
      Differential Revision: https://reviews.llvm.org/D69720
      3d6b5398
    • Pavel Labath's avatar
      DWARFDebugLoclists: Add an api to get the location lists of a DWARF unit · 1eea3fa0
      Pavel Labath authored
      Summary:
      This avoid the need to duplicate the location lists searching logic in
      various users. The "inline location list dumping" code (which is the
      only user actually updated to handle DWARF v5 location lists)  is
      switched to this method. After adding v4 location list support, I'll
      switch other users too.
      
      Reviewers: dblaikie, probinson, JDevlieghere, aprantl, SouraVX
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70084
      1eea3fa0
    • Simon Pilgrim's avatar
    • Simon Pilgrim's avatar
      86f07e82
    • Simon Pilgrim's avatar
      Fix uninitialized variable warning. NFCI. · e1670175
      Simon Pilgrim authored
      e1670175
    • Simon Pilgrim's avatar
      Fix uninitialized variable warning. NFCI. · 29a5a6ee
      Simon Pilgrim authored
      29a5a6ee
    • Simon Pilgrim's avatar
      Fix uninitialized variable warning. NFCI. · 6ebc5089
      Simon Pilgrim authored
      6ebc5089
    • Simon Pilgrim's avatar
      b3be859b
    • Simon Pilgrim's avatar
      PPCReduceCRLogicals - fix static analyzer warnings. NFC · 66f2ed07
      Simon Pilgrim authored
      - Fix uninitialized variable warnings.
      - Fix null dereference warnings.
      66f2ed07
    • Simon Pilgrim's avatar
      SLPVectorizer - make comparison operators + isInSchedulingRegion const · d1bd5e47
      Simon Pilgrim authored
      Fixes cppcheck warnings.
      d1bd5e47
    • Kadir Cetinkaya's avatar
      [clang][Tooling] Filter flags that generate output in SyntaxOnlyAdjuster · 16bdcc80
      Kadir Cetinkaya authored
      Summary:
      Flags that generate output could result in failures when creating
      syntax only actions. This patch introduces initial logic for filtering out
      those. The first such flag is "save-temps", which saves intermediate
      files(bitcode, assembly, etc.) into a specified directory.
      
      Fixes https://github.com/clangd/clangd/issues/191
      
      Reviewers: hokein
      
      Subscribers: ilya-biryukov, usaxena95, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D70173
      16bdcc80
    • Haojian Wu's avatar
      [clangd] Add bool return type to Index::refs API. · 33e882d5
      Haojian Wu authored
      Summary:
      Similar to fuzzyFind, the bool indicates whether there are more xref
      results.
      
      Reviewers: ilya-biryukov
      
      Reviewed By: ilya-biryukov
      
      Subscribers: merge_guards_bot, MaskRay, jkorous, arphaman, kadircet, usaxena95, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D70139
      33e882d5
    • Florian Hahn's avatar
      [InstCombine] Avoid moving ops that do restrict undef across shuffles. · f7499011
      Florian Hahn authored
      I think we have to be a bit more careful when it comes to moving
      ops across shuffles, if the op does restrict undef. For example, without
      this patch, we would move 'and %v, <0, 0, -1, -1>' over a
      'shufflevector %a, undef, <undef, undef, 1, 2>'. As a result, the first
      2 lanes of the result are undef after the combine, but they really
      should be 0, unless I am missing something.
      
      For ops that do fold to undef on undef operands, the current behavior
      should be fine. I've add conservative check OpDoesRestrictUndef, maybe
      there's a better existing utility?
      
      Reviewers: spatel, RKSimon, lebedev.ri
      
      Reviewed By: spatel
      
      Differential Revision: https://reviews.llvm.org/D70093
      f7499011
    • Luís Marques's avatar
      Revert "[RISCV] Fix wrong CFI directives" · c5b56caa
      Luís Marques authored
      test/DebugInfo/RISCV/relax-debug-frame.ll wasn't properly updated.
      c5b56caa
    • Florian Hahn's avatar
      70cc355f
    • Sjoerd Meijer's avatar
      [ARM][MVE] canTailPredicateLoop · d90804d2
      Sjoerd Meijer authored
      This implements TTI hook 'preferPredicateOverEpilogue' for MVE.  This is a
      first version and it operates on single block loops only. With this change, the
      vectoriser will now determine if tail-folding scalar remainder loops is
      possible/desired, which is the first step to generate MVE tail-predicated
      vector loops.
      
      This is disabled by default for now. I.e,, this is depends on option
      -disable-mve-tail-predication, which is off by default.
      
      I will follow up on this soon with a patch for the vectoriser to respect loop
      hint 'vectorize.predicate.enable'. I.e., with this loop hint set to Disabled,
      we don't want to tail-fold and we shouldn't query this TTI hook, which is
      done in D70125.
      
      Differential Revision: https://reviews.llvm.org/D69845
      d90804d2
    • Luís Marques's avatar
      [RISCV] Fix wrong CFI directives · a5ce8bd7
      Luís Marques authored
      Summary: Removes CFI CFA directives that could incorrectly propagate
      beyond the basic block they were inteded for. Specifically it removes
      the epilogue CFI directives. See the branch_and_tail_call test for an
      example of the issue. Should fix the stack unwinding issues caused by
      the incorrect directives.
      
      Reviewers: asb, lenary, shiva0217
      Reviewed By: lenary
      Tags: #llvm
      Differential Revision: https://reviews.llvm.org/D69723
      a5ce8bd7
    • Simon Tatham's avatar
      [ARM,MVE] Add intrinsics for contiguous load/stores. · a12f588e
      Simon Tatham authored
      This patch adds the ACLE intrinsics for all the MVE load and store
      instructions not already handled by D69791. These ones don't need new
      IR intrinsics, because they can be implemented in terms of standard
      LLVM IR constructions.
      
      Some of the load and store instructions access less than 128 bits of
      memory, sign/zero extending each value to a wider vector lane on load
      or truncating it on store. These are represented in IR by a load of a
      shorter vector followed by a zext/sext, and conversely, a trunc
      followed by a short store. Existing ISel patterns already recognize
      those combinations and turn them into the right MVE instructions.
      
      The predicated forms of all these instructions are represented in the
      same way, except that the ordinary load/store operation is replaced
      with the existing intrinsics @llvm.masked.{load,store}. These are
      currently only code-generated as predicated MVE load/store
      instructions if you give LLVM the `-enable-arm-maskedldst` option; so
      I've done that in the LLVM codegen test. When we make that the
      default, that option can be removed.
      
      In the Tablegen backend, I've had to add a handful of extra support
      features:
      
      * We need to be able to make clang::Address objects out of a
        pointer and an alignment (previously we only needed these when the
        user passed us an existing one).
      
      * We can now specify vector types that aren't 128 bits wide (for use
        in those intermediate values in IR), the parametrized type system
        can make one starting from two existing vector types (using the lane
        count of one and the element type of the other).
      
      * I've added support for code generation of pointer casts, and for
        specifying LLVM types as operands to IRBuilder operations (for zext
        and sext, though I think they'll come in useful again).
      
      * Now not all IR construction operations need to be specified as
        Builder.CreateFoo; some don't involve a Builder at all, and one
        passes it as a parameter to a tiny static helper function in
        CGBuiltin.cpp.
      
      Reviewers: ostannard, MarkMurrayARM, dmgreen
      
      Subscribers: kristof.beyls, cfe-commits, llvm-commits
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D70088
      a12f588e
    • Simon Pilgrim's avatar
      [X86][AVX] Add plausible schedule classes to MASKPAIR/VP2INTERSECT/VDPBF16PS instructions · 4d0e7b62
      Simon Pilgrim authored
      These are really just placeholders that use approximately the right resources - once we have CPUs scheduler models that support these instructions they will need revisiting.
      
      In the meantime this means that all instructions have a class of some kind., meaning models can be more easily flagged as complete.
      4d0e7b62
    • JonChesterfield's avatar
      [libomptarget] Move supporti.h to support.cu · fd9fa999
      JonChesterfield authored
      Summary:
      [libomptarget] Move supporti.h to support.cu
      Reimplementation of D69652, without the unity build and refactors.
      Will need a clean build of libomptarget as the cmakelists changed.
      
      Reviewers: ABataev, jdoerfert
      
      Reviewed By: jdoerfert
      
      Subscribers: mgorny, jfb, openmp-commits
      
      Tags: #openmp
      
      Differential Revision: https://reviews.llvm.org/D70131
      fd9fa999