1. Jun 22, 2020
  2. Jun 21, 2020
    • David Green's avatar
      67121d7b
    • Florian Hahn's avatar
      [DSE,MSSA] Move reachability check to main loop. · 40569db7
      Florian Hahn authored
      As we traverse the CFG backwards, we could end up reaching unreachable
      blocks. For unreachable blocks, we won't have computed post order
      numbers and because DomAccess is reachable, unreachable blocks cannot be
      on any path from it.
      
      This fixes a crash with unreachable blocks.
      40569db7
    • Luboš Luňák's avatar
      add option to instantiate templates already in the PCH · a45f713c
      Luboš Luňák authored
      Add -fpch-instantiate-templates which makes template instantiations be
      performed already in the PCH instead of it being done in every single
      file that uses the PCH (but every single file will still do it as well
      in order to handle its own instantiations). I can see 20-30% build
      time saved with the few tests I've tried.
      
      The change may reorder compiler output and also generated code, but
      should be generally safe and produce functionally identical code.
      There are some rare cases that do not compile with it,
      such as test/PCH/pch-instantiate-templates-forward-decl.cpp. If
      template instantiation bailed out instead of reporting the error,
      these instantiations could even be postponed, which would make them
      work.
      
      Enable this by default for clang-cl. MSVC creates PCHs by compiling
      them using an empty .cpp file, which means templates are instantiated
      while building the PCH and so the .h needs to be self-contained,
      making test/PCH/pch-instantiate-templates-forward-decl.cpp to fail
      with MSVC anyway. So the option being enabled for clang-cl matches this.
      
      Differential Revision: https://reviews.llvm.org/D69585
      a45f713c
    • David Green's avatar
      [CGP] Convert phi types · 730ecb63
      David Green authored
      If a collection of interconnected phi nodes is only ever loaded, stored
      or bitcast then we can convert the whole set to the bitcast type,
      potentially helping to reduce the number of register moves needed as the
      phi's are passed across basic block boundaries. This has to be done in
      CodegenPrepare as it naturally straddles basic blocks.
      
      The alorithm just looks from phi nodes, looking at uses and operands for
      a collection of nodes that all together are bitcast between float and
      integer types. We record visited phi nodes to not have to process them
      more than once. The whole subgraph is then replaced with a new type.
      Loads and Stores are bitcast to the correct type, which should then be
      folded into the load/store, changing it's type.
      
      This comes up in the biquad testcase due to the way MVE needs to keep
      values in integer registers. I have also seen it come up from aarch64
      partner example code, where a complicated set of sroa/inlining produced
      integer phis, where float would have been a better choice.
      
      I also added undef and extract element handling which increased the
      potency in some cases.
      
      This adds it with an option that defaults to off, and disabled for 32bit
      X86 due to potential issues around canonicalizing NaNs.
      
      Differential Revision: https://reviews.llvm.org/D81827
      730ecb63
    • David Green's avatar
      [CGP][AArch64] Convert Phi type tests. NFC · 0ee21cdb
      David Green authored
      0ee21cdb
    • Nikita Popov's avatar
      [ValueTracking, BasicAA] Don't simplify instructions · 37d30307
      Nikita Popov authored
      GetUnderlyingObject() (and by required symmetry
      DecomposeGEPExpression()) will call SimplifyInstruction() on the
      passed value if other checks fail. This simplification is very
      expensive, but has little effect in practice. This patch removes
      the SimplifyInstruction call(), and replaces it with a check for
      single-argument phis (which can occur in canonical IR in LCSSA
      form), which is the only useful simplification case I was able to
      identify.
      
      At O3 the geomean CTMark improvement is -1.7%. The largest
      improvement is SPASS with ThinLTO at -6%.
      
      In test-suite, I see only two tests with a hash difference and
      no code size difference (PAQ8p, Ptrdist), which indicates that
      the simplification only ends up being useful very rarely. (I would
      have liked to figure out which simplification is responsible here,
      but wasn't able to spot it looking at transformation logs.)
      
      The AMDGPU test case that is update was using two selects with
      undef condition, in which case GetUnderlyingObject will return
      the first select operand as the underlying object. This will of
      course not happen with non-undef conditions, so this was not
      testing anything realistic. Additionally this illustrates potential
      unsoundness: While GetUnderlyingObject will pick the first operand,
      the select might be later replaced by the second operand, resulting
      in inconsistent assumptions about the undef value.
      
      Differential Revision: https://reviews.llvm.org/D82261
      37d30307
    • Bruno Ricci's avatar
      Revert "Add --hot-func-list to llvm-profdata show for sample profiles" · 5342dd6b
      Bruno Ricci authored
      This reverts commit 7348b951.
      It is causing Asan failures.
      5342dd6b
    • Sanjay Patel's avatar
      [ValueTracking] improve analysis for fdiv with same operands · 2ad42c26
      Sanjay Patel authored
      (The 'nnan' variant of this pattern is already tested to produce '1.0'.)
      
      https://alive2.llvm.org/ce/z/D4hPBy
      
      define i1 @src(float %x, i32 %y) {
      %0:
        %d = fdiv float %x, %x
        %uge = fcmp uge float %d, 0.000000
        ret i1 %uge
      }
      =>
      define i1 @tgt(float %x, i32 %y) {
      %0:
        ret i1 1
      }
      Transformation seems to be correct!
      2ad42c26
    • Sanjay Patel's avatar
      97c02326
    • Bruno Ricci's avatar
      [clang][test][NFC] Also test for serialization in AST dump tests, part 3/n. · cddc9993
      Bruno Ricci authored
      The outputs between the direct ast-dump test and the ast-dump test after
      deserialization should match modulo a few differences.
      
      For hand-written tests, strip the "<undeserialized declarations>"s and
      the "imported"s with sed.
      
      For tests generated with "make-ast-dump-check.sh", regenerate the output.
      
      Part 3/n.
      cddc9993
    • Bruno Ricci's avatar
      [clang][test][NFC] Also test for serialization in AST dump tests, part 2/n. · ecbf2f5f
      Bruno Ricci authored
      The outputs between the direct ast-dump test and the ast-dump test after
      deserialization should match modulo a few differences.
      
      For hand-written tests, strip the "<undeserialized declarations>"s and
      the "imported"s with sed.
      
      For tests generated with "make-ast-dump-check.sh", regenerate the
      output.
      
      Part 2/n.
      ecbf2f5f
    • Bruno Ricci's avatar
    • Bruno Ricci's avatar
      [clang][utils] Minor tweak to make-ast-dump-check.sh · 0dbeffdd
      Bruno Ricci authored
      Remove the space after the "CHECK:" on each line. This space makes the use
      of FileCheck --match-full-lines impossible.
      0dbeffdd
    • Bruno Ricci's avatar
      [clang][Serialization] Fix the serialization of ConstantExpr. · e7ce0528
      Bruno Ricci authored
      The serialization of ConstantExpr has currently a number of problems:
      
      - Some fields are just not serialized (ConstantExprBits.APValueKind and
        ConstantExprBits.IsImmediateInvocation).
      
      - ASTStmtReader::VisitConstantExpr forgets to add the trailing APValue
        to the list of objects to be destroyed when the APValue needs cleanup.
      
      While we are at it, bring the serialization of ConstantExpr more in-line
      with what is done with the other expressions by doing the following NFCs:
      
      - Get rid of ConstantExpr::DefaultInit. It is better to not initialize
        the fields of an empty ConstantExpr since this will allow msan to
        detect if a field was not deserialized.
      
      - Move the initialization of the fields of ConstantExpr to the constructor;
        ConstantExpr::Create allocates the memory and ConstantExpr::ConstantExpr
        is responsible for the initialization.
      
      Review after commit since this is a straightforward mechanical fix
      similar to the other serialization fixes.
      e7ce0528
    • Bruno Ricci's avatar
      [clang][NFC] Fix typos/wording in the comments of ConstantExpr. · ef3adbfc
      Bruno Ricci authored
      It is "trailing objects" and "tail-allocated storage".
      ef3adbfc
    • Nikita Popov's avatar
      [LangRef] Fix sphinx warnings · 93a0f0e4
      Nikita Popov authored
      93a0f0e4
    • Nikita Popov's avatar
      f26b4201
    • Simon Pilgrim's avatar
      [X86][SSE] Add SimplifyDemandedVectorEltsForTargetShuffle to handle target shuffle variable masks · fb9f9dc3
      Simon Pilgrim authored
      Pulled out from the ongoing work on D66004, currently we don't do a good job of simplifying variable shuffle masks that have already lowered to constant pool entries.
      
      This patch adds SimplifyDemandedVectorEltsForTargetShuffle (a custom x86 helper) to first try SimplifyDemandedVectorElts (which we already do) and then constant pool simplification to help mark undefined elements.
      
      To prevent lowering/combines infinite loops, we only handle basic constant pool loads instead of creating new BUILD_VECTOR nodes for lowering - e.g. we don't try to convert them to broadcast/vzext_load - there might be some benefit to this but if so I'd rather we come up with some way to reuse existing code than reimplement a lot of BUILD_VECTOR code.
      
      Differential Revision: https://reviews.llvm.org/D81791
      fb9f9dc3
    • clfbbn's avatar
      [Attributor][NFC] Fix indentation · 10b05397
      clfbbn authored
      Summary: The patch D81022 seems to break the indentation of the `cleanupIR()` function. This patch fixes this problem
      
      Reviewers: jdoerfert, sstefan1, uenoku
      
      Reviewed By: jdoerfert
      
      Subscribers: hiraditya, uenoku, kuter, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D82260
      10b05397
    • Wenlei He's avatar
      [Remarks] Add callsite locations to inline remarks · 7c8a6936
      Wenlei He authored
      Summary:
      Add call site location info into inline remarks so we can differentiate inline sites.
      This can be useful for inliner tuning. We can also reconstruct full hierarchical inline
      tree from parsing such remarks. The messege of inline remark is also tweaked so we can
      differentiate SampleProfileLoader inline from CGSCC inline.
      
      Reviewers: wmi, davidxl, hoy
      
      Subscribers: hiraditya, cfe-commits, llvm-commits
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D82213
      7c8a6936
    • Jonas Devlieghere's avatar
      6e3faaeb
    • Jonas Devlieghere's avatar
      e13fca4f
    • Amy Kwan's avatar
      [PowerPC][Power10] Implement Vector Clear Left/Rightmost Bytes Builtins in LLVM/Clang · cc95635b
      Amy Kwan authored
      This patch implements builtins for the following prototypes:
      ```
      vector signed char vec_clrl (vector signed char a, unsigned int n);
      vector unsigned char vec_clrl (vector unsigned char a, unsigned int n);
      vector signed char vec_clrr (vector signed char a, unsigned int n);
      vector signed char vec_clrr (vector unsigned char a, unsigned int n);
      ```
      
      Differential Revision: https://reviews.llvm.org/D81707
      cc95635b
    • Eric Christopher's avatar
      [clang/llvm] As part of using inclusive language within · 0861889b
      Eric Christopher authored
      the llvm project, migrate away from the use of blacklist and whitelist.
      0861889b
    • Craig Topper's avatar
      [X86] Set the cpu_vendor in __cpu_indicator_init to VENDOR_OTHER if cpuid... · 35f7d583
      Craig Topper authored
      [X86] Set the cpu_vendor in __cpu_indicator_init to VENDOR_OTHER if cpuid isn't supported on the CPU.
      
      We need to set the cpu_vendor to a non-zero value to indicate
      that we already called __cpu_indicator_init once.
      
      This should only happen on a 386 or 486 CPU.
      35f7d583
    • Eric Christopher's avatar
      [clang-tidy] As part of using inclusive language within · da6332f5
      Eric Christopher authored
      the llvm project, migrate away from the use of blacklist and whitelist.
      da6332f5