1. Mar 30, 2021
    • Craig Topper's avatar
      [RISCV] When custom iseling masked loads/stores, copy the mask into V0 instead of virtual register. · 3dd4aa7d
      Craig Topper authored
      This matches what we do in our isel patterns. In our internal
      testing we've found this is needed to make the fast register
      allocator happy at -O0. Otherwise it may assign V0 to an earlier
      operand and find itself with no registers left when it reaches
      the mask operand. By using V0 explicitly, the fast register allocator
      will see it when it checks for phys register usages before it
      starts allocating vregs. I'll try to update this with a test case.
      
      Unfortunately, this does appear to prevent some instruction reordering
      by the pre-RA scheduler which leads to the increased spills seen in
      some tests. I suspect that problem could already occur for other
      instructions that already used V0 directly.
      
      There's a lot of repeated code here that could do with some
      wrapper functions. Not sure if that should be at the level of the
      new code that deals with V0. That would require multiple output
      parameters to pass the glue, chain and register back. Maybe it
      should be at a higher level over the entire set of push_backs.
      
      Reviewed By: frasercrmck, HsiangKai
      
      Differential Revision: https://reviews.llvm.org/D99367
      3dd4aa7d
    • Peter Steinfeld's avatar
      [flang] Fix CHECK() calls on erroneous procedure declarations · a7afc8a5
      Peter Steinfeld authored
      When writing tests for a previous problem, I ran across situations where the
      compiler was failing calls to CHECK().  In these situations, the compiler had
      inconsistent semantic information because the programs were erroneous.  This
      inconsistent information was causing the calls to CHECK().
      
      I fixed this by avoiding the code that ended up making the failed calls to
      CHECK() and making sure that we were only avoiding these situations when the
      associated symbols were erroneous.
      
      I also added tests that would cause the calls to CHECK() without these changes.
      
      Differential Revision: https://reviews.llvm.org/D99342
      a7afc8a5
    • Craig Topper's avatar
      [X86] Always use rip-relative addressing on 64-bit when rematerializing all... · 54bacaf3
      Craig Topper authored
      [X86] Always use rip-relative addressing on 64-bit when rematerializing all zeros/ones registers using a folded load.
      
      Previously we only used RIP relative when PIC was enabled. But
      we know we're in small/kernel code model here so we should
      be able to always use RIP-relative which will give a smaller
      encoding.
      
      Here's a godbolt link that demonstrates the current codegen https://godbolt.org/z/j3158o
      Note in the non-PIC version the load from .LCPI0_0 doesn't use
      RIP-relative addressing, but if you change the constant in the
      source from 0.0 to 1.0 it will become RIP-relative.
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D97208
      54bacaf3
    • Roger Ferrer Ibanez's avatar
      [RISCV] Fix offset computation for RVV · ef76a333
      Roger Ferrer Ibanez authored
      In D97111 we changed the RVV frame layout when using sp or bp to address
      the stack slots so we could address the emergency stack slot. The idea
      is to put the RVV objects as far as possible (in offset terms) from the
      frame reference register (sp / fp / bp).
      
      When using fp this happens naturally because the RVV objects are already
      the top of the stack and due to the constraints of RVV (VLENB being a
      power of two >= 128) the stack remains aligned. The rest of this summary
      does not apply to this case.
      
      When using sp / bp we need to skip the non-RVV stack slots. The size of
      the the non-RVV objects is computed subtracting the callee saved
      register size (whose computation is added in D97111 itself) to the total
      size of the stack (which does not account for RVV stack slots). However,
      when doing so we round to 16 bytes when computing that size and we end
      emitting a smaller offset that may belong to a scalar stack slot (see
      D98801). So this change removes that rounding.
      
      Also, because we want the RVV objects be between the non-RVV stack slots
      and the callee-saved register slots, we need to make sure the RVV
      objects are properly aligned to 8 bytes. Adding a padding of 8 would
      render the stack unaligned. So when allocating space for RVV (only when
      we don't use fp) we need to have extra padding that preserves the stack
      alignment. This way we can round to 8 bytes the offset that skips the
      non-RVV objects and we do not misalign the whole stack in the way. In
      some circumstances this means that the RVV objects may have padding
      before (=lower offsets from sp/bp) and after (before the CSR stack
      slots).
      
      Differential Revision: https://reviews.llvm.org/D98802
      ef76a333
    • Roger Ferrer Ibanez's avatar
      [NFC][RISCV] Add test showing wrong stack slot for GPR and RVV spilled registers · 3abd0bac
      Roger Ferrer Ibanez authored
      This testcase shows that we attempt to assign the same offset sp + 16 to
      two different stack objects.
      
      The fix will come in a later change.
      
      Differential Revision: https://reviews.llvm.org/D98801
      3abd0bac
    • Roger Ferrer Ibanez's avatar
      [NFC][RISCV] Pass file through update_llc_tests to fix whitespace issues · 96d14ff5
      Roger Ferrer Ibanez authored
      While addressing RVV frame layout issues I found this file had
      whitespace differences that made diffs noisier than they should be.
      
      Differential Revision: https://reviews.llvm.org/D98800
      96d14ff5
    • Wenlei He's avatar
      [CSSPGO][llvm-profgen] Context-sensitive global pre-inliner · 30b02323
      Wenlei He authored
      This change sets up a framework in llvm-profgen to estimate inline decision and adjust context-sensitive profile based on that. We call it a global pre-inliner in llvm-profgen.
      
      It will serve two purposes:
        1) Since context profile for not inlined context will be merged into base profile, if we estimate a context will not be inlined, we can merge the context profile in the output to save profile size.
        2) For thinLTO, when a context involving functions from different modules is not inined, we can't merge functions profiles across modules, leading to suboptimal post-inline count quality. By estimating some inline decisions, we would be able to adjust/merge context profiles beforehand as a mitigation.
      
      Compiler inline heuristic uses inline cost which is not available in llvm-profgen. But since inline cost is closely related to size, we could get an estimate through function size from debug info. Because the size we have in llvm-profgen is the final size, it could also be more accurate than the inline cost estimation in the compiler.
      
      This change only has the framework, with a few TODOs left for follow up patches for a complete implementation:
        1) We need to retrieve size for funciton//inlinee from debug info for inlining estimation. Currently we use number of samples in a profile as place holder for size estimation.
        2) Currently the thresholds are using the values used by sample loader inliner. But they need to be tuned since the size here is fully optimized machine code size, instead of inline cost based on not yet fully optimized IR.
      
      Differential Revision: https://reviews.llvm.org/D99146
      30b02323
    • Florian Hahn's avatar
      [Clang] Fix line numbers in CHECK lines. · d3ff65dc
      Florian Hahn authored
      d3ff65dc
    • Wei Mi's avatar
      [SampleFDO] Do not scale the magic number NOMORE_ICP_MAGICNUM in value profile · 3cbf4419
      Wei Mi authored
      during profile update.
      
      When we inline a function and update the profile, the value profiles of the
      indirect call in the inliner and inlinee will be scaled. In
      https://reviews.llvm.org/D96806 and https://reviews.llvm.org/D97350, we start
      using the magic number NOMORE_ICP_MAGICNUM (-1) to mark targets which have
      been promoted. The magic number shouldn't be scaled during the profile update.
      
      Although the problem has been suppressed by https://reviews.llvm.org/D98187
      for SampleFDO, which stops profile update for inlining in sampleFDO, the patch
      is still wanted since it will be more consistent to handle the magic number
      properly in profile update.
      
      Differential Revision: https://reviews.llvm.org/D99394
      3cbf4419
    • Florian Hahn's avatar
      [Clang] Only run test when X86 backend is built. · 9320ac9b
      Florian Hahn authored
      After c773d0f9 the remark is only emitted if the loop is profitable
      to vectorize, but cannot be vectorized. Hence, it depends on
      X86-specific cost-modeling.
      9320ac9b
    • Jonas Devlieghere's avatar
      [lldb] Move UpdateISAToDescriptorMap into ClassInfoExtractor (NFC) · bf8cbfa6
      Jonas Devlieghere authored
      Move UpdateISAToDescriptorMap into ClassInfoExtractor so that all the
      formerly public functions can be private and remain an implementation
      detail of the extractor.
      
      Differential revision: https://reviews.llvm.org/D99448
      bf8cbfa6
    • Joseph Huber's avatar
      [OpenMP] Trim error messages in CUDA plugin · 29338459
      Joseph Huber authored
      Summary:
      Remove some of the error messages printed when the CUDA plugin fails. The current error messages can be confusing because they are the first error messages printed after the async stream finds an error. This means that the printed values aren't related to what caused the issue, but are simply the last asyncronous operation that succeeded on the device. Remove these as they can be misleading.
      
      Reviewers: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D99510
      29338459
    • MaheshRavishankar's avatar
      [mlir][Linalg] Rewrite SubTensors that take a slice out of a unit-extend dimension. · f0a2fe7f
      MaheshRavishankar authored
      Subtensor operations that are taking a slice out of a tensor that is
      unit-extent along a dimension can be rewritten to drop that dimension.
      
      Differential Revision: https://reviews.llvm.org/D99226
      f0a2fe7f
    • Asher Mancinelli's avatar
      [flang] Update output format test to use GTest · e8515ca8
      Asher Mancinelli authored
      Better document each test in output formatting tests. Use GTest primitives and infrastructure in same
      spirit as [[ https://reviews.llvm.org/D97403 | D97403 ]]. [[ https://github.com/flang-compiler/f18/issues/995#issuecomment-790737912 | See legacy github issue linked here ]] for additional context. Reorganize long test cases to be more readable.
      
      Reviewed By: awarzynski, klausler
      
      Differential Revision: https://reviews.llvm.org/D98303
      e8515ca8
    • MaheshRavishankar's avatar
      [mlir][Linalg] Drop spurious error message · 7d8b478c
      MaheshRavishankar authored
      Drop usage of `emitRemark` and use `notifyMatchFailure` instead to
      avoid unnecessary spew during compilation.
      
      Differential Revision: https://reviews.llvm.org/D99485
      7d8b478c
    • Christopher Di Bella's avatar
      [libcxx] adds std::identity to <functional> · 24c44c37
      Christopher Di Bella authored
      Implements parts of:
          - P0898R3 Standard Library Concepts
      
      Differential Revision: https://reviews.llvm.org/D98151
      24c44c37
    • Fanbo Meng's avatar
      [SystemZ][z/OS] Add test of leading zero length bitfield in const/volatile struct · f1e0c7fd
      Fanbo Meng authored
      Reviewed By: abhina.sreeskantharajan
      
      Differential Revision: https://reviews.llvm.org/D99508
      f1e0c7fd
  2. Mar 29, 2021