1. Sep 22, 2023
    • Jeff Niu's avatar
    • Matthew Devereau's avatar
      [AArch64] Separate PNR into its own Register Class (#65306) · b967f3a1
      Matthew Devereau authored
      This patch separates PNR registers into their own register class instead
      of sharing a register class with PPR registers. This primarily allows us
      to return more accurate register classes when applying assembly
      constraints, but also more protection from supplying an incorrect
      predicate type to an invalid register operand.
      b967f3a1
    • michaelrj-google's avatar
      [libc] Fix Off By One Errors In Printf Long Double (#66957) · 5bd34e0a
      michaelrj-google authored
      Two major off-by-one errors are fixed in this patch. The first is in
      float_to_string.h with length_for_num, which wasn't accounting for the
      implicit leading bit when calculating the length of a number, causing
      a missing digit on 80 bit float max. The other off-by-one is the
      ryu_long_double_constants.h (a.k.a the Mega Table) not having any
      entries for the last POW10_OFFSET in POW10_SPLIT. This was also found on
      80 bit float max. Finally, the integer calculation mode was using a
      slightly too short integer, again on 80 bit float max, not accounting
      for the mantissa width. All of these are fixed in this patch.
      5bd34e0a
    • Alexey Bataev's avatar
      [SLP]Use source vector type as the original vector type instead of · 9a99944d
      Alexey Bataev authored
      artificial for better cost estimation.
      
      Need to use original source vector type, not the one artificially
      constructed, based on the number of vectorized scalars. It affect the
      cost significantly.
      9a99944d
    • Andres Villegas's avatar
      Reland: [sanitizer_symbolizer] Add StackTracePrinter virtual class (#66689) · f8ae2e42
      Andres Villegas authored
      Introduce a new virtual class StackTracePrinter and an implementation
      FormattedStackTracePrinter in preparation of enabling symbolizer markup
      for linux.
      This change allows us to implement other behaviour under the same api
      for StackTracePrinter, for example, MarkupStackTracePrinter.
      
      Reason for revert: A missing header file for the
      sanitizer_symbolizer_markup.cpp files.
      This was not caught in local builds or pre-merge checks given that to
      trigger the error, the code
      has to be compiled for Fuchsia.
      For this reland I've build for the fuchsia targets as well as linux.
      f8ae2e42
    • Arthur Eubanks's avatar
      [clang-format] Don't split "DPI"/"DPI-C" in Verilog imports (#66951) · e0388e0e
      Arthur Eubanks authored
      The spec doesn't allow splitting these strings and we're seeing compile
      issues with splitting it.
      
      String splitting was enabled for Verilog in
      https://reviews.llvm.org/D154093.
      e0388e0e
    • Kristof Beyls's avatar
      [BOLT] Fix data race in MCPlusBuilder::getOrCreateAnnotationIndex (#67004) · 8fb28e45
      Kristof Beyls authored
      
      
      MCPlusBuilder::getOrCreateAnnotationIndex(Name) can be called from
      different threads, for example when making use of
      ParallelUtilities::runOnEachFunctionWithUniqueAllocId.
      
      The race occurs when an Index for a particular annotation Name needs to
      be created for the first time.
      
      For example, this can easily happen when multiple "copies" of an
      analysis pass run on different BinaryFunctions, and the analysis pass
      creates a new Annotation Index to be able to store analysis results as
      annotations.
      
      This was found by using the ThreadSanitizer.
      
      No regression test was added; I don't think there is good way to write
      regression tests that verify the absence of data races?
      
      ---------
      
      Co-authored-by: default avatarAmir Ayupov <fads93@gmail.com>
      8fb28e45
    • Momchil Velikov's avatar
      [AArch64] Pre-commit some tests for D152828 (NFC) · 3769aaaf
      Momchil Velikov authored
      Generate a few of the relevant tests with `update_llc_test_checks.py`
      and pre-commit. Makes it easier to spot the differences in D152828.
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D157116
      3769aaaf
    • Momchil Velikov's avatar
      [AArch64] Correctly determine if {ADD,SUB}{W,X}rs instructions are cheap · 0eb0a65d
      Momchil Velikov authored
      These are marked to be "as cheap as a move".
      
      According to publicly available Software Optimization Guides, they
      have one cycle latency and maximum throughput only on some
      microarchitectures, only for `LSL` and only for some shift amounts.
      
      This patch uses the subtarget feature `FeatureALULSLFast` to determine
      how cheap the instructions are.
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D152827
      
      Change-Id: I8f0d7e79bcf277ebf959719991c29a1bc7829486
      0eb0a65d
    • Louis Dionne's avatar
    • Louis Dionne's avatar
    • Momchil Velikov's avatar
      [AArch64] Refactor AArch64InstrInfo::isAsCheapAsAMove (NFC) · ededcb00
      Momchil Velikov authored
          - remove `FeatureCustomCheapAsMoveHandling`: when you have target
            features affecting `isAsCheapAsAMove` that can be given on command
            line or passed via attributes, then every sub-target effectively has
            custom handling
      
          - remove special handling of `FMOVD0`/etc: `FVMOV` with an immediate
            zero operand is never[1] more expensive tha an `FMOV` with a
            register operand.
      
          - remove special handling of `COPY` - copy is trivially as cheap as
            itself
      
          - make the function default to the `MachineInstr` attribute
            `isAsCheapAsAMove`
      
          - remove special handling of `ANDWrr`/etc and of `ANDWri`/etc: the
            fallback `MachineInstr` attribute is already non-zero.
      
          - remove special handling of `ADDWri`/`SUBWri`/`ADDXri`/`SUBXri` -
            there are always[1] one cycle latency with maximum (for the
            micro-architecture) throughput
      
          - check if `MOVi32Imm`/`MOVi64Imm` can be expanded into a "cheap"
            sequence of instructions
      
            There is a little twist with determining whether a
            MOVi32Imm`/`MOVi64Imm` is "as-cheap-as-a-move". Even if one of these
            pseudo-instructions needs to be expanded to more than one MOVZ,
            MOVN, or MOVK instructions, materialisation may be preferrable to
            allocating a register to hold the constant. For the moment a cutoff
            at two instructions seems like a reasonable compromise.
      
          [1] according to 19 software optimisation manuals
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D154722
      ededcb00
    • Kazu Hirata's avatar
      Revert "[InlineCost] Check for conflicting target attributes early" · b4301df6
      Kazu Hirata authored
      This reverts commit d6f994ac.
      
      Several people have reported breakage resulting from this patch:
      
      - https://github.com/llvm/llvm-project/issues/65152
      - https://github.com/llvm/llvm-project/issues/65205
      b4301df6
    • Alexey Bataev's avatar
      [SLp]Fix a crash because of wrong deps between vectorized nodes. · 3dc28e6c
      Alexey Bataev authored
      Need to change the order of the nodes vectorization to avoid too early
      insertion of the first node.
      3dc28e6c
    • Ramkumar Ramachandra's avatar
      ISel/RISCV: remove dead code corresponding to VP_FSH[L|R] (#67035) · 3b3ff5c1
      Ramkumar Ramachandra authored
      70de0e ([VP][RISCV] Add vp.fshl/fshr and RISC-V support.) introduced
      VP_FSHL and VP_FSHR, by using a generic expansion for all targets: the
      core of this change is in TargetLowering. However, the commit
      erroneously introduced dead code in RISCVISelLowering. Remove this dead
      code.
      3b3ff5c1
    • Justin Lebar's avatar
      Set -rpath-link only if the path is nonempty. · 36b87d80
      Justin Lebar authored
      This cmake rule is used by external clients, who may or may not have
      the LLVM_LIBRARY_OUTPUT_INTDIR variable set.
      
      If it is not set, then we pass `-Wl,-rpath-link,` to the compiler.  It
      turns out that gcc and clang interpret this differently.
      
        * gcc passes `-rpath-link ""` to the linker, which is what we want.
      
        * clang passes `-rpath-link` to the linker.  This is not what we want,
          because then the linker gobbles the next command-line argument,
          whatever it happens to be, and uses it as the -rpath-link target.
      
      Fix this by passing -rpath-link only if we actually have a path we want.
      36b87d80
    • Joseph Huber's avatar
      [libc][NFC] Remove unused function from the RPC server · e2bc0f92
      Joseph Huber authored
      Summary:
      I missed removing this now-unused function in the previous patch. Remove
      it to clean up the interface.
      e2bc0f92
    • jeanPerier's avatar
      [flang] Centralize automatic deallocation code in lowering (#67003) · 2cb31fe8
      jeanPerier authored
      There are currently several places that automatically deallocate
      allocatble if they are allocated:
       - INTENT(OUT) allocatable are deallocated on entry in the callee
      - INTENT(OUT) allocatable are also deallocated on the caller side of
      BIND(C) function in case the implementation is in C.
      - Results of function returning allocatable are deallocated after usage.
      - OPENMP privatized allocatable are deallocated at the end of OPENMP
      region.
      
      Introduce genDeallocateIfAllocated that centralize all this code, except
      for the function return that use genFreememIfAllocated since
      finalization is done separately currently.
      
      `fir::factory::genFinalization` and
      `fir::factory::genInlinedDeallocation` are removed and replaced by
      genFreemem since their name were misleading: finalization was not
      called.
      
      There is a fallout in the tests because previous generated code did not
      check the allocated status when doing inline deallocation. This was OK
      since free(null) is guaranteed to be a no-op, but this makes compiler
      code more complex, is a bit surprising in the generated IR IMHO, and it
      relied on knowing when genDeallocateBox inserts runtime calls or uses
      inlined code.
      2cb31fe8
    • Amir Bishara's avatar
      [mlir][tosa] Constant optimizations for reduce operations · f5f7e2a3
      Amir Bishara authored
      Replace the different reduce operations which is getting
      a constant tensor as an input argument with a constant
      tensor.
      
      As the arguement of the reduce operation is constant tensor
      and has only a single user we could calculate the resulted
      constant tensor in compilation time and replace it
      with reduced memory tensor
      
      This optimization has been implemented for:
      tosa.reduce_sum
      tosa.reduce_prod
      tosa.reduce_any
      tosa.reduce_all
      tosa.reduce_max
      tosa.reduce_min
      
      Reviewed By: rsuderman
      
      Differential Revision: https://reviews.llvm.org/D154832
      f5f7e2a3
    • Dávid Ferenc Szabó's avatar
      [CodeGen] Improve compilation time with VLIWMachineScheduler (#66942) · a7612e2e
      Dávid Ferenc Szabó authored
      A straight forward improvement which can already achieve 2x speed up in
      some cases like the one here:
      https://github.com/llvm/llvm-project/issues/65946.
      a7612e2e
    • Matthias Springer's avatar
      [mlir][Interfaces][NFC] Better documentation for `RegionBranchOpInterface` (#66920) · d56537a5
      Matthias Springer authored
      Update outdated documentation and add an example.
      d56537a5
    • Ingo Müller's avatar
      [mlir][memref][transform] Add new alloca_to_global op. (#66511) · 991cb147
      Ingo Müller authored
      This PR adds a new transform op that replaces `memref.alloca`s with
      `memref.get_global`s to newly inserted `memref.global`s. This is useful,
      for example, for allocations that should reside in the shared memory of
      a GPU, which have to be declared as globals.
      991cb147
    • Joseph Huber's avatar
      [libc] Remove the 'rpc_reset' routine from the RPC implementation (#66700) · 59896c16
      Joseph Huber authored
      Summary:
      This patch removes the `rpc_reset` function. This was previously used to
      initialize the RPC client on the device by setting up the pointers to
      communicate with the server. The purpose of this was to make it easier
      to initialize the device for testing. However, this prevented us from
      enforcing an invariant that the buffers are all read-only from the
      client side.
      
      The expected way to initialize the server is now to copy it from the
      host runtime. This will allow us to maintain that the RPC client is in
      the constant address space on the GPU, potentially through inference,
      and improving caching behaviour.
      59896c16
    • Ramkumar Ramachandra's avatar
      CostModel/RISCV: fix typos in fround test, vector length (#67025) · 88800f79
      Ramkumar Ramachandra authored
      There are several typos in fround.ll, persumably caused by copy-pasting,
      where there is a strange nvx5* type. From the surrounding code, it is
      clear that this was intended to be nvx4*. Fix these typos.
      88800f79
    • Matthias Springer's avatar
      [mlir][Interfaces] Clean up `DestinationStyleOpInterface` (#67015) · 0b2197b0
      Matthias Springer authored
      * "init" operands are specified with `MutableOperandRange` (which gives
      access to the underlying `OpOperand *`). No more magic numbers.
      * Remove most interface methods and make them helper functions. Only
      `getInitsMutable` should be implemented.
      * Provide separate helper functions for accessing mutable/immutable
      operands (`OpOperand`/`Value`, in line with #66515): `getInitsMutable`
      and `getInits` (same naming convention as auto-generated op accessors).
      `getInputOperands` was not renamed because this function cannot return a
      `MutableOperandRange` (because the operands are not necessarily
      consecutive). `OpOperandVector` is no longer needed.
      * The new `getDpsInits`/`getDpsInitsMutable` is more efficient than the
      old `getDpsInitOperands` because no `SmallVector` is created. The new
      functions return a range of operands.
      * Fix a bug in `getDpsInputOperands`: out-of-bounds operands were
      potentially returned.
      0b2197b0
  2. Sep 21, 2023