1. Sep 15, 2023
    • Christopher Bate's avatar
      [mlir][VectorToGPU] Update memref stride preconditions on `nvgpu.mma.sync` path · cafb6284
      Christopher Bate authored
      This change removes the requirement that the row stride be statically known when
      converting `vector.transfer_read` and `vector.transfer_write` to distributed
      SIMT operations in the `nvgpu` lowering path. It also adds a check to verify
      that the last dimension of the source memref is statically known to have stride
      1 since this is assumed in the conversion logic.  No other change should be
      required since the generated `vector.load` operations are never created across
      dimensions other than the last. The routines for checking preconditions on
      `vector.transfer_read/write` are moved to under nvgpu utilities.
      
      The change is NFC with respect to the GPU dialect lowering path.
      
      Reviewed By: ThomasRaoux
      
      Differential Revision: https://reviews.llvm.org/D155753
      cafb6284
    • Adrian Prantl's avatar
      e8ad9b0e
    • Alexey Bataev's avatar
      [SLP]Do not account non-instructions for external use. · c15c1e5d
      Alexey Bataev authored
      If the non-instruction gets vectorized, no need to account its extract
      cost, it won't be removed and replaced by extractelement instruction.
      c15c1e5d
    • Alexey Bataev's avatar
    • Christopher Ferris's avatar
      [scudo] Add -Wconversion for tests and clean-up warnings. (#66147) · fd1721d8
      Christopher Ferris authored
      Fix all the places where the tests are doing implicit conversions.
      fd1721d8
    • Louis Dionne's avatar
      [libc++] Make sure LWG2070 is implemented as a DR (#65998) · 85f27d12
      Louis Dionne authored
      When we implemented C++20's P0674R1, we didn't enable the part of
      P0674R1 that was resolving LWG2070 as a DR. This patch fixes that and
      makes sure that we consistently go through the allocator when
      constructing and destroying the underlying object in
      std::allocate_shared.
      
      Fixes #54365.
      85f27d12
    • Amir Ayupov's avatar
      [BOLT][NFC] Simplify RI::selectFunctionsToProcess · 4a6426a8
      Amir Ayupov authored
      Reviewed By: #bolt, maksfb
      
      Differential Revision: https://reviews.llvm.org/D159516
      4a6426a8
    • Kinuko Yasuda's avatar
      [clang][dataflow] Ignore assignment where base class's operator is used (#66364) · 0612c9b0
      Kinuko Yasuda authored
      In C++ it seems it is legit to use base class's operator (e.g. `using
      Base::operator=`) to perform copy if the base class is the common
      ancestor of the source and destination object. In such a case we
      shouldn't try to access fields beyond that of the base class, however
      such a case seems to be very rare (typical code would implement a copy
      constructor instead), and could add complexities, so in this patch we
      simply bail if the method operator's parent class is different from the
      type of the destination object that this framework recognizes.
      0612c9b0
    • Leonard Chan's avatar
      Reland "[clang] Add experimental option to omit the RTTI component from the... · f45f1c35
      Leonard Chan authored
      Reland "[clang] Add experimental option to omit the RTTI component from the vtable when -fno-rtti is used"
      
      This reverts commit 070493dd (and
      relands the original change). This removes a test run that makes an
      assumption of RTTI being on by default for a given target.
      f45f1c35
    • LLVM GN Syncbot's avatar
      [gn build] Port 71e36426 · a126d61e
      LLVM GN Syncbot authored
      a126d61e
    • Kuba (Brecka) Mracek's avatar
      [AArch64] Relax binary format switch in AArch64MCInstLower::LowerSymbolOperand... · 454cc366
      Kuba (Brecka) Mracek authored
      [AArch64] Relax binary format switch in AArch64MCInstLower::LowerSymbolOperand to allow non-Darwin Mach-O files (#66011)
      
      Trying to use a arm64-apple-none-macho target triple today crashes with
      an assertion, this patch fixes that.
      454cc366
    • Matthias Braun's avatar
      Avoid BlockFrequency overflow problems (#66280) · b0c8c454
      Matthias Braun authored
      Multiplying raw block frequency with an integer carries a high risk
      of overflow.
      
      - Add `BlockFrequency::mul` return an std::optional with the product
        or `nullopt` to indicate an overflow.
      - Fix two instances where overflow was likely.
      b0c8c454
    • Danila Malyutin's avatar
      [NFC] Add test for #66382 · e80a8b4a
      Danila Malyutin authored
      e80a8b4a
    • Justin Bogner's avatar
      [github] Simplify DirectX backend labeling (#66407) · 11f0a632
      Justin Bogner authored
      Based on how AMDGPU and RISCV set up their labels - this way does a
      better job of catching everything with less maintenance burden.
      11f0a632
    • Justin Bogner's avatar
      [Transforms][DXIL] Wire up a basic DXILUpgrade pass (#66275) · 71e36426
      Justin Bogner authored
      This pass will upgrade DXIL-style llvm constructs (which are mostly
      metadata) into the representations we use in LLVM for the same concepts.
      
      For now we just strip the valver metadata, which we don't need. Later
      changes will make this pass more useful, and then we should be able to
      wire it into clang and possibly the DirectX backend's AsmParser.
      71e36426
    • Alex Langford's avatar
      [lldb][NFCI] Remove use of ConstString in StructuredData · a5a2a5a3
      Alex Langford authored
      The remaining use of ConstString in StructuredData is the Dictionary
      class. Internally it's backed by a `std::map<ConstString, ObjectSP>`.
      I propose that we replace it with a `llvm::StringMap<ObjectSP>`.
      
      Many StructuredData::Dictionary objects are ephemeral and only exist for
      a short amount of time. Many of these Dictionaries are only produced
      once and are never used again. That leaves us with a lot of string data
      in the ConstString StringPool that is sitting there never to be used
      again. Even if the same string is used many times for keys of different
      Dictionary objects, that is something we can measure and adjust for
      instead of assuming that every key may be reused at some point in the
      future.
      
      Quick comparisons of key data is likely not a concern with Dictionary,
      but the use of `llvm::StringMap` means that lookups should be fast with
      its hashing strategy.
      
      Switching to a llvm::StringMap meant that the iteration order may be
      different. To account...
      a5a2a5a3
    • Christopher Bate's avatar
      [mlir][Transform] Add `updateConversionTarget` to `ConversionPatternDescriptorOpInterface` · e2d39f79
      Christopher Bate authored
      This change adds a method to modify the ConversionTarget used during
      `transform.apply_conversion_patterns` to the
      `ConversionPatternDescriptorOpInterface`. This is needed when the TypeConverter
      is used to dictate the dynamic legality of operations, as in "structural"
      conversion patterns present in, for example, the SCF and func dialects.
      
      As a first use case/test, this change also adds a
      `transform.apply_patterns.scf.structural_conversions` operation to the SCF
      dialect.
      
      Reviewed By: springerm
      
      Differential Revision: https://reviews.llvm.org/D158672
      e2d39f79
    • Fangrui Song's avatar
      [ELF] Align the end of PT_GNU_RELRO associated PT_LOAD to a common-page-size boundary (#66042) · 5a58e98c
      Fangrui Song authored
      Close #57618: currently we align the end of PT_GNU_RELRO to a
      common-page-size
      boundary, but do not align the end of the associated PT_LOAD. This is
      benign
      when runtime_page_size >= common-page-size.
      
      However, when runtime_page_size < common-page-size, it is possible that
      `alignUp(end(PT_LOAD), page_size) < alignDown(end(PT_GNU_RELRO),
      page_size)`.
      In this case, rtld's mprotect call for PT_GNU_RELRO will apply to
      unmapped
      regions and lead to an error, e.g.
      
      ```
      error while loading shared libraries: cannot apply additional memory protection after relocation: Cannot allocate memory
      ```
      
      To fix the issue, add a padding section .relro_padding like mold, which
      is contained in the PT_GNU_RELRO segment and the associated PT_LOAD
      segment. The section also prevents strip from corrupting PT_LOAD program
      headers.
      
      .relro_padding has the largest `sortRank` among RELRO sections.
      Therefore, it is naturally placed at the end of `PT_GNU_RELRO` segment
      in the absence of `PHDRS`/`SECTIONS` commands.
      
      In the presence of `SECTIONS` commands, we place .relro_padding
      immediately before a symbol assignment using DATA_SEGMENT_RELRO_END (see
      also https://reviews.llvm.org/D124656), if present.
      DATA_SEGMENT_RELRO_END is changed to align to max-page-size instead of
      common-page-size.
      
      Some edge cases worth mentioning:
      
      * ppc64-toc-addis-nop.s: when PHDRS is present, do not append
      .relro_padding
      * avoid-empty-program-headers.s: when the only RELRO section is .tbss,
      it is not part of PT_LOAD segment, therefore we do not append
      .relro_padding.
      
      ---
      
      Close #65002: GNU ld from 2.39 onwards aligns the end of PT_GNU_RELRO to
      a
      max-page-size boundary (https://sourceware.org/PR28824) so that the last
      page is
      protected even if runtime_page_size > common-page-size.
      
      In my opinion, losing protection for the last page when the runtime page
      size is
      larger than common-page-size is not really an issue. Double mapping a
      page of up
      to max-common-page for the protection could cause undesired VM waste.
      Internally
      we had users complaining about 2MiB max-page-size applying to shared
      objects.
      
      Therefore, the end of .relro_padding is padded to a common-page-size
      boundary. Users who are really anxious can set common-page-size to match
      their runtime page size.
      
      ---
      
      17 tests need updating as there are lots of change detectors.
      5a58e98c
    • Sam McCall's avatar
      [dataflow] Add global invariant condition to DataflowAnalysisContext (#65949) · 21ab252f
      Sam McCall authored
      This records facts that are not sensitive to the current flow condition,
      and should apply to all environments.
      
      The motivating case is recording information about where a Value
      originated, such as nullability:
       - we may see the same Value for multiple expressions (e.g. reads of the
         same field) in multiple environments (multiple blocks or iterations)
       - we want to record information only when we first see the Value
         (e.g. Nullability annotations on fields only add information if we
         don't know where the value came from)
       - this information should be expressible as a SAT condition
       - we must add this SAT condition to every environment where the
         Value may appear
      
      We solve this by recording the information in the global condition.
      This doesn't seem particularly elegant, but solves the problem and is
      a fairly small and natural extension of the Environment.
      
      Alternatives considered:
       - store the constraint directly as a property on the Value.
         But it's more composable for such properties to always be variables
         (AtomicBoolValue), and constrain them with SAT conditions.
       - add a hook whenever values are created, giving the analysis the
         chance to populate them.
         However the framework relies on/provides the ability to construct
         values in arbitrary places without providing the context such a hook
         would need, this would be a very invasive change.
      21ab252f
    • Alex Langford's avatar
      [lldb][NFCI] Remove use of ConstString from UnixSignals · 2f377c5b
      Alex Langford authored
      The majority of UnixSignals strings are static in the sense that they do
      not change. The overwhelming majority of these strings are string
      literals. Using ConstString to manage their lifetime does not make
      sense. The only exception to this is one of the subclasses of
      UnixSignals, for which I have created a StringSet local to that file
      which will guarantee the lifetimes of these StringRefs.
      
      As for the other benefits of ConstString, string uniqueness is not a
      concern (as many of them are already string literals) and comparing
      signal names and aliases should not be a hot path.
      
      Differential Revision: https://reviews.llvm.org/D159011
      2f377c5b
    • Ye Luo's avatar
    • Jingu Kang's avatar
      [MachineLICM] Handle Subloops · 5ec9699c
      Jingu Kang authored
      Following discussion on https://reviews.llvm.org/D154205, make MachineLICM pass
      handle subloops with only visiting outermost loop's blocks once.
      
      Differential Revision: https://reviews.llvm.org/D154205
      5ec9699c
    • Aart Bik's avatar
      [mlir][sparse] deprecate the convert{To,From}MLIRSparseTensor methods (#66304) · 156a4ba9
      Aart Bik authored
      Rationale:
      These libraries provided COO input and output at external boundaries
      which, since then, has been generalized to the much more powerful pack
      and unpack operations of the sparse tensor dialect.
      156a4ba9
    • Adrian Prantl's avatar
      Clean up test case (#66400) · 9dfc6d37
      Adrian Prantl authored
      9dfc6d37
    • Matt Arsenault's avatar
    • Shubham Sandeep Rastogi's avatar
      0d0ab760
    • Shubham Sandeep Rastogi's avatar
      [Dexter] Fix test failures on greendragon (#66299) · e6cc7b72
      Shubham Sandeep Rastogi authored
      The issue with these test failures is that the dSYM was not being found
      by lldb, which is why setting breakpoints was failing and lldb quit
      without performing any steps. This change copies the dSYM to the same
      temp directory that the executable is copied to.
      e6cc7b72
    • Yinying Li's avatar
      [mlir][sparse] Migrate more tests to new syntax (#66309) · e2e429d9
      Yinying Li authored
      CSR:
      `lvlTypes = [ "dense", "compressed" ]` to `map = (d0, d1) -> (d0 :
      dense, d1 : compressed)`
      
      CSC:
      `lvlTypes = [ "dense", "compressed" ], dimToLvl = affine_map<(d0, d1) ->
      (d1, d0)>` to `map = (d0, d1) -> (d1 : dense, d0 : compressed)`
      
      This is an ongoing effort: #66146
      e2e429d9
    • Aaron Ballman's avatar
      [C23] Remove N2713 from the list · 1db6b127
      Aaron Ballman authored
      This paper was obsoleted by the changes in N3138 and US-045
      1db6b127
  2. Sep 14, 2023
  3. Sep 15, 2023
  4. Sep 14, 2023
    • David Green's avatar
      [AArch64] Split Ampere1Write_Arith into rr/ri and rs/rx InstRWs. (#66384) · 74724902
      David Green authored
      The ampere1 scheduling model uses IsCheapLSL predicates for ADDXri and
      ADDWrr instructions, which only have 3 operands. In attempting to check
      that the third is a shift, the predicate can attempt to access an out of
      bounds operand, hitting an assert. This splits the rr/ri instructions
      (which can never have shifts) from the rs/rx instructions to ensure they
      both work correctly. Ampere1Write_1cyc_1AB was chosen for the rr/ir
      instructions to match the cheap case.
      
      This also sets CompleteModel = 0 for the ampere1 scheduling model, as at
      runtime under debug it will attempt to check that as well as all
      instructions having scheduling info, there is information for each
      output operand.
      
      DefIdx 1 exceeds machine model writes for
        renamable $w9, renamable $w8 = LDPWi renamable $x8, 0
      (Try with MCSchedModel.CompleteModel set to false)incomplete machine
      model
      74724902
    • Ingo Müller's avatar
      [mlir][linalg][transform][python] Drop _get_op_result... from mix-ins. (#65726) · 360c6290
      Ingo Müller authored
      `_get_op_result_or_value` was used in mix-ins to unify the handling of
      op results and values. However, that function is now called in the
      generated constructors, such that doing so in the mix-ins is not
      necessary anymore.
      360c6290
    • Matt Arsenault's avatar
    • Duncan P. N. Exon Smith's avatar
      Resign as code owner of branch weights and block frequency · 7976bdb5
      Duncan P. N. Exon&nbsp;Smith authored
      Somewhat overdue... it has been a few years since I stopped watching block frequency / branch weight patches actively, so I effectively stopped acting as code owner a while ago. Reflect the reality.
      
      Still happy to help out; feel free to pull me in if you think I might have useful context!
      7976bdb5
    • Manos Anagnostakis's avatar
      [AArch64] New subtarget features to control ldp and stp formation (#66098) · 008f26b1
      Manos Anagnostakis authored
      On some AArch64 cores, including Ampere's ampere1 and ampere1a
      architectures, load and store pair instructions are faster compared to
      simple loads/stores only when the alignment of the pair is at least
      twice that of the individual element being loaded.
      
      Based on that, this patch introduces four new subtarget features, two
      for controlling ldp and two for controlling stp, to cover the ampere1
      and ampere1a alignment needs and to enable optional fine-grained control
      over ldp and stp generation in general. The latter can be utilized by
      another cpu, if there are possible benefits
      with a different policy than the default provided by the compiler.
      
      More specifically, for each of the ldp and stp respectively we have:
      
      - disable-ldp/disable-stp: Do not emit ldp/stp.
      - ldp-aligned-only/stp-aligned-only: Emit ldp/stp only if the source
      pointer is aligned to at least double the alignment of the type.
      
      Therefore, for -mcpu=ampere1 and -mcpu=ampere1a
      ldp-aligned-only/stp-aligned-only become the defaults, because of the
      benefit from the alignment, whereas for the rest of the cpus the default
      behaviour of the compiler is maintained.
      008f26b1
    • Slava Zakharin's avatar
      [flang] Select proper library APIs for derived type io. (#66327) · b7d02d7e
      Slava Zakharin authored
      This patch syncs the logic inside `getInputFunc` that selects
      the library API and the logic in `createIoRuntimeCallForItem`
      that creates the input arguments for the library call.
      There were cases where we selected `InputDerivedType` API
      and passed only two arguments, and also we selected `InputDescriptor`
      and passed three arguments.
      It turns out we also were incorrectly selecting `OutputDescriptor`
      in `getOutputFunc` (`test4` case in the new LIT test),
      which caused runtime issues for output of a derived type
      with descriptor components (due to the missing non-type-bound table).
      b7d02d7e
    • Slava Zakharin's avatar
    • Björn Pettersson's avatar
      [LICM] Simplify isLoadInvariantInLoop given opaque pointers (#65597) · a0ce4384
      Björn Pettersson authored
      Since we no longer support typed pointers in LLVM IR, the PtrASXTy
      in isLoadInvariantInLoop was set to be equal to Addr->getType() (an
      opaque ptr in the same address space). That made the loop looking
      through bitcasts redundant.
      a0ce4384
    • Matthias Springer's avatar
      [mlir][transform] Check for invalidated iterators on payload IR mappings (#66369) · aca9019b
      Matthias Springer authored
      Add extra error checking (in debug mode) to detect cases where an
      iterator on "direct" payload IR mappings is invalidated (due to elements
      being removed). Such errors are hard to debug: they are often
      non-deterministic; sometimes the program crashes, sometimes it produces
      wrong results. Even when it crashes, the stack trace often points to
      completely unrelated code locations.
      
      Store a timestamp with each "direct" mapping. The timestamp is increased
      whenever an operation is performed that invaldiates an iterator on that
      mapping. A debug iterator is added that checks the timestamp as payload
      IR is enumerated.
      aca9019b