1. Dec 05, 2021
  2. Dec 04, 2021
    • Matt Arsenault's avatar
      AMDGPU: Enable fixed function ABI by default · 729bf9b2
      Matt Arsenault authored
      Code using indirect calls is broken without this, and there isn't
      really much value in supporting the old attempt to vary the argument
      placement based on uses. This resulted in more argument shuffling code
      anyway.
      
      Also have the option stop implying all inputs need to be passed. This
      will no rely on the amdgpu-no-* attributes to avoid passing
      unnecessary values.
      729bf9b2
    • Florian Hahn's avatar
      [BasicAA] Add atomic mem intrinsic tests. · 89f0f277
      Florian Hahn authored
      89f0f277
    • Matt Arsenault's avatar
      AMDGPU: Assume all amdhsa kernarg passed implicit arguments by default · 2959e082
      Matt Arsenault authored
      Previously we would require adding an attribute to kernels to enable
      the inputs passed in the kernarg segment, accessed by
      llvm.amdgcn.implicitarg.ptr. This violates the principle of being
      correct by default. Some OpenMP testcases were broken recently since
      it wasn't correctly setting this attribute, and no known frontends are
      setting this to anything other than the maximum.
      
      Most of the test changes are from load widening of argument loads
      since there now more implied dereferenceable bytes.
      2959e082
    • Matt Arsenault's avatar
      AMDGPU: Optimize out implicit kernarg argument allocation if unused · ae0ba7de
      Matt Arsenault authored
      We already annotate whether llvm.amdgcn.implicitarg.ptr is known to be
      unused. Start using it to avoid allocating the implicit arguments if
      unneeded.
      ae0ba7de
    • Kristina Bessonova's avatar
      [DwarfDebug] Support emitting function-local declaration for a lexical block · ee691970
      Kristina Bessonova authored
      This is another attempt to make function-local declarations
      (like static variables, structs/classes and other) be correctly
      emitted within a lexical (bracketed) block.
      
      Fixes https://bugs.llvm.org/show_bug.cgi?id=19238.
      
      Differential Revision: https://reviews.llvm.org/D113741
      ee691970
    • Hugo Pompougnac's avatar
      Apply the permutation map on each affine nest · 5d49511b
      Hugo Pompougnac authored
      When using -test-loop-permutation="permutation-map=...", applies the
      permutation map on each affine nest in the function (and not only the
      first one). If the size of the permutation map and the size of a nest
      are not consistent, do nothing on this particular nest (instead of
      making MLIR crash).
      
      Differential Revision: https://reviews.llvm.org/D112947
      5d49511b
    • Kristina Bessonova's avatar
      [DwarfDebug] Move emission of global vars, types and imports to endModule() · 79d31329
      Kristina Bessonova authored
      This patch proposes to move emission of global variables, types,
      imported entities, etc from DwarfDebug::beginModule() to DwarfDebug::endModule().
      Effectively, this changes nothing but the order of debug entities which
      will be as follows:
      * subprograms (including related context, local variables/labels,
        local imported entities; related types can be created as a part of
        the emission of local entities of an abstract subprogram);
      * global variables (including related context and types);
      * retained types and enums;
      * non-local-scoped imported entities;
      * basic types;
      * other types left (as a part of local variables attributes emission).
      
      Note that the order of emitted compile units may also be changed as now we emit
      units that contain subprograms first and then all other non-empty units.
      
      The motivation behind this change is the following:
      (1) DwarfDebug::beginModule() is run at the very beginning of backend's pipeline,
          from this time IR can be significantly changed by target-specific passes.
          If it happens for debug metadata of global entities, those changes will not
          be reflected in the emitted DWARF.
      (2) imported subprogram names should refer to an abstract subprogram if it exists,
          but it isn't known in DwarfDebug::beginModule() (it's possible to make some
          guesses based on location info, but it's not quite reliable);
      (3) aforementioned entities if they are scoped within a bracketed block
          (subject of D113741) couldn't be emitted in DwarfDebug::beginModule()
          (they need parent emitted first). Another problem is if to try to gather
          some information about local entities and defer their emission
          (till subprogram's processing or DwarfDebug::endModule()) all the gathered
          details might be irrelevant / invalid by the time the entities are being
          emitted (because of (1)).
      
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D114705
      79d31329
    • Dmitry Vyukov's avatar
    • Anton Afanasyev's avatar
      [Passes] Move AggressiveInstCombine after InstCombine · c34d157f
      Anton Afanasyev authored
      Swap AIC and IC neighbouring in pipeline. This looks more natural and even
      almost has no effect for now (three slightly touched tests of test-suite). Also
      this could be the first step towards merging AIC (or its part) to -O2 pipeline.
      
      After several changes in AIC (like D108091, D108201, D107766, D109515, D109236)
      there've been observed several regressions (like PR52078, PR52253, PR52289)
      that were fixed in different passes (see D111330, D112721) by extending their
      functionality, but these regressions were exposed since changed AIC prevents IC
      from making some of early optimizations.
      
      This is common problem and it should be fixed by just moving AIC after IC
      which looks more logically by itself: make aggressive instruction combining
      only after failed ordinary one.
      
      Fixes PR52289
      
      Reviewed By: spatel, RKSimon
      
      Differential Revision: https://reviews.llvm.org/D113179
      c34d157f
    • Jay Foad's avatar
      [AMDGPU] Change llvm.amdgcn.image.bvh.intersect.ray to take vec3 args · 2774bad1
      Jay Foad authored
      The ray_origin, ray_dir and ray_inv_dir arguments should all be vec3 to
      match how the hardware instruction works.
      
      Don't change the API of the corresponding OpenCL builtins.
      
      Differential Revision: https://reviews.llvm.org/D115032
      2774bad1
    • Jay Foad's avatar
      [IR,TableGen] Add support for vec3 intrinsic arguments · c8e84c7a
      Jay Foad authored
      Add generic support for vec3 types, and in particular define
      llvm_v3f32_ty which will be used by AMDGPU's
      llvm.amdgcn.image.bvh.intersect.ray intrinsic.
      
      Differential Revision: https://reviews.llvm.org/D114956
      c8e84c7a
    • Jay Foad's avatar
      bc7dacf5
    • Nikita Popov's avatar
      [PhaseOrdering] Add test for incorrect merge function scheduling · 5b94037a
      Nikita Popov authored
      Add an -enable-merge-functions option to allow testing of function
      merging as it will actually happen in the optimization pipeline.
      Based on that add a test where we currently produce two identical
      functions without merging them due to incorrect pass scheduling
      under the new pass manager.
      5b94037a
    • Carlos Galvez's avatar
      [clang-tidy][NFC] Move CachedGlobList to GlobList.h · 946eb7a0
      Carlos Galvez authored
      Currently it's hidden inside ClangTidyDiagnosticConsumer,
      so it's hard to know it exists.
      
      Given that there are multiple uses of globs in clang-tidy,
      it makes sense to have these classes publicly available
      for other use cases that might benefit from it.
      
      Also, add unit test by converting the existing tests
      for GlobList into typed tests.
      
      Reviewed By: salman-javed-nz
      
      Differential Revision: https://reviews.llvm.org/D113422
      946eb7a0