1. Dec 06, 2022
    • serge-sans-paille's avatar
      Store OptTable::Info::Name as a StringRef · 8ae18303
      serge-sans-paille authored
      This avoids implicit conversion to StringRef at several points, which in
      turns avoid redundant calls to strlen.
      
      As a side effect, this greatly simplifies the implementation of
      StrCmpOptionNameIgnoreCase.
      
      It also eventually gives a consistent, humble speedup in compilation
      time.
      
      https://llvm-compile-time-tracker.com/compare.php?from=5f5b942823474e98e43a27d515a87ce140396c53&to=60e13b778119fc32d50dc38ff1a564a87146e9c6&stat=instructions:u
      
      Differential Revision: https://reviews.llvm.org/D139274
      8ae18303
    • Matt Arsenault's avatar
    • Haojian Wu's avatar
      6f12281d
    • Jean Perier's avatar
      [flang] Allow conversion from hlfir.expr to fir::ExtendedValue · 788960d6
      Jean Perier authored
      For now at least, the plan is to keep hlfir.expr usage limited as
      sub-expression operand, assignment rhs, and a few other contexts (
      e.g. Associate statements). The rest of lowering (statements lowering
      in the bridge) will still expect to get and manipulate characters and
      arrays in memory. That means that hlfir.expr must be converted to
      variable in converter.genExprAddr/converter.genExprBox.
      
      This is done using an hlfir.associate, and generating the related
      hlfir.end_associate in the statement context.
      
      hlfir::getFirBase of is updated to avoid bringing in the HLFIR
      fir.boxchar/fir.box into FIR when the entity was created with
      hlfir::AssociateOp.
      
      Differential Revision: https://reviews.llvm.org/D139328
      788960d6
    • bipmis's avatar
      Add tests which can be matched to umull · bda1f0b9
      bipmis authored
      bda1f0b9
    • Manuel Brito's avatar
    • Adrian Kuegel's avatar
      [mlir][SparseTensor] Apply ClangTidyLegacy finding (NFC). · f083c9bd
      Adrian Kuegel authored
      Converting integer literal to bool, use bool literal instead.
      f083c9bd
    • Archibald Elliott's avatar
      [AArch64] Implement __arm_rsr128/__arm_wsr128 · 83b3304d
      Archibald Elliott authored
      This only contains the SelectionDAG implementation. GlobalISel to
      follow.
      
      The broad approach is:
      - Introduce new builtins for 128-bit wide instructions.
      - Lower these to @llvm.read_register.i128/@llvm.write_register.i128
      - Introduce target-specific ISD nodes which have legal operands (two
        i64s rather than an i128). These are named AArch64::{MRRS, MSRR} to
        match the instructions they are for. These are a little complex as
        they need to match the "shape" of what they're replacing or the
        legaliser complains.
      - Select these using the existing tryReadRegister/tryWriteRegister to
        share the MDString parsing code, and introduce additional code to
        ensure these are selected into the right MRRS/MSRR instructions. What
        makes this hard is ensuring that the two i64s end up in an XSeqPair
        register pair, because SelectionDAG doesn't care that much about
        register classes if it can avoid doing so.
      
      The main change to existing code is the reorganisation of
      tryReadRegister and tryWriteRegister to try to keep the string parsing
      code separate from the instruction creating code.
      
      This also includes the changes to clang to define and use the ACLE
      feature macro named `__ARM_FEATURE_SYSREG128`.
      
      Contributors:
        Sam Elliott
        Lucas Prates
      
      Differential Revision: https://reviews.llvm.org/D139086
      83b3304d
    • Ties Stuij's avatar
      [AArch64] implement GPR (U/S)(MIN/MAX) instruction SDag support · 94e7e58f
      Ties Stuij authored
      Using SelectionDag, lower umin, umax, smin, smax intrinsics to corresponding
      UMIN, UMAX, SMIN, SMAX instructions when feat CSSC is available.
      
      See specs for corresponding immediate and register versions in:
      https://developer.arm.com/documentation/ddi0602/2022-09/Base-Instructions/
      
      Reviewed By: lenary
      
      Differential Revision: https://reviews.llvm.org/D138813
      94e7e58f
    • Nikita Popov's avatar
      [IR] Don't assume readnone/readonly intrinsics are willreturn · 0b20c303
      Nikita Popov authored
      This removes our "temporary" hack to assume that readnone/readonly
      intrinsics are also willreturn. An explicit willreturn annotation,
      usually via default intrinsic attributes, is now required.
      
      Differential Revision: https://reviews.llvm.org/D137630
      0b20c303
    • Ties Stuij's avatar
      [AArch64] lower abs intrinsic to new ABS instruction in SelDag · eaea4608
      Ties Stuij authored
      When feature CSSC is available, the SelectionDag abs intrinsic should map to the
      new scalar ABS instruction.
      
      Additionally, the SIMDTwoScalarD tablegen defm includes a pattern match for
      scalar i64, which we don't want to use when CSSC is enabled.
      
      spec:
      https://developer.arm.com/documentation/ddi0602/2022-09/Base-Instructions/ABS--Absolute-value-
      
      Reviewed By: lenary
      
      Differential Revision: https://reviews.llvm.org/D138812
      eaea4608
    • David Spickett's avatar
      [LLVM][ARM] Correct llvm feature for vfpv3d16 host feature · 9f85af54
      David Spickett authored
      d16 was removed in https://reviews.llvm.org/D60691.
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D139304
      9f85af54
    • David Spickett's avatar
      [lld-macho] Fix map file test on 32 bit hosts · 7c7e39db
      David Spickett authored
      The test added in https://reviews.llvm.org/D137368 has been failing
      on our 32 bit arm bots:
      https://lab.llvm.org/buildbot/#/builders/178/builds/3460
      
      You get this for the strings:
      <<dead>> 0x883255000000003 [ 10] literal string: Hello, it's me
      Instead of the expected:
      <<dead>>  0x0000000F  [  3] literal string: Hello, it's me
      
      This is because unlike symbols whose size is a uint64_t, strings
      use a StringRef whose size is size_t. size_t changes size between
      32 and 64 bit platforms.
      
      This fixes the test by using %z to print the size of the strings,
      this works for 32 and 64 bit.
      7c7e39db
    • Ties Stuij's avatar
      [AArch64] SelectionDag codegen for gpr CTZ instruction · 2f778e60
      Ties Stuij authored
      When feature CSSC is available we should use instruction CTZ in SelectionDag
      where applicable:
      
      - CTTZ intrinsics are lowered to using the gpr CTZ instruction
      - BITREVERSE -> CTLZ instruction pattern gets replaced by CTZ
      
      spec:
      https://developer.arm.com/documentation/ddi0602/2022-09/Base-Instructions/CTZ--Count-Trailing-Zeros-
      
      Reviewed By: lenary
      
      Differential Revision: https://reviews.llvm.org/D138811
      2f778e60
    • Kristina Bessonova's avatar
      [llvm-objdump] Avoid using mapping symbols as branch target labels · 4e958b4d
      Kristina Bessonova authored
      The main motivation for this change is to avoid ambiguity because
      mapping symbol names may not be unique across a binary and do not allow uniquely
      identifying target address. So that mapping symbols used as branch target
      labels make llvm-objdump output less readable.
      
      Another point is that mapping symbols sometimes appear in
      non-allocatable sections, like debug info sections which make objdump
      output even more confusing.
      
      For example, a small AArch64 executable may contain plenty of `$d[.*]`
      symbols and none of them would be useful as a label for resolving
      a branch or a memory operand target address:
      
      ```
        0000000000000254 l       .note.ABI-tag	0000000000000000 $d
        00000000000008d4 l       .eh_frame            0000000000000000 $d
        0000000000000868 l       .rodata              0000000000000000 $d
        0000000000011028 l       .data                0000000000000000 $d
        0000000000010db8 l       .fini_array          0000000000000000 $d
        0000000000010db0 l       .init_array          0000000000000000 $d
        00000000000008e8 l       .eh_frame            0000000000000000 $d
        0000000000011034 l       .bss                 0000000000000000 $d
      ```
      
      Note that GNU objdump doesn't use mapping symbols as branch target
      labels for all targets that support such symbols (ARM, AArch64, CSKY).
      
      Differential Revision: https://reviews.llvm.org/D139131
      4e958b4d
    • Javier Setoain's avatar
      [mlir] Add hoisting of transfer ops in affine loops · 825da072
      Javier Setoain authored
      The only way to do this with the current hoisting strategy is by
      lowering Affine to Scf first, but that prevents further passes on
      Affine.
      
      Differential Revision: https://reviews.llvm.org/D137600
      825da072
    • Max Kazantsev's avatar
      [SCEVExpander] Support cost evaluation of several SCEVs with same budget · 6dac1701
      Max Kazantsev authored
      This is a follow-up from discussion in D138412. Sometimes we want to evaluate
      the cost of expansion of several SCEVs together with same budget. For example,
      if one of them is a bit above cheap limit, and the second one is free, then
      we still want to expand. Checking each of them with "cheap" limit is a bit more
      pessimistic.
      
      Differential Revision: https://reviews.llvm.org/D138475
      Reviewed By: lebedev.ri
      6dac1701
    • HanSheng Zhang's avatar
    • HanSheng Zhang's avatar
      [CMake]Allow user specified CPack Options · 27c4b509
      HanSheng Zhang authored
      This should allow downstream vendors to install multiple LLVM distributions in parallel.
      
      Should we also patch the default values to allow multiple upstream llvm distribution?
      
      Reviewed By: thieta
      
      Differential Revision: https://reviews.llvm.org/D138632
      27c4b509
    • Juan Manuel MARTINEZ CAAMAÑO's avatar
    • Sergey Kachkov's avatar
      [RISCV] Generate .cfi_def_cfa_expression for RVV stack adjustment · 132dc442
      Sergey Kachkov authored
      Cannonical frame address after RVV stack adjustment is sp + StackSize +
      RVVStackSize * vlenb, and since vlenb is unknown at compile-time (but it
      is a constant for particular HW implementation), emit
      .cfi_def_cfa_expression so libunwind can read VLENB CSR register at
      run-time and obtain correct frame address.
      
      Fixes https://github.com/llvm/llvm-project/issues/58356 (but additional
      run-time support for reading CSR may be required)
      
      Differential Revision: https://reviews.llvm.org/D136263
      132dc442
    • Vlad Serebrennikov's avatar
      [clang] Add test for CWG600 · 7e31d072
      Vlad Serebrennikov authored
      P1787: //CWG600 is resolved by explaining that accessibility affects naming a member in the sense of the ODR.//
      Wording: see changes to [class.access] p1 and p4.
      Additional references: [[ http://eel.is/c++draft/basic.def.odr#8.sentence-2 | basic.def.odr/8 ]]: //A function is odr-used if it is named by a potentially-evaluated expression or conversion.//
      
      Reviewed By: #clang-language-wg, aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D139173
      7e31d072
    • Corentin Jabot's avatar
      [Clang] make_cxx_dr_status download the issue list automatically · 2fbe3f9e
      Corentin Jabot authored
      if none is provided
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D139212
      2fbe3f9e
    • Vlad Serebrennikov's avatar
      [clang] Mark CWG554 as N/A · 6d971cb8
      Vlad Serebrennikov authored
      P1787: //CWG554 is resolved by using the word “scope” instead of “declarative region”, consistent with its very common use in phrases like “namespace scope”.//
      
      Reviewed By: #clang-language-wg, cor3ntin, aaron.ballman, shafik
      
      Differential Revision: https://reviews.llvm.org/D139172
      6d971cb8
    • David Spickett's avatar
      [LLVM][Release] Prevent empty runtime name in release script · 500587e2
      David Spickett authored
      Unlike projects, runtimes doesn't have a default set of names.
      This means you get a leading space at the start, which gets converted
      to a ';' giving ";<runtime name>;<runtime name>".
      
      CMake then errors because the "" before the first ';' is treated
      as a runtime name and of course it's not a valid name.
      
      Fix this by removing the leading spaces from runtimes before we
      insert the ';'.
      
      Reviewed By: ldionne
      
      Differential Revision: https://reviews.llvm.org/D139306
      500587e2
    • chenglin.bi's avatar
    • Vlad Serebrennikov's avatar
      [clang] Add test for CWG405 · 80bae9aa
      Vlad Serebrennikov authored
      P1787: //CWG405 is resolved by stating that argument-dependent lookup (sometimes) occurs after an ordinary unqualified lookup (making statements like “finding a variable prevents argument-dependent lookup” formally correct).//
      Wording: see changes to [basic.lookup.argdep] p1 and p3
      
      This issue seems a duplicate of CWG218, even though it is not officially recognized. A part of a test for CWG218 is reused here, adding cross-references.
      
      Reviewed By: #clang-language-wg, aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D139095
      80bae9aa
    • Tobias Hieta's avatar
      [CodeView] Add support for local S_CONSTANT records · 2298a44c
      Tobias Hieta authored
      CodeView doesn't have the ability to represent variables
      in other ways than as in registers or memory values, but
      LLVM very often transforms simple values into constants,
      consider this program:
      
      int f () { int i = 123; return i; }
      
      LLVM will transform `i` into a constant value and just
      leave behind a llvm.dbg.value, this can't be represented
      as a S_LOCAL record in CodeView. But we can represent it
      as a S_CONSTANT record.
      
      This patch checks if the location of a debug value is null,
      then we will insert a S_CONSTANT record instead of a S_LOCAL
      value with the flag "OptimizedAway".
      
      In lld we then output the S_CONSTANT in the right scope, before
      they where always inserted in the global stream, now we check
      the scope before inserting it.
      
      This has shown to improve debugging for our developers
      internally.
      
      Fixes to llvm/llvm-project#55958
      
      Reviewed By: aganea
      
      Differential Revision: https://reviews.llvm.org/D138995
      2298a44c
    • Vlad Serebrennikov's avatar
      [clang] Add test for CWG952 · 5b22c512
      Vlad Serebrennikov authored
      P1787: // [[ https://wg21.link/cwg952 | CWG952 ]] is resolved by refining the definition of “naming class” per Richard’s suggestion in [[ https://lists.isocpp.org/core/2020/09/9963.php | “CWG1621 and [class.static/2”]].//
      Wording:
      - [class.static]/2 removed;
      - [class.access.base]/5 rephrased.
      
      Currently behavior is the following: unqualified names undergo //unqualified name lookup// [1], which perform //unqualified search// in immediate scope [2]. This scope is the scope the definition of //naming class// [3] refers to. `A::I` is not //accessible// when named in classes `C` and `D` per [3]. In particular, the last item regarding base class ([class.access.base]/5.4) is not applicable, because class `A` is not //accessible// in both classes `C` and `D` per [4].
      
      References:
      1. [[ https://eel.is/c++draft/basic.lookup#unqual-4.sentence-2 | basic.lookup.unqual/4 ]]
      2. [[ https://eel.is/c++draft/basic.lookup#unqual-3 | basic.lookup.unqual/3 ]]
      3. [[ https://eel.is/c++draft/class.access#base-5.sentence-4 | class.access.base/5 ]]
      4. [[ https://eel.is/c++draft/class.access#base-4 | class.access.base/4 ]]
      
      Reviewed By: #clang-language-wg, erichkeane, aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D139326
      5b22c512
    • Nikita Popov's avatar
      [MemCpyOpt] Use BatchAA when processing one instruction (NFCI) · b7ede701
      Nikita Popov authored
      While we can't use a single BatchAA instance for the entire
      MemCpyOpt run without further justification, we can use BatchAA
      while performing the queries related to a single instruction
      (these will first perform some AA-based checks, and then modify
      the IR only afterwards).
      b7ede701
    • Gedare Bloom's avatar
      [clang-format] Avoid breaking )( with BlockIndent · b40e9dce
      Gedare Bloom authored
      The BracketAlignmentStyle BAS_BlockIndent was forcing breaks before a
      closing right parenthesis yielding strange-looking results in case of
      code structures that have a left parens immediately following a right
      parens ")(" such as is seen with indirect function calls via function
      pointers and with type casting.
      
      Fixes 57250.
      Fixes 58496.
      
      Differential Revision: https://reviews.llvm.org/D137762
      b40e9dce
    • Sergey Kachkov's avatar
      [libunwind][RISCV] Support reading of VLENB CSR register · ca0b4d58
      Sergey Kachkov authored
      Support reading of VLENB (vector byte length) control register, that can be
      required for correct unwinding of RVV objects on stack.
      
      Differential Revision: https://reviews.llvm.org/D136264
      ca0b4d58
    • Nikita Popov's avatar
      [DSE] Reuse BatchAA for MSSA clobber queries · 330ee040
      Nikita Popov authored
      This is not NFC because the DSE BatchAA is more powerful than the
      default one due to EarliestEscape CaptureInfo, so this might
      improve results in some cases.
      330ee040
    • Jean Perier's avatar
      [flang] do not generate padding/truncation code when character length are equals · 9b9a8475
      Jean Perier authored
      When generating character assignment operations, the generic code
      generates some code to handle truncation and padding when the length
      differ at runtime. A bypass already exists when the length are compile
      time constant and match, but it was not used for the trivial case where
      the RHS and LHS length is the same SSA value. In such case, even though,
      the length is not know at compile time, it is known to be the same.
      
      This will simplify the code creating character temporaries from a
      variable in HLFIR that will use this assignment code.
      
      Note that this probably has little impact on performance (llvm may be clever enough
      to later catch that for us). But it makes the generated IR a lot more readable at
      little cost.
      
      Differential Revision: https://reviews.llvm.org/D139330
      9b9a8475
    • Valery Pykhtin's avatar
      [AMDGPU] Fix GCNSubtarget::getMinNumVGPRs, add unit test to check consistency... · d09d834b
      Valery Pykhtin authored
      [AMDGPU] Fix GCNSubtarget::getMinNumVGPRs, add unit test to check consistency between GCNSubtarget's getMinNumVGPRs, getMaxNumVGPRs and getOccupancyWithNumVGPRs.
      
      ```
        /// \returns Minimum number of VGPRs that meets given number of waves per
        /// execution unit requirement supported by the subtarget.
        unsigned getMinNumVGPRs(unsigned WavesPerEU) const;
      
        /// \returns Maximum number of VGPRs that meets given number of waves per
        /// execution unit requirement supported by the subtarget.
        unsigned getMaxNumVGPRs(unsigned WavesPerEU) const;
      
        /// Return the maximum number of waves per SIMD for kernels using \p VGPRs
        /// VGPRs
        unsigned getOccupancyWithNumVGPRs(unsigned VGPRs) const;
      ```
      
      While working on RP tracking issues I noticed that getMinNumVGPRs return incorrect
      values: the problem is large VGPR granule sizes on GFX10+ architectures. Some of the
      occupancies aren't reachable because require the same amount of VGPR granules as others.
      For example 19 waves occupancy on gfx1010 require the same amount of granules as 20 waves
      so the resultng occupancy would be 20.
      
      SGPRs have the same issue and even have inconsistency between getMaxNumSGPRs and getOccupancyWithNumSGPRs.
      It will be addressed in the next patch.
      
      Legend:
        # MinVGPR and MaxVGPR are values returned by getMinNumVGPRs and getMaxNumVGPRs for a given Occ.
        # (ONumber) is the value returned by getOccupancyWithNumVGPRs for a given MinVGPR or MaxVGPR.
        # R means range problem: MinVGPR should be less than MaxVGPR and both should refer to the same occupancy.
      
      Unit test output without the fix:
      ```
      ./build/unittests/Target/AMDGPU/AMDGPUTests --gtest_filter=AMDGPU.TestVGPRLimitsPerOccupancy --print-cpu-reg-limits
      
       gfx90a gfx940:
      Occ    MinVGPR        MaxVGPR
        8        0 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     80  (O6)
        5       81 (O5)     96  (O5)
        4       97 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      257 (O1)     512 (O1)
      
       gfx600 gfx600 gfx601 gfx601 gfx601 gfx602 gfx602 gfx602 gfx700 gfx700 gfx701 gfx701 gfx702 gfx703 gfx703 gfx703 gfx704 gfx704 gfx705 gfx801 gfx801 gfx802 gfx802 gfx802 gfx803 gfx803 gfx803 gfx803 gfx805 gfx805 gfx810 gfx810 gfx900 gfx902 gfx904 gfx906 gfx908 gfx909 gfx90c:
      Occ    MinVGPR        MaxVGPR
       10        0 (O10)    24  (O10)
        9       25 (O9)     28  (O9)
        8       29 (O8)     32  (O8)
        7       33 (O7)     36  (O7)
        6       37 (O6)     40  (O6)
        5       41 (O5)     48  (O5)
        4       49 (O4)     64  (O4)
        3       65 (O3)     84  (O3)
        2       85 (O2)     128 (O2)
        1      129 (O1)     256 (O1)
      
       gfx1030w64 gfx1031w64 gfx1032w64 gfx1033w64 gfx1034w64 gfx1035w64 gfx1036w64 gfx1102w64 gfx1103w64:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    32  (O16)
       15       33 (O12) R  32  (O16)
       14       33 (O12) R  32  (O16)
       13       33 (O12) R  32  (O16)
       12       33 (O12)    40  (O12)
       11       41 (O10) R  40  (O12)
       10       41 (O10)    48  (O10)
        9       49 (O9)     56  (O9)
        8       57 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     80  (O6)
        5       81 (O5)     96  (O5)
        4       97 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      256 (O2) R   256 (O2)
      
       gfx1100w64 gfx1101w64:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    48  (O16)
       15       49 (O12) R  48  (O16)
       14       49 (O12) R  48  (O16)
       13       49 (O12) R  48  (O16)
       12       49 (O12)    60  (O12)
       11       61 (O10) R  60  (O12)
       10       61 (O10)    72  (O10)
        9       73 (O9)     84  (O9)
        8       85 (O8)     96  (O8)
        7       97 (O7)     108 (O7)
        6      109 (O6)     120 (O6)
        5      121 (O5)     144 (O5)
        4      145 (O4)     192 (O4)
        3      193 (O3)     252 (O3)
        2      253 (O2)     256 (O2)
        1      256 (O2) R   256 (O2)
      
       gfx1030w32 gfx1031w32 gfx1032w32 gfx1033w32 gfx1034w32 gfx1035w32 gfx1036w32 gfx1102w32 gfx1103w32:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    64  (O16)
       15       65 (O12) R  64  (O16)
       14       65 (O12) R  64  (O16)
       13       65 (O12) R  64  (O16)
       12       65 (O12)    80  (O12)
       11       81 (O10) R  80  (O12)
       10       81 (O10)    96  (O10)
        9       97 (O9)     112 (O9)
        8      113 (O8)     128 (O8)
        7      129 (O7)     144 (O7)
        6      145 (O6)     160 (O6)
        5      161 (O5)     192 (O5)
        4      193 (O4)     256 (O4)
        3      256 (O4) R   256 (O4)
        2      256 (O4) R   256 (O4)
        1      256 (O4) R   256 (O4)
      
       gfx1100w32 gfx1101w32:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    96  (O16)
       15       97 (O12) R  96  (O16)
       14       97 (O12) R  96  (O16)
       13       97 (O12) R  96  (O16)
       12       97 (O12)    120 (O12)
       11      121 (O10) R  120 (O12)
       10      121 (O10)    144 (O10)
        9      145 (O9)     168 (O9)
        8      169 (O8)     192 (O8)
        7      193 (O7)     216 (O7)
        6      217 (O6)     240 (O6)
        5      241 (O5)     256 (O5)
        4      256 (O5) R   256 (O5)
        3      256 (O5) R   256 (O5)
        2      256 (O5) R   256 (O5)
        1      256 (O5) R   256 (O5)
      
       gfx1010w64 gfx1011w64 gfx1012w64 gfx1013w64:
      Occ    MinVGPR        MaxVGPR
       20        0 (O20)    24  (O20)
       19       25 (O18) R  24  (O20)
       18       25 (O18)    28  (O18)
       17       29 (O16) R  28  (O18)
       16       29 (O16)    32  (O16)
       15       33 (O14) R  32  (O16)
       14       33 (O14)    36  (O14)
       13       37 (O12) R  36  (O14)
       12       37 (O12)    40  (O12)
       11       41 (O11)    44  (O11)
       10       45 (O10)    48  (O10)
        9       49 (O9)     56  (O9)
        8       57 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     84  (O6)
        5       85 (O5)     100 (O5)
        4      101 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      256 (O2) R   256 (O2)
      
       gfx1010w32 gfx1011w32 gfx1012w32 gfx1013w32:
      Occ    MinVGPR        MaxVGPR
       20        0 (O20)    48  (O20)
       19       49 (O18) R  48  (O20)
       18       49 (O18)    56  (O18)
       17       57 (O16) R  56  (O18)
       16       57 (O16)    64  (O16)
       15       65 (O14) R  64  (O16)
       14       65 (O14)    72  (O14)
       13       73 (O12) R  72  (O14)
       12       73 (O12)    80  (O12)
       11       81 (O11)    88  (O11)
       10       89 (O10)    96  (O10)
        9       97 (O9)     112 (O9)
        8      113 (O8)     128 (O8)
        7      129 (O7)     144 (O7)
        6      145 (O6)     168 (O6)
        5      169 (O5)     200 (O5)
        4      201 (O4)     256 (O4)
        3      256 (O4) R   256 (O4)
        2      256 (O4) R   256 (O4)
        1      256 (O4) R   256 (O4)
      ```
      
      After the fix:
      ```
       gfx90a gfx940:
      Occ    MinVGPR        MaxVGPR
        8        0 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     80  (O6)
        5       81 (O5)     96  (O5)
        4       97 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      257 (O1)     512 (O1)
      
       gfx600 gfx600 gfx601 gfx601 gfx601 gfx602 gfx602 gfx602 gfx700 gfx700 gfx701 gfx701 gfx702 gfx703 gfx703 gfx703 gfx704 gfx704 gfx705 gfx801 gfx801 gfx802 gfx802 gfx802 gfx803 gfx803 gfx803 gfx803 gfx805 gfx805 gfx810 gfx810 gfx900 gfx902 gfx904 gfx906 gfx908 gfx909 gfx90c:
      Occ    MinVGPR        MaxVGPR
       10        0 (O10)    24  (O10)
        9       25 (O9)     28  (O9)
        8       29 (O8)     32  (O8)
        7       33 (O7)     36  (O7)
        6       37 (O6)     40  (O6)
        5       41 (O5)     48  (O5)
        4       49 (O4)     64  (O4)
        3       65 (O3)     84  (O3)
        2       85 (O2)     128 (O2)
        1      129 (O1)     256 (O1)
      
       gfx1030w64 gfx1031w64 gfx1032w64 gfx1033w64 gfx1034w64 gfx1035w64 gfx1036w64 gfx1102w64 gfx1103w64:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    32  (O16)
       15        0 (O16)    32  (O16)
       14        0 (O16)    32  (O16)
       13        0 (O16)    32  (O16)
       12       33 (O12)    40  (O12)
       11       33 (O12)    40  (O12)
       10       41 (O10)    48  (O10)
        9       49 (O9)     56  (O9)
        8       57 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     80  (O6)
        5       81 (O5)     96  (O5)
        4       97 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      169 (O2)     256 (O2)
      
       gfx1100w64 gfx1101w64:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    48  (O16)
       15        0 (O16)    48  (O16)
       14        0 (O16)    48  (O16)
       13        0 (O16)    48  (O16)
       12       49 (O12)    60  (O12)
       11       49 (O12)    60  (O12)
       10       61 (O10)    72  (O10)
        9       73 (O9)     84  (O9)
        8       85 (O8)     96  (O8)
        7       97 (O7)     108 (O7)
        6      109 (O6)     120 (O6)
        5      121 (O5)     144 (O5)
        4      145 (O4)     192 (O4)
        3      193 (O3)     252 (O3)
        2      253 (O2)     256 (O2)
        1      253 (O2)     256 (O2)
      
       gfx1030w32 gfx1031w32 gfx1032w32 gfx1033w32 gfx1034w32 gfx1035w32 gfx1036w32 gfx1102w32 gfx1103w32:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    64  (O16)
       15        0 (O16)    64  (O16)
       14        0 (O16)    64  (O16)
       13        0 (O16)    64  (O16)
       12       65 (O12)    80  (O12)
       11       65 (O12)    80  (O12)
       10       81 (O10)    96  (O10)
        9       97 (O9)     112 (O9)
        8      113 (O8)     128 (O8)
        7      129 (O7)     144 (O7)
        6      145 (O6)     160 (O6)
        5      161 (O5)     192 (O5)
        4      193 (O4)     256 (O4)
        3      193 (O4)     256 (O4)
        2      193 (O4)     256 (O4)
        1      193 (O4)     256 (O4)
      
       gfx1100w32 gfx1101w32:
      Occ    MinVGPR        MaxVGPR
       16        0 (O16)    96  (O16)
       15        0 (O16)    96  (O16)
       14        0 (O16)    96  (O16)
       13        0 (O16)    96  (O16)
       12       97 (O12)    120 (O12)
       11       97 (O12)    120 (O12)
       10      121 (O10)    144 (O10)
        9      145 (O9)     168 (O9)
        8      169 (O8)     192 (O8)
        7      193 (O7)     216 (O7)
        6      217 (O6)     240 (O6)
        5      241 (O5)     256 (O5)
        4      241 (O5)     256 (O5)
        3      241 (O5)     256 (O5)
        2      241 (O5)     256 (O5)
        1      241 (O5)     256 (O5)
      
       gfx1010w64 gfx1011w64 gfx1012w64 gfx1013w64:
      Occ    MinVGPR        MaxVGPR
       20        0 (O20)    24  (O20)
       19        0 (O20)    24  (O20)
       18       25 (O18)    28  (O18)
       17       25 (O18)    28  (O18)
       16       29 (O16)    32  (O16)
       15       29 (O16)    32  (O16)
       14       33 (O14)    36  (O14)
       13       33 (O14)    36  (O14)
       12       37 (O12)    40  (O12)
       11       41 (O11)    44  (O11)
       10       45 (O10)    48  (O10)
        9       49 (O9)     56  (O9)
        8       57 (O8)     64  (O8)
        7       65 (O7)     72  (O7)
        6       73 (O6)     84  (O6)
        5       85 (O5)     100 (O5)
        4      101 (O4)     128 (O4)
        3      129 (O3)     168 (O3)
        2      169 (O2)     256 (O2)
        1      169 (O2)     256 (O2)
      
       gfx1010w32 gfx1011w32 gfx1012w32 gfx1013w32:
      Occ    MinVGPR        MaxVGPR
       20        0 (O20)    48  (O20)
       19        0 (O20)    48  (O20)
       18       49 (O18)    56  (O18)
       17       49 (O18)    56  (O18)
       16       57 (O16)    64  (O16)
       15       57 (O16)    64  (O16)
       14       65 (O14)    72  (O14)
       13       65 (O14)    72  (O14)
       12       73 (O12)    80  (O12)
       11       81 (O11)    88  (O11)
       10       89 (O10)    96  (O10)
        9       97 (O9)     112 (O9)
        8      113 (O8)     128 (O8)
        7      129 (O7)     144 (O7)
        6      145 (O6)     168 (O6)
        5      169 (O5)     200 (O5)
        4      201 (O4)     256 (O4)
        3      201 (O4)     256 (O4)
        2      201 (O4)     256 (O4)
        1      201 (O4)     256 (O4)
      ```
      
      Reviewed By: #amdgpu, arsenm
      
      Differential Revision: https://reviews.llvm.org/D138443
      d09d834b
    • Kazu Hirata's avatar
      [mlir] Use std::nullopt instead of None in comments (NFC) · e823abab
      Kazu Hirata authored
      This is part of an effort to migrate from llvm::Optional to
      std::optional:
      
      https://discourse.llvm.org/t/deprecating-llvm-optional-x-hasvalue-getvalue-getvalueor/63716
      e823abab
    • Vladislav Khmelevsky's avatar
      [BOLT] Fix blocks layout reverse iterators · 7bb0cbfc
      Vladislav Khmelevsky authored
      Use container's reverse iterators
      
      Differential Revision: https://reviews.llvm.org/D139335
      7bb0cbfc
    • Kazu Hirata's avatar
      [clang-tools-extra] Use std::nullopt instead of llvm::None (NFC) · 2402c46b
      Kazu Hirata authored
      This is part of an effort to migrate from llvm::Optional to
      std::optional:
      
      https://discourse.llvm.org/t/deprecating-llvm-optional-x-hasvalue-getvalue-getvalueor/63716
      2402c46b
    • Kazu Hirata's avatar
      [llvm] Use std::nullopt instead of llvm::None (NFC) · 1ea9dd32
      Kazu Hirata authored
      This is part of an effort to migrate from llvm::Optional to
      std::optional:
      
      https://discourse.llvm.org/t/deprecating-llvm-optional-x-hasvalue-getvalue-getvalueor/63716
      1ea9dd32
    • Diego Caballero's avatar
      [mlir] Add `replaceAllUsesExcept` to rewriter · 77603e28
      Diego Caballero authored
      This patch adds `replaceAllUsesExcept` to the rewriter class.
      The implementation is copy-pasted from Value + calling
      `updateRootInPlace` to notify the listeners about the
      corresponding IR changes.
      
      Reviewed By: Mogball
      
      Differential Revision: https://reviews.llvm.org/D139382
      77603e28