1. May 06, 2024
    • Ulrich Weigand's avatar
      [SystemZ] Simplify f128 atomic load/store (#90977) · 0a0cac6d
      Ulrich Weigand authored
      Change definition of expandBitCastI128ToF128 and expandBitCastF128ToI128
      to allow for simplified use in atomic load/store.
      
      Update logic to split 128-bit loads and stores in DAGCombine to also
      handle the f128 case where appropriate. This fixes the regressions
      introduced by recent atomic load/store patches.
      0a0cac6d
    • Simon Pilgrim's avatar
      [DAG] Fold bitreverse(shl/srl(bitreverse(x),y)) -> srl/shl(x,y) (#89897) · 522b4bfe
      Simon Pilgrim authored
      Noticed while investigating GFNI per-element vector shifts (we can form SHL but not SRL/SRA)
      
      Alive2: https://alive2.llvm.org/ce/z/fSH-rf
      522b4bfe
    • WANG Rui's avatar
      0933a7a1
    • Timm Bäder's avatar
    • WANG Rui's avatar
      d98a7859
    • hev's avatar
      [LoongArch] Optimize *W Instructions at MI level (#90463) · e9bcd2bf
      hev authored
      Referring to RISC-V, adding an MI level pass to optimize *W instructions
      for LoongArch.
      
      First it removes unneeded sext(addi.w rd, rs, 0) instructions. Either
      because the sign extended bits aren't consumed or because the input was
      already sign extended by an earlier instruction.
      
      Then:
      1. Unless explicit disabled or the target prefers instructions with W
      suffix, it removes the -w suffix from opw instructions whenever all
      users are dependent only on the lower word of the result of the
      instruction. The cases handled are:
      * addi.w because it helps reduce test differences between LA32 and LA64
      w/o being a pessimization.
      
      2. Or if explicit enabled or the target prefers instructions with W
      suffix, it adds the W suffix to the instruction whenever all users are
      dependent only on the lower word of the result of the instruction. The
      cases handled are:
         * add.d/addi.d/sub.d/mul.d.
         * slli.d with imm < 32.
         * ld.d/ld.wu.
      e9bcd2bf
    • Timm Bäder's avatar
      [clang][Interp] Fix primitive lambda capture defaults · 9a521e27
      Timm Bäder authored
      We need to use InitField here, not SetField.
      9a521e27
    • Sameer Sahasrabuddhe's avatar
      [AMDGPU] don't mark control-flow intrinsics as convergent (#90026) · 8a65ee8b
      Sameer Sahasrabuddhe authored
      This is really a workaround to allow control flow lowering in the
      presence of convergence control tokens. Control-flow intrinsics in LLVM
      IR are convergent because they indirectly represent the wave CFG, i.e.,
      sets of threads that are "converged" or "execute in lock-step". But they
      exist during a small window in the lowering process, inserted after the
      structurizer and then translated to equivalent MIR pseudos. So rather
      than create convergence tokens for these builtins, we simply mark them
      as not convergent.
      
      The corresponding MIR pseudos are marked as having side effects, which
      is sufficient to prevent optimizations without having to mark them as
      convergent.
      8a65ee8b
    • Yingwei Zheng's avatar
    • Pavel Labath's avatar
      [lldb] Add SBType::GetByteAlign (#90960) · 30367cb5
      Pavel Labath authored
      lldb already mostly(*) tracks this information. This just makes it
      available to the SB users.
      
      (*) It does not do that for typedefs right now see llvm.org/pr90958
      30367cb5
    • Matt Arsenault's avatar
      Reapply "SystemZ: Fold copy of vector immediate to gr128" (#91099) · eb75af22
      Matt Arsenault authored
      This reverts commit a415b4df.
      
      Modify the instruction in place to transform it into a REG_SEQUENCE,
      which is what other implementations of foldImmediate do. Also start
      erasing the def instruction if there are no other uses.
      
      Fixes #91110.
      eb75af22
    • Jay Foad's avatar
      [AMDGPU] Fix typo in function name · e2c89254
      Jay Foad authored
      e2c89254
    • Matt Arsenault's avatar
      4b61d046
    • Matt Arsenault's avatar
      181e8214
    • Mehdi Amini's avatar
      Revert "Remove redundant move in return statement" (#91169) · ef8d8148
      Mehdi Amini authored
      Reverts llvm/llvm-project#90546
      
      This broke some bots, seems like some toolchain don’t consider the
      implicit move here.
      ef8d8148
    • Serge Pavlov's avatar
      [clang] Enable FPContract with optnone (#91061) · 0140ba03
      Serge Pavlov authored
      Previously treatment of the attribute `optnone` was modified in
      https://github.com/llvm/llvm-project/pull/85605 ([clang] Set correct
      FPOptions if attribute 'optnone' presents). As a side effect FPContract
      was disabled for optnone. It created unneeded divergence with the
      behavior of -O0, which enables this optimization.
      
      In the discussion
      https://github.com/llvm/llvm-project/pull/85605#issuecomment-2089350379
      it was pointed out that FP contraction should be enabled even if all
      optimizations are turned off, otherwise results of calculations would be
      different. This change enables FPContract at optnone.
      0140ba03
    • Matt Arsenault's avatar
      Reapply "AMDGPU: Implement llvm.set.rounding (#88587)" series (#91113) · d654278b
      Matt Arsenault authored
      Revert "Revert 4 last AMDGPU commits to unbreak Windows bots"
      
      This reverts commit 0d493ed2.
      
      MSVC does not like constexpr on the definition after an extern
      declaration of a global.
      d654278b
    • xiaoleis-nv's avatar
      Remove redundant move in return statement (#90546) · db532ff9
      xiaoleis-nv authored
      
      
      This pull request removes unnecessary move in the return statement to
      suppress compilation warnings.
      
      Co-authored-by: default avatarXiaolei Shi <xiaoleis@nvidia.com>
      db532ff9
    • Luke Lau's avatar
      [RISCV] Use virtual registers for AVL instrs in coalesce-vsetvli.mir. NFC · 1500dc0a
      Luke Lau authored
      All GPR registers will still be virtual at this stage, so update the test
      to reflect that.
      1500dc0a
    • martinboehme's avatar
      [clang][dataflow] Fix crash when `operator=` result type is not destination type. (#90898) · 0348e718
      martinboehme authored
      The existing code was full of comments about how we assume this is
      always the
      case, but it's not mandated by the standard, and there is code out there
      that
      returns a different type. So check that the result type is in fact the
      same as
      the destination type before attempting to copy to the result.
      
      To make sure that we don't bail out in more cases than intended, I've
      extended
      existing tests to verify that in the common case, we do return the
      destination
      object (by reference or value, as the case may be).
      0348e718
    • Yeting Kuo's avatar
      [RISCV] Teach .option arch to support experimental extensions. (#89727) · d70267fb
      Yeting Kuo authored
      Previously `.option arch` denied extenions are not belongs to RISC-V
      features. But experimental features have experimental- prefix, so
      `.option arch` can not serve for experimental extension.
      This patch uses the features of extensions to identify extension
      existance.
      d70267fb
    • Chuanqi Xu's avatar
      Reland "[Modules] No transitive source location change (#86912)" · 947b0628
      Chuanqi Xu authored
      This relands 6c311046.
      
      The patch was reverted due to incorrectly introduced alignment. And the
      patch was re-commited after fixing the alignment issue.
      
      Following off are the original message:
      
      This is part of "no transitive change" patch series, "no transitive
      source location change". I talked this with @Bigcheese in the tokyo's
      WG21 meeting.
      
      The idea comes from @jyknight posted on LLVM discourse. That for:
      
      ```
      // A.cppm
      export module A;
      ...
      
      // B.cppm
      export module B;
      import A;
      ...
      
      //--- C.cppm
      export module C;
      import C;
      ```
      
      Almost every time A.cppm changes, we need to recompile `B`. Due to we
      think the source location is significant to the semantics. But it may be
      good if we can avoid recompiling `C` if the change from `A` wouldn't
      change the BMI of B.
      
      This patch only cares source locations. So let's focus on source
      location's example. We can see the full example from the attached test.
      
      ```
      //--- A.cppm
      export module A;
      export template <class T>
      struct C {
          T func() {
              return T(43);
          }
      };
      export int funcA() {
          return 43;
      }
      
      //--- A.v1.cppm
      export module A;
      
      export template <class T>
      struct C {
          T func() {
              return T(43);
          }
      };
      export int funcA() {
          return 43;
      }
      
      //--- B.cppm
      export module B;
      import A;
      
      export int funcB() {
          return funcA();
      }
      
      //--- C.cppm
      export module C;
      import A;
      export void testD() {
          C<int> c;
          c.func();
      }
      ```
      
      Here the only difference between `A.cppm` and `A.v1.cppm` is that
      `A.v1.cppm` has an additional blank line. Then the test shows that two
      BMI of `B.cppm`, one specified `-fmodule-file=A=A.pcm` and the other
      specified `-fmodule-file=A=A.v1.pcm`, should have the bit-wise same
      contents.
      
      However, it is a different story for C, since C instantiates templates
      from A, and the instantiation records the source information from module
      A, which is different from `A` and `A.v1`, so it is expected that the
      BMI `C.pcm` and `C.v1.pcm` can and should differ.
      
      To fully understand the patch, we need to understand how we encodes
      source locations and how we serialize and deserialize them.
      
      For source locations, we encoded them as:
      
      ```
      |
      |
      | _____ base offset of an imported module
      |
      |
      |
      |_____ base offset of another imported module
      |
      |
      |
      |
      | ___ 0
      ```
      
      As the diagram shows, we encode the local (unloaded) source location
      from 0 to higher bits. And we allocate the space for source locations
      from the loaded modules from high bits to 0. Then the source locations
      from the loaded modules will be mapped to our source location space
      according to the allocated offset.
      
      For example, for,
      
      ```
      // a.cppm
      export module a;
      ...
      
      // b.cppm
      export module b;
      import a;
      ...
      ```
      
      Assuming the offset of a source location (let's name the location as
      `S`) in a.cppm is 45 and we will record the value `45` into the BMI
      `a.pcm`. Then in b.cppm, when we import a, the source manager will
      allocate a space for module 'a' (according to the recorded number of
      source locations) as the base offset of module 'a' in the current source
      location spaces. Let's assume the allocated base offset as 90 in this
      example. Then when we want to get the location in the current source
      location space for `S`, we can get it simply by adding `45` to `90` to
      `135`. Finally we can get the source location for `S` in module B as
      `135`.
      
      And when we want to write module `b`, we would also write the source
      location of `S` as `135` directly in the BMI. And to clarify the
      location `S` comes from module `a`, we also need to record the base
      offset of module `a`, 90 in the BMI of `b`.
      
      Then the problem comes. Since the base offset of module 'a' is computed
      by the number source locations in module 'a'. In module 'b', the
      recorded base offset of module 'a' will change every time the number of
      source locations in module 'a' increase or decrease. In other words, the
      contents of BMI of B will change every time the number of locations in
      module 'a' changes. This is pretty sensitive. Almost every change will
      change the number of locations. So this is the problem this patch want
      to solve.
      
      Let's continue with the existing design to understand what's going on.
      Another interesting case is:
      
      ```
      // c.cppm
      export module c;
      import whatever;
      import a;
      import b;
      ...
      ```
      
      In `c.cppm`, when we import `a`, we still need to allocate a base
      location offset for it, let's say the value becomes to `200` somehow.
      Then when we reach the location `S` recorded in module `b`, we need to
      translate it into the current source location space. The solution is
      quite simple, we can get it by `135 + (200 - 90) = 245`. In another
      word, the offset of a source location in current module can be computed
      as `Recorded Offset + Base Offset of the its module file - Recorded Base
      Offset`.
      
      Then we're almost done about how we handle the offset of source
      locations in serializers.
      
      From the abstract level, what we want to do is to remove the hardcoded
      base offset of imported modules and remain the ability to calculate the
      source location in a new module unit. To achieve this, we need to be
      able to find the module file owning a source location from the encoding
      of the source location.
      
      So in this patch, for each source location, we will store the local
      offset of the location and the module file index. For the above example,
      in `b.pcm`, the source location of `S` will be recorded as `135`
      directly. And in the new design, the source location of `S` will be
      recorded as `<1, 45>`. Here `1` stands for the module file index of `a`
      in module `b`. And `45` means the offset of `S` to the base offset of
      module `a`.
      
      So the trade-off here is that, to make the BMI more independent, we need
      to record more abstract information. And I feel it is worthy. The
      recompilation problem of modules is really annoying and there are still
      people complaining this. But if we can make this (including stopping
      other changes transitively), I think this may be a killer feature for
      modules. And from @Bigcheese , this should be helpful for clang explicit
      modules too.
      
      And the benchmarking side, I tested this patch against
      https://github.com/alibaba/async_simple/tree/CXX20Modules. No
      significant change on compilation time. The size of .pcm files becomes
      to 204M from 200M. I think the trade-off is pretty fair.
      
      I didn't use another slot to record the module file index. I tried to
      use the higher 32 bits of the existing source location encodings to
      store that information. This design may be safe. Since we use `unsigned`
      to store source locations but we use uint64_t in serialization. And
      generally `unsigned` is 32 bit width in most platforms. So it might not
      be a safe problem. Since all the bits we used to store the module file
      index is not used before. So the new encodings may be:
      
      ```
         |-----------------------|-----------------------|
         |           A           |         B         | C |
      
        * A: 32 bit. The index of the module file in the module manager + 1.
        * The +1
                here is necessary since we wish 0 stands for the current
      module file.
        * B: 31 bit. The offset of the source location to the module file
        * containing it.
        * C: The macro bit. We rotate it to the lowest bit so that we can save
        * some
                space in case the index of the module file is 0.
      ```
      
      (The B and C is the existing raw encoding for source locations)
      
      Another reason to reuse the same slot of the source location is to
      reduce the impact of the patch. Since there are a lot of places assuming
      we can store and get a source location from a slot. And if I tried to
      add another slot, a lot of codes breaks. I don't feel it is worhty.
      
      Another impact of this decision is that, the existing small
      optimizations for encoding source location may be invalided. The key of
      the optimization is that we can turn large values into small values then
      we can use VBR6 format to reduce the size. But if we decided to put the
      module file index into the higher bits, then maybe it simply doesn't
      work. An example may be the `SourceLocationSequence` optimization.
      
      This will only affect the size of on-disk .pcm files. I don't expect
      this impact the speed and memory use of compilations. And seeing my
      small experiments above, I feel this trade off is worthy.
      
      The mental model for handling source location offsets is not so complex
      and I believe we can solve it by adding module file index to each stored
      source location.
      
      For the practical side, since the source location is pretty sensitive,
      and the patch can pass all the in-tree tests and a small scale projects,
      I feel it should be correct.
      
      I'll continue to work on no transitive decl change and no transitive
      identifier change (if matters) to achieve the goal to stop the
      propagation of unnecessary changes. But all of this depends on this
      patch. Since, clearly, the source locations are the most sensitive
      thing.
      
      ---
      
      The release nots and documentation will be added seperately.
      947b0628
    • Luke Lau's avatar
    • Owen Pan's avatar
      [clang-format] Don't remove parentheses of fold expressions (#91045) · db0ed553
      Owen Pan authored
      Fixes #90966.
      db0ed553
    • Emilia Kond's avatar
      [clang-format] Don't allow comma in front of structural enum (#91056) · c609043d
      Emilia Kond authored
      Assume that a comma in front of `enum` means it is actually a part of an
      elaborated type in a template parameter list.
      
      Fixes https://github.com/llvm/llvm-project/issues/47782
      c609043d
    • Kazu Hirata's avatar
      [ADT] Reimplement operator==(StringRef, StringRef) (NFC) (#91139) · 774b7eb7
      Kazu Hirata authored
      I'm planning to deprecate and eventually remove StringRef::equals in
      favor of operator==.  This patch reimplements operator== without using
      StringRef::equals.
      
      I'm not sure if there is a good way to make StringRef::compareMemory
      available to operator==, which is not a member function.  "friend"
      works to some extent but breaks corner cases, which is why I've chosen
      to "inline" compareMemory.
      774b7eb7
    • Phoebe Wang's avatar
      [X86][FP16] Do not create VBROADCAST_LOAD for f16 without AVX2 (#91125) · f7bfb078
      Phoebe Wang authored
      AVX doesn't provide 16-bit BROADCAST instruction.
      
      Fixes #91005
      f7bfb078
    • Jeremy Kun's avatar
      fix formatting issues with ODS docs around assembly format directives (#91149) · 3d6cf533
      Jeremy Kun authored
      
      
      - Some sentences are incorrectly split across list items.
      - Some pre-formatted syntax is left in plaintext
      - Some lines end in spaces
      
      Co-authored-by: default avatarJeremy Kun <j2kun@users.noreply.github.com>
      3d6cf533
    • Doug Wyatt's avatar
      [clang backend] In AArch64's DataLayout, specify a minimum function alignment of 4. (#90702) · ddecadab
      Doug Wyatt authored
      This addresses an issue where the explicit alignment of 2 (for C++ ABI
      reasons) was being propagated to the back end and causing under-aligned
      functions (in special sections).
      
      This is an alternate approach suggested by @efriedma-quic in PR #90415.
      
      Fixes #90358
      ddecadab
    • Allen's avatar
      [AArch64][SelectionDAG] Lower multiplication by a constant to shl+sub+shl+sub (#90199) · e1236430
      Allen authored
      Change the costmodel to lower a = b * C where C = 1 - (1 - 2^m) * 2^n to
                    sub  w8, w0, w0, lsl #m
                    sub  w0, w0, w8, lsl #n
      Fix https://github.com/llvm/llvm-project/issues/89430
      e1236430
    • Fangrui Song's avatar
      X86FixupBWInsts: Remove redundant code. NFC · 2aaec48d
      Fangrui Song authored
      2aaec48d
    • Eric's avatar
      [NFC] Remove BLOCKLIT workaround. (#91001) · 2574cabd
      Eric authored
      Lit already has support for stopping LIT from parsing further test
      directives. It is
      
      // END.
      
      After that directive, LIT will stop parsing.
      
      This change removes the BLOCKLIT hack and replaces it with END.
      2574cabd
    • Kazu Hirata's avatar
      [Target] Use StringRef::operator== instead of StringRef::equals (NFC) (#91072) (#91138) · c18bcd0a
      Kazu Hirata authored
      I'm planning to remove StringRef::equals in favor of
      StringRef::operator==.
      
      - StringRef::operator==/!= outnumber StringRef::equals by a factor of
        38 under llvm/ in terms of their usage.
      
      - The elimination of StringRef::equals brings StringRef closer to
        std::string_view, which has operator== but not equals.
      
      - S == "foo" is more readable than S.equals("foo"), especially for
        !Long.Expression.equals("str") vs Long.Expression != "str".
      c18bcd0a
    • Florian Hahn's avatar
      [LAA] Directly pass DepChecker to getSource/getDestination (NFC). · 3219c0ed
      Florian Hahn authored
      Instead of passing LoopAccessInfo only to fetch the MemoryDepChecker,
      directly pass MemoryDepChecker. This simplifies the code and also allows
      new uses in places where no LAI is available.
      3219c0ed
    • Fangrui Song's avatar
      [HLSL] Remove overridden -S · 57f13b51
      Fangrui Song authored
      The cc1 option -S (https://reviews.llvm.org/D124983) is overridden by
      the latter -emit-llvm.
      57f13b51
    • Fangrui Song's avatar
      [test] %clang_cc1: remove redundant actions · d33937b6
      Fangrui Song authored
      ParseFrontendArgs takes the last OPT_Action_Group option. The other
      actions are overridden.
      d33937b6
    • David Blaikie's avatar
      Add new BuiltinType introduced in 7a484d3a · 41574f5a
      David Blaikie authored
      I don't think this is one lldb would encounter when building ASTs from
      DWARF.
      41574f5a
    • Fangrui Song's avatar
      [test] %clang_cc1: remove redundant actions · 7e59223a
      Fangrui Song authored
      ParseFrontendArgs takes the last OPT_Action_Group option. The other
      actions are overridden.
      7e59223a
    • Jeremy Kun's avatar
      Upstream polynomial.ntt and polynomial.intt (#90992) · 624c9fc8
      Jeremy Kun authored
      
      
      These two ops represent a number-theoretic transform of a polynomial to
      a tensor of evaluations of the polynomial at a list of powers of
      primitive roots of the polynomial.
      
      To support this, a new optional attribute is added to the ring attribute
      to specify the primitive root of unity used for the NTT. A verifier for
      the op is added to ensure the chosen root is a primitive nth root of
      unity.
      
      ---------
      
      Co-authored-by: default avatarJeremy Kun <j2kun@users.noreply.github.com>
      Co-authored-by: default avatarOleksandr "Alex" Zinenko <ftynse@gmail.com>
      624c9fc8
  2. May 05, 2024