1. May 05, 2024
  2. May 04, 2024
    • Karl-Johan Karlsson's avatar
      [clang][CodeGen] Propagate pragma set fast-math flags to floating point builtins (#90377) · cb015b9e
      Karl-Johan Karlsson authored
      This is a fix for the issue #87758 where fast-math flags are not
      propagated all builtins.
      
      It seems like pragmas with fast math flags was only propagated to calls
      of unary floating point builtins. This patch propagate them also for
      binary and ternary floating point builtins.
      cb015b9e
    • Kazu Hirata's avatar
      [Support] Use StringRef::operator== instead of StringRef::equals (NFC) (#91042) · 7ee62883
      Kazu Hirata authored
      I'm planning to remove StringRef::equals in favor of
      StringRef::operator==.
      
      - StringRef::operator== outnumbers StringRef::equals by a factor of 25
        under llvm/ in terms of their usage.
      
      - The elimination of StringRef::equals brings StringRef closer to
        std::string_view, which has operator== but not equals.
      
      - S == "foo" is more readable than S.equals("foo"), especially for
        !Long.Expression.equals("str") vs Long.Expression != "str".
      7ee62883
    • Matt Stephanson's avatar
      [libc++] Adjust some of the [rand.dist] critical values that are too strict (#88669) · 76aa042d
      Matt Stephanson authored
      Adjust some of the [rand.dist] critical values that are too strict
      
      - Most critical values are determined empirically by running each test
      51
      times with a different PRNG seed and finding the smallest symmetric
      interval
      around the median that contains 90% of the sample means, variances, etc.
      
      - For the Kolmogorov-Smirnov tests, the alpha=0.1 critical value for
      large N
         is 1.224/sqrt(N).
      
      - For normally distributed variates, the sample kurtosis is distributed
      as
         Normal(0, 24/N). For N=1e5, this gives a 90% confidence interval of
      0+/-0.0255. For Binomial(40, 0.25), which is approximately normal, the
         kurtosis is -0.0167, so the relative 90% CI is large, on the order of
      0.0255/0.0167 = 153%. In most cases the distribution of the sample
      kurtosis
      isn't known analytically, but similarly large relative tolerances can be
         expected if the kurtosis is near zero.
      76aa042d
    • Simon Pilgrim's avatar
      [DAG] Fold freeze(shuffle(x,y,m)) -> shuffle(freeze(x),freeze(y),m) (#90952) · caacf868
      Simon Pilgrim authored
      If the shuffle mask contains no undef elements, then we can move the freeze through a shuffle node.
      
      This requires special case handling to create a new ShuffleVectorSDNode.
      
      Includes VECTOR_SHUFFLE support for isGuaranteedNotToBeUndefOrPoison  / canCreateUndefOrPoison.
      caacf868
    • orbiri's avatar
      [MLIR] Extend floating point parsing support (#90442) · 1e3c630f
      orbiri authored
      Parsing support for floating point types was missing a few features:
      1. Parsing floating point attributes from integer literals was supported
      only for types with bitwidth smaller or equal to 64.
      2. Downstream users could not use `AsmParser::parseFloat` to parse float
      types which are printed as integer literals.
      
      This commit addresses both these points. It extends
      `Parser::parseFloatFromIntegerLiteral` to support arbitrary bitwidth,
      and exposes a new API to parse arbitrary floating point given an
      fltSemantics as input. The usage of this new API is introduced in the
      Test Dialect.
      1e3c630f
    • Nikita Kniazev's avatar
    • Kristof Beyls's avatar
      [BOLT] Fix runOnEachFunctionWithUniqueAllocId (#90039) · 554459a0
      Kristof Beyls authored
      When runOnEachFunctionWithUniqueAllocId is invoked with
      ForceSequential=true, then the current implementation runs the function
      with AllocId==0, which is the Id for the shared, non-unique, default
      AnnotationAllocator.
      
      However, the documentation for runOnEachFunctionWithUniqueAllocId
      states:
      ```
      /// Perform the work on each BinaryFunction except those that are rejected
      /// by SkipPredicate, and create a unique annotation allocator for each
      /// task. This should be used whenever the work function creates annotations to
      /// allow thread-safe annotation creation.
      ```
      
      Therefore, even when ForceSequential==true, a unique AllocId should be
      used, i.e. different from 0.
      
      In the current upstream BOLT this is presumably not depended on, but it
      is needed to reduce memory usage for analyses that use a lot of
      memory/annotations. Examples are the pac-ret and stack-clash analyses
      that currently have prototype implementations as described in
      https://discourse.llvm.org/t/rfc-bolt-based-binary-analysis-tool-to-verify-correctness-of-security-hardening/78148
      
      
      These analyses use the DataFlowAnalysis framework to sometimes store
      quite a lot of information on each MCInst. They run in parallel on each
      function. When the dataflow analysis is finished, the annotations on
      each MCInst can be removed, hugely saving on memory consumption. The
      only annotations that need to remain are those that indicate some
      unexpected properties somewhere in the binary.
      
      Fixing this bug enables implementing the deletion of the memory used by
      those huge number of DataFlowAnalysis annotations (by invoking
      BC.MIB->freeValuesAllocator(AllocatorId)), even when run with
      --no-threads. Without this bug fixed, the invocation of
      BC.MIB->freeValuesAllocator(AllocatorId) results in also the memory for
      all other annotations to be deleted, as AllocatorId is 0.
      
      ---------
      
      Co-authored-by: default avatarMaksim Panchenko <maks@meta.com>
      554459a0
    • Andreas Jonson's avatar
    • Nikita Popov's avatar
      [InstCombine] Do not request non-splat vector support in code reviews (NFC) (#90709) · f16e234f
      Nikita Popov authored
      The InstCombine contributor guide already says:
      
      > Handle non-splat vector constants if doing so is free, but do
      > not add handling for them if it adds any additional complexity
      > to the code.
      
      This change strengthens this guideline to explicitly discourage
      asking (new) contributors to implement non-splat support during code
      reviews. Doing so will almost certainly increase the number of
      necessary review iterations, or result in outright contradictory review
      feedback, as different people are willing to accept a different degree
      of complexity for non-splat vector support.
      f16e234f
    • Patrick O'Neill's avatar
      [lld] Error on unsupported split stack (#88063) · 96aac679
      Patrick O'Neill authored
      Targets with no `-fstack-split` support now emit `ld.lld: error: target
      doesn't support split stacks` instead of `UNREACHABLE executed` with a
      backtrace asking the user to report a bug.
      
      Resolves #88061
      96aac679
    • Rafael Ubal's avatar
      Avoid buffer hoisting from parallel loops (#90735) · a42a2ca1
      Rafael Ubal authored
      This change corrects an invalid behavior in pass
      `--buffer-loop-hoisting`. The pass is in charge of extracting buffer
      allocations (e.g., `memref.alloca`) from loop regions (e.g., `scf.for`)
      when possible. This works OK for looks with sequential execution
      semantics. However, a buffer allocated in the body of a parallel loop
      may be concurrently accessed by multiple thread to store its local data.
      Extracting such buffer from the loop causes all threads to wrongly share
      the same memory region.
      
      In the following example, dimension 1 of the input tensor is reversed.
      Dimension 0 is traversed with a parallel loop.
      
      ```
      func.func @f(%input: memref<2x3xf32>) -> memref<2x3xf32> {
        %c0 = index.constant 0
        %c1 = index.constant 1
        %c2 = index.constant 2
        %c3 = index.constant 3
      
        %output = memref.alloc() : memref<2x3xf32>
        scf.parallel (%index) = (%c0) to (%c2) step (%c1) {
          // Create subviews for working input and output slices
          %input_slice = memref.subview %input[%index, 2][1, 3][1, -1] : memref<2x3xf32> to memref<1x3xf32, strided<[3, -1], offset: ?>>
          %output_slice = memref.subview %output[%index, 0][1, 3][1, 1] : memref<2x3xf32> to memref<1x3xf32, strided<[3, 1], offset: ?>>
      
          // Copy the input slice into this temporary buffer. This intermediate
          // copy is unnecessary, but is used for illustration purposes.
          %temp = memref.alloc() : memref<1x3xf32>
          memref.copy %input_slice, %temp : memref<1x3xf32, strided<[3, -1], offset: ?>> to memref<1x3xf32>
      
          // Copy temporary buffer into output slice
          memref.copy %temp, %output_slice : memref<1x3xf32> to memref<1x3xf32, strided<[3, 1], offset: ?>>
          scf.reduce
        }
      
        return %output : memref<2x3xf32>
      }
      ```
      
      The patch submitted here prevents `%temp = memref.alloc() :
      memref<1x3xf32>` from being hoisted when the containing op is
      `scf.parallel` or `scf.forall`. A new op trait called
      `HasParallelRegion` is introduced and assigned to these two ops to
      indicate that their regions have parallel execution semantics.
      
      @joker-eph @ftynse @nicolasvasilache @sabauma
      a42a2ca1
    • Joseph Huber's avatar
      [libc] Fix assert dependency on macro header (#91036) · 1022636b
      Joseph Huber authored
      Summary:
      This file was missing a dependency so it wasn't being installed.
      1022636b
    • paperchalice's avatar
      [Instrumentation] Support verifying machine function (#90931) · e7939d0d
      paperchalice authored
      We need it to test isel related passes. Currently
      `verifyMachineFunction` is incomplete (no LiveIntervals support), but is
      enough for testing isel pass, will migrate to complete
      `MachineVerifierPass` in future.
      e7939d0d
    • Maksim Levental's avatar
      Update GettingInvolved.rst (#91008) · b958ef19
      Maksim Levental authored
      b958ef19
    • Fangrui Song's avatar
      666679a5
    • Johannes Doerfert's avatar
      cd3a4c31
    • Craig Topper's avatar
      [AArch64] Pre-commit another test case for #90936. NFC · 0c7e706c
      Craig Topper authored
      Another similar problem was added to the ticket after the first fix.
      0c7e706c
    • Vitaly Buka's avatar
      [tsan] Don't crash on vscale (#91018) · a441645f
      Vitaly Buka authored
      
      
      Co-authored-by: default avatarHeejin Ahn <aheejin@gmail.com>
      a441645f
    • Teresa Johnson's avatar
      [MemProf] Optionally match profiles on to manually hinted hot/cold new (#91027) · e5cbe8fd
      Teresa Johnson authored
      While we don't currently rewrite the hints on manually hot/cold hinted
      allocations, enable optionally matching profiles onto those allocations
      as a first step to being able to do this.
      
      By explicitly checking whether the library function is in the list of
      operator new also fixes one limitation of the prior call to isNewLikeFn.
      Some operator new calls (those that specify nothrow) are considered
      Malloc-like because they may return null. We want to be able to match
      and rewrite these. Therefore the new test uses a nothrow variant to test
      the fix for this as well.
      e5cbe8fd
    • Benoit Jacob's avatar
      Let `memref.expand_shape` implement `ReifyRankedShapedTypeOpInterface` (#90975) · b05a12e9
      Benoit Jacob authored
      This is a new take on #89111. Now that #90040 is merged, this has become
      trivial to implement. The added test shows the kind of benefit that we
      get from this: now dim-of-expand-shape naturally folds without us
      needing to implement an ad-hoc folding rewrite.
      b05a12e9
    • Heejin Ahn's avatar
      [WebAssembly] Add all remaining features to bleeding-edge (#90875) · 5d81b1c5
      Heejin Ahn authored
      I'm not entirely sure what the criteria for 'bleeding-edge' used to be,
      but at this point it seems to be the set of all added features in LLVM.
      This adds remaining features to bleeding-edge config.
      5d81b1c5
    • Krystian Stasiowski's avatar
      [Clang][Sema] Fix template name lookup for operator= (#90999) · 3191e0b5
      Krystian Stasiowski authored
      This fixes a bug in #90152 where `operator=` was never looked up in the
      current instantiation, resulting in `<` never being interpreted as the
      start of a template argument list.
      
      Since function templates are not copy/move assignment operators, the fix
      is accomplished by allowing lookup in the current instantiation for
      `operator=` when looking up a template name.
      3191e0b5
    • Alexey Bataev's avatar
      [SLP]Fix PR90892: do a correct sign analysis of the entries elements in gather shuffles. · 03972261
      Alexey Bataev authored
      Need to do extra analysis of the scalar elements of the tree entry to be
      shuffled instead of the vectorized value to correctly deduce signedness
      info.
      03972261
    • Reid Kleckner's avatar
      Revert "[gn] port 2d4acb08 (LLVM_ENABLE_CURL)" · 48039b19
      Reid Kleckner authored
      This reverts commit 0558c7e0 to match
      the revert of 2d4acb08 in  327bfc97
      48039b19
    • Reid Kleckner's avatar
      [ARM/X86] Standardize the isEligibleForTailCallOptimization prototypes (#90688) · 385faf9c
      Reid Kleckner authored
      Pass in CallLoweringInfo (CLI) instead of passing in the various fields
      directly. Also pass in CCState (CCInfo), which is computed in both the
      caller and the callee for a minor efficiency saving. There may also be a
      small correctness improvement for sibcalls with vectorcall, which has an
      odd way of recomputing argument locations.
      
      This is a step towards improving the handling of musttail on armv7,
      which we have numerous issues filed about in our tracker.
      
      I took inspiration for this from the RISCV tail call eligibility check,
      which uses a similar prototype.
      385faf9c
    • Alexey Bataev's avatar
    • Reid Kleckner's avatar
      [clang] Note that optnone and target attributes do not apply to nested functions (#82815) · 4e6d30e2
      Reid Kleckner authored
      This behavior is true for all attributes, but this behavior can be
      surprising for attributes which have function-wide effects, such as
      `optnone` and `target`. Most other function attributes affect the
      prototype or semantics, but do not affect code generation in the
      function body. I believe it is worth calling this out in the
      documentation of these function-wide attributes. There may be more,
      these were the two that came to mind.
      4e6d30e2
    • jeffreytan81's avatar
      Fix dap variable value format issue (#90799) · b8d38bb5
      jeffreytan81 authored
      
      
      While adding a UI feature in VSCode to toggle hex/dec in variables view
      window. I noticed that it does not work after second toggle. Then I
      noticed that there is a bug that we only explicitly set hex format not
      reset back to default during further toggle. The new test demonstrates
      the bug.
      
      This PR resets the format back to default if not using hex. One
      complexity is that, we explicitly set registers value format to
      AddressInfo, which shouldn't be overridden by default or hex settings.
      
      ---------
      
      Co-authored-by: default avatarjeffreytan81 <jeffreytan@fb.com>
      b8d38bb5
    • Jorge Gorbe Moya's avatar
      Revert "[BasicBlockUtils] Remove redundant llvm.dbg instructions after blocks... · 2cde0e2f
      Jorge Gorbe Moya authored
      Revert "[BasicBlockUtils] Remove redundant llvm.dbg instructions after blocks to reduce compile time (#89069)"
      
      This reverts commit 2e3e0868. It caused
      quadratic slowdown at compilation time in some cases. See the comments
      in the original PR: https://github.com/llvm/llvm-project/pull/89069
      2cde0e2f
    • Chris B's avatar
      [DirectX] Remove unneccary check lines (#90979) · 9299a136
      Chris B authored
      These check lines break as of 91446e2a due to changes in how LLVM
      handles debug information. Since debug informaiton isn't important to
      what this test is verifying we can remove the check lines.
      9299a136
    • Matt Arsenault's avatar
      AMDGPU: Add tests for minimum and maximum intrinsics (#90997) · 7ec698e6
      Matt Arsenault authored
      Baseline tests for new expansion. I think we can do better and avoid the
      classes.
      7ec698e6
    • whisperity's avatar
    • Jonas Devlieghere's avatar
      Revert "[lldb] Unify CalculateMD5 return types" (#90998) · ca8b0649
      Jonas Devlieghere authored
      Reverts llvm/llvm-project#90921
      ca8b0649
    • Noah Goldstein's avatar
      [InstCombine] Add example usage for new Checked matcher API · f561daf9
      Noah Goldstein authored
      There is no real motivation for this change other than to highlight a
      case where the new `Checked` matcher API can handle non-splat-vecs
      without increasing code complexity.
      
      Closes #85676
      f561daf9
    • Noah Goldstein's avatar
    • Noah Goldstein's avatar
      [PatternMatching] Add generic API for matching constants using custom conditions · d8428dfe
      Noah Goldstein authored
      The new API is:
          `m_CheckedInt(Lambda)`/`m_CheckedFp(Lambda)`
              - Matches non-undef constants s.t `Lambda(ele)` is true for all
                elements.
          `m_CheckedIntAllowUndef(Lambda)`/`m_CheckedFpAllowUndef(Lambda)`
              - Matches constants/undef s.t `Lambda(ele)` is true for all
                elements.
      
      The goal with these is to be able to replace the common usage of:
      ```
          match(X, m_APInt(C)) && CustomCheck(C)
      ```
      with
      ```
          match(X, m_CheckedInt(C, CustomChecks);
      ```
      
      The rationale if we often ignore non-splat vectors because there are
      no good APIs to handle them with and its not worth increasing code
      complexity for such cases.
      
      The hope is the API creates a common method handling
      scalars/splat-vecs/non-splat-vecs to essentially make this a
      non-issue.
      d8428dfe
    • Noah Goldstein's avatar
      [Inliner] Propagate callee argument memory access attributes before inlining · 285dbed1
      Noah Goldstein authored
      To avoid losing information, we can propagate some access attribute
      from the to-be-inlined callee to its callsites.
      
      We can propagate argument memory access attributes to callsite
      parameters if they are from the same underlying object.
      
      Closes #89024
      285dbed1