1. May 05, 2024
  2. May 04, 2024
    • Karl-Johan Karlsson's avatar
      [clang][CodeGen] Propagate pragma set fast-math flags to floating point builtins (#90377) · cb015b9e
      Karl-Johan Karlsson authored
      This is a fix for the issue #87758 where fast-math flags are not
      propagated all builtins.
      
      It seems like pragmas with fast math flags was only propagated to calls
      of unary floating point builtins. This patch propagate them also for
      binary and ternary floating point builtins.
      cb015b9e
    • Kazu Hirata's avatar
      [Support] Use StringRef::operator== instead of StringRef::equals (NFC) (#91042) · 7ee62883
      Kazu Hirata authored
      I'm planning to remove StringRef::equals in favor of
      StringRef::operator==.
      
      - StringRef::operator== outnumbers StringRef::equals by a factor of 25
        under llvm/ in terms of their usage.
      
      - The elimination of StringRef::equals brings StringRef closer to
        std::string_view, which has operator== but not equals.
      
      - S == "foo" is more readable than S.equals("foo"), especially for
        !Long.Expression.equals("str") vs Long.Expression != "str".
      7ee62883
    • Matt Stephanson's avatar
      [libc++] Adjust some of the [rand.dist] critical values that are too strict (#88669) · 76aa042d
      Matt Stephanson authored
      Adjust some of the [rand.dist] critical values that are too strict
      
      - Most critical values are determined empirically by running each test
      51
      times with a different PRNG seed and finding the smallest symmetric
      interval
      around the median that contains 90% of the sample means, variances, etc.
      
      - For the Kolmogorov-Smirnov tests, the alpha=0.1 critical value for
      large N
         is 1.224/sqrt(N).
      
      - For normally distributed variates, the sample kurtosis is distributed
      as
         Normal(0, 24/N). For N=1e5, this gives a 90% confidence interval of
      0+/-0.0255. For Binomial(40, 0.25), which is approximately normal, the
         kurtosis is -0.0167, so the relative 90% CI is large, on the order of
      0.0255/0.0167 = 153%. In most cases the distribution of the sample
      kurtosis
      isn't known analytically, but similarly large relative tolerances can be
         expected if the kurtosis is near zero.
      76aa042d
    • Simon Pilgrim's avatar
      [DAG] Fold freeze(shuffle(x,y,m)) -> shuffle(freeze(x),freeze(y),m) (#90952) · caacf868
      Simon Pilgrim authored
      If the shuffle mask contains no undef elements, then we can move the freeze through a shuffle node.
      
      This requires special case handling to create a new ShuffleVectorSDNode.
      
      Includes VECTOR_SHUFFLE support for isGuaranteedNotToBeUndefOrPoison  / canCreateUndefOrPoison.
      caacf868
    • orbiri's avatar
      [MLIR] Extend floating point parsing support (#90442) · 1e3c630f
      orbiri authored
      Parsing support for floating point types was missing a few features:
      1. Parsing floating point attributes from integer literals was supported
      only for types with bitwidth smaller or equal to 64.
      2. Downstream users could not use `AsmParser::parseFloat` to parse float
      types which are printed as integer literals.
      
      This commit addresses both these points. It extends
      `Parser::parseFloatFromIntegerLiteral` to support arbitrary bitwidth,
      and exposes a new API to parse arbitrary floating point given an
      fltSemantics as input. The usage of this new API is introduced in the
      Test Dialect.
      1e3c630f
    • Nikita Kniazev's avatar
    • Kristof Beyls's avatar
      [BOLT] Fix runOnEachFunctionWithUniqueAllocId (#90039) · 554459a0
      Kristof Beyls authored
      When runOnEachFunctionWithUniqueAllocId is invoked with
      ForceSequential=true, then the current implementation runs the function
      with AllocId==0, which is the Id for the shared, non-unique, default
      AnnotationAllocator.
      
      However, the documentation for runOnEachFunctionWithUniqueAllocId
      states:
      ```
      /// Perform the work on each BinaryFunction except those that are rejected
      /// by SkipPredicate, and create a unique annotation allocator for each
      /// task. This should be used whenever the work function creates annotations to
      /// allow thread-safe annotation creation.
      ```
      
      Therefore, even when ForceSequential==true, a unique AllocId should be
      used, i.e. different from 0.
      
      In the current upstream BOLT this is presumably not depended on, but it
      is needed to reduce memory usage for analyses that use a lot of
      memory/annotations. Examples are the pac-ret and stack-clash analyses
      that currently have prototype implementations as described in
      https://discourse.llvm.org/t/rfc-bolt-based-binary-analysis-tool-to-verify-correctness-of-security-hardening/78148
      
      
      These analyses use the DataFlowAnalysis framework to sometimes store
      quite a lot of information on each MCInst. They run in parallel on each
      function. When the dataflow analysis is finished, the annotations on
      each MCInst can be removed, hugely saving on memory consumption. The
      only annotations that need to remain are those that indicate some
      unexpected properties somewhere in the binary.
      
      Fixing this bug enables implementing the deletion of the memory used by
      those huge number of DataFlowAnalysis annotations (by invoking
      BC.MIB->freeValuesAllocator(AllocatorId)), even when run with
      --no-threads. Without this bug fixed, the invocation of
      BC.MIB->freeValuesAllocator(AllocatorId) results in also the memory for
      all other annotations to be deleted, as AllocatorId is 0.
      
      ---------
      
      Co-authored-by: default avatarMaksim Panchenko <maks@meta.com>
      554459a0
    • Andreas Jonson's avatar
    • Nikita Popov's avatar
      [InstCombine] Do not request non-splat vector support in code reviews (NFC) (#90709) · f16e234f
      Nikita Popov authored
      The InstCombine contributor guide already says:
      
      > Handle non-splat vector constants if doing so is free, but do
      > not add handling for them if it adds any additional complexity
      > to the code.
      
      This change strengthens this guideline to explicitly discourage
      asking (new) contributors to implement non-splat support during code
      reviews. Doing so will almost certainly increase the number of
      necessary review iterations, or result in outright contradictory review
      feedback, as different people are willing to accept a different degree
      of complexity for non-splat vector support.
      f16e234f
    • Patrick O'Neill's avatar
      [lld] Error on unsupported split stack (#88063) · 96aac679
      Patrick O'Neill authored
      Targets with no `-fstack-split` support now emit `ld.lld: error: target
      doesn't support split stacks` instead of `UNREACHABLE executed` with a
      backtrace asking the user to report a bug.
      
      Resolves #88061
      96aac679
    • Rafael Ubal's avatar
      Avoid buffer hoisting from parallel loops (#90735) · a42a2ca1
      Rafael Ubal authored
      This change corrects an invalid behavior in pass
      `--buffer-loop-hoisting`. The pass is in charge of extracting buffer
      allocations (e.g., `memref.alloca`) from loop regions (e.g., `scf.for`)
      when possible. This works OK for looks with sequential execution
      semantics. However, a buffer allocated in the body of a parallel loop
      may be concurrently accessed by multiple thread to store its local data.
      Extracting such buffer from the loop causes all threads to wrongly share
      the same memory region.
      
      In the following example, dimension 1 of the input tensor is reversed.
      Dimension 0 is traversed with a parallel loop.
      
      ```
      func.func @f(%input: memref<2x3xf32>) -> memref<2x3xf32> {
        %c0 = index.constant 0
        %c1 = index.constant 1
        %c2 = index.constant 2
        %c3 = index.constant 3
      
        %output = memref.alloc() : memref<2x3xf32>
        scf.parallel (%index) = (%c0) to (%c2) step (%c1) {
          // Create subviews for working input and output slices
          %input_slice = memref.subview %input[%index, 2][1, 3][1, -1] : memref<2x3xf32> to memref<1x3xf32, strided<[3, -1], offset: ?>>
          %output_slice = memref.subview %output[%index, 0][1, 3][1, 1] : memref<2x3xf32> to memref<1x3xf32, strided<[3, 1], offset: ?>>
      
          // Copy the input slice into this temporary buffer. This intermediate
          // copy is unnecessary, but is used for illustration purposes.
          %temp = memref.alloc() : memref<1x3xf32>
          memref.copy %input_slice, %temp : memref<1x3xf32, strided<[3, -1], offset: ?>> to memref<1x3xf32>
      
          // Copy temporary buffer into output slice
          memref.copy %temp, %output_slice : memref<1x3xf32> to memref<1x3xf32, strided<[3, 1], offset: ?>>
          scf.reduce
        }
      
        return %output : memref<2x3xf32>
      }
      ```
      
      The patch submitted here prevents `%temp = memref.alloc() :
      memref<1x3xf32>` from being hoisted when the containing op is
      `scf.parallel` or `scf.forall`. A new op trait called
      `HasParallelRegion` is introduced and assigned to these two ops to
      indicate that their regions have parallel execution semantics.
      
      @joker-eph @ftynse @nicolasvasilache @sabauma
      a42a2ca1
    • Joseph Huber's avatar
      [libc] Fix assert dependency on macro header (#91036) · 1022636b
      Joseph Huber authored
      Summary:
      This file was missing a dependency so it wasn't being installed.
      1022636b
    • paperchalice's avatar
      [Instrumentation] Support verifying machine function (#90931) · e7939d0d
      paperchalice authored
      We need it to test isel related passes. Currently
      `verifyMachineFunction` is incomplete (no LiveIntervals support), but is
      enough for testing isel pass, will migrate to complete
      `MachineVerifierPass` in future.
      e7939d0d
    • Maksim Levental's avatar
      Update GettingInvolved.rst (#91008) · b958ef19
      Maksim Levental authored
      b958ef19
    • Fangrui Song's avatar
      666679a5
    • Johannes Doerfert's avatar
      cd3a4c31
    • Craig Topper's avatar
      [AArch64] Pre-commit another test case for #90936. NFC · 0c7e706c
      Craig Topper authored
      Another similar problem was added to the ticket after the first fix.
      0c7e706c