1. Apr 30, 2022
  2. Apr 29, 2022
    • Craig Topper's avatar
      [RISCV] Improve constant materialization for cases that can use LUI+ADDI instead of LUI+ADDIW. · 5c383731
      Craig Topper authored
      It's possible that we have a constant that isn't simm32 so we can't
      use LUI+ADDIW, but we can use LUI+ADDI. Because ADDI uses a sign
      extended constant, it's possible that after subtracting it out, we
      end up with a simm32 that maps to LUI.
      
      This patch detects this case after removing Lo12 and before shifting
      the value for SLLI.
      
      Reviewed By: luismarques
      
      Differential Revision: https://reviews.llvm.org/D124222
      5c383731
    • Joseph Huber's avatar
      [OpenMP] Allow CUDA to be linked with OpenMP using the new driver · d9c64d33
      Joseph Huber authored
      After basic support for embedding and handling CUDA files was added to
      the new driver, we should be able to call CUDA functions from OpenMP
      code. This patch makes the necessary changes to successfuly link in CUDA
      programs that were compiled using the new driver. With this patch it
      should be possible to compile device-only CUDA code (no kernels) and
      call it from OpenMP as follows:
      
      ```
      $ clang++ cuda.cu -fopenmp-new-driver -offload-arch=sm_70 -c
      $ clang++ openmp.cpp cuda.o -fopenmp-new-driver -fopenmp -fopenmp-targets=nvptx64 -Xopenmp-target=nvptx64 -march=sm_70
      ```
      
      Currently this requires using a host variant to suppress the generation
      of a CPU-side fallback call.
      
      Depends on D120272
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D120273
      d9c64d33
    • Nikita Popov's avatar
      [InstCombine] Require LoopInfo in test (NFC) · 6aeb2a21
      Nikita Popov authored
      This test case doesn't show what it was intended to without
      require<loops>.
      6aeb2a21
    • Joseph Huber's avatar
      [OpenMP] Add options to only compile the host or device when offloading · 47d66255
      Joseph Huber authored
      OpenMP recently moved to the new offloading driver, this had the effect
      of making it more difficult to inspect intermediate code for the device.
      This patch adds `-foffload-host-only` and `-foffload-device-only` to
      control which sides get compiled. This will allow users to more easily
      inspect output without needing the temp files.
      
      Reviewed By: tra
      
      Differential Revision: https://reviews.llvm.org/D124220
      47d66255
    • Nikita Popov's avatar
    • Simon Pilgrim's avatar
      [X86] lowerShuffleAsRepeatedMaskAndLanePermute - move the sublane split code... · b424055b
      Simon Pilgrim authored
      [X86] lowerShuffleAsRepeatedMaskAndLanePermute - move the sublane split code into a lambda helper. NFC.
      
      This is a NFC cleanup as part of the work on #55066 - the idea being that we will be able to check for multiple sub lane scales.
      b424055b
    • Alexey Bataev's avatar
      [COST]Fix crash for non-power-2 vector shuffle mask. · 371412e0
      Alexey Bataev authored
      Need to normalizize the mask to avoid possible crashes during attempts
      to estimate cost of the very long shuffles with non-power-2 number of
      elements in masks.
      371412e0
    • Florian Hahn's avatar
      [SimplifyCFG] Avoid shifting by a too large exponent. · a8008176
      Florian Hahn authored
      TI->getBitWidth can be > 64 and in those cases the shift will be UB due
      to the exponent being too large.
      
      To fix this, cap the shift at 63. I think this should work out fine,
      because TableSize is itself a 64 bit type and the maximum table size
      must fit in the type. Also, if we would underestimate the size here, at
      most we get an extra ZExt.
      
      Reviewed By: spatel
      
      Differential Revision: https://reviews.llvm.org/D124608
      a8008176
    • David Candler's avatar
      Additionally set f32 mode with denormal-fp-math · 9e7c9967
      David Candler authored
      When the denormal-fp-math option is used, this should set the
      denormal handling mode for all floating point types. However,
      currently 32-bit float types can ignore this setting as there is a
      variant of the option, denormal-fp-math-f32, specifically for that type
      which takes priority when checking the mode based on type and remains
      at the default of IEEE. From the description, denormal-fp-math would
      be expected to set the mode for floats unless overridden by the f32
      variant, and code in the front end only emits the f32 option if it is
      different to the general one, so setting just denormal-fp-math should
      be valid.
      
      This patch changes the denormal-fp-math option to also set the f32
      mode. If denormal-fp-math-f32 is also specified, this is then
      overridden as expected, but if it is absent floats will be set to the
      mode specified by the former option, rather than remain on the default.
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D122589
      9e7c9967
    • Anna Thomas's avatar
      [CompileTime] [Passes] Avoid computing unnecessary analyses. NFC · 205246cb
      Anna Thomas authored
      Similar to c515b2f3, If there are no loops in the function as seen
      through LI, we should avoid computing the remaining expensive analyses
      (such as SCEV, BPI).  Reordered the analyses requests and early return
      if there are no loops.
      
      The logic of avoiding expensive analyses is applied to LoopVectorizer,
      LoopLoadElimination and LoopUnrollPass, i.e. all function passes which operate
      on loops.
      
      This is an NFC with compile time improvement.
      
      Differential Revision: https://reviews.llvm.org/D124529
      205246cb
    • Stefan Pintilie's avatar
      [PowerPC][NFC] Add a function to determine if a call needs to be NOTOC. · f685bce8
      Stefan Pintilie authored
      Add the isNoTOCCallInstr function to PPCInstrInfo to determine if a call opcode
      does not need a TOC restore after the call. All call opcodes should be listed in
      this function. A default unreachable in this function should force future call
      opcodes to also be added.
      
      This is a follow up patch to D122012
      
      Reviewed By: jsji, shchenz
      
      Differential Revision: https://reviews.llvm.org/D124415
      f685bce8
    • Martin Boehme's avatar
      [clang] Eliminate TypeProcessingState::trivial. · 23c10e8d
      Martin Boehme authored
      This flag is redundant -- it's true iff `savedAttrs` is empty.
      
      Querying `savedAttrs.empty()` should not take any more time than querying the
      `trivial` flag, so this should not have a performance impact either.
      
      I noticed this while working on https://reviews.llvm.org/D111548.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D123783
      23c10e8d
    • Paul Walker's avatar
      [DAGCombiner] Stop invalid sign conversion in refineIndexType. · 23c50975
      Paul Walker authored
      When looking through extends of gather/scatter indices it's safe
      to convert a known positive signed index to unsigned, but unsigned
      indices must remain unsigned.
      
      Depends On D123318
      
      Differential Revision: https://reviews.llvm.org/D123326
      23c50975
    • Paul Walker's avatar
      [SVE][ISel] Ensure explicit gather/scatter offset extension isn't lost. · 59588f0a
      Paul Walker authored
      getGatherScatterIndexIsExtended currently looks through all
      SIGN_EXTEND_INREG operations regardless of their input type.  This
      patch restricts the code to only look through i32->i64 extensions,
      which are the ones supported implicitly by SVE addressing modes.
      
      Differential Revision: https://reviews.llvm.org/D123318
      59588f0a
    • Joseph Huber's avatar
      [CUDA] Add driver support for compiling CUDA with the new driver · c5e5b543
      Joseph Huber authored
      This patch adds the basic support for the clang driver to compile and link CUDA
      using the new offloading driver. This requires handling the CUDA offloading kind
      and embedding the generated files into the host. This will allow us to link
      OpenMP code with CUDA code in the linker wrapper. More support will be required
      to create functional CUDA / HIP binaries using this method.
      
      Depends on D120270 D120271 D120934
      
      Reviewed By: tra
      
      Differential Revision: https://reviews.llvm.org/D120272
      c5e5b543
    • Joseph Huber's avatar
      [Clang] Make enabling the new driver more generic · 4e2b5a66
      Joseph Huber authored
      In preparation for allowing other offloading kinds to use the new driver
      a new opt-in flag `-foffload-new-driver` is added. This is distinct from
      the existing `-fopenmp-new-driver` because OpenMP will soon use the new
      driver by default while the others should not.
      
      Reviewed By: yaxunl, tra
      
      Differential Revision: https://reviews.llvm.org/D123325
      4e2b5a66
    • Joseph Huber's avatar
      [OpenMP] Make clang argument handling for the new driver more generic · ca6bbe00
      Joseph Huber authored
      In preparation for accepting other offloading kinds with the new driver,
      this patch makes the way we handle offloading actions more generic. A
      new field to get the associated device action's toolchain is used rather
      than manually iterating a list. This makes building the arguments easier
      and makes sure that we doin't rely on any implicit ordering.
      
      Reviewed By: yaxunl
      
      Differential Revision: https://reviews.llvm.org/D123313
      ca6bbe00
    • Joseph Huber's avatar
      [OpenMP] Make generating offloading entries more generic · 643c9b22
      Joseph Huber authored
      This patch moves the logic for generating the offloading entries to the
      OpenMPIRBuilder. This makes it easier to re-use in other places, such as
      for OpenMP support in Flang or using the same method for generating
      offloading entires for other languages like Cuda.
      
      Reviewed By: tianshilei1992
      
      Differential Revision: https://reviews.llvm.org/D123460
      643c9b22
    • Nikita Popov's avatar
    • Nikita Popov's avatar
      [SelectionDAGBuilder] Don't create MGATHER/MSCATTER with Scale != ElemSize · 027c728f
      Nikita Popov authored
      This is an alternative to D124530. In getUniformBase() only create
      scales that match the gather/scatter element size. If targets also
      support other scales, then they can produce those scales in target
      DAG combines. This is what X86 already does (as long as the
      resulting scale would be 1, 2, 4 or 8).
      
      This essentially restores the pre-opaque-pointer state of things.
      
      Fixes https://github.com/llvm/llvm-project/issues/55021.
      
      Differential Revision: https://reviews.llvm.org/D124605
      027c728f
    • Jean Perier's avatar
      [flang] Handle common block with different sizes in same file · 2c8cb9ac
      Jean Perier authored
      Semantics is not preventing a named common block to appear with
      different size in a same file (named common block should always have
      the same storage size (see Fortran 2018 8.10.2.5), but it is a common
      extension to accept different sizes).
      
      Lowering was not coping with this well, since it just use the first
      common block appearance, starting with BLOCK DATAs to define common
      blocks (this also was an issue with the blank common block, which can
      legally appear with different size in different scoping units).
      
      Semantics is also not preventing named common from being initialized
      outside of a BLOCK DATA, and lowering was dealing badly with this,
      since it only gave an initial value to common blocks Globals if the
      first common block appearance, starting with BLOCK DATAs had an initial
      value.
      
      Semantics is also allowing blank common to be initialized, while
      lowering was assuming this would never happen, and was never creating
      an initial value for it.
      
      Lastly, semantics was not complaining if a COMMON block was initialized
      in several scoping unit in a same file, while lowering can only generate
      one of these initial value.
      
      To fix this, add a structure to keep track of COMMON block properties
      (biggest size, and initial value if any) at the Program level. Once the
      size of a common block appearance is know, the common block appearance
      is checked against this information. It allows semantics to emit an error
      in case of multiple initialization in different scopes of a same common
      block, and to warn in case named common blocks appears with different
      sizes. Lastly, this allows lowering to use the Program level info about
      common blocks to emit the right GlobalOp for a Common Block, regardless
      of the COMMON Block appearances order: It emits a GlobalOp with the
      biggest size, whose lowest bytes are initialized with the initial value
      if any is given in a scope where the common block appears.
      
      Lowering is updated to go emit the common blocks before anything else so
      that the related GlobalOps are available when lowering the scopes where
      common block appear. It is also updated to not assume that blank common
      are never initialized.
      
      Differential Revision: https://reviews.llvm.org/D124622
      2c8cb9ac
    • Nikita Popov's avatar
      [InstCombine] Remove memset of undef value · 1881711f
      Nikita Popov authored
      This removes memset with undef char. We already do this for stores
      of undef value.
      
      This comes with the caveat that this optimization is not, strictly
      speaking, legal for undef values, because we might be overwriting
      a poison value. However, our entire load/store model currently still
      operates on undef values, so we need to support undef here as well
      for internal consistency.
      
      Once https://github.com/llvm/llvm-project/issues/52930 is resolved,
      these and related folds can be limited to poison -- I've added
      FIXMEs to that effect.
      
      Differential Revision: https://reviews.llvm.org/D124173
      1881711f
    • Ricky Zhou's avatar
      [LV] Rename CountRoundDown to VectorTripCount (NFC) · 24a133e1
      Ricky Zhou authored
      The name CountRoundDown is potentially misleading, as the number of
      iterations can be rounded up when folding the tail.
      
      Reviewed By: fhahn
      
      Differential Revision: https://reviews.llvm.org/D119681
      24a133e1
    • Nikita Popov's avatar
      [InstCombine] Fold logical and/or of range icmps with nowrap flags · 982cbed8
      Nikita Popov authored
      This is an edge-case where we don't convert to bitwise and/or based
      on implies poison reasoning, so explicitly try to perform the fold
      in logical form. The transform itself is poison-safe, as both icmps
      are based on the same value and any nowrap flags are discarded as
      part of the fold (https://alive2.llvm.org/ce/z/aCwC8b for the used
      example).
      982cbed8
    • Matthias Springer's avatar
      [mlir][linalg][transform] Add TileOp to transform dialect · 3c2a74a3
      Matthias Springer authored
      This commit adds a tiling op to the transform dialect as an external op.
      
      Differential Revision: https://reviews.llvm.org/D124661
      3c2a74a3
    • Florian Hahn's avatar
      [VPlan] Simplify & adjust code as suggested in D123005. · e66127e6
      Florian Hahn authored
      Improve code as suggested in D123005. Applied separately, because the
      comments where made a diff that has not been rebased to current main.
      e66127e6
    • David Spickett's avatar
      [lldb] Allow EXE or exe in toolchain-msvc.test · f8463da4
      David Spickett authored
      I suspect that one of link or cl is found by shutil.which
      and one isn't, hence the case difference. It doesn't really
      matter for what the test is looking for.
      f8463da4
    • NAKAMURA Takumi's avatar
      llvm/Support/Debug.h: Suppress warnings with -Asserts. [-Wunused-variable] · 2e6657b3
      NAKAMURA Takumi authored
      Re. setCurrentDebugTypes(X,N), the only user is llvm-ml.cpp (exc. DebugTests)
      since llvmorg-15-init-8355-g82ecf9a0.
      
      FIXME: X and N are evaluated regardless of NDEBUG.
      Could we avoid evaluating (but w/o warnings) with NDEBUG?
      2e6657b3