1. Oct 09, 2018
  2. Oct 08, 2018
    • Sanjay Patel's avatar
      [x86] simplify hadd tests; NFC · d48789c0
      Sanjay Patel authored
      The tests from PR39195 don't use 2 parameters. That's the
      root problem for the pattern matching in isHorizontalBinOp().
      
      llvm-svn: 343974
      d48789c0
    • Neil Henning's avatar
      [AMDGPU] Add an AMDGPU specific atomic optimizer. · 66416574
      Neil Henning authored
      This commit adds a new IR level pass to the AMDGPU backend to perform
      atomic optimizations. It works by:
      
      - Running through a function and finding atomicrmw add/sub or uses of
        the atomic buffer intrinsics for add/sub.
      - If all arguments except the value to be added/subtracted are uniform,
        record the value to be optimized.
      - Run through the atomic operations we can optimize and, depending on
        whether the value is uniform/divergent use wavefront wide operations
        (DPP in the divergent case) to calculate the total amount to be
        atomically added/subtracted.
      - Then let only a single lane of each wavefront perform the atomic
        operation, reducing the total number of atomic operations in flight.
      - Lastly we recombine the result from the single lane to each lane of
        the wavefront, and calculate our individual lanes offset into the
        final result.
      
      Differential Revision: https://reviews.llvm.org/D51969
      
      llvm-svn: 343973
      66416574
    • Sid Manning's avatar
      [ELF][HEXAGON] Add R_HEX_GOT_16_X support · 307c7901
      Sid Manning authored
      Differential Revision: https://reviews.llvm.org/D52909
      
      llvm-svn: 343972
      307c7901
    • Zachary Turner's avatar
      Don't use back-quotes in a run line. · affaff8b
      Zachary Turner authored
      This works on Windows, but seems to be breaking tests that
      use an external shell (e.g. bash) because backquote has special
      meaning.
      
      This particular argument wasn't crucial for the test, so I've
      just removed it.
      
      llvm-svn: 343971
      affaff8b
    • Xin Tong's avatar
      [ThinLTO] Keep non-prevailing (linkonce|weak)_odr symbols live · bfdad33b
      Xin Tong authored
      Summary:
      If we have a symbol with (linkonce|weak)_odr linkage, we do not want
      to dead strip it even it is not prevailing.
      
      IR level (linkonce|weak)_odr symbol can become non-prevailing when we mix
      ELF objects and IR objects where the (linkonce|weak)_odr symbol in the ELF
      object is prevailing and the ones in the IR objects are not. Stripping
      them will prevent us from doing optimizations with them.
      
      By not dead stripping them, We will convert these symbols to
      available_externally linkage as a result of non-prevailing and eventually
      dropping them after inlining.
      
      I modified cache-prevailing.ll to use linkonce linkage as it is
      testing whether cache prevailing bit is effective or not, not
      we should treat linkonce_odr alive or not
      
      Reviewers: tejohnson, pcc
      
      Subscribers: mehdi_amini, inglorion, eraman, steven_wu, dexonsmith, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52893
      
      llvm-svn: 343970
      bfdad33b
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Don't create BR instructions in outliner when BTI enabled · 367b4741
      Oliver Stannard authored
      When branch target identification is enabled, we can only do indirect
      tail-calls through x16 or x17. This means that the outliner can't
      transform a BLR instruction at the end of an outlined region into a BR.
      
      Differential revision: https://reviews.llvm.org/D52869
      
      llvm-svn: 343969
      367b4741
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Restrict indirect tail calls to use x16/17 only when using BTI · c922116a
      Oliver Stannard authored
      When branch target identification is enabled, all indirectly-callable
      functions start with a BTI C instruction. this instruction can only be
      the target of certain indirect branches (direct branches and
      fall-through are not affected):
      - A BLR instruction, in either a protected or unprotected page.
      - A BR instruction in a protected page, using x16 or x17.
      - A BR instruction in an unprotected page, using any register.
      
      Without BTI, we can use any non call-preserved register to hold the
      address for an indirect tail call. However, when BTI is enabled, then
      the code being compiled might be loaded into a BTI-protected page, where
      only x16 and x17 can be used for indirect tail calls.
      
      Legacy code withiout this restriction can still indirectly tail-call
      BTI-protected functions, because they will be loaded into an unprotected
      page, so any register is allowed.
      
      Differential revision: https://reviews.llvm.org/D52868
      
      llvm-svn: 343968
      c922116a
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Branch Target Identification code-generation pass · 250e5a5b
      Oliver Stannard authored
      The Branch Target Identification extension, introduced to AArch64 in
      Armv8.5-A, adds the BTI instruction, which is used to mark valid targets
      of indirect branches. When enabled, the processor will trap if an
      instruction in a protected page tries to perform an indirect branch to
      any instruction other than a BTI. The BTI instruction uses encodings
      which were NOPs in earlier versions of the architecture, so BTI-enabled
      code will still run on earlier hardware, just without the extra
      protection.
      
      There are 3 variants of the BTI instruction, which are valid targets for
      different kinds or branches:
      - BTI C can be targeted by call instructions, and is inteneded to be
        used at function entry points. These are the BLR instruction, as well
        as BR with x16 or x17. These BR instructions are allowed for use in
        PLT entries, and we can also use them to allow indirect tail-calls.
      - BTI J can be targeted by BR only, and is intended to be used by jump
        tables.
      - BTI JC acts ab both a BTI C and a BTI J instruction, and can be
        targeted by any BLR or BR instruction.
      
      Note that RET instructions are not restricted by branch target
      identification, the reason for this is that return addresses can be
      protected more effectively using return address signing. Direct branches
      and calls are also unaffected, as it is assumed that an attacker cannot
      modify executable pages (if they could, they wouldn't need to do a
      ROP/JOP attack).
      
      This patch adds a MachineFunctionPass which:
      - Adds a BTI C at the start of every function which could be indirectly
        called (either because it is address-taken, or externally visible so
        could be address-taken in another translation unit).
      - Adds a BTI J at the start of every basic block which could be
        indirectly branched to. This could be either done by a jump table, or
        by taking the address of the block (e.g. the using GCC label values
        extension).
      
      We only need to use BTI JC when a function is indirectly-callable, and
      takes the address of the entry block. I've not been able to trigger this
      from C or IR, but I've included a MIR test just in case.
      
      Using BTI C at function entries relies on the fact that no other code in
      BTI-protected pages uses indirect tail-calls, unless they use x16 or x17
      to hold the address. I'll add that code-generation restriction as a
      separate patch.
      
      Differential revision: https://reviews.llvm.org/D52867
      
      llvm-svn: 343967
      250e5a5b
    • Alexander Ivchenko's avatar
      [GlobalIsel][X86] Support G_UDIV/G_UREM/G_SREM · 1aedf203
      Alexander Ivchenko authored
      Support G_UDIV/G_UREM/G_SREM. The instruction selection
      code is taken from FastISel with only minor tweaks to adapt
      for GlobalISel.
      
      Differential Revision: https://reviews.llvm.org/D49781
      
      llvm-svn: 343966
      1aedf203
    • Sanjay Patel's avatar
      [x86] add 16 missed hadd patterns (PR39195); NFC · 60badd75
      Sanjay Patel authored
      llvm-svn: 343965
      60badd75
    • David Carlier's avatar
      · b07407e6
      David Carlier authored
      [Sanitizer] fix internal_sysctlbyname build for FreeBSD.
      
      llvm-svn: 343964
      b07407e6
    • Haojian Wu's avatar
      [clangd] Update the out-of-date yaml-symbol-file flag in clangd. · 162510f6
      Haojian Wu authored
      Summary:
      The flag is stale due to the recent changes of clangd indexer, this
      patch renames the flag to "index-file".
      
      Reviewers: sammccall
      
      Subscribers: ilya-biryukov, ioeric, MaskRay, jkorous, arphaman, kadircet, cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D52976
      
      llvm-svn: 343963
      162510f6
    • Neil Henning's avatar
      [IRBuilder] Fixup CreateIntrinsic to allow specifying Types to Mangle. · 57f5d0a8
      Neil Henning authored
      The IRBuilder CreateIntrinsic method wouldn't allow you to specify the
      types that you wanted the intrinsic to be mangled with. To fix this
      I've:
      
      - Added an ArrayRef<Type *> member to both CreateIntrinsic overloads.
      - Used that array to pass into the Intrinsic::getDeclaration call.
      - Added a CreateUnaryIntrinsic to replace the most common use of
        CreateIntrinsic where the type was auto-deduced from operand 0.
      - Added a bunch more unit tests to test Create*Intrinsic calls that
        weren't being tested (including the FMF flag that wasn't checked).
      
      This was suggested as part of the AMDGPU specific atomic optimizer
      review (https://reviews.llvm.org/D51969).
      
      Differential Revision: https://reviews.llvm.org/D52087
      
      llvm-svn: 343962
      57f5d0a8
    • Francis Visoiu Mistrih's avatar
      [AsmParser] Return an error in the case of empty symbol ref in an expression · 627d1469
      Francis Visoiu Mistrih authored
      The following instruction:
      
      > str q28, [x0, #1*6*4*@]
      
      contains a @ which is parsed as an empty symbol. The parser returns true
      but has no error, so the assembler continues by ignoring the
      instruction.
      
      Differential Revision: https://reviews.llvm.org/D52645
      
      llvm-svn: 343961
      627d1469
    • Peter Smith's avatar
      [ARM] Account for implicit IT when calculating inline asm size · 6f36cd4d
      Peter Smith authored
      When deciding if it is safe to optimize a conditional branch to a CBZ or
      CBNZ the offsets of the BasicBlocks from the start of the function are
      estimated. For inline assembly the generic getInlineAsmLength() function is
      used to get a worst case estimate of the inline assembly by multiplying the
      number of instructions by the max instruction size of 4 bytes. This
      unfortunately doesn't take into account the generation of Thumb implicit IT
      instructions. In edge cases such as when all the instructions in the block
      are 4-bytes in size and there is an implicit IT then the size is
      underestimated. This can cause an out of range CBZ or CBNZ to be generated.
      
      The patch takes a conservative approach and assumes that every instruction
      in the inline assembly block may have an implicit IT.
      
      Fixes pr31805
      
      Differential Revision: https://reviews.llvm.org/D52834
      
      llvm-svn: 343960
      6f36cd4d
    • Oliver Stannard's avatar
      [AArch64] Fix verifier error when outlining indirect calls · 9ecdac8e
      Oliver Stannard authored
      The MachineOutliner for AArch64 transforms indirect calls into indirect
      tail calls, replacing the call with the TCRETURNri pseudo-instruction.
      This pseudo lowers to a BR, but has the isCall and isReturn flags set.
      
      The problem is that TCRETURNri takes a tcGPR64 as the register argument,
      to prevent indiret tail-calls from using caller-saved registers. The
      indirect calls transformed by the outliner could use caller-saved
      registers. This is fine, because the outliner ensures that the register
      is available at all call sites. However, this causes a verifier failure
      when the register is not in tcGPR64. The fix is to add a new
      pseudo-instruction like TCRETURNri, but which accepts any GPR.
      
      Differential revision: https://reviews.llvm.org/D52829
      
      llvm-svn: 343959
      9ecdac8e
    • Alex Bradbury's avatar
      [RISCV] Update alu8.ll and alu16.ll test cases · 5af6c149
      Alex Bradbury authored
      The srli test in alu8.ll was a no-op, as it shifted by 8 bits. Fix this, and 
      also change the immediate in alu16.ll as shifted by something other than a 
      poewr of 8 is more interesting.
      
      llvm-svn: 343958
      5af6c149
    • Kristina Brooks's avatar
      [DebugInfo][PDB] Fix a signed/unsigned coversion warning · bcc86a95
      Kristina Brooks authored
      Fix the following warning when compiling with clang (caused by commit
      rL343951):
      
      GlobalsStream.cpp:61:33: warning: comparison of integers of different
      signs: 'int' and 'uint32_t'
      
      This also avoids double evaluation of `GlobalsTable.HashBuckets.size()`.
      
      llvm-svn: 343957
      bcc86a95
    • Ewan Crawford's avatar
      [InstCombine] Fix incongruous GEP type addrspace · fa120cbd
      Ewan Crawford authored
      Currently running the @insertelem_after_gep function below through the InstCombine pass with opt produces invalid IR.
      
      Input:
      ```
      define void @insertelem_after_gep(<16 x i32>* %t0) {
         %t1 = bitcast <16 x i32>* %t0 to [16 x i32]*
         %t2 = addrspacecast [16 x i32]* %t1 to [16 x i32] addrspace(3)*
         %t3 = getelementptr inbounds [16 x i32], [16 x i32] addrspace(3)* %t2, i64 0, i64 0
         %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
         call void @extern_vec_pointers_func(<16 x i32 addrspace(3)*> %t4)
         ret void
      }
      ```
      
      Output:
      
      ```
      define void @insertelem_after_gep(<16 x i32>* %t0) {
        %t3 = getelementptr inbounds <16 x i32>, <16 x i32>* %t0, i64 0, i64 0
        %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
        call void @my_extern_func(<16 x i32 addrspace(3)*> %t4)
        ret void
      }
      ```
      
      Which although causes no complaints when produced, isn't valid IR as the insertelement use of the %t3 GEP expects an address space.
      
      ```
      opt: /tmp/bad.ll:52:73: error: '%t3' defined with type 'i32*' but expected 'i32 addrspace(3)*'
        %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
      ```
      
      I've fixed this by adding an addrspacecast after the GEP in the InstCombine pass, and including a check for this type mismatch to the verifier.
      
      Reviewers: spatel, lebedev.ri
      Subscribers: llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52294
      
      llvm-svn: 343956
      fa120cbd
    • Alex Bradbury's avatar
      [SelectionDAGBuilder][NFC] Pass LHSTy to getShiftAmountTy rather than RHSTy · f27c67af
      Alex Bradbury authored
      r126518 introduced a a type parameter to the getShiftAmountTy target hook. It 
      produces the type of the shift (RHSTy), parameterised by the type of the value 
      being shifted (LHSTy). SelectionDAGBuilder::visitShift passed RHSTy rather 
      than LHSTy and this patch corrects this. The change is a no-op because in LLVM 
      IR the LHS and RHS types for a shift must be equal anyway.
      
      llvm-svn: 343955
      f27c67af
    • Max Kazantsev's avatar
      [LV] Do not create SCEVs on broken IR in emitTransformedIndex. PR39160 · b0736965
      Max Kazantsev authored
      At the point when we perform `emitTransformedIndex`, we have a broken IR (in
      particular, we have Phis for which not every incoming value is properly set). On
      such IR, it is illegal to create SCEV expressions, because their internal
      simplification process may try to prove some predicates and break when it
      stumbles across some broken IR.
      
      The only purpose of using SCEV in this particular place is attempt to simplify
      the generated code slightly. It seems that the result isn't worth it, because
      some trivial cases (like addition of zero and multiplication by 1) can be
      handled separately if needed, but more generally InstCombine is able to achieve
      the goals we want to achieve by using SCEV.
      
      This patch fixes a functional crash described in PR39160, and as side-effect it
      also generates a bit smarter code in some simple cases. It also may cause some
      optimality loss (i.e. we will now generate `mul` by power of `2` instead of
      shift etc), but there is nothing what InstCombine could not handle later. In
      case of dire need, we can support more trivial cases just in place.
      
      Note that this patch only fixes one particular case of the general problem that
      LV misuses SCEV, attempting to create SCEVs or prove predicates on invalid IR.
      The general solution, however, seems complex enough.
      
      Differential Revision: https://reviews.llvm.org/D52881
      Reviewed By: fhahn, hsaito
      
      llvm-svn: 343954
      b0736965
    • Zachary Turner's avatar
      Fix a -Wsign-compare warning. · ba73a914
      Zachary Turner authored
      llvm-svn: 343953
      ba73a914
    • Zachary Turner's avatar
      Fix a compilation failure on non-MSVC compilers. · 9f6ac4c2
      Zachary Turner authored
      llvm-svn: 343952
      9f6ac4c2
    • Zachary Turner's avatar
      [PDB] Add the ability to lookup global symbols by name. · 94926a6d
      Zachary Turner authored
      The Globals table is a hash table keyed on symbol name, so
      it's possible to lookup symbols by name in O(1) time.  Add
      a function to the globals stream to do this, and add an option
      to llvm-pdbutil to exercise this, then use it to write some
      tests to verify correctness.
      
      llvm-svn: 343951
      94926a6d
    • Craig Topper's avatar
      Revert r343948 "[LegalizeDAG] Make one of the ReplaceNode signatures take an... · 98dd9d68
      Craig Topper authored
      Revert r343948 "[LegalizeDAG] Make one of the ReplaceNode signatures take an ArrayRef instead a pointer to an array. Add assert on size of array. NFC"
      
      The assert is failing some asan tests on the bots.
      
      llvm-svn: 343950
      98dd9d68
    • Brian Gesiak's avatar
      [coro]Pass rvalue reference for named local variable to return_value · 0b568300
      Brian Gesiak authored
      Summary:
      Addressing https://bugs.llvm.org/show_bug.cgi?id=37265.
      
      Implements [class.copy]/33 of coroutines TS.
      
      When the criteria for elision of a copy/move operation are met, but not
      for an exception-declaration, and the object to be copied is designated by an
      lvalue, or when the expression in a return or co_return statement is a
      (possibly parenthesized) id-expression that names an object with automatic
      storage duration declared in the body or parameter-declaration-clause of the
      innermost enclosing function or lambda-expression, overload resolution to select
      the constructor for the copy or the return_value overload to call is first
      performed as if the object were designated by an rvalue.
      
      Patch by Tanoy Sinha!
      
      Reviewers: modocache, GorNishanov
      
      Reviewed By: modocache, GorNishanov
      
      Subscribers: cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D51741
      
      llvm-svn: 343949
      0b568300
    • Craig Topper's avatar
      [LegalizeDAG] Make one of the ReplaceNode signatures take an ArrayRef instead... · c058a687
      Craig Topper authored
      [LegalizeDAG] Make one of the ReplaceNode signatures take an ArrayRef instead a pointer to an array. Add assert on size of array. NFC
      
      llvm-svn: 343948
      c058a687
    • Craig Topper's avatar
      [LegalizeDAG] Move legalization of scatter and masked store from LegalizeVectorOps to LegalizeDAG. · cd38de8b
      Craig Topper authored
      This is where we legalize gather and masked load so this is consistent.
      
      Since these ops are always on vectors I've chosen to go with LegalizeDAG since that's what we do for other vector only ops like BUILD_VECTOR, VECTOR_SHUFFLE, etc. The ScalarizeMaskedMemIntrinsic pass should take care of scalarizing these before SelectionDAG so hopefully we don't need to worry about illegally typed scalar ops being emitted in the legalizing. If we did we would need to do this in LegalizeVectorOps so we could get the second type legalization that runs between LegalizeVectorOps and LegalizeDAG.
      
      llvm-svn: 343947
      cd38de8b
    • Fangrui Song's avatar
      [clangd] Migrate to LLVM STLExtras range API · 8380c9e9
      Fangrui Song authored
      llvm-svn: 343946
      8380c9e9
    • Sanjay Patel's avatar
      [DAGCombiner] allow undef elts in vector fadd matching · ecc8af61
      Sanjay Patel authored
      llvm-svn: 343945
      ecc8af61
    • Sanjay Patel's avatar
      [x86] add vector fadd with undef elts test; NFC · f956840d
      Sanjay Patel authored
      llvm-svn: 343944
      f956840d
    • Sanjay Patel's avatar
      [x86] remove redundant tests; NFC · 6c02c6a3
      Sanjay Patel authored
      The equivalent tests were added to the file with related folds in rL343941.
      
      llvm-svn: 343943
      6c02c6a3
    • Sanjay Patel's avatar
      [DAGCombiner] allow undefs when matching vector splats for fmul folds · ef76e279
      Sanjay Patel authored
      llvm-svn: 343942
      ef76e279
    • Sanjay Patel's avatar
      [x86] add vector fmul with undef elts tests; NFC · fcb1061c
      Sanjay Patel authored
      llvm-svn: 343941
      fcb1061c
  3. Oct 07, 2018