1. Oct 09, 2018
    • Rong Xu's avatar
      [X86] condition branches folding for three-way conditional codes · 67b1b328
      Rong Xu authored
      This patch implements a pass that optimizes condition branches on x86 by
      taking advantage of the three-way conditional code generated by compare
      instructions.
      
      Currently, it tries to hoisting EQ and NE conditional branch to a dominant
      conditional branch condition where the same EQ/NE conditional code is
      computed. An example:
      bb_0:
        cmp %0, 19
        jg bb_1
        jmp bb_2
      bb_1:
        cmp %0, 40
        jg bb_3
        jmp bb_4
      bb_4:
        cmp %0, 20
        je bb_5
        jmp bb_6
      Here we could combine the two compares in bb_0 and bb_4 and have the
      following code:
      
      bb_0:
        cmp %0, 20
        jg bb_1
        jl bb_2
        jmp bb_5
      bb_1:
        cmp %0, 40
        jg bb_3
        jmp bb_6
      
      For the case of %0 == 20 (bb_5), we eliminate two jumps, and the control height
      for bb_6 is also reduced. bb_4 is gone after the optimization.
      
      This optimization is motivated by the branch pattern generated by the switch
      lowering: we always have pivot-1 compare for the inner nodes and we do a pivot
      compare again the leaf (like above pattern).
      
      This pass currently is enabled on Intel's Sandybridge and later arches. Some
      reviewers pointed out that on some arches (like AMD Jaguar), this pass may
      increase branch density to the point where it hurts the performance of the
      branch predictor.
      
      Differential Revision: https://reviews.llvm.org/D46662
      
      llvm-svn: 343993
      67b1b328
    • Scott Linder's avatar
      [AMDGPU] Legalize VGPR Rsrc operands for MUBUF instructions · 823549a6
      Scott Linder authored
      Emit a waterfall loop in the general case for a potentially-divergent Rsrc
      operand. When practical, avoid this by using Addr64 instructions.
      
      Recommits r341413 with changes to update the MachineDominatorTree when present.
      
      Differential Revision: https://reviews.llvm.org/D51742
      
      llvm-svn: 343992
      823549a6
    • Simon Pilgrim's avatar
      [X86][AVX2] Enable ZERO_EXTEND_VECTOR_INREG lowering of 256-bit vectors · 6fc8d055
      Simon Pilgrim authored
      Some necessary yak shaving before lowering *_EXTEND_VECTOR_INREG 256-bit vectors on AVX1 targets as suggested by D52964.
      
      Differential Revision: https://reviews.llvm.org/D52970
      
      llvm-svn: 343991
      6fc8d055
    • Charles Davis's avatar
      [CMake] Link to compiler-rt if LIBUNWIND_USE_COMPILER_RT is ON. · 869e5d1f
      Charles Davis authored
      Summary:
      If `-nodefaultlibs` is given, we weren't actually linking to it. This
      was true irrespective of passing `-rtlib=compiler-rt` (see previous
      patch). Now we explicitly link it to handle that case.
      
      I wonder if we should be linking these libraries only if we're using
      `-nodefaultlibs`...
      
      Reviewers: beanz
      
      Subscribers: dberris, mgorny, christof, chrib, cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D51657
      
      llvm-svn: 343990
      869e5d1f
    • Sanjay Patel's avatar
      [x86] make horizontal binop matching clearer; NFCI · 43bf9917
      Sanjay Patel authored
      The instructions are complicated, so this code will
      probably never be very obvious, but hopefully this
      makes it better. 
      
      As shown in PR39195:
      https://bugs.llvm.org/show_bug.cgi?id=39195
      ...we need to improve the matching to not miss cases
      where we're h-opping on 1 source vector, and that
      should be a small patch after this rearranging.
      
      llvm-svn: 343989
      43bf9917
    • Kamil Rytarowski's avatar
      Remove remnant code of using indirect syscall on NetBSD · 88e545ec
      Kamil Rytarowski authored
      Summary:
      The NetBSD version of internal routines no longer call
      the indirect syscall interfaces, as these functions were
      switched to lib calls.
      
      Remove the remnant code complication that is no
      longer needed after this change. Remove the variations
      of internal_syscall, as they were NetBSD specific.
      
      No functional change intended.
      
      Reviewers: vitalybuka, joerg, javed.absar
      
      Reviewed By: vitalybuka
      
      Subscribers: kubamracek, fedor.sergeev, llvm-commits, #sanitizers
      
      Tags: #sanitizers
      
      Differential Revision: https://reviews.llvm.org/D52955
      
      llvm-svn: 343988
      88e545ec
    • Kamil Rytarowski's avatar
      Don't harcode -ldl test/sanitizer_common/TestCases · bfd14ca6
      Kamil Rytarowski authored
      Summary:
      The dl library does not exist on all system and in particular
      this breaks build on NetBSD. Make it conditional and
      enable only for Linux, following the approach from other
      test suites in the same repository.
      
      Reviewers: joerg, vitalybuka
      
      Reviewed By: vitalybuka
      
      Subscribers: kubamracek, llvm-commits, #sanitizers
      
      Tags: #sanitizers
      
      Differential Revision: https://reviews.llvm.org/D52994
      
      llvm-svn: 343987
      bfd14ca6
    • Robert Lougher's avatar
      [TailCallElim] Enable marking of calls with byval as tails · 0c93ea26
      Robert Lougher authored
      In r339636 the alias analysis rules were changed with regards to tail calls
      and byval arguments. Previously, tail calls were assumed not to alias
      allocas from the current frame. This has been updated, to not assume this
      for arguments with the byval attribute.
      
      This patch aligns TailCallElim with the new rule. Tail marking can now be
      more aggressive and mark more calls as tails, e.g.:
      
      define void @test() {
        %f = alloca %struct.foo
        call void @bar(%struct.foo* byval %f)
        ret void
      }
      
      define void @test2(%struct.foo* byval %f) {
        call void @bar(%struct.foo* byval %f)
        ret void
      }
      
      define void @test3(%struct.foo* byval %f) {
        %agg.tmp = alloca %struct.foo
        %0 = bitcast %struct.foo* %agg.tmp to i8*
        %1 = bitcast %struct.foo* %f to i8*
        call void @llvm.memcpy.p0i8.p0i8.i64(i8* %0, i8* %1, i64 40, i1 false)
        call void @bar(%struct.foo* byval %agg.tmp)
        ret void
      }
      
      The problematic case where a byval parameter is captured by a call is still
      handled correctly, and will not be marked as a tail (see PR7272).
      
      llvm-svn: 343986
      0c93ea26
    • Tom Stellard's avatar
      AMDGPU/GlobalISel: Select amdgcn.cvt.pkrtz to 64-bit instructions · 14d8807d
      Tom Stellard authored
      Summary: The 32-bit variants do not exist on VI+.
      
      Reviewers: arsenm
      
      Reviewed By: arsenm
      
      Subscribers: kzhuravl, jvesely, wdng, nhaehnle, yaxunl, rovka, kristof.beyls, dstuttard, tpr, t-tye, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52958
      
      llvm-svn: 343985
      14d8807d
    • Kristina Brooks's avatar
      Fix incorrect Twine usage in CFGPrinter · 4f197cd7
      Kristina Brooks authored
      CFGPrinter (-view-cfg, -dot-cfg) invokes an undefined behaviour (dangling
      pointer to rvalue) on IR files with branch weights. This patch fixes the
      problem caused by Twine initialization and string conversion split into
      two statements.
      
      This change fixes the bug 37019. A similar patch to this problem was
      provided in the llvmlite project
      
      Patch by mcopik (Marcin Copik).
      
      Differential Revision: https://reviews.llvm.org/D52933
      
      llvm-svn: 343984
      4f197cd7
    • Rui Ueyama's avatar
      Fix a broken buildbot. · 9b5a495d
      Rui Ueyama authored
      llvm-svn: 343983
      9b5a495d
    • Eric Liu's avatar
      [clang-move] Dump whether a declaration is templated. · bce181c6
      Eric Liu authored
      llvm-svn: 343982
      bce181c6
    • Kamil Rytarowski's avatar
      Disable TestCases/pthread_mutexattr_get on NetBSD · 0fbf3e99
      Kamil Rytarowski authored
      The pshared feature is unsupported on NetBSD as of today.
      
      llvm-svn: 343981
      0fbf3e99
    • Kamil Rytarowski's avatar
      Fix Posix/devname_r for NetBSD · 73214e31
      Kamil Rytarowski authored
      NetBSD returns a different type as a return value of
      devname_r(3) than FreeBSD and Darwin (int vs char*).
      
      This implies that checking for successful completion of this
      function has to be handled differently.
      
      This test used to work well, but was switched to fix Darwin,
      which broke NetBSD.
      
      Add a dedicated ifdef for NetBSD and make it functional again
      for this OS.
      
      llvm-svn: 343980
      73214e31
    • Rui Ueyama's avatar
      Avoid unnecessary buffer allocation and memcpy for compressed sections. · e28c1464
      Rui Ueyama authored
      Previously, we uncompress all compressed sections before doing anything.
      That works, and that is conceptually simple, but that could results in
      a waste of CPU time and memory if uncompressed sections are then
      discarded or just copied to the output buffer.
      
      In particular, if .debug_gnu_pub{names,types} are compressed and if no
      -gdb-index option is given, we wasted CPU and memory because we
      uncompress them into newly allocated bufers and then memcpy the buffers
      to the output buffer. That temporary buffer was redundant.
      
      This patch changes how to uncompress sections. Now, compressed sections
      are uncompressed lazily. To do that, `Data` member of `InputSectionBase`
      is now hidden from outside, and `data()` accessor automatically expands
      an compressed buffer if necessary.
      
      If no one calls `data()`, then `writeTo()` directly uncompresses
      compressed data into the output buffer. That eliminates the redundant
      memory allocation and redundant memcpy.
      
      This patch significantly reduces memory consumption (20 GiB max RSS to
      15 Gib) for an executable whose .debug_gnu_pub{names,types} are in total
      5 GiB in an uncompressed form.
      
      Differential Revision: https://reviews.llvm.org/D52917
      
      llvm-svn: 343979
      e28c1464
    • Nicolai Haehnle's avatar
      AMDGPU: Future-proof {raw,struct}.buffer.atomic intrinsics · ea36cd59
      Nicolai Haehnle authored
      Summary:
      The ISA is really supposed to support 64-bit atomics as well,
      so the data type should be an overload.
      
      Mesa doesn't use these atomics yet, in fact I noticed this
      issue while trying to use the atomics from Mesa.
      
      Change-Id: I77f58317a085a0d3eb933cc7e99308c48a19f83e
      
      Reviewers: tpr
      
      Subscribers: kzhuravl, jvesely, wdng, yaxunl, dstuttard, t-tye, jfb, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52291
      
      llvm-svn: 343978
      ea36cd59
    • Nicolai Haehnle's avatar
      TableGen/CodeGenDAGPatterns: addPredicateFn only once · 46c91fd2
      Nicolai Haehnle authored
      Summary:
      The predicate function is added in InlinePatternFragments, no need to
      do it here. As a result, all uses of addPredicateFn are located in
      InlinePatternFragments.
      
      Test confirmed that there are no changes to generated files when
      building all (non-experimental) targets.
      
      Change-Id: I720e42e045ca596eb0aa339fb61adf6fe71034d5
      
      Reviewers: arsenm, rampitec, RKSimon, craig.topper, hfinkel, uweigand
      
      Subscribers: wdng, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D51993
      
      llvm-svn: 343977
      46c91fd2
    • Xin Tong's avatar
      Fix test case for @r343970 · 8dd92482
      Xin Tong authored
      op2 for weakodr symbols is 101 from bcanalyzer.
      
      llvm-svn: 343976
      8dd92482
    • Sanjay Patel's avatar
      [x86] add hadd test with no undefs, remove duplicate tests; NFC · 8459465a
      Sanjay Patel authored
      llvm-svn: 343975
      8459465a
  2. Oct 08, 2018
    • Sanjay Patel's avatar
      [x86] simplify hadd tests; NFC · d48789c0
      Sanjay Patel authored
      The tests from PR39195 don't use 2 parameters. That's the
      root problem for the pattern matching in isHorizontalBinOp().
      
      llvm-svn: 343974
      d48789c0
    • Neil Henning's avatar
      [AMDGPU] Add an AMDGPU specific atomic optimizer. · 66416574
      Neil Henning authored
      This commit adds a new IR level pass to the AMDGPU backend to perform
      atomic optimizations. It works by:
      
      - Running through a function and finding atomicrmw add/sub or uses of
        the atomic buffer intrinsics for add/sub.
      - If all arguments except the value to be added/subtracted are uniform,
        record the value to be optimized.
      - Run through the atomic operations we can optimize and, depending on
        whether the value is uniform/divergent use wavefront wide operations
        (DPP in the divergent case) to calculate the total amount to be
        atomically added/subtracted.
      - Then let only a single lane of each wavefront perform the atomic
        operation, reducing the total number of atomic operations in flight.
      - Lastly we recombine the result from the single lane to each lane of
        the wavefront, and calculate our individual lanes offset into the
        final result.
      
      Differential Revision: https://reviews.llvm.org/D51969
      
      llvm-svn: 343973
      66416574
    • Sid Manning's avatar
      [ELF][HEXAGON] Add R_HEX_GOT_16_X support · 307c7901
      Sid Manning authored
      Differential Revision: https://reviews.llvm.org/D52909
      
      llvm-svn: 343972
      307c7901
    • Zachary Turner's avatar
      Don't use back-quotes in a run line. · affaff8b
      Zachary Turner authored
      This works on Windows, but seems to be breaking tests that
      use an external shell (e.g. bash) because backquote has special
      meaning.
      
      This particular argument wasn't crucial for the test, so I've
      just removed it.
      
      llvm-svn: 343971
      affaff8b
    • Xin Tong's avatar
      [ThinLTO] Keep non-prevailing (linkonce|weak)_odr symbols live · bfdad33b
      Xin Tong authored
      Summary:
      If we have a symbol with (linkonce|weak)_odr linkage, we do not want
      to dead strip it even it is not prevailing.
      
      IR level (linkonce|weak)_odr symbol can become non-prevailing when we mix
      ELF objects and IR objects where the (linkonce|weak)_odr symbol in the ELF
      object is prevailing and the ones in the IR objects are not. Stripping
      them will prevent us from doing optimizations with them.
      
      By not dead stripping them, We will convert these symbols to
      available_externally linkage as a result of non-prevailing and eventually
      dropping them after inlining.
      
      I modified cache-prevailing.ll to use linkonce linkage as it is
      testing whether cache prevailing bit is effective or not, not
      we should treat linkonce_odr alive or not
      
      Reviewers: tejohnson, pcc
      
      Subscribers: mehdi_amini, inglorion, eraman, steven_wu, dexonsmith, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52893
      
      llvm-svn: 343970
      bfdad33b
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Don't create BR instructions in outliner when BTI enabled · 367b4741
      Oliver Stannard authored
      When branch target identification is enabled, we can only do indirect
      tail-calls through x16 or x17. This means that the outliner can't
      transform a BLR instruction at the end of an outlined region into a BR.
      
      Differential revision: https://reviews.llvm.org/D52869
      
      llvm-svn: 343969
      367b4741
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Restrict indirect tail calls to use x16/17 only when using BTI · c922116a
      Oliver Stannard authored
      When branch target identification is enabled, all indirectly-callable
      functions start with a BTI C instruction. this instruction can only be
      the target of certain indirect branches (direct branches and
      fall-through are not affected):
      - A BLR instruction, in either a protected or unprotected page.
      - A BR instruction in a protected page, using x16 or x17.
      - A BR instruction in an unprotected page, using any register.
      
      Without BTI, we can use any non call-preserved register to hold the
      address for an indirect tail call. However, when BTI is enabled, then
      the code being compiled might be loaded into a BTI-protected page, where
      only x16 and x17 can be used for indirect tail calls.
      
      Legacy code withiout this restriction can still indirectly tail-call
      BTI-protected functions, because they will be loaded into an unprotected
      page, so any register is allowed.
      
      Differential revision: https://reviews.llvm.org/D52868
      
      llvm-svn: 343968
      c922116a
    • Oliver Stannard's avatar
      [AArch64][v8.5A] Branch Target Identification code-generation pass · 250e5a5b
      Oliver Stannard authored
      The Branch Target Identification extension, introduced to AArch64 in
      Armv8.5-A, adds the BTI instruction, which is used to mark valid targets
      of indirect branches. When enabled, the processor will trap if an
      instruction in a protected page tries to perform an indirect branch to
      any instruction other than a BTI. The BTI instruction uses encodings
      which were NOPs in earlier versions of the architecture, so BTI-enabled
      code will still run on earlier hardware, just without the extra
      protection.
      
      There are 3 variants of the BTI instruction, which are valid targets for
      different kinds or branches:
      - BTI C can be targeted by call instructions, and is inteneded to be
        used at function entry points. These are the BLR instruction, as well
        as BR with x16 or x17. These BR instructions are allowed for use in
        PLT entries, and we can also use them to allow indirect tail-calls.
      - BTI J can be targeted by BR only, and is intended to be used by jump
        tables.
      - BTI JC acts ab both a BTI C and a BTI J instruction, and can be
        targeted by any BLR or BR instruction.
      
      Note that RET instructions are not restricted by branch target
      identification, the reason for this is that return addresses can be
      protected more effectively using return address signing. Direct branches
      and calls are also unaffected, as it is assumed that an attacker cannot
      modify executable pages (if they could, they wouldn't need to do a
      ROP/JOP attack).
      
      This patch adds a MachineFunctionPass which:
      - Adds a BTI C at the start of every function which could be indirectly
        called (either because it is address-taken, or externally visible so
        could be address-taken in another translation unit).
      - Adds a BTI J at the start of every basic block which could be
        indirectly branched to. This could be either done by a jump table, or
        by taking the address of the block (e.g. the using GCC label values
        extension).
      
      We only need to use BTI JC when a function is indirectly-callable, and
      takes the address of the entry block. I've not been able to trigger this
      from C or IR, but I've included a MIR test just in case.
      
      Using BTI C at function entries relies on the fact that no other code in
      BTI-protected pages uses indirect tail-calls, unless they use x16 or x17
      to hold the address. I'll add that code-generation restriction as a
      separate patch.
      
      Differential revision: https://reviews.llvm.org/D52867
      
      llvm-svn: 343967
      250e5a5b
    • Alexander Ivchenko's avatar
      [GlobalIsel][X86] Support G_UDIV/G_UREM/G_SREM · 1aedf203
      Alexander Ivchenko authored
      Support G_UDIV/G_UREM/G_SREM. The instruction selection
      code is taken from FastISel with only minor tweaks to adapt
      for GlobalISel.
      
      Differential Revision: https://reviews.llvm.org/D49781
      
      llvm-svn: 343966
      1aedf203
    • Sanjay Patel's avatar
      [x86] add 16 missed hadd patterns (PR39195); NFC · 60badd75
      Sanjay Patel authored
      llvm-svn: 343965
      60badd75
    • David Carlier's avatar
      · b07407e6
      David Carlier authored
      [Sanitizer] fix internal_sysctlbyname build for FreeBSD.
      
      llvm-svn: 343964
      b07407e6
    • Haojian Wu's avatar
      [clangd] Update the out-of-date yaml-symbol-file flag in clangd. · 162510f6
      Haojian Wu authored
      Summary:
      The flag is stale due to the recent changes of clangd indexer, this
      patch renames the flag to "index-file".
      
      Reviewers: sammccall
      
      Subscribers: ilya-biryukov, ioeric, MaskRay, jkorous, arphaman, kadircet, cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D52976
      
      llvm-svn: 343963
      162510f6
    • Neil Henning's avatar
      [IRBuilder] Fixup CreateIntrinsic to allow specifying Types to Mangle. · 57f5d0a8
      Neil Henning authored
      The IRBuilder CreateIntrinsic method wouldn't allow you to specify the
      types that you wanted the intrinsic to be mangled with. To fix this
      I've:
      
      - Added an ArrayRef<Type *> member to both CreateIntrinsic overloads.
      - Used that array to pass into the Intrinsic::getDeclaration call.
      - Added a CreateUnaryIntrinsic to replace the most common use of
        CreateIntrinsic where the type was auto-deduced from operand 0.
      - Added a bunch more unit tests to test Create*Intrinsic calls that
        weren't being tested (including the FMF flag that wasn't checked).
      
      This was suggested as part of the AMDGPU specific atomic optimizer
      review (https://reviews.llvm.org/D51969).
      
      Differential Revision: https://reviews.llvm.org/D52087
      
      llvm-svn: 343962
      57f5d0a8
    • Francis Visoiu Mistrih's avatar
      [AsmParser] Return an error in the case of empty symbol ref in an expression · 627d1469
      Francis Visoiu Mistrih authored
      The following instruction:
      
      > str q28, [x0, #1*6*4*@]
      
      contains a @ which is parsed as an empty symbol. The parser returns true
      but has no error, so the assembler continues by ignoring the
      instruction.
      
      Differential Revision: https://reviews.llvm.org/D52645
      
      llvm-svn: 343961
      627d1469
    • Peter Smith's avatar
      [ARM] Account for implicit IT when calculating inline asm size · 6f36cd4d
      Peter Smith authored
      When deciding if it is safe to optimize a conditional branch to a CBZ or
      CBNZ the offsets of the BasicBlocks from the start of the function are
      estimated. For inline assembly the generic getInlineAsmLength() function is
      used to get a worst case estimate of the inline assembly by multiplying the
      number of instructions by the max instruction size of 4 bytes. This
      unfortunately doesn't take into account the generation of Thumb implicit IT
      instructions. In edge cases such as when all the instructions in the block
      are 4-bytes in size and there is an implicit IT then the size is
      underestimated. This can cause an out of range CBZ or CBNZ to be generated.
      
      The patch takes a conservative approach and assumes that every instruction
      in the inline assembly block may have an implicit IT.
      
      Fixes pr31805
      
      Differential Revision: https://reviews.llvm.org/D52834
      
      llvm-svn: 343960
      6f36cd4d
    • Oliver Stannard's avatar
      [AArch64] Fix verifier error when outlining indirect calls · 9ecdac8e
      Oliver Stannard authored
      The MachineOutliner for AArch64 transforms indirect calls into indirect
      tail calls, replacing the call with the TCRETURNri pseudo-instruction.
      This pseudo lowers to a BR, but has the isCall and isReturn flags set.
      
      The problem is that TCRETURNri takes a tcGPR64 as the register argument,
      to prevent indiret tail-calls from using caller-saved registers. The
      indirect calls transformed by the outliner could use caller-saved
      registers. This is fine, because the outliner ensures that the register
      is available at all call sites. However, this causes a verifier failure
      when the register is not in tcGPR64. The fix is to add a new
      pseudo-instruction like TCRETURNri, but which accepts any GPR.
      
      Differential revision: https://reviews.llvm.org/D52829
      
      llvm-svn: 343959
      9ecdac8e
    • Alex Bradbury's avatar
      [RISCV] Update alu8.ll and alu16.ll test cases · 5af6c149
      Alex Bradbury authored
      The srli test in alu8.ll was a no-op, as it shifted by 8 bits. Fix this, and 
      also change the immediate in alu16.ll as shifted by something other than a 
      poewr of 8 is more interesting.
      
      llvm-svn: 343958
      5af6c149
    • Kristina Brooks's avatar
      [DebugInfo][PDB] Fix a signed/unsigned coversion warning · bcc86a95
      Kristina Brooks authored
      Fix the following warning when compiling with clang (caused by commit
      rL343951):
      
      GlobalsStream.cpp:61:33: warning: comparison of integers of different
      signs: 'int' and 'uint32_t'
      
      This also avoids double evaluation of `GlobalsTable.HashBuckets.size()`.
      
      llvm-svn: 343957
      bcc86a95
    • Ewan Crawford's avatar
      [InstCombine] Fix incongruous GEP type addrspace · fa120cbd
      Ewan Crawford authored
      Currently running the @insertelem_after_gep function below through the InstCombine pass with opt produces invalid IR.
      
      Input:
      ```
      define void @insertelem_after_gep(<16 x i32>* %t0) {
         %t1 = bitcast <16 x i32>* %t0 to [16 x i32]*
         %t2 = addrspacecast [16 x i32]* %t1 to [16 x i32] addrspace(3)*
         %t3 = getelementptr inbounds [16 x i32], [16 x i32] addrspace(3)* %t2, i64 0, i64 0
         %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
         call void @extern_vec_pointers_func(<16 x i32 addrspace(3)*> %t4)
         ret void
      }
      ```
      
      Output:
      
      ```
      define void @insertelem_after_gep(<16 x i32>* %t0) {
        %t3 = getelementptr inbounds <16 x i32>, <16 x i32>* %t0, i64 0, i64 0
        %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
        call void @my_extern_func(<16 x i32 addrspace(3)*> %t4)
        ret void
      }
      ```
      
      Which although causes no complaints when produced, isn't valid IR as the insertelement use of the %t3 GEP expects an address space.
      
      ```
      opt: /tmp/bad.ll:52:73: error: '%t3' defined with type 'i32*' but expected 'i32 addrspace(3)*'
        %t4 = insertelement <16 x i32 addrspace(3)*> undef, i32 addrspace(3)* %t3, i32 0
      ```
      
      I've fixed this by adding an addrspacecast after the GEP in the InstCombine pass, and including a check for this type mismatch to the verifier.
      
      Reviewers: spatel, lebedev.ri
      Subscribers: llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D52294
      
      llvm-svn: 343956
      fa120cbd
    • Alex Bradbury's avatar
      [SelectionDAGBuilder][NFC] Pass LHSTy to getShiftAmountTy rather than RHSTy · f27c67af
      Alex Bradbury authored
      r126518 introduced a a type parameter to the getShiftAmountTy target hook. It 
      produces the type of the shift (RHSTy), parameterised by the type of the value 
      being shifted (LHSTy). SelectionDAGBuilder::visitShift passed RHSTy rather 
      than LHSTy and this patch corrects this. The change is a no-op because in LLVM 
      IR the LHS and RHS types for a shift must be equal anyway.
      
      llvm-svn: 343955
      f27c67af
    • Max Kazantsev's avatar
      [LV] Do not create SCEVs on broken IR in emitTransformedIndex. PR39160 · b0736965
      Max Kazantsev authored
      At the point when we perform `emitTransformedIndex`, we have a broken IR (in
      particular, we have Phis for which not every incoming value is properly set). On
      such IR, it is illegal to create SCEV expressions, because their internal
      simplification process may try to prove some predicates and break when it
      stumbles across some broken IR.
      
      The only purpose of using SCEV in this particular place is attempt to simplify
      the generated code slightly. It seems that the result isn't worth it, because
      some trivial cases (like addition of zero and multiplication by 1) can be
      handled separately if needed, but more generally InstCombine is able to achieve
      the goals we want to achieve by using SCEV.
      
      This patch fixes a functional crash described in PR39160, and as side-effect it
      also generates a bit smarter code in some simple cases. It also may cause some
      optimality loss (i.e. we will now generate `mul` by power of `2` instead of
      shift etc), but there is nothing what InstCombine could not handle later. In
      case of dire need, we can support more trivial cases just in place.
      
      Note that this patch only fixes one particular case of the general problem that
      LV misuses SCEV, attempting to create SCEVs or prove predicates on invalid IR.
      The general solution, however, seems complex enough.
      
      Differential Revision: https://reviews.llvm.org/D52881
      Reviewed By: fhahn, hsaito
      
      llvm-svn: 343954
      b0736965