1. Jan 23, 2018
    • Sjoerd Meijer's avatar
      [ARM] Pass _Float16 as int or float · ca8f4e74
      Sjoerd Meijer authored
      Pass and return _Float16 as if it were an int or float for ARM, but with the
      top 16 bits unspecified, similarly like we already do for __fp16.
      
      We will implement proper half-precision function argument lowering in the ARM
      backend soon, but want to use this workaround in the mean time.
      
      Differential Revision: https://reviews.llvm.org/D42318
      
      llvm-svn: 323185
      ca8f4e74
    • Stefan Maksimovic's avatar
      [mips] Properly select abs and sqrt instructions · 98749e02
      Stefan Maksimovic authored
      - Alter abs for micromips to have both AFGR64 and FGR64
        variants, same as sqrt
      - Remove sqrt and abs from MicroMips32r6InstrInfo.td,
        use micromips FGR64 variants
      - Restrict non-micromips abs/sqrt with NotInMicroMips
        predicate
      
      Differential revision: https://reviews.llvm.org/D41439
      
      llvm-svn: 323184
      98749e02
    • Ashutosh Nema's avatar
      This change add's optimization remark in LoopVersioning LICM pass. · 007b425b
      Ashutosh Nema authored
      Summary:
      This patch is adding remark messages to the LoopVersioning LICM pass, 
      which will be useful for optimization remark emitter (ORE) infrastructure.
      
      Patch by: Deepak Porwal
      
      Reviewers: anemet, ashutosh.nema, eastig
      
      Subscribers: eastig, vivekvpandya, fhahn, llvm-commits
      llvm-svn: 323183
      007b425b
    • Anton Bikineev's avatar
      [InstSimplify] (X << Y) % X -> 0 · 82f61151
      Anton Bikineev authored
      llvm-svn: 323182
      82f61151
    • Raphael Isemann's avatar
      Prevent unaligned memory read in parseMinidumpString · acc48f3e
      Raphael Isemann authored
      Summary:
      It's possible to hit an unaligned memory read when reading `source_length` as the `data` array is only aligned with 2 bytes (it's actually a UTF16 array). This patch memcpy's `source_length` into a local variable to prevent this:
      
      ```
      MinidumpTypes.cpp:49:23: runtime error: load of misaligned address 0x7f0f4792692a for type 'const uint32_t' (aka 'const unsigned int'), which requires 4 byte alignment
      ```
      
      Reviewers: dvlahovski, zturner, davide
      
      Reviewed By: davide
      
      Subscribers: davide, lldb-commits
      
      Differential Revision: https://reviews.llvm.org/D42348
      
      llvm-svn: 323181
      acc48f3e
    • Jonas Hahnfeld's avatar
      Fix name of 'macOS' and add asteriks to brands, NFC. · e5762030
      Jonas Hahnfeld authored
      llvm-svn: 323180
      e5762030
    • Justin Bogner's avatar
      update_mir_test_checks: Improve the check for LLVM IR in MIR files · eaae305a
      Justin Bogner authored
      The LLVM IR section of a MIR document can start with "--- |" rather
      than just "---", because "|" is a sigil for a freeform document in
      YAML. We need to handle this so that we don't try to add check lines
      to the LLVM IR functions in a MIR file.
      
      llvm-svn: 323178
      eaae305a
    • Hiroshi Inoue's avatar
      [NFC] fix trivial typos in comments · cde18b4b
      Hiroshi Inoue authored
      "the the" -> "the"
      
      llvm-svn: 323177
      cde18b4b
    • Hiroshi Inoue's avatar
      [NFC] fix trivial typos in comments · 4cf7b885
      Hiroshi Inoue authored
      "the the" -> "the"
      
      llvm-svn: 323176
      4cf7b885
    • Craig Topper's avatar
      [X86] Don't reorder (srl (and X, C1), C2) if (and X, C1) can be matched as a movzx · c92edd99
      Craig Topper authored
      Summary:
      If we can match as a zero extend there's no need to flip the order to get an encoding benefit. As movzx is 3 bytes with independent source/dest registers. The shortest 'and' we could make is also 3 bytes unless we get lucky in the register allocator and its on AL/AX/EAX which have a 2 byte encoding.
      
      This patch was more impressive before r322957 went in. It removed some of the same Ands that got deleted by that patch.
      
      Reviewers: spatel, RKSimon
      
      Reviewed By: spatel
      
      Subscribers: llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D42313
      
      llvm-svn: 323175
      c92edd99
    • Craig Topper's avatar
      [X86] Remove 'NOREX' comment from the printing of _NOREX instructions. · e5aea259
      Craig Topper authored
      Some of the NOREX instructions are used in 32-bit mode making this printing confusing. It also doesn't provide a lot of value since you can see the h-register being used by the instruction.
      
      llvm-svn: 323174
      e5aea259
    • Craig Topper's avatar
      [X86] Various vXi1 insertion improvements. · 26a701f2
      Craig Topper authored
      Add missing patterns for inserting v1i1 into a zero vector. Use insert_subvector to zero upper bits before inserting an element into a vXi1 vector. Replace kshift based isel pattern with insert_subvector based pattern now that code that caused the pattern has been fixed to emit insert_subvector.
      
      llvm-svn: 323173
      26a701f2
    • Rafael Espindola's avatar
      Use 4 as the alignment of .eh_frame_hdr. · 6b2b4502
      Rafael Espindola authored
      It includes 32 bit values and this matches both gold and bfd.
      
      llvm-svn: 323172
      6b2b4502
    • Don Hinton's avatar
      66ddb50e
    • Peter Collingbourne's avatar
      libcxx: Provide overloads for basic_filebuf::open() et al that take wchar_t* filenames on Windows. · 4801624e
      Peter Collingbourne authored
      This is an MSVC standard library extension. It seems like a reasonable
      enough extension to me because wchar_t* is the native format for
      filenames on that platform.
      
      Differential Revision: https://reviews.llvm.org/D42225
      
      llvm-svn: 323170
      4801624e
    • Peter Collingbourne's avatar
      libcxx: Move Windows threading support into a .cpp file. · ac15ae6d
      Peter Collingbourne authored
      This allows us to avoid polluting the namespace of users of <thread>
      with the definitions in windows.h.
      
      Differential Revision: https://reviews.llvm.org/D42214
      
      llvm-svn: 323169
      ac15ae6d
    • Sam Clegg's avatar
      [WebAssembly] Update to match llvm changes to TABLE relocations · ab604a98
      Sam Clegg authored
      TABLE relocations now store the function that is being refered
      to indirectly.
      
      See rL323165.
      
      Also extend the call-indirect.ll a little.
      
      Based on a patch by Nicholas Wilson!
      
      llvm-svn: 323168
      ab604a98
    • David Blaikie's avatar
      NewPM: Improve/fix GCOV - which needs to run early in the pass pipeline. · ac904d0e
      David Blaikie authored
      Using a new extension point in the new PM, register GCOV at the start of
      the pipeline rather than the end.
      
      llvm-svn: 323167
      ac904d0e
    • David Blaikie's avatar
      NewPM: Add an extension point for the start of the pipeline. · 0c64f5a2
      David Blaikie authored
      This applies to most pipelines except the LTO and ThinLTO backend
      actions - it is for use at the beginning of the overall pipeline.
      
      This extension point will be used to add the GCOV pass when enabled in
      Clang.
      
      llvm-svn: 323166
      0c64f5a2
    • Sam Clegg's avatar
      [WebAssembly] Store function index rather than table index in TABLE_INDEX relocations · 60ec3034
      Sam Clegg authored
      Relocations of type R_WEBASSEMBLY_TABLE_INDEX represent places
      where the table index for a given function is needed.  While the
      value stored in this location is a table index, the index in
      the relocation entry itself is a function index (the index of
      the function which is to be called indirectly).
      
      This is how is was spec'd originally but the LLVM implementation
      didn't do this.  This makes things a little simpler in the linker
      since the table in the input file can essentially be ignored that
      the output table can be created purely based on these relocations.
      
      Patch by Nicholas Wilson!
      
      Differential Revision: https://reviews.llvm.org/D42080
      
      llvm-svn: 323165
      60ec3034
    • Bob Haarman's avatar
      [COFF] don't replace import library if contents are unchanged · 4ce341ff
      Bob Haarman authored
      Summary:
      This detects when an import library is about to be overwritten with a
      newly built one with the same contents, and keeps the old library
      instead. The use case for this is to avoid needlessly rebuilding
      targets that depend on the import library in build systems that rely
      on timestamps to determine whether a target requires rebuilding.
      
      This feature was requested in PR35917.
      
      Reviewers: rnk, ruiu, zturner, pcc
      
      Reviewed By: ruiu
      
      Subscribers: llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D42326
      
      llvm-svn: 323164
      4ce341ff
    • Lang Hames's avatar
      [lldb] Fix some C++ virtual method call bugs in LLDB expression evaluation by · 48b32f4c
      Lang Hames authored
      building method override tables for CXXMethodDecls in 
      DWARFASTParserClang::CompleteTypeFromDWARF.
      
      C++ virtual method calls in LLDB expressions may fail if the override table for
      the method being called is not correct as IRGen will produce references to the
      wrong (or a missing) vtable entry.
      
      This patch does not fix calls to virtual methods with covariant return types as
      it mistakenly treats these as overloads, rather than overrides. This will be
      addressed in a future patch.
      
      Review: https://reviews.llvm.org/D41997
      
      Partially fixes <rdar://problem/14205774>
      
      llvm-svn: 323163
      48b32f4c
    • Alex Shlyapnikov's avatar
      Small fixes for detect_invalid_pointer_pairs. · ac8217de
      Alex Shlyapnikov authored
      Summary:
      One test-case uses a wrong operation (should be subtraction).
      Second test-case should declare a global variables before a tested one
      in order to guarantee we will find a red-zone.
      
      Reviewers: kcc, jakubjelinek, alekseyshl
      
      Reviewed By: alekseyshl
      
      Subscribers: kubamracek
      
      Differential Revision: https://reviews.llvm.org/D41481
      
      llvm-svn: 323162
      ac8217de
    • Rui Ueyama's avatar
      Revert r322595: Specify inline for isWhitespace in CommandLine.cpp · 322fcfee
      Rui Ueyama authored
      The original change was made based on a misunderstanding that
      -DCMAKE_BUILD_TYPE=RelWithDebugInfo would produce the same executable
      as -DCMAKE_BUILD_TYPE=Release modulo debug info. Turned out that's not
      true -- it at least disables some optimizations such as function inlining.
      
      llvm-svn: 323161
      322fcfee
    • Marshall Clow's avatar
      Update cxx2a status · d161ac9e
      Marshall Clow authored
      llvm-svn: 323160
      d161ac9e
    • Marshall Clow's avatar
      8da1a487
    • Julie Hockett's avatar
      Add hasTrailingReturn AST matcher · 239d25a1
      Julie Hockett authored
      Adds AST matcher for a FunctionDecl that has a trailing return type.
      
      Differential Revision: https://reviews.llvm.org/D42273
      
      llvm-svn: 323158
      239d25a1
    • Fangrui Song's avatar
      [ASTMatchers] [NFC] Fix code examples · 55942abb
      Fangrui Song authored
      Subscribers: klimek, cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D42213
      
      llvm-svn: 323157
      55942abb
    • Volodymyr Sapsai's avatar
      Reland "[CodeGen] Fix crash when a function taking transparent union is redeclared." · 17ebdb23
      Volodymyr Sapsai authored
      When a function taking transparent union is declared as taking one of
      union members earlier in the translation unit, clang would hit an
      "Invalid cast" assertion during EmitFunctionProlog. This case
      corresponds to function f1 in test/CodeGen/transparent-union-redecl.c.
      We decided to cast i32 to union because after merging function
      declarations function parameter type becomes int,
      CGFunctionInfo::ArgInfo type matches with ABIArgInfo type, so we decide
      it is a trivial case. But these types should also be castable to
      parameter declaration type which is not the case here.
      
      Now the fix is in converting from ABIArgInfo type to VarDecl type and using
      argument demotion when necessary.
      
      Additional tests in Sema/transparent-union.c capture current behavior and make
      sure there are no regressions.
      
      rdar://problem/34949329
      
      Reviewers: rjmccall, rafael
      
      Reviewed By: rjmccall
      
      Subscribers: aemerson, cfe-commits, kristof.beyls, ahatanak
      
      Differential Revision: https://reviews.llvm.org/D41311
      
      llvm-svn: 323156
      17ebdb23
    • Chandler Carruth's avatar
      Introduce the "retpoline" x86 mitigation technique for variant #2 of the... · c58f2166
      Chandler Carruth authored
      Introduce the "retpoline" x86 mitigation technique for variant #2 of the speculative execution vulnerabilities disclosed today, specifically identified by CVE-2017-5715, "Branch Target Injection", and is one of the two halves to Spectre..
      
      Summary:
      First, we need to explain the core of the vulnerability. Note that this
      is a very incomplete description, please see the Project Zero blog post
      for details:
      https://googleprojectzero.blogspot.com/2018/01/reading-privileged-memory-with-side.html
      
      The basis for branch target injection is to direct speculative execution
      of the processor to some "gadget" of executable code by poisoning the
      prediction of indirect branches with the address of that gadget. The
      gadget in turn contains an operation that provides a side channel for
      reading data. Most commonly, this will look like a load of secret data
      followed by a branch on the loaded value and then a load of some
      predictable cache line. The attacker then uses timing of the processors
      cache to determine which direction the branch took *in the speculative
      execution*, and in turn what one bit of the loaded value was. Due to the
      nature of these timing side channels and the branch predictor on Intel
      processors, this allows an attacker to leak data only accessible to
      a privileged domain (like the kernel) back into an unprivileged domain.
      
      The goal is simple: avoid generating code which contains an indirect
      branch that could have its prediction poisoned by an attacker. In many
      cases, the compiler can simply use directed conditional branches and
      a small search tree. LLVM already has support for lowering switches in
      this way and the first step of this patch is to disable jump-table
      lowering of switches and introduce a pass to rewrite explicit indirectbr
      sequences into a switch over integers.
      
      However, there is no fully general alternative to indirect calls. We
      introduce a new construct we call a "retpoline" to implement indirect
      calls in a non-speculatable way. It can be thought of loosely as
      a trampoline for indirect calls which uses the RET instruction on x86.
      Further, we arrange for a specific call->ret sequence which ensures the
      processor predicts the return to go to a controlled, known location. The
      retpoline then "smashes" the return address pushed onto the stack by the
      call with the desired target of the original indirect call. The result
      is a predicted return to the next instruction after a call (which can be
      used to trap speculative execution within an infinite loop) and an
      actual indirect branch to an arbitrary address.
      
      On 64-bit x86 ABIs, this is especially easily done in the compiler by
      using a guaranteed scratch register to pass the target into this device.
      For 32-bit ABIs there isn't a guaranteed scratch register and so several
      different retpoline variants are introduced to use a scratch register if
      one is available in the calling convention and to otherwise use direct
      stack push/pop sequences to pass the target address.
      
      This "retpoline" mitigation is fully described in the following blog
      post: https://support.google.com/faqs/answer/7625886
      
      We also support a target feature that disables emission of the retpoline
      thunk by the compiler to allow for custom thunks if users want them.
      These are particularly useful in environments like kernels that
      routinely do hot-patching on boot and want to hot-patch their thunk to
      different code sequences. They can write this custom thunk and use
      `-mretpoline-external-thunk` *in addition* to `-mretpoline`. In this
      case, on x86-64 thu thunk names must be:
      ```
        __llvm_external_retpoline_r11
      ```
      or on 32-bit:
      ```
        __llvm_external_retpoline_eax
        __llvm_external_retpoline_ecx
        __llvm_external_retpoline_edx
        __llvm_external_retpoline_push
      ```
      And the target of the retpoline is passed in the named register, or in
      the case of the `push` suffix on the top of the stack via a `pushl`
      instruction.
      
      There is one other important source of indirect branches in x86 ELF
      binaries: the PLT. These patches also include support for LLD to
      generate PLT entries that perform a retpoline-style indirection.
      
      The only other indirect branches remaining that we are aware of are from
      precompiled runtimes (such as crt0.o and similar). The ones we have
      found are not really attackable, and so we have not focused on them
      here, but eventually these runtimes should also be replicated for
      retpoline-ed configurations for completeness.
      
      For kernels or other freestanding or fully static executables, the
      compiler switch `-mretpoline` is sufficient to fully mitigate this
      particular attack. For dynamic executables, you must compile *all*
      libraries with `-mretpoline` and additionally link the dynamic
      executable and all shared libraries with LLD and pass `-z retpolineplt`
      (or use similar functionality from some other linker). We strongly
      recommend also using `-z now` as non-lazy binding allows the
      retpoline-mitigated PLT to be substantially smaller.
      
      When manually apply similar transformations to `-mretpoline` to the
      Linux kernel we observed very small performance hits to applications
      running typical workloads, and relatively minor hits (approximately 2%)
      even for extremely syscall-heavy applications. This is largely due to
      the small number of indirect branches that occur in performance
      sensitive paths of the kernel.
      
      When using these patches on statically linked applications, especially
      C++ applications, you should expect to see a much more dramatic
      performance hit. For microbenchmarks that are switch, indirect-, or
      virtual-call heavy we have seen overheads ranging from 10% to 50%.
      
      However, real-world workloads exhibit substantially lower performance
      impact. Notably, techniques such as PGO and ThinLTO dramatically reduce
      the impact of hot indirect calls (by speculatively promoting them to
      direct calls) and allow optimized search trees to be used to lower
      switches. If you need to deploy these techniques in C++ applications, we
      *strongly* recommend that you ensure all hot call targets are statically
      linked (avoiding PLT indirection) and use both PGO and ThinLTO. Well
      tuned servers using all of these techniques saw 5% - 10% overhead from
      the use of retpoline.
      
      We will add detailed documentation covering these components in
      subsequent patches, but wanted to make the core functionality available
      as soon as possible. Happy for more code review, but we'd really like to
      get these patches landed and backported ASAP for obvious reasons. We're
      planning to backport this to both 6.0 and 5.0 release streams and get
      a 5.0 release with just this cherry picked ASAP for distros and vendors.
      
      This patch is the work of a number of people over the past month: Eric, Reid,
      Rui, and myself. I'm mailing it out as a single commit due to the time
      sensitive nature of landing this and the need to backport it. Huge thanks to
      everyone who helped out here, and everyone at Intel who helped out in
      discussions about how to craft this. Also, credit goes to Paul Turner (at
      Google, but not an LLVM contributor) for much of the underlying retpoline
      design.
      
      Reviewers: echristo, rnk, ruiu, craig.topper, DavidKreitzer
      
      Subscribers: sanjoy, emaste, mcrosier, mgorny, mehdi_amini, hiraditya, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D41723
      
      llvm-svn: 323155
      c58f2166
    • Sam Clegg's avatar
      [WebAssembly] Remove --emit-relocs · ff2b1221
      Sam Clegg authored
      This was added to mimic ELF, but maintaining it has cost
      and we currently don't have any use for it outside of the
      test code.
      
      Differential Revision: https://reviews.llvm.org/D42324
      
      llvm-svn: 323154
      ff2b1221
    • Mark Searles's avatar
      [AMDGPU] SI Load Store Optimizer: When merging with offset, use V_ADD_{I|U}32_e64 · 7687d420
      Mark Searles authored
      - Change inserted add ( V_ADD_{I|U}32_e32 ) to _e64 version ( V_ADD_{I|U}32_e64 ) so that the add uses a vreg for the carry; this prevents inserted v_add from killing VCC; the _e64 version doesn't accept a literal in its encoding, so we need to introduce a mov instr as well to get the imm into a register.
      - Change pass name to "SI Load Store Optimizer"; this removes the '/', which complicates scripts.
      
      Differential Revision: https://reviews.llvm.org/D42124
      
      llvm-svn: 323153
      7687d420
    • Marshall Clow's avatar
      Another batch of P0202 constepr algirithms.... · e8ea8296
      Marshall Clow authored
      Another batch of P0202 constepr algirithms. remove/remove_if/remove_copy/remove_copy_if/reverse_copy, and tests (commented out) for rotate_copy, because that depends on std::copy
      
      llvm-svn: 323152
      e8ea8296
    • Sam McCall's avatar
      d2a95925
    • Sam McCall's avatar
      [CodeComplete] Omit templated constructors from member list too. · 63c59720
      Sam McCall authored
      Also avoid printing a 'void' return type for constructor expressions.
      
      llvm-svn: 323148
      63c59720
    • Marshall Clow's avatar
    • Alexander Shaposhnikov's avatar
      [analyzer] Protect against dereferencing a null pointer · d7d991e8
      Alexander Shaposhnikov authored
      The check (inside StackHintGeneratorForSymbol::getMessage)
      if (!N)
          return getMessageForSymbolNotFound()
      is moved to the beginning of the function.
      
      Differential revision: https://reviews.llvm.org/D42388
      
      Test plan: make check-all
      
      llvm-svn: 323146
      d7d991e8
    • Don Hinton's avatar
      [cmake] [libcxxabi] Fix find_path() problems when cross compiling · ae858bf6
      Don Hinton authored
      When CMAKE_SYSROOT or CMAKE_FIND_ROOT_PATH is set, cmake
      recommends setting CMAKE_FIND_ROOT_PATH_MODE_INCLUDE=ONLY
      globally which means find_path() always prepends CMAKE_SYSROOT or
      CMAKE_FIND_ROOT_PATH to all paths used in the search.
      
      However, these find_path() invocations are looking for paths in
      the libcxx and libunwind projects on the host system, not the
      target system, which can be done by passing
      NO_CMAKE_FIND_ROOT_PATH.
      
      Differential Revision: https://reviews.llvm.org/D41623
      
      llvm-svn: 323145
      ae858bf6
    • Jake Ehrlich's avatar
      [llvm-objcopy] Use physical instead of virtual address when aligning and placing sections in binary · 46814bee
      Jake Ehrlich authored
      For sections with different virtual and physical addresses, alignment and
      placement in the output binary should be based on the physical address.
      
      Ran into this problem with a bare metal ARM project where llvm-objcopy added a
      lot of zero-padding before the .data section that had differing addresses. GNU
      objcopy did not add the padding, and after this fix, neither does llvm-objcopy.
      
      Update a test case so a section has different physical and virtual addresses.
      
      Fixes B35708
      
      Authored By: Owen Shaw (owenpshaw)
      
      Differential Revision: https://reviews.llvm.org/D41619
      
      llvm-svn: 323144
      46814bee
    • Don Hinton's avatar
      [cmake] [libcxx] Fix find_path() problems when cross compiling. · 98bb4205
      Don Hinton authored
      When CMAKE_SYSROOT or CMAKE_FIND_ROOT_PATH is set, cmake
      recommends setting CMAKE_FIND_ROOT_PATH_MODE_INCLUDE=ONLY
      globally which means find_path() always prepends CMAKE_SYSROOT or
      CMAKE_FIND_ROOT_PATH to all paths used in the search.
      
      However, this find_path() invocation is looking for a path in the
      libcxxabi project on the host system, not the target system,
      which can be done by passing NO_CMAKE_FIND_ROOT_PATH.
      
      Differential Revision: https://reviews.llvm.org/D41622
      
      llvm-svn: 323143
      98bb4205