1. Jan 15, 2020
    • Cullen Rhodes's avatar
      [AArch64][SVE] Add ptest intrinsics · 93a4dede
      Cullen Rhodes authored
      Summary:
      Implements the following intrinsics:
      
          * @llvm.aarch64.sve.ptest.any
          * @llvm.aarch64.sve.ptest.first
          * @llvm.aarch64.sve.ptest.last
      
      Reviewers: sdesmalen, efriedma, dancgr, mgudim, cameron.mcinally, rengolin
      
      Reviewed By: efriedma
      
      Subscribers: tschuett, kristof.beyls, hiraditya, rkruppe, psnobl, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72398
      93a4dede
    • Djordje Todorovic's avatar
      [llvm-locstats] Add the --draw-plot option · ada96466
      Djordje Todorovic authored
      When using the option, draw the histogram representing the debug
      location buckets. The resulting histogram will be saved in a png
      file.
      
      Differential Revision: https://reviews.llvm.org/D71869
      ada96466
    • Georgii Rymar's avatar
      [yaml2obj/obj2yaml] - Add support for SHT_RELR sections. · 46d11e30
      Georgii Rymar authored
      The encoded sequence of Elf*_Relr entries in a SHT_RELR section looks
      like [ AAAAAAAA BBBBBBB1 BBBBBBB1 ... AAAAAAAA BBBBBB1 ... ]
      i.e. start with an address, followed by any number of bitmaps. The address
      entry encodes 1 relocation. The subsequent bitmap entries encode up to 63(31)
      relocations each, at subsequent offsets following the last address entry.
      
      More information is here:
      https://github.com/llvm-mirror/llvm/blob/master/lib/Object/ELF.cpp#L272
      
      This patch adds a support for these sections.
      
      Differential revision: https://reviews.llvm.org/D71872
      46d11e30
    • Scott Egerton's avatar
      Revert "[RISCV] Add Clang frontend support for Bitmanip extension" · cbe681bd
      Scott Egerton authored
      This reverts commit 57cf6ee9.
      cbe681bd
    • Djordje Todorovic's avatar
      [llvm-locstats][NFC] Support OOP concept · a3ebc406
      Djordje Todorovic authored
      Making these changes, the code becomes more robust and easier for
      adding the new features.
      
        -Introduce the LocationStats class representing the statistics
        -Add the pretty_print() method in the LocationStats class
        -Add additional '-' for the program options
        -Add the verify_program_inputs() function
        -Add the parse_locstats() function
        -Rename 'results' => 'opts'
        -Add more comments
      
      Differential Revision: https://reviews.llvm.org/D71868
      a3ebc406
    • Zakk Chen's avatar
      [RISCV] Support ABI checking with per function target-features · 109e4d12
      Zakk Chen authored
      if users don't specific -mattr, the default target-feature come
      from IR attribute.
      109e4d12
    • Igor Kudrin's avatar
      [DWARF] Fix DWARFDebugAranges to support 64-bit CU offsets. · 2142e20f
      Igor Kudrin authored
      DWARFContext, the only user of this class, can already handle such offsets.
      
      Differential Revision: https://reviews.llvm.org/D71834
      2142e20f
    • LLVM GN Syncbot's avatar
      [gn build] Port 0dc6c249 · 4b1d471f
      LLVM GN Syncbot authored
      4b1d471f
    • Igor Kudrin's avatar
      [MachO] Add a test for detecting reserved unit length. · fcc08aa8
      Igor Kudrin authored
      This is a follow-up for D71546 to add a corresponding unit test.
      
      Differential Revision: https://reviews.llvm.org/D72695
      fcc08aa8
    • cdevadas's avatar
      [AMDGPU] Invert the handling of skip insertion. · 0dc6c249
      cdevadas authored
      The current implementation of skip insertion (SIInsertSkip) makes it a
      mandatory pass required for correctness. Initially, the idea was to
      have an optional pass. This patch inserts the s_cbranch_execz upfront
      during SILowerControlFlow to skip over the sections of code when no
      lanes are active. Later, SIRemoveShortExecBranches removes the skips
      for short branches, unless there is a sideeffect and the skip branch is
      really necessary.
      
      This new pass will replace the handling of skip insertion in the
      existing SIInsertSkip Pass.
      
      Differential revision: https://reviews.llvm.org/D68092
      0dc6c249
    • Kazushi (Jam) Marukawa's avatar
      [VE] Minimal codegen for empty functions · 064859bd
      Kazushi (Jam) Marukawa authored
      Summary:
      This patch implements minimal VE code generation for empty function bodies (no args, no value return).
      
      Contents
      
      * empty function code generation test.
      * Minimal function prologue & epilogue emission
      * Instruction formats and instruction definitions as far as required for the empty function prologue & epilogue.
      * I64 register class definitions.
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D72598
      064859bd
    • Craig Topper's avatar
      [X86] Don't call LowerUINT_TO_FP_i32 for i32->f80 on 32-bit targets with sse2. · be8f217b
      Craig Topper authored
      We were performing an emulated i32->f64 in the SSE registers, then
      storing that value to memory and doing a extload into the X87
      domain.
      
      After this patch we'll now just store the i32 to memory along
      with an i32 0. Then do a 64-bit FILD to f80 completely in the X87
      unit. This matches what we do without SSE.
      be8f217b
    • David Green's avatar
      [ARM] Reegenerate MVE tests. NFC · 1b264a82
      David Green authored
      The mve-phireg.ll test no longer really tests what it was added for,
      but the original case was fairly complex. I've left the test in as a
      general codegen test.
      1b264a82
    • Hideto Ueno's avatar
      [Attributor] AAValueConstantRange: Value range analysis using constant range · 188f9a34
      Hideto Ueno authored
      Summary:
      This patch introduces `AAValueConstantRange`, which answers a possible range for integer value in a specific program point.
      One of the motivations is propagating existing `range` metadata. (I think we need to change the situation that `range` metadata cannot be put to Argument).
      
      The state is a tuple of `ConstantRange` and it is initialized to (known, assumed) = ([-∞, +∞], empty).
      
      Currently, AAValueConstantRange is created in `getAssumedConstant` method when `AAValueSimplify` returns `nullptr`(worst state).
      
      Supported
       - BinaryOperator(add, sub, ...)
       - CmpInst(icmp eq, ...)
       - !range metadata
      
      `AAValueConstantRange` is not intended to extend to polyhedral range value analysis.
      
      Reviewers: jdoerfert, sstefan1
      
      Reviewed By: jdoerfert
      
      Subscribers: phosek, davezarzycki, baziotis, hiraditya, javed.absar, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D71620
      188f9a34
    • David Green's avatar
      [Scheduler] Adjust interface of CreateTargetMIHazardRecognizer to use ScheduleDAGMI. NFC · b891490c
      David Green authored
      All the callers of this function will be ScheduleDAGMI from the
      MachineScheduler. This allows us to use the extra info available in
      ScheduleDAGMI without resorting to awkward casts.
      b891490c
    • Jonas Devlieghere's avatar
      [lldb/test] Add test for CMTime data formatter · 914b551e
      Jonas Devlieghere authored
      Add a test for the CMTime data formatter. The coverage report showed
      that this code path was untested.
      914b551e
    • Jonas Devlieghere's avatar
      [lldb/CommandInterpreter] Remove flag that's always true (NFC) · a6faf851
      Jonas Devlieghere authored
      The 'asynchronously' argument to both GetLLDBCommandsFromIOHandler and
      GetPythonCommandsFromIOHandler is true for all call sites. This commit
      simplifies the API by dropping it and giving the baton a default
      argument.
      a6faf851
    • Reid Kleckner's avatar
      c42116cc
    • Fangrui Song's avatar
      [Driver][X86] Add -malign-branch* and -mbranches-within-32B-boundaries · 5ca24d09
      Fangrui Song authored
      These driver options perform some checking and delegate to MC options -x86-align-branch* and -x86-branches-within-32B-boundaries.
      
      Reviewed By: skan
      
      Differential Revision: https://reviews.llvm.org/D72463
      5ca24d09
    • Weverything's avatar
      [ODRHash] Fix wrong error message with bitfields and mutable. · a60e8927
      Weverything authored
      Add a check to bitfield mismatches that may have caused Clang to
      give an error about the bitfield instead of being mutable.
      a60e8927
    • Justin Hibbits's avatar
      [PowerPC] Fix powerpcspe subtarget enablement in llvm backend · 36eedfcb
      Justin Hibbits authored
      Summary:
      As currently written, -target powerpcspe will enable SPE regardless of
      disabling the feature later on in the command line.  Instead, change
      this to just set a default CPU to 'e500' instead of a generic CPU.
      
      As part of this, add FeatureSPE to the e500 definition.
      
      Reviewed By: MaskRay
      Differential Revision: https://reviews.llvm.org/D72673
      36eedfcb
    • Pierre Habouzit's avatar
      Relax the rules around objc_alloc and objc_alloc_init optimizations. · d18fbfc0
      Pierre Habouzit authored
      Today the optimization is limited to:
      - `[ClassName alloc]`
      - `[self alloc]` when within a class method
      
      However it means that when code is written this way:
      
      ```
          @interface MyObject
          - (id)copyWithZone:(NSZone *)zone
          {
              return [[self.class alloc] _initWith...];
          }
      
          @end
      ```
      
      ... then the optimization doesn't kick in and `+[NSObject alloc]` ends
      up in IMP caches where it could have been avoided. It turns out that
      `+alloc` -> `+[NSObject alloc]` is the most cached SEL/IMP pair in the
      entire platform which is rather silly).
      
      There's two theoretical risks allowing this optimization:
      
      1. if the receiver is nil (which it can't be today), but it turns out
         that `objc_alloc()`/`objc_alloc_init()` cope with a nil receiver,
      
      2. if the `Clas` type for the receiver is a lie. However, for such a
         code to work today (and not fail witn an unrecognized selector
         anyway) you'd have to have implemented the `-alloc` **instance
         method**.
      
         Fortunately, `objc_alloc()` doesn't assume that the receiver is a
         Class, it basically starts with a test that is similar to
      
             `if (receiver->isa->bits & hasDefaultAWZ) { /* fastpath */ }`.
      
         This bit is only set on metaclasses by the runtime, so if an instance
         is passed to this function by accident, its isa will fail this test,
         and `objc_alloc()` will gracefully fallback to `objc_msgSend()`.
      
         The one thing `objc_alloc()` doesn't support is tagged pointer
         instances. None of the tagged pointer classes implement an instance
         method called `'alloc'` (actually there's a single class in the
         entire Apple codebase that has such a method).
      
      Differential Revision: https://reviews.llvm.org/D71682
      Radar-Id: rdar://problem/58058316
      
      
      Reviewed-By: Akira Hatanaka
      Signed-off-by: default avatarPierre Habouzit <phabouzit@apple.com>
      d18fbfc0
    • Tom Stellard's avatar
      CMake: Make most target symbols hidden by default · 0dbcb363
      Tom Stellard authored
      Summary:
      For builds with LLVM_BUILD_LLVM_DYLIB=ON and BUILD_SHARED_LIBS=OFF
      this change makes all symbols in the target specific libraries hidden
      by default.
      
      A new macro called LLVM_EXTERNAL_VISIBILITY has been added to mark symbols in these
      libraries public, which is mainly needed for the definitions of the
      LLVMInitialize* functions.
      
      This patch reduces the number of public symbols in libLLVM.so by about
      25%.  This should improve load times for the dynamic library and also
      make abi checker tools, like abidiff require less memory when analyzing
      libLLVM.so
      
      One side-effect of this change is that for builds with
      LLVM_BUILD_LLVM_DYLIB=ON and LLVM_LINK_LLVM_DYLIB=ON some unittests that
      access symbols that are no longer public will need to be statically linked.
      
      Before and after public symbol counts (using gcc 8.2.1, ld.bfd 2.31.1):
      nm before/libLLVM-9svn.so | grep ' [A-Zuvw] ' | wc -l
      36221
      nm after/libLLVM-9svn.so | grep ' [A-Zuvw] ' | wc -l
      26278
      
      Reviewers...
      0dbcb363
    • Richard Smith's avatar
      PR44540: Prefer an inherited default constructor over an initializer · 1b5404af
      Richard Smith authored
      list constructor when initializing from {}.
      
      We would previously pick between calling an initializer list constructor
      and calling a default constructor unstably in this situation, depending
      on whether the inherited default constructor had already been used
      elsewhere in the program.
      1b5404af
    • Douglas Yung's avatar
      Modify test to use -S instead of -c so that it works when an external... · c6e69880
      Douglas Yung authored
      Modify test to use -S instead of -c so that it works when an external assembler is used that is not present.
      c6e69880
    • Hubert Tong's avatar
      DWARFDebugLine.cpp: Restore LF line endings · aca3e70d
      Hubert Tong authored
      rG7e02406f switched the file to CRLF
      line endings.
      aca3e70d
    • Philip Reames's avatar
      [BranchAlign] Add master --x86-branches-within-32B-boundaries flag · 1a7398ec
      Philip Reames authored
      This flag was originally part of D70157, but was removed as we carved away pieces of the review. Since we have the nop support checked in, and it appears mature(*), I think it's time to add the master flag. For now, it will default to nop padding, but once the prefix padding support lands, we'll update the defaults.
      
      (*) I can now confirm that downstream testing of the changes which have landed to date - nop padding and compiler support for suppressions - is passing all of the functional testing we've thrown at it. There might still be something lurking, but we've gotten enough coverage to be confident of the basic approach.
      
      Note that the new flag can be used either when assembling an .s file, or when using the integrated assembler directly from the compiler. The later will use all of the suppression mechanism and should always generate correct code. We don't yet have assembly syntax for the suppressions, so passing this directly to the assembler w/a raw .s file may result in broken code. Use at your own risk.
      
      Also note that this isn't the wiring for the clang option. I think the most recent review for that is D72227, but I've lost track, so that might be off.
      
      Differential Revision: https://reviews.llvm.org/D72738
      1a7398ec
    • Saar Raz's avatar
      [Concepts] Type Constraints · ff1e0fce
      Saar Raz authored
      Add support for type-constraints in template type parameters.
      Also add support for template type parameters as pack expansions (where the type constraint can now contain an unexpanded parameter pack).
      
      Differential Revision: https://reviews.llvm.org/D44352
      ff1e0fce
    • Reid Kleckner's avatar
      [X86] ABI compat bugfix for MSVC vectorcall · 8e780252
      Reid Kleckner authored
      Summary:
      Before this change, X86_32ABIInfo::classifyArgument would be called
      twice on vector arguments to vectorcall functions. This function has
      side effects to track GPR register usage, and this would lead to
      incorrect GPR usage in some cases.  The specific case I noticed is from
      running out of XMM registers with mixed FP and vector arguments and no
      aggregates of any kind. Consider this prototype:
      
        void __vectorcall vectorcall_indirect_vec(
            double xmm0, double xmm1, double xmm2, double xmm3, double xmm4,
            __m128 xmm5,
            __m128 ecx,
            int edx,
            __m128 mem);
      
      classifyArgument has no effects when called on a plain FP type, but when
      called on a vector type, it modifies FreeRegs to model GPR consumption.
      However, this should not happen during the vector call first pass.
      
      I refactored the code to unify vectorcall HVA logic with regcall HVA
      logic. The conventions pass HVAs in registers differently (expanded vs.
      not expanded), but if they do not fit in registers, they both pass them
      indirectly by address.
      
      Reviewers: erichkeane, craig.topper
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72110
      8e780252
    • Zachary Henkel's avatar
      Allow /D flags absent during PCH creation under msvc-compat · 0f9cf42f
      Zachary Henkel authored
      Summary:
      Before this patch adding a new /D flag when compiling a source file that consumed a PCH with clang-cl would issue a diagnostic and then fail.  With the patch, the diagnostic is still issued but the definition is accepted.  This matches the msvc behavior.  The fuzzy-pch-msvc.c is a clone of the existing fuzzy-pch.c tests with some msvc specific rework.
      
      msvc diagnostic:
        warning C4605: '/DBAR=int' specified on current command line, but was not specified when precompiled header was built
      
      Output of the CHECK-BAR test prior to the code change:
        <built-in>(1,9): warning: definition of macro 'BAR' does not match definition in precompiled header [-Wclang-cl-pch]
        #define BAR int
                ^
        D:\repos\llvm\llvm-project\clang\test\PCH\fuzzy-pch-msvc.c(12,1): error: unknown type name 'BAR'
        BAR bar = 17;
        ^
        D:\repos\llvm\llvm-project\clang\test\PCH\fuzzy-pch-msvc.c(23,4): error: BAR was not defined
        #  error BAR was not defined
           ^
        1 warning and 2 errors generated.
      
      Reviewers: rnk, thakis, hans, zturner
      
      Subscribers: mikerice, aganea, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72405
      0f9cf42f
    • Reid Kleckner's avatar
      [Win64] Handle FP arguments more gracefully under -mno-sse · 40cd26c7
      Reid Kleckner authored
      Pass small FP values in GPRs or stack memory according the the normal
      convention. This is what gcc -mno-sse does on Win64.
      
      I adjusted the conditions under which we emit an error to check if the
      argument or return value would be passed in an XMM register when SSE is
      disabled. This has a side effect of no longer emitting an error for FP
      arguments marked 'inreg' when targetting x86 with SSE disabled. Our
      calling convention logic was already assigning it to FP0/FP1, and then
      we emitted this error. That seems unnecessary, we can ignore 'inreg' and
      compile it without SSE.
      
      Reviewers: jyknight, aemerson
      
      Differential Revision: https://reviews.llvm.org/D70465
      40cd26c7
    • Michael Liao's avatar
      [amdgpu] Fix typos in a test case. · 65c8abb1
      Michael Liao authored
      - There are typos introduced due to merge.
      65c8abb1
    • Craig Topper's avatar
      [X86] Drop an unneeded FIXME. NFC · 76291e11
      Craig Topper authored
      The extload on X87 is free.
      76291e11
    • Craig Topper's avatar
      [X86] Swap the 0 and the fudge factor in the constant pool for the 32-bit mode... · 57eb56b8
      Craig Topper authored
      [X86] Swap the 0 and the fudge factor in the constant pool for the 32-bit mode i64->f32/f64/f80 uint_to_fp algorithm.
      
      This allows us to generate better code for selecting the fixup
      to load.
      
      Previously when the sign was set we had to load offset 0. And
      when it was clear we had to load offset 4. This required a testl,
      setns, zero extend, and finally a mul by 4. By switching the offsets
      we can just shift the sign bit into the lsb and multiply it by 4.
      57eb56b8
    • Ahmed Taei's avatar
      [mlir] : Fix ViewOp shape folder for identity affine maps · ab035647
      Ahmed Taei authored
      Summary: Fix the ViewOpShapeFolder in case of no affine mapping associated with a Memref construct identity mapping.
      
      Reviewers: nicolasvasilache
      
      Subscribers: mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, arpith-jacob, mgester, lucyrfox, liufengdb, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72735
      ab035647
    • Petr Hosek's avatar
      [libcxx] Use C11 thread API on Fuchsia · ab9aefee
      Petr Hosek authored
      On Fuchsia, pthread API is emulated on top of C11 thread API. Using C11
      thread API directly is more efficient.
      
      While this implementation is only used by Fuchsia at the moment, it's
      not Fuchsia specific, and could be used by other platforms that use C11
      threads rather than pthreads in the future.
      
      Differential Revision: https://reviews.llvm.org/D64378
      ab9aefee
    • Rong Xu's avatar
      Fix windows bot failures in c410adb092c9cb51ddb0b55862b70f2aa8c5b16f · c9ee5e99
      Rong Xu authored
      (clang diagnostic handler for IR input files)
      c9ee5e99
    • George Rokos's avatar
      [LIBOMPTARGET] Do not increment/decrement the refcount for "declare target" objects · e244145a
      George Rokos authored
      The reference counter for global objects marked with declare target is INF. This patch prevents the runtime from incrementing /decrementing INF refcounts. Without it, the map(delete: global_object) directive actually deallocates the global on the device. With this patch, such a directive becomes a no-op.
      
      Differential Revision: https://reviews.llvm.org/D72525
      e244145a
    • Michael Liao's avatar
      [codegen,amdgpu] Enhance MIR DIE and re-arrange it for AMDGPU. · 01a4b831
      Michael Liao authored
      Summary:
      - `dead-mi-elimination` assumes MIR in the SSA form and cannot be
        arranged after phi elimination or DeSSA. It's enhanced to handle the
        dead register definition by skipping use check on it. Once a register
        def is `dead`, all its uses, if any, should be `undef`.
      - Re-arrange the DIE in RA phase for AMDGPU by placing it directly after
        `detect-dead-lanes`.
      - Many relevant tests are refined due to different register assignment.
      
      Reviewers: rampitec, qcolombet, sunfish
      
      Subscribers: arsenm, kzhuravl, jvesely, wdng, nhaehnle, yaxunl, dstuttard, tpr, t-tye, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72709
      01a4b831
    • Lei Zhang's avatar
      [mlir][spirv] Properly support SPIR-V conversion target · 47c6ab2b
      Lei Zhang authored
      This commit defines a new SPIR-V dialect attribute for specifying
      a SPIR-V target environment. It is a dictionary attribute containing
      the SPIR-V version, supported extension list, and allowed capability
      list. A SPIRVConversionTarget subclass is created to take in the
      target environment and sets proper dynmaically legal ops by querying
      the op availability interface of SPIR-V ops to make sure they are
      available in the specified target environment. All existing conversions
      targeting SPIR-V is changed to use this SPIRVConversionTarget. It
      probes whether the input IR has a `spv.target_env` attribute,
      otherwise, it uses the default target environment: SPIR-V 1.0 with
      Shader capability and no extra extensions.
      
      Differential Revision: https://reviews.llvm.org/D72256
      47c6ab2b