1. Jan 15, 2020
    • Fangrui Song's avatar
      [Driver][X86] Add -malign-branch* and -mbranches-within-32B-boundaries · 5ca24d09
      Fangrui Song authored
      These driver options perform some checking and delegate to MC options -x86-align-branch* and -x86-branches-within-32B-boundaries.
      
      Reviewed By: skan
      
      Differential Revision: https://reviews.llvm.org/D72463
      5ca24d09
    • Weverything's avatar
      [ODRHash] Fix wrong error message with bitfields and mutable. · a60e8927
      Weverything authored
      Add a check to bitfield mismatches that may have caused Clang to
      give an error about the bitfield instead of being mutable.
      a60e8927
    • Justin Hibbits's avatar
      [PowerPC] Fix powerpcspe subtarget enablement in llvm backend · 36eedfcb
      Justin Hibbits authored
      Summary:
      As currently written, -target powerpcspe will enable SPE regardless of
      disabling the feature later on in the command line.  Instead, change
      this to just set a default CPU to 'e500' instead of a generic CPU.
      
      As part of this, add FeatureSPE to the e500 definition.
      
      Reviewed By: MaskRay
      Differential Revision: https://reviews.llvm.org/D72673
      36eedfcb
    • Pierre Habouzit's avatar
      Relax the rules around objc_alloc and objc_alloc_init optimizations. · d18fbfc0
      Pierre Habouzit authored
      Today the optimization is limited to:
      - `[ClassName alloc]`
      - `[self alloc]` when within a class method
      
      However it means that when code is written this way:
      
      ```
          @interface MyObject
          - (id)copyWithZone:(NSZone *)zone
          {
              return [[self.class alloc] _initWith...];
          }
      
          @end
      ```
      
      ... then the optimization doesn't kick in and `+[NSObject alloc]` ends
      up in IMP caches where it could have been avoided. It turns out that
      `+alloc` -> `+[NSObject alloc]` is the most cached SEL/IMP pair in the
      entire platform which is rather silly).
      
      There's two theoretical risks allowing this optimization:
      
      1. if the receiver is nil (which it can't be today), but it turns out
         that `objc_alloc()`/`objc_alloc_init()` cope with a nil receiver,
      
      2. if the `Clas` type for the receiver is a lie. However, for such a
         code to work today (and not fail witn an unrecognized selector
         anyway) you'd have to have implemented the `-alloc` **instance
         method**.
      
         Fortunately, `objc_alloc()` doesn't assume that the receiver is a
         Class, it basically starts with a test that is similar to
      
             `if (receiver->isa->bits & hasDefaultAWZ) { /* fastpath */ }`.
      
         This bit is only set on metaclasses by the runtime, so if an instance
         is passed to this function by accident, its isa will fail this test,
         and `objc_alloc()` will gracefully fallback to `objc_msgSend()`.
      
         The one thing `objc_alloc()` doesn't support is tagged pointer
         instances. None of the tagged pointer classes implement an instance
         method called `'alloc'` (actually there's a single class in the
         entire Apple codebase that has such a method).
      
      Differential Revision: https://reviews.llvm.org/D71682
      Radar-Id: rdar://problem/58058316
      
      
      Reviewed-By: Akira Hatanaka
      Signed-off-by: default avatarPierre Habouzit <phabouzit@apple.com>
      d18fbfc0
    • Tom Stellard's avatar
      CMake: Make most target symbols hidden by default · 0dbcb363
      Tom Stellard authored
      Summary:
      For builds with LLVM_BUILD_LLVM_DYLIB=ON and BUILD_SHARED_LIBS=OFF
      this change makes all symbols in the target specific libraries hidden
      by default.
      
      A new macro called LLVM_EXTERNAL_VISIBILITY has been added to mark symbols in these
      libraries public, which is mainly needed for the definitions of the
      LLVMInitialize* functions.
      
      This patch reduces the number of public symbols in libLLVM.so by about
      25%.  This should improve load times for the dynamic library and also
      make abi checker tools, like abidiff require less memory when analyzing
      libLLVM.so
      
      One side-effect of this change is that for builds with
      LLVM_BUILD_LLVM_DYLIB=ON and LLVM_LINK_LLVM_DYLIB=ON some unittests that
      access symbols that are no longer public will need to be statically linked.
      
      Before and after public symbol counts (using gcc 8.2.1, ld.bfd 2.31.1):
      nm before/libLLVM-9svn.so | grep ' [A-Zuvw] ' | wc -l
      36221
      nm after/libLLVM-9svn.so | grep ' [A-Zuvw] ' | wc -l
      26278
      
      Reviewers...
      0dbcb363
    • Richard Smith's avatar
      PR44540: Prefer an inherited default constructor over an initializer · 1b5404af
      Richard Smith authored
      list constructor when initializing from {}.
      
      We would previously pick between calling an initializer list constructor
      and calling a default constructor unstably in this situation, depending
      on whether the inherited default constructor had already been used
      elsewhere in the program.
      1b5404af
    • Douglas Yung's avatar
      Modify test to use -S instead of -c so that it works when an external... · c6e69880
      Douglas Yung authored
      Modify test to use -S instead of -c so that it works when an external assembler is used that is not present.
      c6e69880
    • Hubert Tong's avatar
      DWARFDebugLine.cpp: Restore LF line endings · aca3e70d
      Hubert Tong authored
      rG7e02406f switched the file to CRLF
      line endings.
      aca3e70d
    • Philip Reames's avatar
      [BranchAlign] Add master --x86-branches-within-32B-boundaries flag · 1a7398ec
      Philip Reames authored
      This flag was originally part of D70157, but was removed as we carved away pieces of the review. Since we have the nop support checked in, and it appears mature(*), I think it's time to add the master flag. For now, it will default to nop padding, but once the prefix padding support lands, we'll update the defaults.
      
      (*) I can now confirm that downstream testing of the changes which have landed to date - nop padding and compiler support for suppressions - is passing all of the functional testing we've thrown at it. There might still be something lurking, but we've gotten enough coverage to be confident of the basic approach.
      
      Note that the new flag can be used either when assembling an .s file, or when using the integrated assembler directly from the compiler. The later will use all of the suppression mechanism and should always generate correct code. We don't yet have assembly syntax for the suppressions, so passing this directly to the assembler w/a raw .s file may result in broken code. Use at your own risk.
      
      Also note that this isn't the wiring for the clang option. I think the most recent review for that is D72227, but I've lost track, so that might be off.
      
      Differential Revision: https://reviews.llvm.org/D72738
      1a7398ec
    • Saar Raz's avatar
      [Concepts] Type Constraints · ff1e0fce
      Saar Raz authored
      Add support for type-constraints in template type parameters.
      Also add support for template type parameters as pack expansions (where the type constraint can now contain an unexpanded parameter pack).
      
      Differential Revision: https://reviews.llvm.org/D44352
      ff1e0fce
    • Reid Kleckner's avatar
      [X86] ABI compat bugfix for MSVC vectorcall · 8e780252
      Reid Kleckner authored
      Summary:
      Before this change, X86_32ABIInfo::classifyArgument would be called
      twice on vector arguments to vectorcall functions. This function has
      side effects to track GPR register usage, and this would lead to
      incorrect GPR usage in some cases.  The specific case I noticed is from
      running out of XMM registers with mixed FP and vector arguments and no
      aggregates of any kind. Consider this prototype:
      
        void __vectorcall vectorcall_indirect_vec(
            double xmm0, double xmm1, double xmm2, double xmm3, double xmm4,
            __m128 xmm5,
            __m128 ecx,
            int edx,
            __m128 mem);
      
      classifyArgument has no effects when called on a plain FP type, but when
      called on a vector type, it modifies FreeRegs to model GPR consumption.
      However, this should not happen during the vector call first pass.
      
      I refactored the code to unify vectorcall HVA logic with regcall HVA
      logic. The conventions pass HVAs in registers differently (expanded vs.
      not expanded), but if they do not fit in registers, they both pass them
      indirectly by address.
      
      Reviewers: erichkeane, craig.topper
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72110
      8e780252
    • Zachary Henkel's avatar
      Allow /D flags absent during PCH creation under msvc-compat · 0f9cf42f
      Zachary Henkel authored
      Summary:
      Before this patch adding a new /D flag when compiling a source file that consumed a PCH with clang-cl would issue a diagnostic and then fail.  With the patch, the diagnostic is still issued but the definition is accepted.  This matches the msvc behavior.  The fuzzy-pch-msvc.c is a clone of the existing fuzzy-pch.c tests with some msvc specific rework.
      
      msvc diagnostic:
        warning C4605: '/DBAR=int' specified on current command line, but was not specified when precompiled header was built
      
      Output of the CHECK-BAR test prior to the code change:
        <built-in>(1,9): warning: definition of macro 'BAR' does not match definition in precompiled header [-Wclang-cl-pch]
        #define BAR int
                ^
        D:\repos\llvm\llvm-project\clang\test\PCH\fuzzy-pch-msvc.c(12,1): error: unknown type name 'BAR'
        BAR bar = 17;
        ^
        D:\repos\llvm\llvm-project\clang\test\PCH\fuzzy-pch-msvc.c(23,4): error: BAR was not defined
        #  error BAR was not defined
           ^
        1 warning and 2 errors generated.
      
      Reviewers: rnk, thakis, hans, zturner
      
      Subscribers: mikerice, aganea, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72405
      0f9cf42f
    • Reid Kleckner's avatar
      [Win64] Handle FP arguments more gracefully under -mno-sse · 40cd26c7
      Reid Kleckner authored
      Pass small FP values in GPRs or stack memory according the the normal
      convention. This is what gcc -mno-sse does on Win64.
      
      I adjusted the conditions under which we emit an error to check if the
      argument or return value would be passed in an XMM register when SSE is
      disabled. This has a side effect of no longer emitting an error for FP
      arguments marked 'inreg' when targetting x86 with SSE disabled. Our
      calling convention logic was already assigning it to FP0/FP1, and then
      we emitted this error. That seems unnecessary, we can ignore 'inreg' and
      compile it without SSE.
      
      Reviewers: jyknight, aemerson
      
      Differential Revision: https://reviews.llvm.org/D70465
      40cd26c7
    • Michael Liao's avatar
      [amdgpu] Fix typos in a test case. · 65c8abb1
      Michael Liao authored
      - There are typos introduced due to merge.
      65c8abb1
    • Craig Topper's avatar
      [X86] Drop an unneeded FIXME. NFC · 76291e11
      Craig Topper authored
      The extload on X87 is free.
      76291e11
    • Craig Topper's avatar
      [X86] Swap the 0 and the fudge factor in the constant pool for the 32-bit mode... · 57eb56b8
      Craig Topper authored
      [X86] Swap the 0 and the fudge factor in the constant pool for the 32-bit mode i64->f32/f64/f80 uint_to_fp algorithm.
      
      This allows us to generate better code for selecting the fixup
      to load.
      
      Previously when the sign was set we had to load offset 0. And
      when it was clear we had to load offset 4. This required a testl,
      setns, zero extend, and finally a mul by 4. By switching the offsets
      we can just shift the sign bit into the lsb and multiply it by 4.
      57eb56b8
    • Ahmed Taei's avatar
      [mlir] : Fix ViewOp shape folder for identity affine maps · ab035647
      Ahmed Taei authored
      Summary: Fix the ViewOpShapeFolder in case of no affine mapping associated with a Memref construct identity mapping.
      
      Reviewers: nicolasvasilache
      
      Subscribers: mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, arpith-jacob, mgester, lucyrfox, liufengdb, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72735
      ab035647
    • Petr Hosek's avatar
      [libcxx] Use C11 thread API on Fuchsia · ab9aefee
      Petr Hosek authored
      On Fuchsia, pthread API is emulated on top of C11 thread API. Using C11
      thread API directly is more efficient.
      
      While this implementation is only used by Fuchsia at the moment, it's
      not Fuchsia specific, and could be used by other platforms that use C11
      threads rather than pthreads in the future.
      
      Differential Revision: https://reviews.llvm.org/D64378
      ab9aefee
    • Rong Xu's avatar
      Fix windows bot failures in c410adb092c9cb51ddb0b55862b70f2aa8c5b16f · c9ee5e99
      Rong Xu authored
      (clang diagnostic handler for IR input files)
      c9ee5e99
    • George Rokos's avatar
      [LIBOMPTARGET] Do not increment/decrement the refcount for "declare target" objects · e244145a
      George Rokos authored
      The reference counter for global objects marked with declare target is INF. This patch prevents the runtime from incrementing /decrementing INF refcounts. Without it, the map(delete: global_object) directive actually deallocates the global on the device. With this patch, such a directive becomes a no-op.
      
      Differential Revision: https://reviews.llvm.org/D72525
      e244145a
    • Michael Liao's avatar
      [codegen,amdgpu] Enhance MIR DIE and re-arrange it for AMDGPU. · 01a4b831
      Michael Liao authored
      Summary:
      - `dead-mi-elimination` assumes MIR in the SSA form and cannot be
        arranged after phi elimination or DeSSA. It's enhanced to handle the
        dead register definition by skipping use check on it. Once a register
        def is `dead`, all its uses, if any, should be `undef`.
      - Re-arrange the DIE in RA phase for AMDGPU by placing it directly after
        `detect-dead-lanes`.
      - Many relevant tests are refined due to different register assignment.
      
      Reviewers: rampitec, qcolombet, sunfish
      
      Subscribers: arsenm, kzhuravl, jvesely, wdng, nhaehnle, yaxunl, dstuttard, tpr, t-tye, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72709
      01a4b831
    • Lei Zhang's avatar
      [mlir][spirv] Properly support SPIR-V conversion target · 47c6ab2b
      Lei Zhang authored
      This commit defines a new SPIR-V dialect attribute for specifying
      a SPIR-V target environment. It is a dictionary attribute containing
      the SPIR-V version, supported extension list, and allowed capability
      list. A SPIRVConversionTarget subclass is created to take in the
      target environment and sets proper dynmaically legal ops by querying
      the op availability interface of SPIR-V ops to make sure they are
      available in the specified target environment. All existing conversions
      targeting SPIR-V is changed to use this SPIRVConversionTarget. It
      probes whether the input IR has a `spv.target_env` attribute,
      otherwise, it uses the default target environment: SPIR-V 1.0 with
      Shader capability and no extra extensions.
      
      Differential Revision: https://reviews.llvm.org/D72256
      47c6ab2b
    • Rong Xu's avatar
      [remark][diagnostics] Using clang diagnostic handler for IR input files · 60d39479
      Rong Xu authored
      For IR input files, we currently use LLVM diagnostic handler even the
      compilation is from clang. As a result, we are not able to use -Rpass
      to get the transformation reports. Some warnings are not handled
      properly either: We found many mysterious warnings in our ThinLTO backend
      compilations in SamplePGO and CSPGO. An example of the warning:
      "warning: net/proto2/public/metadata_lite.h:51:21: 0.02% (1 / 4999)"
      
      This turns out to be a warning by Wmisexpect, which is supposed to be
      filtered out by default. But since the filter is in clang's
      diagnostic hander, we emit these incomplete warnings from LLVM's
      diagnostic handler.
      
      This patch uses clang diagnostic handler for IR input files. We create
      a fake backendconsumer just to install the diagnostic handler.
      
      With this change, we will have proper handling of all the warnings and we can
      use -Rpass* options in IR input files compilation.
      Also note that with is patch, LLVM's diagnostic options, like
      "-mllvm -pass-remarks=*", are no longer be able to get optimization remarks.
      
      Differential Revision: https://reviews.llvm.org/D72523
      60d39479
    • River Riddle's avatar
      [mlir] Refactor ModuleState into AsmState and expose it to users. · fa9dd833
      River Riddle authored
      Summary:
      This allows for users to cache printer state, which can be costly to recompute. Each of the IR print methods gain a new overload taking this new state class.
      
      Depends On D72293
      
      Reviewed By: jpienaar
      
      Differential Revision: https://reviews.llvm.org/D72294
      fa9dd833
    • Alexey Bataev's avatar
      [OPENMP]Do not use RTTI by default for NVPTX devices. · 23058f9d
      Alexey Bataev authored
      NVPTX does not support RTTI, so disable it by default.
      23058f9d
    • River Riddle's avatar
      [mlir] Enable printing of FuncOp in the generic form. · 20c6e074
      River Riddle authored
      Summary:
      This was previously disabled as FunctionType TypeAttrs could not be roundtripped in the IR. This has been fixed, so we can now generically print FuncOp.
      
      Depends On D72429
      
      Reviewed By: jpienaar, mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D72642
      20c6e074
    • Luboš Luňák's avatar
      make -fmodules-codegen and -fmodules-debuginfo work also with PCHs · cbc9d22e
      Luboš Luňák authored
      Allow to build PCH's (with -building-pch-with-obj and the extra .o file)
      with -fmodules-codegen -fmodules-debuginfo to allow emitting shared code
      into the extra .o file, similarly to how it works with modules. A bit of
      a misnomer, but the underlying functionality is the same. This saves up
      to 20% of build time here.
      
      Differential Revision: https://reviews.llvm.org/D69778
      cbc9d22e
    • Luboš Luňák's avatar
      fix recent -fmodules-codegen fix test · b5b2cf7a
      Luboš Luňák authored
      b5b2cf7a
    • Luboš Luňák's avatar
      -fmodules-codegen should not emit extern templates · 729530f6
      Luboš Luňák authored
      If a header contains 'extern template', then the template should be provided
      somewhere by an explicit instantiation, so it is not necessary to generate
      a copy. Worse, this can lead to an unresolved symbol, because the codegen's
      object file will not actually contain functions from such a template
      because of the GVA_AvailableExternally, but the object file for the explicit
      instantiation will not contain them either because it will be blocked
      by the information provided by the module.
      
      Differential Revision: https://reviews.llvm.org/D69779
      729530f6
    • Nicolas Vasilache's avatar
      [mlir][Linalg] Update the semantics, verifier and test for Linalg with tensors. · f52d7173
      Nicolas Vasilache authored
      Summary:
      This diff fixes issues with the semantics of linalg.generic on tensors that appeared when converting directly from HLO to linalg.generic.
      The changes are self-contained within MLIR and can be captured and tested independently of XLA.
      
      The linalg.generic and indexed_generic are updated to:
      
      To allow progressive lowering from the value world (a.k.a tensor values) to
      the buffer world (a.k.a memref values), a linalg.generic op accepts
      mixing input and output ranked tensor values with input and output memrefs.
      
      ```
      %1 = linalg.generic #trait_attribute %A, %B {other-attributes} :
        tensor<?x?xf32>,
        memref<?x?xf32, stride_specification>
        -> (tensor<?x?xf32>)
      ```
      
      In this case, the number of outputs (args_out) must match the sum of (1) the
      number of output buffer operands and (2) the number of tensor return values.
      The semantics is that the linalg.indexed_generic op produces (i.e.
      allocates and fills) its return values.
      
      Tensor values must be legalized by a buffer allocation pass before most
      transformations can be applied. Such legalization moves tensor return values
      into output buffer operands and updates the region argument accordingly.
      
      Transformations that create control-flow around linalg.indexed_generic
      operations are not expected to mix with tensors because SSA values do not
      escape naturally. Still, transformations and rewrites that take advantage of
      tensor SSA values are expected to be useful and will be added in the near
      future.
      
      Subscribers: bmahjour, mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, arpith-jacob, mgester, lucyrfox, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72555
      f52d7173
    • Michael Liao's avatar
      [DAGCombine] Replace `getIntPtrConstant()` with `getVectorIdxTy()`. · 8d07f8d9
      Michael Liao authored
      - Prefer `getVectorIdxTy()` as the index operand type for
        `EXTRACT_SUBVECTOR` as targets expect different types by overloading
        `getVectorIdxTy()`.
      8d07f8d9
    • Alexey Bataev's avatar
      [OPENMP]Do not emit special virtual function for NVPTX target. · a48600c0
      Alexey Bataev authored
      There are no special virtual function handlers (like __cxa_pure_virtual)
      defined for NVPTX target, so just emit such functions as null pointers
      to prevent issues with linking and unresolved references.
      a48600c0
    • River Riddle's avatar
      [mlir] Use double format when parsing bfloat16 hexadecimal values · 1bd14ce3
      River Riddle authored
      Summary: bfloat16 doesn't have a valid APFloat format, so we have to use double semantics when storing it. This change makes sure that hexadecimal values can be round-tripped properly given this fact.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D72667
      1bd14ce3
    • Michael Liao's avatar
      Remove trailing `;`. NFC. · a3490e3e
      Michael Liao authored
      a3490e3e
    • Amara Emerson's avatar
      [AArch64][GlobalISel]: Support @llvm.{return,frame}address selection. · 6078f2fe
      Amara Emerson authored
      These intrinsics expand to a variable number of instructions so just like in
      ISelLowering.cpp we use custom code to deal with them.
      
      Committing Tim's original patch.
      
      Differential Revision: https://reviews.llvm.org/D65656
      6078f2fe
    • Fangrui Song's avatar
      [Driver][test] Fix Driver/hexagon-toolchain-elf.c for -DCLANG_DEFAULT_LINKER=lld builds · 1ca51c06
      Fangrui Song authored
      Reviewed By: nathanchance, sidneym
      
      Differential Revision: https://reviews.llvm.org/D72668
      1ca51c06
    • Craig Topper's avatar
      [LegalizeTypes] Remove untested code from ExpandIntOp_UINT_TO_FP · 9ee90ea5
      Craig Topper authored
      This code is untested in tree because the "APFloat::semanticsPrecision(sem) >= SrcVT.getSizeInBits() - 1" check is false for most combinations for int and fp types except maybe i32 and f64. For that you would need i32 to be an illegal type, but f64 to be legal and have custom handling for legalizing the split sint_to_fp. The precision check itself was added in 2010 to fix a double rounding issue in the algorithm that would occur if the sint_to_fp was not able to do the conversion without rounding.
      
      Differential Revision: https://reviews.llvm.org/D72728
      9ee90ea5
    • Fedor Sergeev's avatar
    • Jan Korous's avatar
      [clang][test][NFC] Use more widely supported sanitizer for file dependency tests · 986202fa
      Jan Korous authored
      The tests aren't concerned at all by the actual sanitizer - only by blacklist being reported as a dependency.
      We're unfortunately limited by platform support for any particular sanitizer but we can at least use one that is widely supported.
      
      Post-commit review:
      https://reviews.llvm.org/D72729
      986202fa
    • Nikita Popov's avatar
      [InstCombine] Fix worklist management when removing guard intrinsic · 04e58615
      Nikita Popov authored
      When multiple guard intrinsics are merged into one, currently the
      result of eraseInstFromFunction() is returned -- however, this
      should only be done if the current instruction is being removed.
      In this case we're removing a different instruction and should
      instead report that the current one has been modified by returning it.
      
      For this test case, this reduces the number of instcombine iterations
      from 5 to 2 (the minimum possible).
      
      Differential Revision: https://reviews.llvm.org/D72558
      04e58615