1. Jan 27, 2021
    • Valery N Dmitriev's avatar
      c8df2d1b
    • Craig Topper's avatar
      [X86] In shrinkAndImmediate, place the new constant into the topological sort. · 74784a5a
      Craig Topper authored
      Revert the change to use APInt::isSignedIntN from
      5ff5cf8e.
      
      Its clear that the games we were playing to avoid the topological
      sort aren't working. So just fix it once and for all.
      
      Fixes PR48888.
      74784a5a
    • Julian Lettner's avatar
      [NFC][lit] Cleanup code using string interpolation · 63273fc4
      Julian Lettner authored
      LLVM now requires Python 3.6, so we can use string interpolation to make
      code more readable.
      63273fc4
    • Amara Emerson's avatar
      [GlobalISel][IRTranslator] Ignore the llvm.experimental.noalias.scope.decl intrinsic. · cbed865e
      Amara Emerson authored
      These don't generate any code.
      cbed865e
    • Atmn Patel's avatar
      [OpenMP][Libomptarget] Fix cmake error on remote plugin · 810572cc
      Atmn Patel authored
      Requiring 3.15 causes a build breakage, I'm sure none of the contents actually require
      3.15 or above.
      
      Differential Revision: https://reviews.llvm.org/D95474
      810572cc
    • LLVM GN Syncbot's avatar
      [gn build] Port 1e634f39 · da9a3540
      LLVM GN Syncbot authored
      da9a3540
    • Fangrui Song's avatar
      [llvm-elfabi] Fix test after D95140 · 79ce46e2
      Fangrui Song authored
      79ce46e2
    • Jon Chesterfield's avatar
      [libomptarget][cuda] Gracefully handle missing cuda library · 7baff00e
      Jon Chesterfield authored
      [libomptarget][cuda] Gracefully handle missing cuda library
      
      If using dynamic cuda, and it failed to load, it is not safe to call
      cuGetErrorString.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95412
      7baff00e
    • Jon Chesterfield's avatar
      [libomptarget][cuda] Only run tests when sure there is cuda available · fdeffd6f
      Jon Chesterfield authored
      [libomptarget][cuda] Only run tests when sure there is cuda available
      
      Prior to D95155, building the cuda plugin implied cuda was installed locally.
      With that change, every machine can build a cuda plugin, but they won't all have
      cuda and/or an nvptx card installed locally.
      
      This change enables the nvptx tests when either:
      - libcuda is present
      - the user has forced use of the dlopen stub
      
      The default case when there is no cuda detected will no longer attempt to
      run the tests on nvptx hardware, as was the case before D95155.
      
      Reviewed By: jdoerfert, ronlieb
      
      Differential Revision: https://reviews.llvm.org/D95467
      fdeffd6f
    • Atmn Patel's avatar
      [OpenMP][Libomptarget] Introduce Remote Offloading Plugin · ec8f4a38
      Atmn Patel authored
      This introduces a remote offloading plugin for libomptarget. This
      implementation relies on gRPC and protobuf, so this library will only
      build if both libraries are available on the system. The corresponding
      server is compiled to `openmp-offloading-server`.
      
      This is a large change, but the only way to split this up is into RTL/server
      but I fear that could introduce an inconsistency amongst them.
      
      Ideally, tests for this should be added to the current ones that but that is
      problematic for at least one reason. Given that libomptarget registers plugin
      on a first-come-first-serve basis, if we wanted to offload onto a local x86
      through a different process, then we'd have to either re-order the plugin list
      in `rtl.cpp` (which is what I did locally for testing) or find a better
      solution for runtime plugin registration in libomptarget.
      
      Differential Revision: https://reviews.llvm.org/D95314
      ec8f4a38
    • Haowei Wu's avatar
      [llvm-elfabi] Support ELF file that lacks .gnu.hash section · 15313f64
      Haowei Wu authored
      Before this change, when reading ELF file, elfabi determines number of
      entries in .dynsym by reading the .gnu.hash section. This change makes
      elfabi read section headers directly first. This change allows elfabi
      works on ELF files which do not have .gnu.hash sections.
      
      Differential Revision: https://reviews.llvm.org/D93362
      15313f64
    • Louis Dionne's avatar
      [libc++] Fix oss-fuzz build · 4210b870
      Louis Dionne authored
      4210b870
    • Fangrui Song's avatar
      Add -fbinutils-version= to gate ELF features on the specified binutils version · 34b60d8a
      Fangrui Song authored
      There are two use cases.
      
      Assembler
      We have accrued some code gated on MCAsmInfo::useIntegratedAssembler().  Some
      features are supported by latest GNU as, but we have to use
      MCAsmInfo::useIntegratedAs() because the newer versions have not been widely
      adopted (e.g. SHF_LINK_ORDER 'o' and 'unique' linkage in 2.35, --compress-debug-sections= in 2.26).
      
      Linker
      We want to use features supported only by LLD or very new GNU ld, or don't want
      to work around older GNU ld. We currently can't represent that "we don't care
      about old GNU ld".  You can find such workarounds in a few other places, e.g.
      Mips/MipsAsmprinter.cpp PowerPC/PPCTOCRegDeps.cpp X86/X86MCInstrLower.cpp
      AArch64 TLS workaround for R_AARCH64_TLSLD_MOVW_DTPREL_* (PR ld/18276),
      R_AARCH64_TLSLE_LDST8_TPREL_LO12 (https://bugs.llvm.org/show_bug.cgi?id=36727 https://sourceware.org/bugzilla/show_bug.cgi?id=22969)
      
      Mixed SHF_LINK_ORDER and non-SHF_LINK_ORDER components (supported ...
      34b60d8a
    • Petr Hosek's avatar
      Revert "Support for instrumenting only selected files or functions" · 1e634f39
      Petr Hosek authored
      This reverts commit 4edf35f1 because
      the test fails on Windows bots.
      1e634f39
    • Jim Ingham's avatar
      Make SBDebugger::CreateTargetWithFileAndArch work with lldb::LLDB_DEFAULT_ARCH · 7636b1f6
      Jim Ingham authored
      Second try, handling both a bogus arch string and the "null file & arch" used
      to create an empty but valid target.
      Also check in that case before logging (previously the logging would have
      crashed.)
      7636b1f6
    • Valentin Clement's avatar
      [flang][openacc][NFC] Organize clause validity tests by directive · d2abd62b
      Valentin Clement authored
      Split the tests from acc-clause-validity.f90 in dedicated files by directives.
      The file acc-clause-validity.f90 was getting too big to be correctly maintained.
      Tests are identical.
      
      Reviewed By: SouraVX
      
      Differential Revision: https://reviews.llvm.org/D95328
      d2abd62b
    • Fangrui Song's avatar
      CGDebugInfo CreatedLimitedType: Drop file/line for RecordType with invalid location · 189f3111
      Fangrui Song authored
      For Clang synthesized `__va_list_tag` (`CreateX86_64ABIBuiltinVaListDecl`),
      its DW_AT_decl_file/DW_AT_decl_line are arbitrarily set from `CurLoc`.
      
      In a stage 2 `-DCMAKE_BUILD_TYPE=Debug` clang build, I observe that
      in driver.cpp, DW_AT_decl_file/DW_AT_decl_line may be set to an `#include` line
      (the transitively included file uses va_arg (`__builtin_va_arg`)).
      This seems arbitrary. Drop that.
      
      Reviewed By: #debug-info, dblaikie
      
      Differential Revision: https://reviews.llvm.org/D94735
      189f3111
    • Fangrui Song's avatar
      CGDebugInfo: Drop Loc.isInvalid() special case from getLineNumber · 31d375f1
      Fangrui Song authored
      `getLineNumber()` picks CurLoc if the parameter is invalid. This appears to
      mainly work around missing SourceLocation information for some constructs, but
      sometimes adds unintended locations.
      
      * For `CodeGenObjC/debug-info-blocks.m`, `CurLoc` has been advanced to the closing brace. The debug line of `ImplicitVarParameter` is set to the line of `}` because this implicit parameter has an invalid `SourceLocation`. The debug line is a bit arbitrary - perhaps the location of `^{` is better.
      * The file/line of Clang synthesized `__va_list_tag` is arbitrarily attached a `#include` line. D94735
      
      Drop the special case to make getLineNumber less magic and add CurLoc fallback in its callers instead.
      
      Tested with stage 2 -DCMAKE_BUILD_TYPE=Debug clang, byte identical.
      
      Reviewed By: #debug-info, aprantl
      
      Differential Revision: https://reviews.llvm.org/D94391
      31d375f1
    • Austin Kerbow's avatar
      [AMDGPU] Update subtarget features for new target ID support · 2291bd13
      Austin Kerbow authored
      Support for XNACK and SRAMECC is not static on some GPUs. We must be able
      to differentiate between different scenarios for these dynamic subtarget
      features.
      
      The possible settings are:
      
      - Unsupported: The GPU has no support for XNACK/SRAMECC.
      - Any: Preference is unspecified. Use conservative settings that can run anywhere.
      - Off: Request support for XNACK/SRAMECC Off
      - On: Request support for XNACK/SRAMECC On
      
      GCNSubtarget will track the four options based on the following criteria. If
      the subtarget does not support XNACK/SRAMECC we say the setting is
      "Unsupported". If no subtarget features for XNACK/SRAMECC are requested we
      must support "Any" mode. If the subtarget features XNACK/SRAMECC exist in the
      feature string when initializing the subtarget, the settings are "On/Off".
      
      The defaults are updated to be conservatively correct, meaning if no setting
      for XNACK or SRAMECC is explicitly requested, defaults will be used which
      generate code that can be run anywhere. This corresponds to the "Any" setting.
      
      Differential Revision: https://reviews.llvm.org/D85882
      2291bd13
    • Atmn's avatar
      [OpenMP][Libomptarget] Introduce changes to support remote plugin · 683719bc
      Atmn authored
      In order to support remote execution, we need to be able to send the
      target binary description to the remote host for registration (and
      consequent deregistration). To support this, I added these two
      optional new functions to the plugin API:
      - `__tgt_rtl_register_lib`
      - `__tgt_rtl_unregister_lib`
      
      These functions will be called to properly manage the instance of
      libomptarget running on the remote host.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D93293
      683719bc
    • LLVM GN Syncbot's avatar
      [gn build] Port 4edf35f1 · 96f09aa2
      LLVM GN Syncbot authored
      96f09aa2
    • Petr Hosek's avatar
      Support for instrumenting only selected files or functions · 4edf35f1
      Petr Hosek authored
      This change implements support for applying profile instrumentation
      only to selected files or functions. The implementation uses the
      sanitizer special case list format to select which files and functions
      to instrument, and relies on the new noprofile IR attribute to exclude
      functions from instrumentation.
      
      Differential Revision: https://reviews.llvm.org/D94820
      4edf35f1
    • Jon Chesterfield's avatar
    • Nathan James's avatar
      [clangd] FindTarget resolves base specifier · 7730599c
      Nathan James authored
      FindTarget on the virtual keyword or access specifier of a base specifier will now resolve to type of the base specifier.
      
      Reviewed By: sammccall
      
      Differential Revision: https://reviews.llvm.org/D95338
      7730599c
    • Nathan James's avatar
      [clangd] Selection handles CXXBaseSpecifier · d92413a4
      Nathan James authored
      Selection now includes the virtual and access modifier as part of their range for cxx base specifiers.
      
      Reviewed By: sammccall
      
      Differential Revision: https://reviews.llvm.org/D95231
      d92413a4
    • Adhemerval Zanella's avatar
      [ARM] [ELF] Fix ARMMaterializeGV for Indirect calls · dad55c22
      Adhemerval Zanella authored
      Recent shouldAssumeDSOLocal changes (introduced by 961f31d8)
      do not take in consideration the relocation model anymore.  The ARM
      fast-isel pass uses the function return to set whether a global symbol
      is loaded indirectly or not, and without the expected information
      llvm now generates an extra load for following code:
      
      ```
      $ cat test.ll
      @__asan_option_detect_stack_use_after_return = external global i32
      define dso_local i32 @main(i32 %argc, i8** %argv) #0 {
      entry:
        %0 = load i32, i32* @__asan_option_detect_stack_use_after_return,
      align 4
        %1 = icmp ne i32 %0, 0
        br i1 %1, label %2, label %3
      
      2:
        ret i32 0
      
      3:
        ret i32 1
      }
      
      attributes #0 = { noinline optnone }
      
      $ lcc test.ll -o -
      [...]
      main:
              .fnstart
      [...]
              movw    r0, :lower16:__asan_option_detect_stack_use_after_return
              movt    r0, :upper16:__asan_option_detect_stack_use_after_return
              ldr     r0, [r0]
              ldr     r0, [r0]
              cmp     r0, #0
      [...]
      ```
      
      And without 'optnone' it produces:
      ```
      [...]
      main:
              .fnstart
      [...]
              movw    r0, :lower16:__asan_option_detect_stack_use_after_return
              movt    r0, :upper16:__asan_option_detect_stack_use_after_return
              ldr     r0, [r0]
              clz     r0, r0
              lsr     r0, r0, #5
              bx      lr
      
      [...]
      ```
      
      This triggered a lot of invalid memory access in sanitizers for
      arm-linux-gnueabihf.  I checked this patch both a stage1 built with
      gcc and a stage2 bootstrap and it fixes all the Linux sanitizers
      issues.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D95379
      dad55c22
    • Craig Topper's avatar
      [RISCV] Have customLegalizeToWOp truncate to the original type instead of i32... · f9d7f772
      Craig Topper authored
      [RISCV] Have customLegalizeToWOp truncate to the original type instead of i32 now that we use it for i8/i16 as well.
      
      239cfbcc add support for legalizing
      i8/i16 UDIV/UREM/SDIV to use *W instructions. So we need to truncate
      to i8/i16 if we're legalizing one of those.
      f9d7f772
    • Eric Schweitz's avatar
      [mlir] sret and byval now require a type argument when constructed. · 1d6df1fc
      Eric Schweitz authored
      Fixes the LLVM code gen bugs and adds the missing tests.
      
      Reviewed By: ftynse
      
      Differential Revision: https://reviews.llvm.org/D95378
      1d6df1fc
    • Julian Lettner's avatar
      Reland "[lit] Use os.cpu_count() to cleanup TODO" · 302432f7
      Julian Lettner authored
      The initial problem with the remaining bot config was resolved.
      
      We can now use Python3.  Let's use `os.cpu_count()` to cleanup this
      helper.
      
      Differential Revision: https://reviews.llvm.org/D94734
      302432f7
    • Raphael Isemann's avatar
      [lldb][NFC] Another attempt to fix GCC 5.x compilation · 48e09faa
      Raphael Isemann authored
      37510f69 tried to fix GCC 5.x compilation
      by making the enum which is used as a unordered_map key unscoped. However it
      seems that in GCC 5.x, enum keys are not supported *at all* in unordered_maps
      (at least that's what some trial&error on godbolt tells me). This updates the
      workaround to just use an int until GCC 5.x support is dropped.
      48e09faa
    • Christian Sigg's avatar
      [mlir] Set CUDA/ROCm context before creating resources. · 8262cd8a
      Christian Sigg authored
      The current context is thread-local state, and in preparation of GPU async execution (on multiple threads) we need to set the context before calling API that create resources.
      
      Reviewed By: herhut
      
      Differential Revision: https://reviews.llvm.org/D94495
      8262cd8a
    • Matt Arsenault's avatar
      AMDGPU: Fix redundant FP spilling/assert in some functions · 5f9707b7
      Matt Arsenault authored
      If a function has stack objects, and a call, we require an FP. If we
      did not initially have any stack objects, and only introduced them
      during PrologEpilogInserter for CSR VGPR spills, SILowerSGPRSpills
      would end up spilling the FP register as if it were a normal
      register. This would result in an assert in a debug build, or
      redundant handling of the FP register in a release build.
      
      Try to predict that we will have an FP later, although this is ugly.
      5f9707b7
    • Matt Arsenault's avatar
      AMDGPU: Add assertion to determineCalleeSaves · 92d1195b
      Matt Arsenault authored
      Make sure this isn't getting called multiple times. I was surprised we
      were modifying the function here, which I think is a bit questionable.
      92d1195b
    • Shilei Tian's avatar
      [OpenMP][deviceRTLs] Build the deviceRTLs with OpenMP instead of target dependent language · 7c03f7d7
      Shilei Tian authored
      From this patch (plus some landed patches), `deviceRTLs` is taken as a regular OpenMP program with just `declare target` regions. In this way, ideally, `deviceRTLs` can be written in OpenMP directly. No CUDA, no HIP anymore. (Well, AMD is still working on getting it work. For now AMDGCN still uses original way to compile) However, some target specific functions are still required, but they're no longer written in target specific language. For example, CUDA parts have all refined by replacing CUDA intrinsic and builtins with LLVM/Clang/NVVM intrinsics.
      Here're a list of changes in this patch.
      1. For NVPTX, `DEVICE` is defined empty in order to make the common parts still work with AMDGCN. Later once AMDGCN is also available, we will completely remove `DEVICE` or probably some other macros.
      2. Shared variable is implemented with OpenMP allocator, which is defined in `allocator.h`. Again, this feature is not available on AMDGCN, so two macros are redefined properly.
      3. CUDA header `cuda.h` is dropped in the source code. In order to deal with code difference in various CUDA versions, we build one bitcode library for each supported CUDA version. For each CUDA version, the highest PTX version it supports will be used, just as what we currently use for CUDA compilation.
      4. Correspondingly, compiler driver is also updated to support CUDA version encoded in the name of bitcode library. Now the bitcode library for NVPTX is named as `libomptarget-nvptx-cuda_[cuda_version]-sm_[sm_number].bc`, such as `libomptarget-nvptx-cuda_80-sm_20.bc`.
      
      With this change, there are also multiple features to be expected in the near future:
      1. CUDA will be completely dropped when compiling OpenMP. By the time, we also build bitcode libraries for all supported SM, multiplied by all supported CUDA version.
      2. Atomic operations used in `deviceRTLs` can be replaced by `omp atomic` if OpenMP 5.1 feature is fully supported. For now, the IR generated is totally wrong.
      3. Target specific parts will be wrapped into `declare variant` with `isa` selector if it can work properly. No target specific macro is needed anymore.
      4. (Maybe more...)
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D94745
      7c03f7d7
    • Dave Lee's avatar
      [lldb] Remove unused ThreadPlanStack::GetStackOfKind (NFC) · 90b8ae01
      Dave Lee authored
      This function isn't used.
      
      Differential Revision: https://reviews.llvm.org/D95411
      90b8ae01
    • Kadir Cetinkaya's avatar
      [clangd] Add std::size_t to StdSymbol mapping · 9190f17a
      Kadir Cetinkaya authored
      This is a common symbol that's missing from our mapping because
      cppreference yields multiple headers.
      
      Add it manually by picking cstddef to prevent insertion of some stdlib-internal
      headers instead.
      
      Fixes https://github.com/clangd/clangd/issues/666.
      
      Differential Revision: https://reviews.llvm.org/D95423
      9190f17a
    • Alex Zinenko's avatar
      [mlir] Add Python bindings for IntegerSet · b208e5bc
      Alex Zinenko authored
      This follows up on the introduction of C API for the same object and is similar
      to AffineExpr and AffineMap.
      
      Reviewed By: stellaraccident
      
      Differential Revision: https://reviews.llvm.org/D95437
      b208e5bc
    • Sanjay Patel's avatar
      [LoopVectorize] add test for fmin/fmax FMF propagation; NFC · 00773ef7
      Sanjay Patel authored
      The existing test has less FMF than we might expect if
      our FMF was fixed (on all FP values), so this additional
      test is intended to check propagation in a more "normal"
      example.
      00773ef7
    • Sanjay Patel's avatar
      [LoopUtils] do not initialize Cmp predicate unnecessarily; NFC · 09b1c563
      Sanjay Patel authored
      The switch must set the predicate correctly; anything else
      should lead to unreachable/assert.
      
      I'm trying to fix FMF propagation here and the callers,
      so this is a preliminary cleanup.
      09b1c563
    • Simon Pilgrim's avatar
      Fix null dereference static analysis warning. NFCI. · 879c12d9
      Simon Pilgrim authored
      Replace cast_or_null<> with cast<> as we immediately dereference the pointer afterward so we're not expecting a null pointer.
      879c12d9