1. Jul 14, 2022
    • gbreynoo's avatar
      Revert "[llvm-ar][test] Add testing for bitcode file handling" · 8564b2ab
      gbreynoo authored
      This reverts commit 264b9a48.
      
      Due to build bot test failure.
      8564b2ab
    • Simon Moll's avatar
      [VP] Add test to show optimization opportunities · 173d4b84
      Simon Moll authored
      Add vp.add test cases that can are optimized with D92086 to show the
      potential of generalized pattern rewriting.
      
      Reviewed By: frasercrmck
      
      Differential Revision: https://reviews.llvm.org/D129746
      173d4b84
    • Michał Górny's avatar
      [lldb] [gdb-remote] Remove stray GetSupportsThreadSuffix() method (NFC) · c164efb0
      Michał Górny authored
      Remove stray GDBRemoteCommunicationClient::GetSupportsThreadSuffix()
      method that is not implemented nor used anywhere.
      c164efb0
    • gbreynoo's avatar
      [llvm-ar][test] Add testing for bitcode file handling · 264b9a48
      gbreynoo authored
      This change adds testing for handling of bitcode files in archives,
      particularly the creation of symbol tables and through MRI scripts.
      Although there is some testing of bitcode handling in the archive
      library testing, this was not covered.
      
      Differential Revision: https://reviews.llvm.org/D129088
      264b9a48
    • Cullen Rhodes's avatar
      Revert "[ORC] Add a shared-memory based orc::MemoryMapper." · 3e9cc543
      Cullen Rhodes authored
      This reverts commit 5acd4716.
      
      Breaks shared library build with:
      
        ld.lld-12: error: undefined symbol: shm_open
        >>> referenced by ExecutorSharedMemoryMapperService.cpp:68
        (/home/culrho01/llvm-project/llvm/lib/ExecutionEngine/Orc/TargetProcess/ExecutorSharedMemoryMapperService.cpp:68)
        >>>
        lib/ExecutionEngine/Orc/TargetProcess/CMakeFiles/LLVMOrcTargetProcess.dir/ExecutorSharedMemoryMapperService.cpp.o:(llvm::orc::rt_bootstrap::ExecutorSharedMemoryMapperService::reserve[abi:cxx11](unsigned
        long))
        >>> did you mean: sem_open
        >>> defined in:
        /usr/bin/../lib/gcc/aarch64-linux-gnu/9/../../../aarch64-linux-gnu/libpthread.so
      3e9cc543
    • Cullen Rhodes's avatar
      Revert "[ORC] Fix compilation on mingw" · f3eacb4f
      Cullen Rhodes authored
      This reverts commit 46b1a7c5.
      
      Parent commit breaks shared library build, reverting both commits.
      f3eacb4f
    • Jay Foad's avatar
      [AMDGPU] Update LiveVariables after killing an immediate def · e45aa230
      Jay Foad authored
      D114999 added code to kill an immediate def if it was folded into its
      only use by convertToThreeAddress. This patch updates LiveVariables when
      that happens in order to fix verification failures exposed by D129213.
      
      Differential Revision: https://reviews.llvm.org/D129661
      e45aa230
    • Nikita Popov's avatar
      [IndVars] Make sure header phi simplification preserves LCSSA form · 7a43b382
      Nikita Popov authored
      When simplifying instructions, make sure that the replacement
      preserves LCSSA form. This fixes the issue reported at:
      https://reviews.llvm.org/D129293#3650851
      7a43b382
    • Fraser Cormack's avatar
      [RISCV] Add a test showing a miscompilation with subreg liveness · 3b334978
      Fraser Cormack authored
      This patch adds a test which shows that we may incorrectly register
      allocate for RVV instructions which have no-overlap constraints on
      source/dest registers of different LMUL groups.
      
      The particular case shows that a vrgatherei16 instruction writes to a
      LMUL=1 register group v11 and reads from an EMUL=2 register group
      v10/v11. This breaks the overlap constraints of the vrgatherei16
      instruction.
      
      The test also shows that disabling subregister liveness fixes the test.
      
      We use `early-clobber` on the `VR` dest and the `VRM2` source to enforce
      the constraint but with subregister liveness this constraint is not met.
      
      It's unclear to me at this point whether this is per-design of
      early-clobber in conjunction with subregisters (meaning we should find
      another way of expressing this constraint) or whether it's a bug in the
      register allocator somewhere.
      
      Reviewed By: rogfer01
      
      Differential Revision: https://reviews.llvm.org/D129639
      3b334978
    • Cullen Rhodes's avatar
      [AArch64][NFC] Drop 'V' from ASIMD FP convert, other, D/Q-form regex · 055b409c
      Cullen Rhodes authored
      In the Cortex A57 Optimization Guide [1] VCVTAU (AArch32) is incorrectly
      listed in the AArch64 instructions for instruction groups:
      
        - ASIMD FP convert, other, D-form
        - ASIMD FP convert, other, Q-form
      
      It's meant to be FCVTAU, this will be fixed in future releases of the guide.
      
      [1] https://developer.arm.com/documentation/uan0015/b
      055b409c
    • Cullen Rhodes's avatar
      [NFC][SVE] Add tests for zext(cmpeq(x, splat(0))) · a2fe6aa9
      Cullen Rhodes authored
      In preparation for follow up patch folding above to CNOT.
      
      Reviewed By: paulwalker-arm, peterwaller-arm
      
      Differential Revision: https://reviews.llvm.org/D129625
      a2fe6aa9
    • Weining Lu's avatar
      9b87ad33
    • Ingo Müller's avatar
      [mlir][doc] Fix usage of PatternApplicator. · d4a7ca81
      Ingo Müller authored
      The constructor of PatternApplicator doesn't have a constructor that
      accepts only a `RewritePatternSet` as currently used in the example
      code in PatternRewriter.md. Instead, one has to turn it into a
      `FrozenRewritePatternSet`.
      
      Reviewed By: nicolasvasilache
      
      Differential Revision: https://reviews.llvm.org/D125236
      d4a7ca81
    • Martin Storsjö's avatar
      [ORC] Fix compilation on mingw · 46b1a7c5
      Martin Storsjö authored
      Explicitly call the -W suffixed API functions when passing wchar based
      strings.
      46b1a7c5
    • Nikita Popov's avatar
      [SCCP] Make check for unknown/undef in unary op handling more explicit (NFCI) · ebc54e0c
      Nikita Popov authored
      Make the implementation more similar to other functions, by
      explicitly skipping an unknown/undef first, and always falling
      back to overdefined at the end. I don't think it makes a difference
      now, but could make one once the constant evaluation can fail. In
      that case we would directly mark the result as overdefined now,
      rather than keeping it unknown (and later making it overdefined
      because we think it's undef-based).
      ebc54e0c
    • David Green's avatar
      [CodeGen] Move instruction predicate verification to emitInstruction · 3e0bf1c7
      David Green authored
      D25618 added a method to verify the instruction predicates for an
      emitted instruction, through verifyInstructionPredicates added into
      <Target>MCCodeEmitter::encodeInstruction. This is a very useful idea,
      but the implementation inside MCCodeEmitter made it only fire for object
      files, not assembly which most of the llvm test suite uses.
      
      This patch moves the code into the <Target>_MC::verifyInstructionPredicates
      method, inside the InstrInfo.  The allows it to be called from other
      places, such as in this patch where it is called from the
      <Target>AsmPrinter::emitInstruction methods which should trigger for
      both assembly and object files. It can also be called from other places
      such as verifyInstruction, but that is not done here (it tends to catch
      errors earlier, but in reality just shows all the mir tests that have
      incorrect feature predicates). The interface was also simplified
      slightly, moving computeAvailableFeatures into the function so that it
      does not need to be called externally.
      
      The ARM, AMDGPU (but not R600), AVR, Mips and X86 backends all currently
      show errors in the test-suite, so have been disabled with FIXME
      comments.
      
      Recommitted with some fixes for the leftover MCII variables in release
      builds.
      
      Differential Revision: https://reviews.llvm.org/D129506
      3e0bf1c7
    • Fangrui Song's avatar
      [CommandLine] --help: print "-o <xxx>" instead of "-o=<xxx>" · 52cb9725
      Fangrui Song authored
      Accepting -o= is a quirk of CommandLine. For --help, we should print the
      conventional "-o <xxx>".
      52cb9725
    • Amara Emerson's avatar
      Revert "[llvm] add zstd to llvm::compression namespace" · 6e6be5f9
      Amara Emerson authored
      This reverts commit d449c600.
      
      Breaks macOS builds with this:
      llvm/lib/Support/Compression.cpp:24:10: fatal error: 'zstd.h' file not found
      6e6be5f9
    • Nikita Popov's avatar
      [SCCP] Don't check for UndefValue before calling markConstant() · 6db3edc8
      Nikita Popov authored
      The value lattice explicitly represents undef, and markConstant()
      internally checks for UndefValue and will create an undef rather
      than constant lattice element in that case.
      
      This is mostly a code simplification, it has little practical impact
      because we usually get undef results from undef operands, and those
      don't get processed.
      
      Only leave the check behind for the CmpInst case, because it
      currently goes through this incorrect code in the getCompare()
      implementation: https://github.com/llvm/llvm-project/blob/f98697642cea761448dc0f84f750d3f5def8af6b/llvm/include/llvm/Analysis/ValueLattice.h#L456-L457
      
      Differential Revision: https://reviews.llvm.org/D128330
      6db3edc8
    • Amara Emerson's avatar
      [GlobalISel] Re-generate some checks. · cef349a3
      Amara Emerson authored
      cef349a3
    • Jason Molenda's avatar
      jGetLoadedDynamicLibrariesInfos can inspect machos not yet loaded · ac49e902
      Jason Molenda authored
      jGetLoadedDynamicLibrariesInfos normally checks with dyld to find
      the list of binaries loaded in the inferior, and getting the filepath,
      before trying to parse the Mach-O binary in inferior memory.
      This allows for debugserver to parse a Mach-O binary present in memory,
      but not yet registered with dyld.  This patch also adds some simple
      sanity checks that we're reading a Mach-O header before we begin
      stepping through load commands, because we won't have the sanity check
      of consulting dyld for the list of loaded binaries before parsing.
      Also adds a testcase.
      
      [This patch was reverted after causing a testsuite failure on a CI bot;
      I haven't been able to repro the failure outside the CI, but I have a
      theory that my sanity check on cputype which only matched arm64 and
      x86_64 - and the CI machine may have a watch simulator that is still
      using i386.]
      
      Differential Revision: https://reviews.llvm.org/D128956
      rdar://95737734
      ac49e902
    • Matthias Springer's avatar
      [mlir][sparse] Switch to One-Shot Bufferize · c66303c2
      Matthias Springer authored
      This change removes the partial bufferization passes from the sparse compilation pipeline and replaces them with One-Shot Bufferize. One-Shot Analysis (and TensorCopyInsertion) is used to resolve all out-of-place bufferizations, dense and sparse. Dense ops are then bufferized with BufferizableOpInterface. Sparse ops are still bufferized in the Sparsification pass.
      
      Details:
      * Dense allocations are automatically deallocated, unless they are yielded from a block. (In that case the alloc would leak.) All test cases are modified accordingly. E.g., some funcs now have an "out" tensor argument that is returned from the function. (That way, the allocation happens at the call site.)
      * Sparse allocations are *not* automatically deallocated. They must be "released" manually. (No change, this will be addressed in a future change.)
      * Sparse tensor copies are not supported yet. (Future change)
      * Sparsification no longer has to consider inplacability. If necessary, allocations and/or copies are inserted during TensorCopyInsertion. All tensors are inplaceable by the time Sparsification is running. Instead of marking a tensor as "not inplaceable", it can be marked as "not writable", which will trigger an allocation and/or copy during TensorCopyInsertion.
      
      Differential Revision: https://reviews.llvm.org/D129356
      c66303c2
    • Jannik Silvanus's avatar
      [AMDGPU] SIMachineScheduler: Add support for several MachineScheduler features · e5c4cde4
      Jannik Silvanus authored
      The SI machine scheduler inherits from ScheduleDAGMI.
      This patch adds support for a few features that are implemented
      in ScheduleDAGMI (or its base classes) that were missing so far
      because their support is implemented in overridden functions.
      
      * Support cl::opt -view-misched-dags
        This option allows to open a graphical window of the scheduling DAG.
      
      * Support cl::opt -misched-print-dags
        This option allows to print the scheduling DAG in text form.
      
      * After constructing the scheduling DAG, call postprocessDAG()
        to apply any registered DAG mutations.
        Note that currently there are no mutations defined in AMDGPUTargetMachine.cpp
        in case SIScheduler is used.
        Still add this to avoid surprises in the future in case mutations are added.
      
      Differential Revision: https://reviews.llvm.org/D128808
      e5c4cde4
    • Fangrui Song's avatar
      [obj2yaml] Add -o to specify output filename · cfec2080
      Fangrui Song authored
      -o is very common among tools. yaml2obj supports -o and it surprised me that
      obj2yaml doesn't support -o. Just add it which doesn't take much code.
      
      Differential Revision: https://reviews.llvm.org/D129713
      cfec2080
    • Shoaib Meenai's avatar
      [clang] Add missing header include · 42b3a5fb
      Shoaib Meenai authored
      With my version of the MSVC tools (14.11.25503), this was failing to
      build because of missing declarations of `std::isalnum` and
      `std::isdigit`. Include `<cctype>` to get these.
      42b3a5fb
    • Kazu Hirata's avatar
      [mlir] Use value instead of getValue (NFC) · c27d8152
      Kazu Hirata authored
      c27d8152
    • Balázs Kéri's avatar
      [clang-tidy] Improve check cert-dcl58-cpp. · 0e95921b
      Balázs Kéri authored
      Detect template specializations that should be handled specially.
      In some cases it is allowed to extend the `std` namespace with
      template specializations.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D129353
      0e95921b
    • Kazu Hirata's avatar
      [clang] Use value instead of getValue (NFC) · cb2c8f69
      Kazu Hirata authored
      cb2c8f69
    • Huan Nguyen's avatar
      [BOLT] Support multiple parents for split jump table · 05523dc3
      Huan Nguyen authored
      There are two assumptions regarding jump table:
      (a) It is accessed by only one fragment, say, Parent
      (b) All entries target instructions in Parent
      
      For (a), BOLT stores jump table entries as relative offset to Parent.
      For (b), BOLT treats jump table entries target somewhere out of Parent
      as INVALID_OFFSET, including fragment of same split function.
      
      In this update, we extend (a) and (b) to include fragment of same split
      functinon. For (a), we store jump table entries in absolute offset
      instead. In addition, jump table will store all fragments that access
      it. A fragment uses this information to only create label for jump table
      entries that target to that fragment.
      
      For (b), using absolute offset allows jump table entries to target
      fragments of same split function, i.e., extend support for split jump
      table. This can be done using relocation (fragment start/size) and
      fragment detection heuristics (e.g., using symbol name pattern for
      non-stripped binaries).
      
      For jump table targets that can only be reached by one fragment, we
      mark them as local label; otherwise, they would be the secondary
      function entry to the target fragment.
      
      Test Plan
      ```
      ninja check-bolt
      ```
      
      Reviewed By: Amir
      
      Differential Revision: https://reviews.llvm.org/D128474
      05523dc3
    • Kazu Hirata's avatar
      [llvm] Use value instead of getValue (NFC) · 611ffcf4
      Kazu Hirata authored
      611ffcf4
    • Corentin Jabot's avatar
      [Clang] Adjust extension warnings for delimited sequences · 6882ca9a
      Corentin Jabot authored
      WG21 approved delimited escape sequences and named escape
      sequences.
      Adjust the extension warnings accordingly, and update
      the release notes.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D129664
      6882ca9a
    • owenca's avatar
      [llvm] Make lib/Target/BPF/BTF.h self-contained · 18a910df
      owenca authored
      18a910df
    • Zi Xuan Wu (Zeson)'s avatar
      [CSKY] Fix the br target operand type in td · 033324db
      Zi Xuan Wu (Zeson) authored
      br target operand should be Operand<OtherVT> type instead of Operand<iPTR>
      033324db
    • Cole Kissane's avatar
      [llvm] add zstd to llvm::compression namespace · d449c600
      Cole Kissane authored
      - add `FindZSTD.cmake`
      - add zstd to `llvm::compression` namespace
      - add a CMake option `LLVM_ENABLE_ZSTD` with behavior mirroring that of `LLVM_ENABLE_ZLIB`
      - add tests for zstd to `llvm/unittests/Support/CompressionTest.cpp`
      
      Reviewed By: leonardchan, MaskRay
      
      Differential Revision: https://reviews.llvm.org/D128465
      d449c600
    • Cole Kissane's avatar
      Revert "[llvm] add zstd to `llvm::compression` namespace" · 5ecb161c
      Cole Kissane authored
      This reverts commit cef07169.
      5ecb161c
    • Alexander Potapenko's avatar
      [compiler-rt][hwasan] Support for new Intel LAM API · b191056f
      Alexander Potapenko authored
      New version of Intel LAM patches
      (https://lore.kernel.org/linux-mm/20220712231328.5294-1-kirill.shutemov@linux.intel.com/)
      uses a different interface based on arch_prctl():
       - arch_prctl(ARCH_GET_UNTAG_MASK, &mask) returns the current mask for
         untagging the pointers. We use it to detect kernel LAM support.
       - arch_prctl(ARCH_ENABLE_TAGGED_ADDR, nr_bits) enables pointer tagging
         for the current process.
      
      Because __NR_arch_prctl is defined in different headers, and no other
      platforms need it at the moment, we only declare internal_arch_prctl()
      on x86_64.
      
      Reviewed By: vitalybuka
      
      Differential Revision: https://reviews.llvm.org/D129645
      b191056f
    • Cole Kissane's avatar
      [llvm] add zstd to `llvm::compression` namespace · cef07169
      Cole Kissane authored
      - add `FindZSTD.cmake`
      - add zstd to `llvm::compression` namespace
      - add a CMake option `LLVM_ENABLE_ZSTD` with behavior mirroring that of `LLVM_ENABLE_ZLIB`
      - add tests for zstd to `llvm/unittests/Support/CompressionTest.cpp`
      
      Reviewed By: leonardchan, MaskRay
      
      Differential Revision: https://reviews.llvm.org/D128465
      cef07169
    • Florian Hahn's avatar
    • Joseph Huber's avatar
      [CUDA] Allow the new driver to compile CUDA in non-RDC mode · b370be37
      Joseph Huber authored
      The new driver primarily allows us to support RDC-mode compilations with
      proper linking. This is not needed for non-RDC mode compilation, but we
      still would like the new driver to be able to handle this mode so we can
      transition away from the old driver in the future. This patch adds the
      necessary code to support creating a fatbinary for CUDA code generation
      as well as removing old assumptions and errors about RDC-mode with the
      new driver.
      
      Reviewed By: tra
      
      Differential Revision: https://reviews.llvm.org/D129655
      b370be37
    • Jez Ng's avatar
      [lld-macho] Enable EH frame relocation / pruning · 403d61ae
      Jez Ng authored
      This just removes the code that gates the logic. The main issue here is
      perf impact: without {D122258}, LLD takes a significant perf hit because
      it now has to do a lot more work in the input parsing phase. But with
      that change to eliminate unnecessary EH frames from input object files,
      the perf overhead here is minimal. Concretely, here are the numbers for
      some builds as measured on my 16-core Mac Pro:
      
      **chromium_framework**
      
      This is without the use of `-femit-dwarf-unwind=no-compact-unwind`:
      
                   base           diff           difference (95% CI)
        sys_time   1.826 ± 0.019  1.962 ± 0.034  [  +6.5% ..   +8.4%]
        user_time  9.306 ± 0.054  9.926 ± 0.082  [  +6.2% ..   +7.1%]
        wall_time  8.225 ± 0.068  8.947 ± 0.128  [  +8.0% ..   +9.6%]
        samples    15             22
      
      With that flag enabled, the regression mostly disappears, as hoped:
      
                   base           diff           difference (95% CI)
        sys_time   1.839 ± 0.062  1.866 ± 0.068  [  -0.9% ..   +3.8%]
        user_time  9.452 ± 0.068  9.490 ± 0.067  [  -0.1% ..   +0.9%]
        wall_time  8.383 ± 0.127  8.452 ± 0.114  [  -0.1% ..   +1.8%]
        samples    17             21
      
      **Unnamed internal app**
      
      Without `-femit-dwarf-unwind`, this is the perf hit:
      
                   base           diff           difference (95% CI)
        sys_time   1.372 ± 0.029  1.317 ± 0.024  [  -4.6% ..   -3.5%]
        user_time  2.835 ± 0.028  2.980 ± 0.027  [  +4.8% ..   +5.4%]
        wall_time  3.205 ± 0.079  3.383 ± 0.066  [  +4.9% ..   +6.2%]
        samples    102            83
      
      With `-femit-dwarf-unwind`, the perf hit almost disappears:
      
                   base           diff           difference (95% CI)
        sys_time   1.274 ± 0.026  1.270 ± 0.025  [  -0.9% ..   +0.3%]
        user_time  2.812 ± 0.023  2.822 ± 0.035  [  +0.1% ..   +0.7%]
        wall_time  3.166 ± 0.047  3.174 ± 0.059  [  -0.2% ..   +0.7%]
        samples    95             97
      
      Just for fun, I measured the impact of `-femit-dwarf-unwind` on ld64
      (`base` has the extra DWARF unwind info in the input object files,
      `diff` doesn't):
      
                   base           diff           difference (95% CI)
        sys_time   1.128 ± 0.010  1.124 ± 0.023  [  -1.3% ..   +0.6%]
        user_time  7.176 ± 0.030  7.106 ± 0.094  [  -1.5% ..   -0.4%]
        wall_time  7.874 ± 0.041  7.795 ± 0.121  [  -1.7% ..   -0.3%]
        samples    16             25
      
      And for LLD:
      
                   base           diff           difference (95% CI)
        sys_time   1.315 ± 0.019  1.280 ± 0.019  [  -3.2% ..   -2.0%]
        user_time  2.980 ± 0.022  2.822 ± 0.016  [  -5.5% ..   -5.0%]
        wall_time  3.369 ± 0.038  3.175 ± 0.033  [  -6.2% ..   -5.3%]
        samples    47             47
      
      So parsing the extra EH frames is a lot more expensive for us than for
      ld64. But given that we are quite a lot faster than ld64 to begin with,
      I guess this isn't entirely unexpected...
      
      Reviewed By: #lld-macho, oontvoo
      
      Differential Revision: https://reviews.llvm.org/D129540
      403d61ae