1. Jun 24, 2023
    • Matt Arsenault's avatar
    • Matt Arsenault's avatar
      clang: Add __builtin_elementwise_rint and nearbyint · 9d84f8dc
      Matt Arsenault authored
      These are basically the same thing and only differ for strictfp,
      so add both for future proofing. Note all the elementwise functions are
      currently broken for strictfp, and use non-constrained ops. Add a test
      that demonstrates this, but doesn't attempt to fix it.
      9d84f8dc
    • Hongtao Yu's avatar
      [Pseudo Probe] Remove the assert of allowing only one call probe for a callsite. · abe34ce4
      Hongtao Yu authored
      Compiler-generated static symbols, such as the global initializers, can shared the same name and can coexist in the binary. As a result, their pseudo probes are all kept in the binary too. This could cause multiple call probes decoded against one callsite, as probes are decoded against there owning functions by name. I'm temporarily disabling an assert to keep the debug build green until we have a better fix.
      
      Reviewed By: wenlei
      
      Differential Revision: https://reviews.llvm.org/D153588
      abe34ce4
    • Matt Arsenault's avatar
    • Johannes Doerfert's avatar
      [Attributor][NFC] Allow to restrict the Attributor to cached passes · 767e429a
      Johannes Doerfert authored
      If the user wants to avoid running additional passes, they can now
      initialize the AnalysisGetter accordingly.
      767e429a
    • Johannes Doerfert's avatar
      [Attributor] Remove the iteration count verification · 23dafbb1
      Johannes Doerfert authored
      It was never really useful to track #iterations, though it helped during
      the initial development. What we should track, in a follow up, are
      potentially #updates. That is also what we should restrict instead of
      the #iterations.
      23dafbb1
    • Matt Arsenault's avatar
      ValueTracking: Use new version of cannotBeOrderedLessThanZero · 13cf479d
      Matt Arsenault authored
      Pass all arguments so now assumes work.
      13cf479d
    • Matt Arsenault's avatar
      ValueTracking: Prepare to phase out CannotBeOrderedLessThanZero · 95209e37
      Matt Arsenault authored
      Introduce a full featured wrapper around computeKnownFPClass
      to start replacing the uses with.
      95209e37
    • Paul Kirth's avatar
      Reland [llvm] Preliminary fat-lto-objects support · 44265dc3
      Paul Kirth authored
      Fat LTO objects contain both LTO compatible IR, as well as generated
      object code. This allows users to defer the choice of whether to use LTO
      or not to link-time. This is a feature available in GCC for some time,
      and makes the existing -ffat-lto-objects flag functional in the same
      way as GCC's.
      
      Within LLVM, we add a new EmbedBitcodePass that serializes the module to
      the object file, and expose a new pass pipeline for compiling fat
      objects. The new pipeline initially clones the module and runs the
      selected (Thin)LTOPrelink pipeline, after which it will serialize the
      module into a `.llvm.lto` section of an ELF file. When compiling for
      (Thin)LTO, this normally the point at which the compiler would emit a
      object file containing the bitcode and metadata.
      
      After that point we compile the original module using the
      PerModuleDefaultPipeline used for non-LTO compilation. We generate
      standard object files at the end of this pipeline, which contain machine
      code and the new `.llvm.lto` section containing bitcode.
      
      Since the two pipelines operate on different copies of the module, we
      can be sure that the bitcode in the `.llvm.lto` section and object code
      in  `.text` are congruent with the existing output produced by the
      default and LTO pipelines.
      
      Original RFC: https://discourse.llvm.org/t/rfc-ffat-lto-objects-support/63977
      
      Earlier versions of this patch were missing REQUIRES lines for llc
      related tests in Transforms/EmbedBitcode. Those tests are now under
      CodeGen/X86, which should avoid running the check on unsupported
      platforms.
      
      Reviewed By: tejohnson, MaskRay, nikic
      
      Differential Revision: https://reviews.llvm.org/D146776
      44265dc3
    • Sami Tolvanen's avatar
      [RISCV] Implement KCFI operand bundle lowering · 83835e22
      Sami Tolvanen authored
      With `-fsanitize=kcfi` (Kernel Control-Flow Integrity), Clang emits
      "kcfi" operand bundles to indirect call instructions. Similarly to
      the target-specific lowering added in D119296, implement KCFI operand
      bundle lowering for RISC-V.
      
      This patch disables the generic KCFI pass for RISC-V in Clang, and
      adds the KCFI machine function pass in `RISCVPassConfig::addPreSched`
      to emit target-specific `KCFI_CHECK` pseudo instructions before calls
      that have KCFI operand bundles. The machine function pass also bundles
      the instructions to ensure we emit the checks immediately before the
      calls, which is not possible with the generic pass.
      
      `KCFI_CHECK` instructions are lowered in `RISCVAsmPrinter` to a
      contiguous code sequence that traps if the expected hash in the
      operand bundle doesn't match the hash before the target function
      address. This patch emits an `ebreak` instruction for error handling
      to match the Linux kernel's `BUG()` implementation. Just like for X86,
      we also emit trap locations to a `.kcfi_traps` section to support
      error handling, as we cannot embed additional information to the trap
      instruction itself.
      
      Relands commit 62fa708c with fixed
      tests.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D148385
      83835e22
    • Mehdi Amini's avatar
      Fix missing return introduced when adding support for properties in MLIR in 5e118f93 · 4d60c654
      Mehdi Amini authored
      This is a rarely used API: `matchAndRewrite()` is the usual way of using
      patterns.
      4d60c654
    • Matt Arsenault's avatar
      ValueTracking: Handle cannotBeOrderedLessThanZero for fadd · 5af973d5
      Matt Arsenault authored
      Move cannotBeOrderedLessThanZero logic into computeKnownFPClass.
      5af973d5
    • Kazu Hirata's avatar
      [llvm-profdata] Fix mixed-sign comparison warnings · a2e7f261
      Kazu Hirata authored
      This patch fixes -Wsign-compare warnings from:
      
        llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:117:3: note:
        in instantiation of function template specialization
        'testing::internal::EqHelper::Compare<unsigned long, int, nullptr>'
        requested here
      
        llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:134:3: note:
        in instantiation of function template specialization
        'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>'
        requested here
      
        llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:144:3: note:
        in instantiation of function template specialization
        'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>'
        requested here
      
        llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:160:3: note:
        in instantiation of function template specialization
        'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>'
        requested here
      a2e7f261
    • Vitaly Buka's avatar
      [test][msan] Add test for dc4d9d61 · 6fc0e548
      Vitaly Buka authored
      6fc0e548
    • Vitaly Buka's avatar
      [msan] Optimize zeroing allocated memory · ca9c18a2
      Vitaly Buka authored
      Reviewed By: thurston
      
      Differential Revision: https://reviews.llvm.org/D153599
      ca9c18a2
    • LLVM GN Syncbot's avatar
      [gn build] Port 31af18bc · 7425e770
      LLVM GN Syncbot authored
      7425e770
    • William Huang's avatar
      [llvm-profdata] Refactoring Sample Profile Reader to increase FDO build speed... · 31af18bc
      William Huang authored
      [llvm-profdata] Refactoring Sample Profile Reader to increase FDO build speed using MD5 as key to Sample Profile map
      
      This is phase 1 of multiple planned improvements on the sample profile loader.   The major change is to use MD5 hash code ((instead of the function itself) as the key to look up the function offset table and the profiles, which significantly reduce the time it takes to construct the map.
      
      The optimization is based on the fact that many practical sample profiles are using MD5 values for function names to reduce profile size, so we shouldn't need to convert the MD5 to a string and then to a SampleContext and use it as the map's key, because it's extremely slow.
      
      Several changes to note:
      
      (1) For non-CS SampleContext, if it is already MD5 string, the hash value will be its integral value, instead of hashing the MD5 again. In phase 2 this is going to be optimized further using a union to represent MD5 function (without converting it to string) and regular function names.
      
      (2) The SampleProfileMap is a wrapper to *map<uint64_t, FunctionSamples>, while providing interface allowing using SampleContext as key, so that existing code still work. It will check for MD5 collision (unlikely but not too unlikely, since we only takes the lower 64 bits) and handle it to at least guarantee compilation correctness (conflicting old profile is dropped, instead of returning an old profile with inconsistent context). Other code should not try to use MD5 as key to access the map directly, because it will not be able to handle MD5 collision at all. (see exception at (5) )
      
      (3) Any SampleProfileMap::emplace() followed by SampleContext assignment if newly inserted, should be replaced with SampleProfileMap::Create(), which does the same thing.
      
      (4) Previously we ensure an invariant that in SampleProfileMap, the key is equal to the Context of the value, for profile map that is eventually being used for output (as in llvm-profdata/llvm-profgen). Since the key became MD5 hash, only the value keeps the context now, in several places where an intermediate SampleProfileMap is created, each new FunctionSample's context is set immediately after insertion, which is necessary to "remember" the context otherwise irretrievable.
      
      (5) When reading a profile, we cache the MD5 values of all functions, because they are used at least twice (one to index into FuncOffsetTable, the other into SampleProfileMap, more if there are additional sections), in this case the SampleProfileMap is directly accessed with MD5 value so that we don't recalculate it each time (expensive)
      
      Performance impact:
      When reading a ~1GB extbinary profile (fixed length MD5, not compressed) with 10 million function names and 2.5 million top level functions (non CS functions, each function has varying nesting level from 0 to 20), this patch improves the function offset table loading time by 20%, and improves full profile read by 5%.
      
      Reviewed By: davidxl, snehasish
      
      Differential Revision: https://reviews.llvm.org/D147740
      31af18bc
    • Sami Tolvanen's avatar
      Revert "[RISCV] Implement KCFI operand bundle lowering" · e809ebeb
      Sami Tolvanen authored
      This reverts commit 62fa708c.
      
      Reverting to investigate -verify-machineinstrs errors in MIR tests.
      e809ebeb
    • xortoast's avatar
      [WebAssembly] Add lowering for llvm.rint and llvm.roundeven · bb648c91
      xortoast authored
      WebAssembly doesn't expose inexact exceptions, so frint can be mapped to
      fnearbyint. Likewise, WebAssembly always rounds ties-to-even, so
      froundeven can be mapped to fnearbyint.
      
      Differential Revision: https://reviews.llvm.org/D153451
      bb648c91
    • Amara Emerson's avatar
      Darwin: Use the GOT to reference ___stack_chk_guard. · 1ec30106
      Amara Emerson authored
      e018cbf7 changed the default behaviour for Darwin, and this breaks some
      existing software.
      
      rdar://110350601
      1ec30106
    • Joseph Huber's avatar
      [libc] Add basic utility support for timing functions on the GPU · 6fb57f73
      Joseph Huber authored
      This patch adds the utilities for the clocks on the GPU. This is done
      prior to exporting it via some other interface and is mainly just done
      so they are availible if we wish to do internal testing.
      
      Reviewed By: lntue
      
      Differential Revision: https://reviews.llvm.org/D153388
      6fb57f73
    • Joseph Huber's avatar
      [ClangPackager] Add an option to extract inputs to an archive · 869baa91
      Joseph Huber authored
      Currently we simply overwrite the output file if we get muliple matches
      in the fatbinary. This patch introduces the `--archive` option which
      allows us to combine all of the files into a static archive instead.
      This is usefuly for creating a device specific static archive library
      from a fatbinary.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D153568
      869baa91
    • Florian Hahn's avatar
      [ConstraintElim] Add tests to check negated OR simplifications. · c52268a9
      Florian Hahn authored
      Additional test coverage for D151799.
      c52268a9
    • Matt Arsenault's avatar
      6e94a9bf
    • Joseph Huber's avatar
      [libc][Obvious] Use the current binary dir instead of the base one · a24a1e04
      Joseph Huber authored
      Summary:
      We include things off of `libc/include` so we need to use the current
      binary dir when setting up the directory.
      a24a1e04
    • Robert Suderman's avatar
      [mlir][math] Modified the 'math.exp' lowering for higher precision · 710dc728
      Robert Suderman authored
      The existing lowering has lower precision for certain use cases, e.g.
      tanh. Improved version should demonstrate an overall higher level of precision.
      
      Reviewed By: cota, jpienaar
      
      Differential Revision: https://reviews.llvm.org/D153592
      710dc728
    • Matt Arsenault's avatar
      OpenMP/cmake: Use DEPFILE instead of IMPLICIT_DEPENDS · a2f5bcc7
      Matt Arsenault authored
      IMPLICIT_DEPENDS doesn't actually work with ninja and this does.
      a2f5bcc7
    • Matt Arsenault's avatar
    • Zahira Ammarguellat's avatar
      When float_t and double_t types are used inside a scope with · 63b0b82f
      Zahira Ammarguellat authored
      a '#pragma clang fp eval_method, it can lead to ABI breakage.
      See https://godbolt.org/z/56zG4Wo91
      This patch prevents this.
      
      Differential Revision: https://reviews.llvm.org/D153590
      63b0b82f
    • LLVM GN Syncbot's avatar
      [gn build] Port a3800ad9 · 55d04119
      LLVM GN Syncbot authored
      55d04119
    • Jonas Devlieghere's avatar
      [lldb] Use format specific for unprintabe char in DumpDataExtractor · c0045a8e
      Jonas Devlieghere authored
      Addresses Jason's post-commit feedback in D153644.
      c0045a8e
    • Chia-hung Duan's avatar
      [scudo] PopBatch after populateFreeList() · d0290b2f
      Chia-hung Duan authored
      Ensure the thread that refills freelist will get the Batch without
      contending the lock in SizeClassAllocator64.
      
      Reviewed By: cferris
      
      Differential Revision: https://reviews.llvm.org/D152419
      d0290b2f
    • Chia-hung Duan's avatar
      [scudo] update Pushedblocks/PoppedBlocks in Impl functions · 18207dbc
      Chia-hung Duan authored
      Reviewed By: cferris
      
      Differential Revision: https://reviews.llvm.org/D152420
      18207dbc
    • Joseph Huber's avatar
      [libc] Fix installing GPU headers · 3368a92b
      Joseph Huber authored
      The patch in D152592 changed the logic for this. We could never check if
      we were on the GPU as this was before the variable was defined so I
      moved it later. Secondly, we cannot use the `LLVM_BINARY_DIR` here, and
      I do not know if that works in general. The problem is that it will
      isntall the headers under a normal path outside of the
      `LLVM_ENABLE_RUNTIMES` build. I don't know if that's correct for the
      other targets, but for the GPU I need to set it back to the
      CMAKE_BINARY_DIR so it works.
      
      Reviewed By: phosek
      
      Differential Revision: https://reviews.llvm.org/D153637
      3368a92b
    • Paul Kirth's avatar
      Revert "[llvm] Preliminary fat-lto-objects support" · a3800ad9
      Paul Kirth authored
      There seems to be a problem on arm buildbots. Reverting until I can
      investigate.  https://lab.llvm.org/buildbot#builders/245/builds/10184
      
      This reverts commit a67208e1
      and dependent commit e54a3112.
      a3800ad9
    • Sami Tolvanen's avatar
      [RISCV] Implement KCFI operand bundle lowering · 62fa708c
      Sami Tolvanen authored
      With `-fsanitize=kcfi` (Kernel Control-Flow Integrity), Clang emits
      "kcfi" operand bundles to indirect call instructions. Similarly to
      the target-specific lowering added in D119296, implement KCFI operand
      bundle lowering for RISC-V.
      
      This patch disables the generic KCFI pass for RISC-V in Clang, and
      adds the KCFI machine function pass in `RISCVPassConfig::addPreSched`
      to emit target-specific `KCFI_CHECK` pseudo instructions before calls
      that have KCFI operand bundles. The machine function pass also bundles
      the instructions to ensure we emit the checks immediately before the
      calls, which is not possible with the generic pass.
      
      `KCFI_CHECK` instructions are lowered in `RISCVAsmPrinter` to a
      contiguous code sequence that traps if the expected hash in the
      operand bundle doesn't match the hash before the target function
      address. This patch emits an `ebreak` instruction for error handling
      to match the Linux kernel's `BUG()` implementation. Just like for X86,
      we also emit trap locations to a `.kcfi_traps` section to support
      error handling, as we cannot embed additional information to the trap
      instruction itself.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D148385
      62fa708c
    • LLVM GN Syncbot's avatar
      [gn build] Port a67208e1 · fd65b8da
      LLVM GN Syncbot authored
      fd65b8da
    • Benjamin Kramer's avatar
      Remove unused include. NFC · e54a3112
      Benjamin Kramer authored
      e54a3112
    • Jonas Devlieghere's avatar
      [lldb] Print unprintable characters as unsigned · 85f40fc6
      Jonas Devlieghere authored
      When specifying the C-string format for dumping memory, we treat
      unprintable characters as signed. Whether a character is signed or not
      is implementation defined, but all printable characters are signed.
      Therefore it's fair to assume that unprintable characters are unsigned.
      
      Before this patch, "\xcf\xfa\xed\xfe\f" would be printed as
      "\xffffffcf\xfffffffa\xffffffed\xfffffffe\f". Now we correctly print the
      original string.
      
      rdar://111126134
      
      Differential revision: https://reviews.llvm.org/D153644
      85f40fc6
    • Artem Belevich's avatar
      [NVPTX] Lower v2f16 and v2bf16 stores as 32-bit scalars. · 60941f1d
      Artem Belevich authored
      This avoids unnecessary vector splitting that was needed for vectorized store
      instruction.
      
      Differential Revision: https://reviews.llvm.org/D152593
      60941f1d