- May 25, 2023
-
-
Mehdi Amini authored
The move of the bytecode serialization to be tablegen driven in https://reviews.llvm.org/D144820 added a new condition in the reading path that forbid 0-sized integer, even though we still produce them. Fix #62920 Differential Revision: https://reviews.llvm.org/D151372
-
AdityaK authored
[libc++, std::vector] call the optimized version of __uninitialized_allocator_copy for trivial types See: https://github.com/llvm/llvm-project/issues/61987 Fix suggested by: @philnik and @var-const Reviewers: philnik, ldionne, EricWF, var-const Differential Revision: https://reviews.llvm.org/D147741 Testing: ninja check-cxx check-clang check-llvm Benchmark Testcases (BM_CopyConstruct, and BM_Assignment) added. performance improvement: Run on (8 X 4800 MHz CPU s) CPU Caches: L1 Data 48 KiB (x4) L1 Instruction 32 KiB (x4) L2 Unified 1280 KiB (x4) L3 Unified 12288 KiB (x1) Load Average: 1.66, 3.02, 2.43 Comparing build-runtimes-base/libcxx/benchmarks/vector_operations.libcxx.out to build-runtimes/libcxx/benchmarks/vector_operations.libcxx.out Benchmark Time CPU Time Old Time New CPU Old CPU New ---------------------------------------------------------------------------------------------------------------------------------------- BM_ConstructSize/vector_byte/5140480 +0.0362 +0.0362 116906 121132 116902 121131 BM_CopyConstruct/vector_int/5140480 -0.4563 -0.4577 1755224 954241 1755330 951987 BM_Assignment/vector_int/5140480 -0.0222 -0.0220 990045 968095 989917 968125 BM_ConstructSizeValue/vector_byte/5140480 +0.0308 +0.0307 116970 120567 116977 120573 BM_ConstructIterIter/vector_char/1024 -0.0831 -0.0831 19 17 19 17 BM_ConstructIterIter/vector_size_t/1024 +0.0129 +0.0131 88 89 88 89 BM_ConstructIterIter/vector_string/1024 -0.0064 -0.0018 54455 54109 54208 54112 OVERALL_GEOMEAN -0.0845 -0.0842 0 0 0 0 FYI, the perf improvements for BM_CopyConstruct due to this patch is mostly subsumed by the https://reviews.llvm.org/D149826. However this patch still adds value by converting copy to memmove (the second testcase). Before the patch: ``` define linkonce_odr dso_local void @_ZNSt3__16vectorIiNS_9allocatorIiEEE18__construct_at_endIPiS5_EEvT_T0_m(ptr noundef nonnull align 8 dereferenceable(24) %0, ptr noundef %1, ptr noundef %2, i64 noundef %3) local_unnamed_addr #4 comdat align 2 { %5 = getelementptr inbounds %"class.std::__1::vector", ptr %0, i64 0, i32 1 %6 = load ptr, ptr %5, align 8, !tbaa !12 %7 = icmp eq ptr %1, %2 br i1 %7, label %16, label %8 8: ; preds = %4, %8 %9 = phi ptr [ %13, %8 ], [ %1, %4 ] %10 = phi ptr [ %14, %8 ], [ %6, %4 ] %11 = icmp ne ptr %10, null tail call void @llvm.assume(i1 %11) %12 = load i32, ptr %9, align 4, !tbaa !14 store i32 %12, ptr %10, align 4, !tbaa !14 %13 = getelementptr inbounds i32, ptr %9, i64 1 %14 = getelementptr inbounds i32, ptr %10, i64 1 %15 = icmp eq ptr %13, %2 br i1 %15, label %16, label %8, !llvm.loop !16 16: ; preds = %8, %4 %17 = phi ptr [ %6, %4 ], [ %14, %8 ] store ptr %17, ptr %5, align 8, !tbaa !12 ret void } ``` After the patch: ``` define linkonce_odr dso_local void @_ZNSt3__16vectorIiNS_9allocatorIiEEE18__construct_at_endIPiS5_EEvT_T0_m(ptr noundef nonnull align 8 dereferenceable(24) %0, ptr noundef %1, ptr noundef %2, i64 noundef %3) local_unnamed_addr #4 comdat align 2 { %5 = getelementptr inbounds %"class.std::__1::vector", ptr %0, i64 0, i32 1 %6 = load ptr, ptr %5, align 8, !tbaa !12 %7 = ptrtoint ptr %2 to i64 %8 = ptrtoint ptr %1 to i64 %9 = sub i64 %7, %8 %10 = ashr exact i64 %9, 2 tail call void @llvm.memmove.p0.p0.i64(ptr align 4 %6, ptr align 4 %1, i64 %9, i1 false) %11 = getelementptr inbounds i32, ptr %6, i64 %10 store ptr %11, ptr %5, align 8, !tbaa !12 ret void } ``` This is due to the optimized version of uninitialized_allocator_copy function.
-
Vitaly Buka authored
-
Vitaly Buka authored
__sanitizer_get_current_allocated_bytes had as body, but allocator caches were not registered to collect stats. It's done by SizeClassAllocator64LocalCache::Init(). Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D151355
-
Ramkumar Ramachandra authored
This patch modifies the description in TosaOps.td, taking into account all the arguments, and supplying examples. Signed-off-by:
Ramkumar Ramachandra <r@artagnon.com> Differential Revision: https://reviews.llvm.org/D139089
-
Eugene Burmako authored
https://reviews.llvm.org/D151104 moved PDL-related transform ops into an extension and updated the Bazel build, but one tiny thing fell through the cracks - TransformOpsPyFiles also needs to include the newly introduced `mlir/python/mlir/dialects/_transform_pdl_extension_ops_ext.py`. Reviewed By: saugustine, bkramer Differential Revision: https://reviews.llvm.org/D151368
-
Craig Topper authored
Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D150996
-
Thurston Dang authored
This is a very simple test that calls wsclen. There are currently no other HWASan tests that call wsclen, which is why the wcslen interceptor issue (triggered by https://reviews.llvm.org/D150708 and fixed in https://reviews.llvm.org/D150909) was only detected by stage2/hwasan check on the buildbots. With this test, the issue would have been caught by stage1 check-sanitizer, with a more obvious diagnosis. Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D151000
-
Florian Hahn authored
If there is an infinite cycle in the IR, the loop will never exit. Keep track of visited basic blocks in a set and return nullptr if a block is visited again. Fixes #62830. Reviewed By: rjmccall Differential Revision: https://reviews.llvm.org/D151076
-
Slava Zakharin authored
Even though the constant expression actual argument is not definable, and the associated dummy argument is not definable, the compiler may produce implicit copies into the memory storage associated with the constant expression. For example, a constant expression storage passed by reference to a subprogram may be used for implicit copy-out: ``` subroutine sub(i, n) interface subroutine sub2(i) integer :: i(*) end subroutine sub2 end interface integer :: i(n) call sub2(i(3::2)) ! copy-out after the call will write to 'i' end subroutine sub subroutine test call sub((/1,2,3,4,5/), 5) end subroutine test ``` If we pass a reference to constant expression storage to 'sub' directly, the copy-out inside 'sub' will try to write into readonly memory. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D151271 -
Artem Belevich authored
Previous implementation provided wrappers for the internal implementations used by CUDA headers. However, clang does not include those, so we need to provide the public functions instead. Differential Revision: https://reviews.llvm.org/D151243
-
Erich Keane authored
Clang 16 changed to consider dereferencing a void* to be a warning-as-error, plus made this an error in SFINAE contexts, since this resulted in incorrect template instantiation. When doing so, the Clang 16 documentation was updated to reflect that this was likely to change again to a non-disablable error in the next version. As there has been no response to changing from a warning to an error, I believe this is a non-controversial change. This patch changes this to be an Error, consistent with the standard and other compilers. This was discussed in this RFC: https://discourse.llvm.org/t/rfc-can-we-stop-the-extension-to-allow-dereferencing-void-in-c/65708 Differential Revision: https://reviews.llvm.org/D150875
-
Michael Liao authored
-
Sterling Augustine authored
-
Med Ismail Bennani authored
This test started failing on the green-dragon bot, but after some investigation, it doesn't have anything to do with Lua. If we use a variable watchpoint with a condition using a scope variable, if we go out-of-scope, the watpoint remains active which can the expression evaluator to fail to parse the watchpoint condition (because of the missing varible bindings). For now, we should disable this test until we come up with a fix for it. rdar://109574319 Signed-off-by:
Med Ismail Bennani <ismail@bennani.ma>
-
Vitaly Buka authored
-
Valentin Clement authored
Use the new reduction design in acc.loop operation. Depends on D151146 Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D151164
-
Vitaly Buka authored
__sanitizer_get_current_allocated_bytes had as body, but allocator caches were not registered to collect stats. It's done by SizeClassAllocator64LocalCache::Init(). Reviewed By: kstoimenov Differential Revision: https://reviews.llvm.org/D151352
-
Philip Reames authored
This adds the vfslide1down (and vfslide1up for consistency) nodes. These mostly parallel the existing vslide1down/up nodes. (See note below on instruction semantics.) We then use the vfslide1down in build_vector lowering instead of going through the stack. The specification is more than a bit vague on the meaning of these instructions. All we're given is "The vfslide1down instruction is defined analogously, but sources its scalar argument from an f register." We have to combine this with a general note at the beginning of section 10. Vector Arithmetic Instruction Formats which reads: "For floating-point operations, the scalar can be taken from a scalar f register. If FLEN > SEW, the value in the f registers is checked for a valid NaN-boxed value, in which case the least-signicant SEW bits of the f register are used, else the canonical NaN value is used. Vector instructions where any floating-point vector operand’s EEW is not a supported floating-point type width (which includes when FLEN < SEW) are reserved.". Note that floats are NaN-boxed when D is implemented. Combining that all together, we're fine as long as the element type matches the vector type - which is does by construction. We shouldn't have legal vectors which hit the reserved encoding case. An assert is included, just to be careful. Differential Revision: https://reviews.llvm.org/D151347
-
Valentin Clement authored
Add the missing check on private list information. The check is the same than the one done for acc.parallel. Depends on D151146 Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D151149
-
Valentin Clement authored
Add the missing check on private list information. The check is the same than the one done for acc.parallel. Depends on D151146 Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D151149
-
Aaron Ballman authored
We've been hosting these meetings regularly for a while now, so this begins advertising the meetings more widely.
-
Valentin Clement authored
After D150818 the reduction clause is represented with a acc.reduction.recipe operation and an operand. This patch updates the acc.parallel op for the new design. Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D151146
-
Philip Reames authored
The immediate field on the vsetivli is fairly limited. For larger vectors, we end up having to materialize a constant in a register. We hadn't plumbed the infrastructure to treat such materialized constants as constants for purpose of vsetvli elimination. I only bothered to handle LI. We could extend this to LUI sequences, but well, 2048 elements is probably enough for all practical fixed length vector codegen. :) The test delta does point out a related problem. At LMUL8, we see increased register allocation pressure, and we should probably either a) address register allocation remat, or b) be less aggressive about eliminating vsetvlis at high lmul. Note that high LMUL code is not generated much by default. Differential Revision: https://reviews.llvm.org/D151212
-
Kazu Hirata authored
This patch fixes: mlir/lib/Dialect/GPU/IR/GPUDialect.cpp:175:2: error: extra ';' outside of a function is incompatible with C++98 [-Werror,-Wc++98-compat-extra-semi]
-
Amy Kwan authored
Fix the shared library build failure on clang-ppc64le-rhel from 1c9a8004 as seen in: https://lab.llvm.org/buildbot/#/builders/57/builds/27080/steps/6/logs/stdio
-
Stefan Pintilie authored
My previous patch had added a couple of asserts to the disassembler. The problem with this is that the disassembler is not just used for the text section it is also used to disassemble the data section of an object where the bytes do not necessarily represent instructions. If the data in the data section happens to look like an illegal instruction then llvm-objdump will assert on data because it is finding an illegal instruction that is not actually an instruction at all. Reviewed By: nemanjai, #powerpc Differential Revision: https://reviews.llvm.org/D149711
-
Vitaly Buka authored
-
Alex Langford authored
I landed D151001 before it had gotten sign-off from all the reviewers. This is a follow-up to address the additional feedback. Differential Revision: https://reviews.llvm.org/D151233
-
Kun Wu authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D151014
-
Kelvin Li authored
The following PowerPC vector type syntax is added: VECTOR ( element-type-spec ) where element-type-sec is integer-type-spec, real-type-sec or unsigned-type-spec. Two opaque types (__VECTOR_PAIR and __VECTOR_QUAD) are also added. A finite set of functionalities are implemented in order to support the new types: 1. declare objects 2. declare function result 3. declare type dummy arguments 4. intrinsic assignment between the new type objects (e.g. v1=v2) 5. reference functions that return the new types Submit on behalf of @tislam @danielcchen Authors: @tislam @danielcchen Differential Revision: https://reviews.llvm.org/D150876
-
Vitaly Buka authored
Breaks https://lab.llvm.org/buildbot/#/builders/18/builds/9118 This reverts commit 8064caf8.
-
https://reviews.llvm.org/D144552Sterling Augustine authored
Differential Revision: https://reviews.llvm.org/D151346
-
Jay Foad authored
Verify the LiveIntervals analysis after a pass that claims to preserve it, even if there are no further passes (apart from the verifier itself) that would use the analysis. Fixes https://github.com/llvm/llvm-project/issues/46217 Differential Revision: https://reviews.llvm.org/D129208
-
John Brawn authored
Fix several instances of macros being defined multiple times in several targets. Most of these are just simple duplication in a TargetInfo or OSTargetInfo of things already defined in InitializePredefinedMacros or InitializeStandardPredefinedMacros, but there are a few that aren't: * AArch64 defines a couple of feature macros for armv8.1a that are handled generically by getTargetDefines. * CSKY needs to take care when CPUName and ArchName are the same. * Many os/target combinations result in __ELF__ being defined twice. Instead define __ELF__ just once in InitPreprocessor based on the Triple, which already knows what the object format is based on os and target. These changes shouldn't change the final result of which macros are defined, with the exception of the changes to __ELF__ where if you explicitly specify the object type in the triple then this affects if __ELF__ is defined, e.g. --target=i686-windows-elf results in it being defined where it wasn't before, but this is more accurate as an ELF file is in fact generated. Differential Revision: https://reviews.llvm.org/D150966
-
Marco Elver authored
Fix typo introduced in 2f1e2a6b. Reported-by: RamNalamothu
-
Tom Eccles authored
This reverts commit 74c2ec50. This caused a regression building spec2017 with -Ofast.
-
Matthias Braun authored
This bumps the "large-interval-freq-threshold" limit in the register coalescer to 256. The limit was introduced in https://reviews.llvm.org/D59143 without much justify for the particular value "100", so I hope bumping it is ok. This change is motivated by bad codegen for the popular crc32c algorithm; the code is often based/copied from this implementation: https://github.com/htot/crc32c/blob/master/crc32c/crc32intelc.cc which uses a duffs-device pattern with 128 switch-cases. There are examples in RocksDB (https://github.com/facebook/rocksdb/blob/main/util/crc32c.cc) and Folly (https://github.com/facebook/folly/blob/main/folly/hash/detail/Crc32cDetail.cpp) which are important use cases for us. Differential Revision: https://reviews.llvm.org/D150994
-
Nemanja Ivanovic authored
Commit 8064caf8 added a call to a function that performs this combine without checking whether the target supports FPCVT. This caused asserts to trip on BE bots as the default target does not have this feature.
-
Mark de Wever authored
When using with clang-tidy 17 Node.getAttrName() sometimes returns a nullptr. This caused segfaults in the CI. Reviewed By: philnik, #libc Differential Revision: https://reviews.llvm.org/D151224
-