- Feb 21, 2024
-
-
Rajveer Singh Bharadwaj authored
Resolves Issue #82249 As described in the issue, any deallocation function for a `class X` is a static member (even if not explicitly declared static).
-
Timm Bäder authored
This function tried to be smart about the dereferenced value, but it ended up hurting more than it helped. At least in the current state, where we still try get the correct output. I might add something similar back later.
-
harishch4 authored
Name mangling is invoked for a bind(C) procedure contained in a block in a context that does not have access to block ID mapping. Relaxing an assert to account for this. Fixes #79408
-
Momchil Velikov authored
Certain stack probing sequences might clobber flags, then we can't use a block as a prologue if the flags register is a live-in on entry to that block.
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
Simon Pilgrim authored
Use the 3 or 4 active bits as a shift amount into a i32/i64 constant representing the number of set bits. In future, it might be worthwhile to move this into a generic location in case other targets want to make use of them. Another expansion pulled from #79823
-
Simon Pilgrim authored
If we only have 2 active bits then we can avoid the i8 CTPOP multiply expansion entirely Another expansion pulled from #79823
-
Hui authored
For the mutex vs atomic test: Old: `unique_lock<mutex>` New: a lock implemented with `atomic::wait` On 10 years old Intel Macbook, `atomic::wait` is 50% slower than `mutex` ``` Benchmark Time CPU Time Old Time New CPU Old CPU New ---------------------------------------------------------------------------------------------------------------------------------- BM_multi_thread_lock_unlock/1024 +0.3735 +2.4497 1724726 2368935 153159 528354 BM_multi_thread_lock_unlock/2048 +0.4174 +1.2487 3410538 4834012 435062 978311 BM_multi_thread_lock_unlock/4096 +0.5256 +1.9824 6903783 10532681 590266 1760405 BM_multi_thread_lock_unlock/8192 +0.5415 +0.4578 14536391 22408399 ...
-
chuongg3 authored
Ensure BITCAST is only legal for types with the same amount of bits. Enable BITCAST to work with non-legal vector types as well.
-
Adrian Kuegel authored
Use const reference for loop variable.
-
Timm Baeder authored
The Float print type is backed by the Floating class, which in turn uses APFloat, which might heap-allocate memory, so might be expensive to copy. Add an 'AsRef' bit to the ArgType tablegen class, which defines whether we pass the argument around by copy or by reference.
-
hev authored
This PR indicates that `addrspacecasts` are always no-ops on LoongArch. Fixes #82330
-
Alex Zinenko authored
Fix a leak of the root operation not being deleted in the recently introduced transform_interpreter.c.
-
Chia authored
Extend D133739 and #76785 to support vector widening floating-point add/sub/mul instructions. Specifically, this patch works for the below optimization case: ### Source code ``` define void @vfwmul_v2f32_multiple_users(ptr %x, ptr %y, ptr %z, <2 x float> %a, <2 x float> %b, <2 x float> %b2) { %c = fpext <2 x float> %a to <2 x double> %d = fpext <2 x float> %b to <2 x double> %d2 = fpext <2 x float> %b2 to <2 x double> %e = fmul <2 x double> %c, %d %f = fadd <2 x double> %c, %d2 %g = fsub <2 x double> %d, %d2 store <2 x double> %e, ptr %x store <2 x double> %f, ptr %y store <2 x double> %g, ptr %z ret void } ``` ### Before this patch [Compiler Explorer](https://godbolt.org/z/aaEMs5s9h) ``` vfwmul_v2f32_multiple_users: vsetivli zero, 2, e32, mf2, ta, ma vfwcvt.f.f.v v11, v8 vfwcvt.f.f.v v8, v9 vfwcvt.f.f.v v9, v10 vsetvli zero, zero, e64, m1, ta, ma vfmul.vv v10, v11, v8 vfadd.vv v11, v11, v9 vfsub.vv v8, v8, v9 vse64.v v10, (a0) vse64.v v11, (a1) vse64.v v8, (a2) ret ``` ### After this patch ``` vfwmul_v2f32_multiple_users: vsetivli zero, 2, e32, mf2, ta, ma vfwmul.vv v11, v8, v9 vfwadd.vv v12, v8, v10 vfwsub.vv v8, v9, v10 vse64.v v11, (a0) vse64.v v12, (a1) vse64.v v8, (a2) ``` -
Paul Walker authored
This is another step in the direction of fixing the `Fixed(0) != Scalable(0)` bugbear, although whilst weird I don't believe it's causing us any real issues.
-
Lukacma authored
This patches adds missing target-feature dependencies for SME2.1
-
Vedant Paranjape authored
In LoopUnroll, peelLoop is called on the loop. After the loop is peeled it calls simplifyLoopAfterUnroll on the loop. This call to simplifyLoopAfterUnroll doesn't preserve the LCSSA form of the parent loop and thus during the next call to peelLoop the LCSSA form is already broken. LoopPeel util takes in the PreserveLCSSA argument and it passes on the same argument to simplifyLoop which checks if the loop is in a valid LCSSA form, when (PreserveLCSSA = true). This causes an assert in simplifyLoop when (PreserveLCSSA = true), as during the last call LCSSA for the loop wasn't preserved, and thus crashes at the following assert. assert(L->isRecursivelyLCSSAForm(*DT, *LI) && "Requested to preserve LCSSA, but it's already broken."); Upon debugging, it is evident that simplifyLoopIVs call inside simplifyLoopAfterUnroll breaks the LCSSA form. This patch fixes llvm#77118, it checks if the replacement of IV Users with Loop Invariant preserves the LCSSA form. If it does not, it emits the required LCSSA Phi instructions. -
harishch4 authored
When a do loop with a construct-name is used inside OpenMP construct with default(none), an incorrect error will be raised as below. ``` program cn_and_default implicit none integer :: i !$omp parallel default(none) loop: do i = 1, 10 end do loop !$omp end parallel end program ``` > The DEFAULT(NONE) clause requires that 'loop' must be listed in a data-sharing attribute clause This patch fixes this by adding a condition to check and skip processing construct-names. -
Timm Bäder authored
We internalle handle these via pointers, but we need to return them as RValues in initializers.
-
Sergei Lebedev authored
_SubClassValueT is only useful when it is has >1 usage in a signature. This was not true for the signatures produced by tblgen. For example def call(result, callee, operands_, *, loc=None, ip=None) -> _SubClassValueT: ... here a type checker does not have enough information to infer a type argument for _SubClassValueT, and thus effectively treats it as Any. -
Ivan Kosarev authored
Saves generating ~1200 instances of the PredConcat TableGen class. Also removes the default predicates from resulting predicate lists.
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
Simon Pilgrim authored
Adds missing avx512 constant broadcast comments
-
David Spickett authored
With some missing config options and a link to the test suite docs that explain how to setup `ISO_FORTRAN_C_HEADER` and set the stop message variable.
-
Kadir Cetinkaya authored
This reverts commit 50373506. Broke include-cleaner tests
-
Sergei Lebedev authored
The two forms are equivalent, so there is no reason to use the longer one.
-
John Brawn authored
The address register should be surrounded by square brackets, like in all the other str instructions. Fixes https://github.com/llvm/llvm-project/issues/81846
-
Oleksandr "Alex" Zinenko authored
Transform interpreter functionality can be used standalone without going through the interpreter pass, make it available in Python.
-
Nick Anderson authored
follow up patch to #78673 @Pierre-vh @jayfoad @arsenm Could you review when you have a chance.
-
David Green authored
In certain case "extreme" values like Nan, Inf and 0xffffffff could lead to generating different code via the inline-generated intrinsics vs the versions in the runtimes (and other compilers like gfortran). There are some examples I was using for testing in https://godbolt.org/z/x4EfqEss5. This changes the generation for the intrinsics to be more like the runtimes, using a condition that is similar to: isFirst || (prev != prev && elem == elem) || elem < prev The middle part is only used for floating point operations, and checks if the values are Nan. This should then hopefully make the logic closer to - return the first element with the lowest value, with Nans ignored unless there are only Nans. The initial limit value for floats are also changed from the largest float to Inf, to make sure it is handled correctly. The integer reductions are also changed to use a similar scheme to make sure they work with masked values. This means that the preamble after the loop can be removed.
-
Tuan Chuong Goh authored
-
Nikita Popov authored
SCEVExpanderCleaner will currently remove instructions created by SCEVExpander, but not restore poison generating flags that it may have dropped. As such, running LIR can currently spuriously drop flags without performing any transforms. Fix this by keeping track of original instruction flags in SCEVExpander. Fixes https://github.com/llvm/llvm-project/issues/82337.
-
Kohei Yamaguchi authored
- Fixed OpenACC's spec link format - Add missed `OpenACCPasses.md` into Passes.md - Add missed `MyExtensionCh4.md` into Ch4.md of tutorial of transform
-
martinboehme authored
-
LLVM GN Syncbot authored
-
kadir çetinkaya authored
-
Clement Courbet authored
…table. All data is derived from a single table rather than being spread out over an enum, a table and the main entry point. This is intended as a replacement for #82092.
-
Yingwei Zheng authored
This patch lowers select of constants if `TrueV == ~FalseV`. Address the comment in https://github.com/llvm/llvm-project/pull/82456#discussion_r1496881603.
-