- May 05, 2023
-
-
Serguei Katkov authored
After applying FMIN/FMAX, if any of operands is NaN, the second operand will be the result. So all we need is to check whether first operand is NaN and return it or result of FMIN/FMAX. So we avoid usage of constant NaN in the lowering. Additionally we can avoid handling NaN after FMIN/FMAX if we are sure that first operand is not NaN. Reviewed By: e-kud Differential Revision: https://reviews.llvm.org/D149729
-
Matt Devereau authored
-
Matt Devereau authored
Emit FNMADD instead of FNEG(FMADD) for optimization levels above Oz when fast-math flags (nsz+contract) permit it. Differential Revision: https://reviews.llvm.org/D149260
-
Nikita Popov authored
We are simplifying the loop and all its children. Each time, we invalidate the top-most loop. The top-most loop is going to be the same every time. The cost of SCEV invalidation is largely independent from how data about the loop is actually cached, so we should avoid redundant invalidations.
-
Florian Hahn authored
VPWidenRecipes should not be generated for scalar VFs. Replace check with an assert. Suggested in preparation for D149081.
-
Christian Ulmann authored
This commit introduces support for locations as part of the loop annotation attribute. These locations indicate the start and the end of the loop. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D149858
-
Lang Hames authored
The __objc_imageinfo section may be deleted (leaving dangling references to any symbols that it contains), and shouldn't have any dependencies anyway. This patch verifies that the section has no dependencies and then skips the section. rdar://108469243
-
Jean Perier authored
I plan to implement lowering from parse tree to HLFIR first for forall and where to ease testing of the rewrite pass while writing it. To avoid cryptic errors in ConvertToFir pass about unhandled operations, this patch already defines the pass that will further lower these operations and make it throw clear TODO messages. Differential Revision: https://reviews.llvm.org/D149852
-
Jean Perier authored
This is the last piece required to lower Forall (except pointer assignments, where an operation may be needed to deal with bounds remapping). Lowering requires symbols to be mapped to memory SSA values produced by a fir_FortranVariableOpInterface operation. This applies to forall index-values, that are symbols. fir.alloca/fir.store/hlfir.declare are not allowed inside the body of an hlfir.forall that only accept operations with the hlfir_OrderedAssignmentTreeOpInterface so that the forall structure is well defined and easy to transform. Allowing such operations in the forall body would open the doors to generating ill-formed programs where such operation would be used for non index-values. Instead, add an hlfir.forall_index with both required interface to produce a memory address for a forall index. As a bonus, since forall index-value are by nature read-only, the loads of hlfir.forall_index can be canonicalized, which will help simplifying the hlfir.forall nested code (it is unclear we will be able to tell MLIR enough about hlfir.forall and hlfir.where structure so that it could safely do a generic mem-to-reg inside it, and getting rid of read-effect operations will benefit the forall rewrite pass). Differential Revision: https://reviews.llvm.org/D149836
-
Aviad Cohen authored
This pass is useful to legalize rankless and dynamic shapes towards static using operands' shapes & types. Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D148998
-
Hristo Hristov authored
Implements parts of **P1614R2**: https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1614r2.html - Implemented `operator<=>` for `optional` - Updated "optional synopsis" to match the current draft https://eel.is/c++draft/optional closer - Implemented https://cplusplus.github.io/LWG/issue3566 - Implemented https://cplusplus.github.io/LWG/issue3746 Reviewed By: #libc, philnik, ldionne Differential Revision: https://reviews.llvm.org/D146392
-
Fangrui Song authored
-
Jie Fu authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D149927
-
Kazu Hirata authored
This is part of an effort to migrate from llvm::Optional to std::optional: https://discourse.llvm.org/t/deprecating-llvm-optional-x-hasvalue-getvalue-getvalueor/63716
-
Kazu Hirata authored
This is part of an effort to migrate from llvm::Optional to std::optional: https://discourse.llvm.org/t/deprecating-llvm-optional-x-hasvalue-getvalue-getvalueor/63716
-
Fangrui Song authored
-
Jonas Devlieghere authored
Use the templated GetPropertyAtIndexAs helper for FileSpecList.
-
Craig Topper authored
Instead of passing a constant to DecodeGPRRegisterClass, just create the X2 register directly.
-
Fangrui Song authored
-
Jonas Devlieghere authored
This patch is a continuation of 6f8b33f6 and eliminates the {Get,Set}PropertyAtIndexAsFileSpec functions.
-
LLVM GN Syncbot authored
-
Jonas Devlieghere authored
There's no reason for FileSpecList to live in lldb/Core while FileSpec lives in lldb/Utility. Move FileSpecList next to FileSpec.
-
LiaoChunyu authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D149925
-
Jonas Devlieghere authored
After 6f8b33f6 this function has no callers anymore.
-
Zhenkai Weng authored
This revision makes ShuffleBlockStrategy deterministic by replacing SmallPtrSet with other data structures that has a deterministic iteration order. Reviewed By: Peter Differential Revision: https://reviews.llvm.org/D149676
-
Fangrui Song authored
-
Brad Smith authored
Make sure that the upper bits of the offset is placed in bits 20-21 of the instruction word. This fixes the encoding of backwards (negative offset) BPr branches. (Previously, the upper two bits of the offset would overwrite parts of the rs1 field, causing it to branch on the wrong register, with the wrong offset) Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D144012
-
Craig Topper authored
Otherwise I think we extract and use a build_vector. There may be some more improvements that can be made and there might be some cases that we should do something different for, but this seemed like a decent starting point. Reviewed By: luke Differential Revision: https://reviews.llvm.org/D149724
-
Joseph Huber authored
Previously this wasn't implemented because it's effectively a no-op. However, this should be safe to emit on sm_60 architectures. It's important because it carries semantic importance for whether or not something can be moved. So we should always emit this instrinsic. Differential Revision: https://reviews.llvm.org/D149923
-
Luo, Yuanke authored
-
Tom Stellard authored
Fix bug constants and sub instructions When finding constants in a chain starting with the RHS operator of sub instructions, we were negating the constant before zero extending it, which is incorrect. Unfortunately, I was unable to find a simple way to implement this transformation correctly, so for now I just disabled this optimization for constants that feed into the RHS of a sub. Resolves #62379 Transformation from alive2.llvm.org: define i16 @src(i8 %a, i8 %b, i8 %c) { entry: %0 = sub nuw nsw i8 %c, %a %1 = sub nuw nsw i8 %b, %0 %2 = zext i8 %1 to i16 ret i16 %2 } Before/Bad: define i16 @tgt(i8 %a, i8 %b, i8 %c) { entry: %0 = zext i8 %a to i16 %1 = zext i8 %b to i16 %c_neg = sub i8 0, %c %c_zext = zext i8 %c_neg to i16 %2 = sub i16 0, %0 %3 = sub i16 %1, %2 %4 = add i16 %3, %c_zext ret i16 %4 } Correct: define i16 @tgt(i8 %a, i8 %b, i8 %c) { entry: %0 = zext i8 %a to i16 %1 = zext i8 %b to i16 %c_zext = zext i8 %c to i16 %c_neg = sub i16 0, %c_zext %2 = sub i16 0, %0 %3 = sub i16 %1, %2 %4 = add i16 %3, %c_neg ret i16 %4 } Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D149507 -
Luo, Yuanke authored
-
Craig Topper authored
Reviewed By: jrtc27 Differential Revision: https://reviews.llvm.org/D149901
-
Advenam Tacet authored
This revision is a part of a series of patches extending AddressSanitizer C++ container overflow detection capabilities by adding annotations, similar to those existing in std::vector, to std::string and std::deque collections. These changes allow ASan to detect cases when the instrumented program accesses memory which is internally allocated by the collection but is still not in-use (accesses before or after the stored elements for std::deque, or between the size and capacity bounds for std::string). The motivation for the research and those changes was a bug, found by Trail of Bits, in a real code where an out-of-bounds read could happen as two strings were compared via a std::equals function that took iter1_begin, iter1_end, iter2_begin iterators (with a custom comparison function). When object iter1 was longer than iter2, read out-of-bounds on iter2 could happen. Container sanitization would detect it. In revision D132522, support for non-aligned memory buffers (sharing first/last granule with other objects) was added, therefore the check for standard allocator is not necessary anymore. This patch removes the check in std::vector annotation member function (__annotate_contiguous_container) to support different allocators. Additionally, this revision fixes unpoisoning in std::vector. It guarantees that __alloc_traits::deallocate may access returned memory. Originally suggested in D144155 revision. If you have any questions, please email: - advenam.tacet@trailofbits.com - disconnect3d@trailofbits.com Reviewed By: #libc, #sanitizers, philnik, vitalybuka, ldionne Spies: mikhail.ramalho, manojgupta, ldionne, AntonBikineev, ayzhao, hans, EricWF, philnik, #sanitizers, libcxx-commits Differential Revision: https://reviews.llvm.org/D136765
-
Joseph Huber authored
The GPU has a different execution model to standard `_start` implementations. On the GPU, all threads are active at the start of a kernel. In order to correctly intitialize and call the constructors we want single threaded semantics. Previously, this was done using a makeshift global barrier with atomics. However, it should be easier to simply put the portions of the code that must be single threaded in separate kernels and then call those with only one thread. Generally, mixing global state between kernel launches makes optimizations more difficult, similarly to calling a function outside of the TU, but for testing it is better to be correct. Depends on D149527 D148943 Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D149581
-
Joseph Huber authored
The execution model of the GPU expects that groups of threads will execute in lock-step in SIMD fashion. It's both important for performance and correctness that we treat this as the smallest possible granularity for an RPC operation. Thus, we map multiple threads to a single larger buffer and ship that across the wire. This patch makes the necessary changes to support executing the RPC on the GPU with multiple threads. This requires some workarounds to mimic the model when handling the protocol from the CPU. I'm not completely happy with some of the workarounds required, but I think it should work. Uses some of the implementation details from D148191. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D148943
-
Craig Topper authored
Reviewed By: fakepaper56 Differential Revision: https://reviews.llvm.org/D149911
-
Alexander Shaposhnikov authored
This reverts commit 3a540229. A new regression is discovered and needs to be investigated.
-
Craig Topper authored
-
Alex Langford authored
This reverts commit 04aa943b. This broke the debian buildbot and I'm not sure why. Reverting so I can investigate.
-