- Jun 21, 2022
-
-
Mehdi Chinoune authored
This turns off a bunch of non-standard behaviors in MSVC. LLVM, as a portable codebase, should build correctly without those behaviors. Note that `/permissive-` implies `/Zc:strictStrings` and `/Zc:rvalueCast`. See also: https://docs.microsoft.com/en-us/cpp/build/reference/permissive-standards-conformance Differential Revision: https://reviews.llvm.org/D125263
-
Amir Ayupov authored
This reverts commit 4cd41619.
-
Florian Hahn authored
-
Nemanja Ivanovic authored
There are instances where using paired vector stores leads to significant performance degradation due to issues with store forwarding.To avoid falling into this trap with compiler - generated code, we will not emit these instructions unless the user requests them explicitly(with a builtin or by specifying the option). Reviewed By : lei, amyk, saghir Differential Revision: https://reviews.llvm.org/D127218
-
Amir Ayupov authored
Make Offsets and OpcodeOperandTypes tables human-readable by printing the instruction name before the operand list. In effect, this makes debugging generated `getOperandType` possible. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D127931
-
Jakob Johnson authored
Add trace load functionality to SBDebugger via the `LoadTraceFromFile` method. Update intelpt test case class to have `testTraceLoad` method so we can take advantage of the testApiAndSB decorator to test both the CLI and SB without duplicating code. Differential Revision: https://reviews.llvm.org/D128107
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
David Green authored
An AArch64ISD::DUP is just a splat, where the known bits for each lane are the same as the input. This teaches that to computeKnownBitsForTargetNode. Problems arise for constants though, as a constant BUILD_VECTOR can be lowered to an AArch64ISD::DUP, which SimplifyDemandedBits would then turn back into a constant BUILD_VECTOR leading to an infinite cycle. This has been prevented by adding a isTargetCanonicalConstantNode node to prevent the conversion back into a BUILD_VECTOR. Differential Revision: https://reviews.llvm.org/D128144
-
Kazu Hirata authored
-
Simon Pilgrim authored
v32i8/v16i16 blend shuffles on AVX1 will expand to OR(AND,ANDN) patterns which can be easily broken by other combines
-
Michał Górny authored
Fix test_platform_file_fstat to correctly truncate/max out the expected value when GDB Remote Serial Protocol specifies a value as an unsigned integer but the underlying platform type uses a signed integer. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.llvm.org/D128042
-
Michał Górny authored
Make the AVX/MPX register tests more robust by checking for the presence of actual registers rather than register sets. Account for the option that the respective registers are defined but not available, as is the case on FreeBSD and NetBSD. This fixes test regression on these platforms. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.llvm.org/D128041
-
Michał Górny authored
The -gmodule tests currently fail on FreeBSD due to include bugs: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=264730 Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.llvm.org/D128034
-
Michał Górny authored
Refactor GDBRemoteCommunicationServerLLGS::SendStopReasonForState() to accept process as an argument rather than hardcoding m_current_process, in order to make it work correctly for multiprocess scenarios. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.llvm.org/D127497
-
Michał Górny authored
Refactor SendStopReplyPacketForThread() to accept process instance as a parameter rather than use m_current_process. This future-proofs it for multiprocess support. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.llvm.org/D127289
-
Philip Reames authored
This change removes an explicit scalable vector bailout for fshl and fshr. This bailout was added in 60e4698b, when sinking a unconditional bailout for all intrinsics into selected cases. Its not clear if the bailout was originally unneeded, or if our cost model infrastructure has simply matured in the meantime. Either way, the generic code appears to handle scalable vectors without issue. Note that the RISC-V cost model changes here aren't particularly interesting. They do probably better match the current lowering, but the main point is to have coverage of the BasicTTI path and simply show lack of crashing. AArch64 costing was changed to preserve legacy behavior. There will most likely be an upcoming change to use the generic costs there too, but I didn't want to make that change not being particularly familiar with the target. Differential Revision: https://reviews.llvm.org/D127680
-
Kazu Hirata authored
-
Stanislav Gatev authored
Extend flow condition in the body of a do/while loop. Differential Revision: https://reviews.llvm.org/D128183 Reviewed-by: gribozavr2, xazax.hun
-
Arthur Eubanks authored
This reverts commit 6f348b14. Am seeing internal test failures plus a linux kernel breakage reported due to this.
-
Arthur Eubanks authored
This reverts commit cc65f3e1. Causes crashes: https://github.com/llvm/llvm-project/issues/56131
-
Philip Reames authored
The code being removed is technically correct; if we end up with two VL=0 instructions next to each other, we can avoid a state transition if the second is a scalar move. However, since both ops are also nops, we should simply delete them instead. As such, this compatibility rule simply complicates the code for no purpose.
-
- Jun 20, 2022
-
-
David Candler authored
Depending on the environment, a floating point instruction should treat denormal inputs as zero, and/or flush a denormal output to zero. Denormals are not currently accounted for when an instruction gets folded to a constant, which can lead to differences in output between a folded and a unfolded instruction when running on the target. The denormal handling mode can be set by the function level attribute denormal-fp-math, which this patch uses to determine whether any denormal inputs to or outputs from folding should be zero, and that the sign is set appropriately. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D116952
-
Fraser Cormack authored
The example wouldn't compile, and used an invalid case style for a function. Reviewed By: MatzeB Differential Revision: https://reviews.llvm.org/D128176
-
Guillaume Chatelet authored
-
Guillaume Chatelet authored
-
Guillaume Chatelet authored
-
Guillaume Chatelet authored
-
Florian Hahn authored
-
Krzysztof Drewniak authored
In order to support newer hardware, define wrappers around MFMA intrinsics that have not previously been exposed in the ROCDL dialect. A `amdgpu.mfma` wrapper around these instructions is in development and will provide a more user-friendly interface to them. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D128079
-
Krzysztof Drewniak authored
Reviewed By: Mogball Differential Revision: https://reviews.llvm.org/D128096
-
Philip Reames authored
When working through correctness issues in this pass, I moved a number of transforms which were phrased as mutating prior vsetvli instructions out of the main data flow because mutating prior instructions can invalidate the running dataflow results in subtle ways. We ended up creating both a prepass and a post-pass. After consideration, I believe the prepass to be redundant, and this change removes it by folding it back into the data flow via a key conceptual change. Instead of phrasing the mutations on instructions, we can phrase them on abstract states. This avoids the dataflow inconsistency problem mentioned above by simply propagating the potential change forward, and thus reflecting its results in the dataflow. Critically, we do so without modifying existing VSETVLI instructions; some of the data flow steps include non-local IR analysis. Compile time wise, this removes a linear pass, but has the potential to increase the number of iterations for the data flow to converge. That's not a algorithmic complexity change, the needVSETVLI mechanism has the same effect. In practice, I don't see this triggering more iterations, so I think it's likely to be a net win overall. (I didn't do any careful analysis here; just an impression from glancing at a couple tests.) This has the potential to produce better results, so this isn't strictly speaking NFC. Differential Revision: https://reviews.llvm.org/D127870
-
Jan Svoboda authored
This is to fix the following error on https://green.lab.llvm.org/green/job/clang-stage2-Rthinlto: BranchProbability.h:236:34: error: declaration of 'distance' must be imported from module 'std.iterator.__iterator.distance' before it is required
-
Philip Reames authored
In D127983, I had flipped from using the computed EEW to using the SEW value pulled from the VSETVLI when checking compatibility. This wasn't intentional, though thankfully it appears to be a non-functional difference. The new code does make a unchecked assumption that the initial SEW operand on the load/store is the EEW. This patch clarifies the assumption, and adds an assert to make sure this remains true. Differential Revision: https://reviews.llvm.org/D128085
-
Kadir Cetinkaya authored
Differential Revision: https://reviews.llvm.org/D128197
-
Florian Hahn authored
-
David Candler authored
These tests demonstrate cases where the constant produced by folding a floating point instruction should differ based on the denormal handling mode set in function attributes. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D125807
-
Guillaume Chatelet authored
-
Valentin Clement authored
This patch is part of the upstreaming effort from fir-dev branch. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D128186 Co-authored-by:
Peter Steinfeld <psteinfeld@nvidia.com>
-