- Oct 13, 2020
-
-
Martin Storsjö authored
This unfortunately means that we don't execute C++ destructors when unwinding past such frames for a different SEH unwind purpose (e.g. as part of setjmp/longjmp), but that case isn't handled properly at the moment (the original unwind intent is lost and we end up with an unhandled exception). This patch makes sure the foreign unwind terminates as intended. After executing a handler, _Unwind_Resume doesn't have access to the target frame parameter of the original foreign unwind. We also currently blindly set ExceptionCode to STATUS_GCC_THROW - we could set that correctly by storing the original code in _GCC_specific_handler, but we don't have access to the original target frame value. This also matches what libgcc's SEH unwinding code does in this case. Differential Revision: https://reviews.llvm.org/D89231
-
JonChesterfield authored
[libomptarget][amdgcn] Implement partial barrier named_sync is used to coordinate non-spmd kernels. This uses bar.sync on nvptx. There is no corresponding ISA support on amdgcn, so this is implemented using shared memory, one word initialized to zero. Each wave increments the variable by one. Whichever wave is last is responsible for resetting the variable to zero, at which point it and the others continue. The race condition on a wave reaching the barrier before another wave has noticed that it has been released is handled with a generation counter, packed into the same word. Uses a shared variable that is not needed on nvptx. Introduces a new hook, kmpc_impl_target_init, to allow different targets to do extra initialization. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D88602
-
Nicolas Vasilache authored
The TensorConstantOp bufferize conversion pattern has a bug that makes it incorrect in the case of vectors whose alignment is not the natural alignment. Circumvent it temporarily by using a power of 2. Differential Revision: https://reviews.llvm.org/D89265
-
Nico Weber authored
It's built with just-built clang, like all other compiler-rt parts in the GN build. This requires adding some cross build support to the mac toolchain. Also add explicit mmacosx-version-min and miphoneos-version-min flags to the build. ios.a is only built with the arm64 slice, iossim.a only with the x86_64 slice for now. (The latter should maybe become host_cpu when Arm Macs become a common iOS development platform.) With this, it's possible to build chromium/iOS with a GN-built LLVM. Differential Revision: https://reviews.llvm.org/D89260
-
Tony authored
Change-Id: Ie409f86876b0437d0b0405aff42872963708d926 Differential Revision: https://reviews.llvm.org/D89259
-
Roman Lebedev authored
Reland "[SCEV] Model ptrtoint(SCEVUnknown) cast not as unknown, but as zext/trunc/self of SCEVUnknown" This relands commit 1c021c64 which was reverted in commit 17cec6a1 because an assertion was being triggered, since `BuildConstantFromSCEV()` wasn't updated to handle the case where the constant we want to truncate is actually a pointer. I was unsuccessful in coming up with a test case where we'd end there with constant zext/sext of a pointer, so i didn't handle those cases there until there is a test case. Original commit message: While we indeed can't treat them as no-ops, i believe we can/should do better than just modelling them as `unknown`. `inttoptr` story is complicated, but for `ptrtoint`, it seems straight-forward to model it just as a zext-or-trunc of unknown. This may be important now that we track towards making inttoptr/ptrtoint casts not no-op, and towards preventing folding them into loads/etc (see D88979/D88789/D88788) Reviewed By: mkazantsev Differential Revision: https://reviews.llvm.org/D88806
-
Roman Lebedev authored
Reduced from the https://reviews.llvm.org/D88806#2325340
-
Arthur Eubanks authored
This reverts commit 9dcd96f7. See https://crbug.com/1134762.
-
Cameron McInally authored
This seems to be a typo that propagated to a number of tests. Replace VBITS_GE_256 with CHECK. There is no VBITS_GE_256.
-
Walter Erquinigo authored
Depends on D88841 As per the discussion in the RFC, we'll implement both thread trace dump [instructions | functions] This is the first step in implementing the "instructions" dumping command. It includes: - A minimal ProcessTrace plugin for representing processes from a trace file. I noticed that it was a required step to mimic how core-based processes are initialized, e.g. ProcessElfCore and ProcessMinidump. I haven't had the need to create ThreadTrace yet, though. So far HistoryThread seems good enough. - The command handling itself in CommandObjectThread, which outputs a placeholder text instead of the actual instructions. I'll do that part in the next diff. - Tests {F13132325} Differential Revision: https://reviews.llvm.org/D88769 -
Valentin Clement authored
This patch upstream the lowering of Data construct that was initially done in https://github.com/flang-compiler/f18-llvm-project/pull/460. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D88918
-
Xun Li authored
In https://reviews.llvm.org/D87470 I added the change to tighten the lifetime of the expression awaiter.await_suspend().address. Howver it was incorrect. ExprWithCleanups will call the dtor and end the lifetime for all the temps created in the current full expr. When this is called on a normal await call, we don't want to do that. We only want to do this for the call on the final_awaiter, to avoid writing into the frame after the frame is destroyed. This change fixes it, by checking IsImplicit. Differential Revision: https://reviews.llvm.org/D89066
-
brett koonce authored
-
Ben Vanik authored
Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D89255
-
Adrian McCarthy authored
A Windows-style LLDB_PYTHON_HOME path in a Cmake template didn't have the backslashes escaped, which led to a garbled paths derived from it. Fixed by expanding the environment variable as a raw string literal. Differential Revision: https://reviews.llvm.org/D89256
-
Arthur Eubanks authored
alloca-dbgdeclare-merge.ll: alloca-merge-align.ll: array_merge.ll: NPM inliner does not merge allocas delete-call.ll: NPM inliner does not delete readonly calls externally_available.ll: NPM inliner does not delete available_externally functions inline-cold-callee.ll: inline-hot-callee.ll: inline-hot-callee.ll has a comment saying it only applies to legacy PM, I assume same for inline-cold-callee.ll devirtualize-2.ll: inline-hot-callsite: monster_scc.ll: pr22285.ll: already has legacy and new PM RUN lines inline-cold.ll: profile-summary required to see callee as cold prof-update-sample.ll: profile-summary required to update branch_weights Reviewed By: davidxl Differential Revision: https://reviews.llvm.org/D89093
-
Simon Pilgrim authored
There's no need to create constant vector splats manually - missed this one in rG24dd0cd1
-
Nathan Ridge authored
Fixes https://github.com/clangd/clangd/issues/543 Differential Revision: https://reviews.llvm.org/D88469
-
Adhemerval Zanella authored
ARM thumb/thumb2 frame pointer is inconsistent on GCC and Clang [1] and fast-unwider is also unreliable when mixing arm and thumb code [2]. The fast unwinder on ARM tries to probe and compare the frame-pointer at different stack layout positions and it works reliable only on systems where all the libraries were built in arm mode (either with gcc or clang) or with clang in thmb mode (which uses the same stack frame pointer layout in arm and thumb). However when mixing objects built with different abi modes the fast unwinder is still problematic as shown by the failures on the AddressSanitizer.ThreadStackReuseTest. For these failures, the malloc is called by the loader itself and since it has been built with a thum enabled gcc, the stack frame is not correctly obtained and the suppression rule is not applied (resulting in a leak warning). The check for fast-unwinder-works is also changed: instead of checking f it is explicit enabled in the compiler flags, it now checks if compiler defined thumb pre-processor. This should fix BZ#44158. [1] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=92172 [2] https://bugs.llvm.org/show_bug.cgi?id=44158 Reviewed By: eugenis Differential Revision: https://reviews.llvm.org/D88958
-
Fangrui Song authored
PR47686. These micro-architecture levels are defined in the x86-64 psABI: https://gitlab.com/x86-psABIs/x86-64-ABI/-/commit/77566eb03bc6a326811cb7e9 GCC 11 will support these levels. Note, -mtune=x86-64-v[234] are invalid and __builtin_cpu_is cannot be used on them. Reviewed By: craig.topper, RKSimon Differential Revision: https://reviews.llvm.org/D89197
-
Valentin Clement authored
This patch upstream the lowering of Parallel construct that was initially done in https://github.com/flang-compiler/f18-llvm-project/pull/460. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D88917
-
Valentin Clement authored
This patch update the loop construct lowring to match fir-dev changes. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D88914
-
Simon Pilgrim authored
There's no need to create constant vector splats manually.
-
Simon Pilgrim authored
Consistently use the original shift instruction's Type/BitWidth instead of the operands, casted values etc.
-
Teresa Johnson authored
This restores commit ab1b4810 which was reverted in 01b9deba, with a fix for the issue it caused. We should use a temporary BitstreamCursor when loading the global decl attachment records so that the abbrev ids held in the lazy loading IndexCursor are not clobbered. Enhanced the test so that the issue is exposed there. Original description: When performing ThinLTO importing, the metadata loader attempts to lazy load, by building an index. However, module level global decl attachment metadata was being parsed early while building the index, since the associated (module level) global values aren't materialized on demand. This results in the creation of forward reference temporary metadatas, which are expensive. Normally, these module level global values don't have much attached metadata. However, in the case of -fwhole-program-vtables (e.g. for whole program devirtualization), the vtables may have many attached type metadatas. This was resulting in very slow performance when performing ThinLTO importing with the default lazy loading. This patch restructures the handling of these global decl attachment records, delaying their parsing until after the lazy loading index has been built. Then the parser can use the interface that loads from the index, which resolves forward references immediately instead of creating expensive temporaries. For one ThinLTO backend that imports from modules containing huge numbers of vtables and associated types, I measured the following compile times for the metadata materialization during function importing, rounded to nearest second: No -fwhole-program-vtables: Lazy loading on (head): 1s Lazy loading off (head): 3s Lazy loading on (patch): 1s With -fwhole-program-vtables: Lazy loading on (head): 440s Lazy loading off (head): 4s Lazy loading on (patch): 2s Differential Revision: https://reviews.llvm.org/D87970
-
Florian Hahn authored
This patch turns VPMemoryInstructionRecipe into a VPValue and uses it during VPlan construction and codegeneration instead of the plain IR reference where possible. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D84680
-
Mark de Wever authored
Jeremy Morse discovered an issue with the lit test introduced in D88363. The test gives different results for Sony's `-O1`. The test needs to run at `-O1` otherwise the likelihood attribute will be ignored. Instead of running all `-O1` passes it only runs the lower-expect pass which is needed to lower `__builtin_expect`. Differential Revision: https://reviews.llvm.org/D89204
-
Fangrui Song authored
Noticed by Peter Foley. In glibc, ::write is declared as __attribute__((__warn_unused_result__)) when __USE_FORTIFY_LEVEL is larger than 0.
-
Hans Wennborg authored
Revert 1c021c64 "[SCEV] Model ptrtoint(SCEVUnknown) cast not as unknown, but as zext/trunc/self of SCEVUnknown" > While we indeed can't treat them as no-ops, i believe we can/should > do better than just modelling them as `unknown`. `inttoptr` story > is complicated, but for `ptrtoint`, it seems straight-forward > to model it just as a zext-or-trunc of unknown. > > This may be important now that we track towards > making inttoptr/ptrtoint casts not no-op, > and towards preventing folding them into loads/etc > (see D88979/D88789/D88788) > > Reviewed By: mkazantsev > > Differential Revision: https://reviews.llvm.org/D88806 It caused the following assert during Chromium builds: llvm/lib/IR/Constants.cpp:1868: static llvm::Constant *llvm::ConstantExpr::getTrunc(llvm::Constant *, llvm::Type *, bool): Assertion `C->getType()->isIntOrIntVectorTy() && "Trunc operand must be integer"' failed. See code review for a link to a reproducer. This reverts commit 1c021c64.
-
Konstantin Schwarz authored
If the known shift amount is bigger than or equal to the bitwidth of the type of the value to be shifted, the result is target dependent, so don't try to infer any bits. This fixes a crash we've seen in one of our internal test suites. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D89232
-
- Oct 12, 2020
-
-
Dávid Bolvanský authored
-
Mircea Trofin authored
The change starts from LiveRangeMatrix and also checks the users of the APIs are typed accordingly. Differential Revision: https://reviews.llvm.org/D89145
-
Florian Hahn authored
Now that operands of the recipe are managed through VPUser, we can simplify the printing by just using the operands.
-
Mircea Trofin authored
It's never null - the reason it's modeled as a pointer is because the pass can't init it in its ctor. Passing by ref simplifies the code, too, as the null checks were unnecessary complexity. Differential Revision: https://reviews.llvm.org/D89171
-
Sebastian Neubauer authored
If the metadata is valid yaml, we can print it, even if it failed validation. That makes it easier to debug any wrong metadata. Differential Revision: https://reviews.llvm.org/D89243
-
Florian Hahn authored
60b85209 introduced SCEV verification to deleteDeadLoop, but it appears this check is currently a bit over-eager and some users of deleteDeadLoop appear to only patch up SE after calling it (e.g. PR47753). Remove the extra check for now. We can consider adding it back after we tracked down the source of the inconsistency for PR47753.
-
Sebastian Neubauer authored
Extend loadSRsrcFromVGPR to allow moving a range of instructions into the loop. The call instruction is surrounded by copies into physical registers which should be part of the waterfall loop. Differential Revision: https://reviews.llvm.org/D88291
-
Cameron McInally authored
Differential Revision: https://reviews.llvm.org/D88974
-
Jay Foad authored
-
Simon Pilgrim authored
[InstCombine] matchFunnelShift - fold or(shl(a,x),lshr(b,sub(bw,x))) -> fshl(a,b,x) iff x < bw (REAPPLIED) If value tracking can confirm that a shift value is less than the type bitwidth then we can more confidently fold general or(shl(a,x),lshr(b,sub(bw,x))) patterns to a funnel/rotate intrinsic pattern without causing bad codegen regressions in the backend (see D89139). Reapplied after the shift canonicalization in rG02295e6d which removed the need to flip the shift values. Differential Revision: https://reviews.llvm.org/D88783
-