- Apr 12, 2024
-
-
Francis Visoiu Mistrih authored
lr.[wd]: mayLoad = 1, mayStore = 0 sc.[wd]: mayLoad = 0, mayStore = 1 all other AMOs: mayLoad = 1, mayStore = 1
-
Craig Topper authored
This allows us to support larger stack offsets for FrameLowering. Fixes #88365.
-
Nathan Sidwell authored
Use the std `if () continue;` idiom before falling into the processing.
-
Vincent Belliard authored
FormatManager::GetCategoryForLanguage and FormatManager::GetCategory(can_create = true) can be called concurrently and they both take the TypeCategory::m_map_mutex and the FormatManager::m_language_categories_mutex but in reverse order. On one thread, GetCategoryForLanguage takes m_language_categories_mutex and then ends calling TypeCategoryMap::Get which takes m_map_mutex On another thread GetCategory calls TypeCategoryMap::Add which takes m_map_mutex and then calls FormatManager::Changed() which takes m_language_categories_mutex If both threads are running concurrently, we have a dead lock. The patch releases the m_map_mutex before calling Changed which avoids the dead lock. --------- Co-authored-by:Vincent Belliard <v-bulle@github.com>
-
Dominik Adamski authored
ROCm installation path is used for finding and automatically linking required bitcode libraries for OpenMP AMDGPU offload. Reported issue: https://github.com/llvm/llvm-project/issues/82553
-
yronglin authored
This PR fix a AST dump issue since https://github.com/llvm/llvm-project/pull/80001 When Clang dumps `CXXDefaultArgExpr`/`CXXDefaultInitExpr`, there has no recursively dump the complete `CXXDefaultArgExpr`/`CXXDefaultInitExpr`. Since this PR, Clang will recursively dump a `CXXDefaultArgExpr`/`CXXDefaultInitExpr` node, even if the node has no rewritten init. *Consider*: ``` struct A { int arr[1]; }; struct B { const A &a = A{{0}}; }; void test() { B b{}; } ``` *Before*: ``` `-FunctionDecl <line:9:1, line:11:1> line:9:6 test 'void ()' `-CompoundStmt <col:13, line:11:1> `-DeclStmt <line:10:3, col:8> `-VarDecl <col:3, col:7> col:5 b 'B' listinit `-InitListExpr <col:6, col:7> 'B' `-CXXDefaultInitExpr <col:7> 'const A' lvalue has rewritten init `-ExprWithCleanups <line:6:16, col:21> 'const A' lvalue ``` *After*: ``` `-FunctionDecl 0x15a9455a8 <line:9:1, line:11:1> line:9:6 test 'void ()' `-CompoundStmt 0x15a945850 <col:13, line:11:1> `-DeclStmt 0x15a945838 <line:10:3, col:8> `-VarDecl 0x15a945708 <col:3, col:7> col:5 b 'B' listinit `-InitListExpr 0x15a9457b0 <col:6, col:7> 'B' `-CXXDefaultInitExpr 0x15a9457f8 <col:7> 'const A' lvalue has rewritten init `-ExprWithCleanups 0x15a945568 <line:6:16, col:21> 'const A' lvalue `-MaterializeTemporaryExpr 0x15a945500 <col:16, col:21> 'const A' lvalue extended by Field 0x15a945160 'a' 'const A &' `-ImplicitCastExpr 0x15a9454e8 <col:16, col:21> 'const A' <NoOp> `-CXXFunctionalCastExpr 0x15a9454c0 <col:16, col:21> 'A' functional cast to A <NoOp> `-InitListExpr 0x15a9452c0 <col:17, col:21> 'A' `-InitListExpr 0x15a945308 <col:18, col:20> 'int[1]' `-IntegerLiteral 0x15a945210 <col:19> 'int' 0 ``` --------- Signed-off-by:
yronglin <yronglin777@gmail.com>
-
Adam Fowler authored
This includes the lldb-dap executable in the MacOS and Linux distributions of Swift. Currently there is a commit in the Apple repo to do this for just MacOS https://github.com/apple/llvm-project/pull/8176. This PR extends this to both Linux and MacOS and brings the change upstream. @JDevlieghere @adrian-prantl
-
Björn Pettersson authored
The load narrowing part of TargetLowering::SimplifySetCC is updated according to this: 1) The offset calculation (for big endian) did not work properly for non byte-sized types. This is basically solved by an early exit if the memory type isn't byte-sized. But the code is also corrected to use the store size when calculating the offset. 2) To still allow some optimizations for non-byte-sized types the TargetLowering::isPaddedAtMostSignificantBitsWhenStored hook is added. By default it assumes that scalar integer types are padded starting at the most significant bits, if the type needs padding when being stored to memory. 3) Allow optimizing when isPaddedAtMostSignificantBitsWhenStored is true, as that hook makes it possible for TargetLowering to know how the non byte-sized value is aligned in memory. 4) Update the algorithm to always search for a narrowed load with a power-of-2 byte-sized type. In the past the algorithm started with the the width of the original load, and then divided it by two for each iteration. But for a type such as i48 that would just end up trying to narrow the load into a i24 or i12 load, and then we would fail sooner or later due to not finding a newVT that fulfilled newVT.isRound(). With this new approach we can narrow the i48 load into either an i8, i16 or i32 load. By checking if such a load is allowed (e.g. alignment wise) for any "multiple of 8 offset", then we can find more opportunities for the optimization to trigger. So even for a byte-sized type such as i32 we may now end up narrowing the load into loading the 16 bits starting at offset 8 (if that is allowed by the target). The old algorithm did not even consider that case. 5) Also start using getObjectPtrOffset instead of getMemBasePlusOffset when creating the new ptr. This way we get "nsw" on the add.
-
Bjorn Pettersson authored
These test cases show some miscomplies for big-endian when dealing with non byte-sized loads. One part of the problem is that LLVM IR isn't really telling where the padding goes for non byte-sized loads/stores. So currently TargetLowering::SimplifySetCC can't assume anything about it. But the implementation also do not consider that the TypeStoreSize could be larger than the TypeSize, resulting in the offset calculation being wrong for big-endian. Pre-commit for https://github.com/llvm/llvm-project/pull/87646
-
Michael Buch authored
[lldb][test] Add tests for evaluating local variables whose name clashes with Objective-C types (#87807) Depends on https://github.com/llvm/llvm-project/pull/87767
-
Xiaoyang Liu authored
This pull request implements LWG3736: move_iterator missing disable_sized_sentinel_for specialization.
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Michael Maitland authored
Co-authored-by:Wang Pengcheng <wangpengcheng.pp@bytedance.com>
-
Xiaoyang Liu authored
This pull request implements LWG3643: Missing constexpr in std::counted_iterator. Specifically, one overload of std::counted_operator::operator++ was not marked as constexpr, despite being eligible for it after the introduction of try-block support in constexpr functions in C++20.
-
Shilei Tian authored
When the alloca is too big for vectorization, the function could have already been modified in previous iteration of the `for` loop.
-
Brandon Wu authored
This reverts commit 29e8bfc1. This patch didn't handle vector return type correctly.
-
Simon Pilgrim authored
[VectorCombine] foldShuffleOfCastops - ensure we can scale shuffle masks between bitcasted vector types Don't just assert that the src/dst vector element counts are multiples of one another - in general IR this can actually happen. Reported by @mikaelholmen
-
Tom Eccles authored
Reverts llvm/llvm-project#88395 This broke the powerpc buildbot. That build doesn't support using `std::filesystem` in flang unit tests.
-
Tom Eccles authored
This is a GNU extension: https://gcc.gnu.org/onlinedocs/gfortran/ACCESS.html Used in SALMON: https://salmon-tddft.jp/download.html Unfortunately the intrinsic takes a file path to operate on so there isn't an easy way to make the test robust. The unit test expects to be able to create, set read write and execute permissions, and delete files called `std::filesystem::temp_directory_path() / <test_name>.<pid>` The test will fail if a file already exists with that name. I have not implemented the intrinsic on Windows because this is wrapping a POSIX system call and Windows doesn't support all of the permission bits tested by the intrinsic. I don't have a Windows machine easily available to check if Gfortran implements this intrinsic on Windows.
-
Nathan Sidwell authored
This originally dealt with tbss, but now handles any bss-like section. So the comment is inaccurate. Also, the `{}` on the messaging seem unnecessary. -
SahilPatidar authored
Resolve #84905
-
Vlad Serebrennikov authored
This patch covers [CWG393](https://cplusplus.github.io/CWG/issues/393.html) "Pointer to array of unknown bound in template argument list in parameter", [CWG528](https://cplusplus.github.io/CWG/issues/528.html) "Why are incomplete class types not allowed with `typeid`?", [CWG550](https://cplusplus.github.io/CWG/issues/550.html) "Pointer to array of unknown bound in parameter declarations", [CWG553](https://cplusplus.github.io/CWG/issues/553.html) "Problems with friend allocation and deallocation functions", [CWG555](https://cplusplus.github.io/CWG/issues/555.html) "Pseudo-destructor name lookup", [CWG560](https://cplusplus.github.io/CWG/issues/560.html) "Use of the `typename` keyword in return types". CWG393 is on this list, because CWG550 is marked as a duplicate of CWG393. Test for CWG553 has been already written, but it was missing a status comment. As a drive-by fix, I'm adding missing status comments to CWG1584 and CWG1903 as well. CWG555 used CWG466 test, and also a variation of that test to test references. CWG466 is now testing non-reference non-pointer case. CWG560 showcases again that converting warnings to errors in DR tests via `-pedantic-errors` doesn't make things more clear. By default that test is accepted with an extension warning since we implemented [P0634R3](https://wg21.link/p0634r3) "Down with `typename`!" in Clang 16.
-
Aaron Ballman authored
This diagnostic used to talk about complex integer types but is issued for use of the increment or decrement operators on any complex type, not just integral ones. The diagnostic also had zero test coverage, so new coverage is added along with the rewording.
-
Sergio Afonso authored
This patch updates Flang lowering to use the new set of OpenMP clause operand structures and their groupings into directive-specific sets of clause operands. It simplifies the passing of information from the clause processor and the creation of operations. The `DataSharingProcessor` is slightly modified to not hold delayed privatization state. Instead, optional arguments are added to `processStep1` which are only passed when delayed privatization is used. This enables using the clause operand structure for `private` and removes the need for the ad-hoc `DelayedPrivatizationInfo` structure. The processing of the `schedule` clause is updated to process the `chunk` modifier rather than requiring two separate calls to the `ClauseProcessor`. Lowering of a block-associated `ordered` construct is updated to emit a TODO error if the `simd` clause is specified, since it is not currently supported by the `ClauseProcessor` or later compilation stages. Removed processing of `schedule` from `omp.simdloop`, as it doesn't apply to `simd` constructs.
-
Simon Pilgrim authored
If the extract_subvector is cheap, attempt to extract directly from an inserted subvector
-
Simon Pilgrim authored
-
Vlad Serebrennikov authored
Refactor `CUDAFunctionTarget` into a scoped enum at namespace scope, so that it can be forward declared. This is done in preparation for `SemaCUDA`.
-
wanglei authored
-
David Green authored
By default the scheduling info of instructions into a BUNDLE are given a latency of 0 as they operate on the implicit register of the bundle. This modifies that for AArch64 so that the latency is adjusted to use the latency from the instruction in the bundle instead. This essentially assumes that the bundled instructions are executed in a single cycle, which for AArch64 is probably OK considering they are mostly used for MOVPFX bundles, where this can help create slightly better scheduling especially for in-order cores.
-
Benjamin Kramer authored
-
Sergio Afonso authored
This patch introduces an operation intended to hold loop information associated to the `omp.distribute`, `omp.simdloop`, `omp.taskloop` and `omp.wsloop` operations. This is a stopgap solution to unblock work on transitioning these operations to becoming wrappers, as discussed in [this RFC](https://discourse.llvm.org/t/rfc-representing-combined-composite-constructs-in-the-openmp-dialect/76986). Long-term, this operation will likely be replaced by `omp.canonical_loop`, which is being designed to address missing support for loop transformations, etc.
-
AtariDreams authored
-
Andreas Jonson authored
-
Vlad Serebrennikov authored
In preparation for `SemaCUDA`, which requires this enum to be forward-declarable.
-
Beal Wang authored
This diff causes the `tblgen`-erated printProperties() function to skip printing a `DefaultValuedAttr` property when the value is equal to the default. Co-authored-by:Biao Wang <biaow@nvidia.com>
-
jeanPerier authored
Fortran mandates "CHARACTER(1), VALUE" be passed as a C "char" in calls to BIND(C) procedures (F'2023 18.3.7 (4)). Lowering passed them by memory instead. Update call interface lowering code to pass them by register. Fix related test and update it to use HLFIR.
-