- Mar 29, 2023
-
-
Johannes de Fine Licht authored
This is a subset of the full LLVM functionality to detect whether realignment is necessary, conservatively copying byval arguments whenever we cannot prove that the alignment requirement is met. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D147049
-
Christian Ulmann authored
This commit makes the name of a DINamespace optional to enable modeling of anonymous namespaces. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D147125
-
David Sherwood authored
This one-line patch just tightens up the code added in 1c4fedfa where we try to avoid tail-folding if we know the runtime VF will always be a multiple of the trip count.
-
Paul Osmialowski authored
This commit extends D134719 "[AArch64] Enable libm vectorized functions via SLEEF" with the mappings for the scalable functions. It also introduces all the necessary changes needed to support masked interfaces. Signed-off-by:Paul Osmialowski <pawel.osmialowski@arm.com>
-
Paul Osmialowski authored
Signed-off-by:Paul Osmialowski <pawel.osmialowski@arm.com>
-
LLVM GN Syncbot authored
-
-
Michael Halkenhaeuser authored
Adds a warning, issued by the clang semantic analysis, if HIP and OpenMP target offloading is requested concurrently. That is, if HIP language mode is active but OpenMP target directives are encountered. Previously, a user might not have been aware that target directives are ignored in such a case. Generation of this warning is (lit-)tested via "make check-clang-semaopenmp". The warning can be ignored via "-Wno-hip-omp-target-directives". Differential Revision: https://reviews.llvm.org/D145591
-
Serguei Katkov authored
Guard widening optimization is able to move the condition from one guard to the previous one. As a result if the condition is poison and orginal second guard is never executed but the first one does, we introduce undefined behavior which was not observed in original program. To resolve the issue we must freeze the condition we are moving. However optimization itself does not know how to work with freeze. Additionally optimization is written in incremental way. For example we have three guards G1(base + 8 < L) G2(base + 16 < L) G3(base + 24 < L) On the first step GW will combine G1 and G2 as G1(base + 8 < L && freeze(base + 16 < L)) G2(true) G3(base + 24 < L) while combining G1 and G3 base appears to be different. To keep optimization enabled after freezing the moving condition, the freeze instruction is pushed as much as possible and later all uses of freezed values are replaced with frozen version. This is similar what instruction combining does but more aggressevely. Reviewed By: mkazantsev Differential Revision: https://reviews.llvm.org/D146699
-
Graham Hunter authored
Precommit for D145163
-
Haojian Wu authored
Reviewed By: kadircet Differential Revision: https://reviews.llvm.org/D146717
-
Simi Pallipurath authored
Changes: - Adding BE32 big endian Support for Arm. - Replace the writele and readle with their endian-aware versions. - Adding test cases for the big-endian be32 arm configuration. Patch by: Milosz Plichta. This patch merges all the changes from this patch https://reviews.llvm.org/D140203 as well. Reviewed By: peter.smith, MaskRay Differential Revision: https://reviews.llvm.org/D140202 -
Matthias Springer authored
This should have been part of D147039.
-
Matthias Springer authored
This change makes it possible to use a greedy pattern rewrite as part of a transform op, even if the transform op does not invalidate the target handle (in particular transform ops without `FunctionalStyleTransformOpTrait`) and the targeted op is not isolated from above. The listener API allows us to track replacements of ops with values, but not ops with ops. Therefore, the TrackingListener is conservative: If an op is replaced with values that all have the same defining op and the defining op is of the same type as the original op, it is safe to assume that the op was replaced with an equivalent op. Otherwise, the op mapping is dropped. When this is not good enough, transforms can track values instead or provide a custom `findReplacementOp` function. Differential Revision: https://reviews.llvm.org/D147039
-
Matthias Springer authored
Differential Revision: https://reviews.llvm.org/D147038
-
Marco Elver authored
Preserve !pcsections metadata on X86-only atomic intrinsics when expanding higher-level atomics. Differential Revision: https://reviews.llvm.org/D147123
-
Serguei Katkov authored
LoopPredication introduces the use of possibly posion value in branch (guard) instruction, so to avoid introducing undefined behavior it should be frozen. Reviewed By: mkazantsev Differential Revision: https://reviews.llvm.org/D146685
-
Yeting Kuo authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D147120
-
Johannes Reifferscheid authored
This header can't be built standalone. Making it textual will prevent blaze from attempting to do so.
-
Dominik Adamski authored
Scope of changes: 1) Extract common code between Clang and Flang for parsing AMDGPU features 2) Add function which adds implicit target features for AMDGPU as Clang does 3) Add AMDGPU target as one of valid targets for Flang Differential Revision: https://reviews.llvm.org/D145579 Reviewed By: yaxunl, awarzynski
-
Matthias Springer authored
This helper function is used for both ExtractSliceOp and InsertSliceOp. Also fixes a bug in the implementation of `InsertSliceOp::getDroppedDims`. Differential Revision: https://reviews.llvm.org/D147048
-
Henry Yu authored
[AsmPrinter] Fix Crash when Emitting Global Constant of small bit width when targeting Big Endian arch For Big Endian, the function `emitGlobalConstantLargeInt` tries to right shift `Realigned` by an amount `ExtraBitSize` in place. However, if the constant to emit has a bit width less than 64 and the bit width is not a multiple of 8, the shift amount will be greater than the bit width of `Realigned`, which causes assertion error described in issue [[ https://github.com/llvm/llvm-project/issues/59055 | issue #59055 ]]. This patch fixes the issue by avoiding right shift when bit width is under 64 to avoid the assertion error. Reviewed By: Peter Differential Revision: https://reviews.llvm.org/D138246
-
Sameer Sahasrabuddhe authored
Use a SetVector to store blocks in a cycle to ensure a quick loop-up when querying whether the cycle contains a given block. This is along the same lines as the SmallPtrSet in LoopBase, introduced by commit be640b28. To make this work, we also enhance SetVector to support vector operations with pointers and set operations with const pointers in the same container. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D146136
-
Sergei Barannikov authored
Without this all checks fail because CMake passes the flags like this: `... -nodefaultlibs -D-Warray-bounds -Werror -std=c++17 ...` Note the `-D` before the `-W`. Reviewed By: ahatanak Differential Revision: https://reviews.llvm.org/D146920
-
Iain Sandoe authored
We need to be able to distinguish individual TUs from the same module in cases where TU-local entities either need to be hidden (or, for some cases of ADL in template instantiation, need to be detected as exposures). This creates a module type for the implementation which implicitly imports its primary module interface per C++20: [module.unit/8] 'A module-declaration that contains neither an export-keyword nor a module-partition implicitly imports the primary module interface unit of the module as if by a module-import-declaration. Implementation modules are never serialized (-emit-module-interface for an implementation unit is diagnosed and rejected). Differential Revision: https://reviews.llvm.org/D126959
-
Chuanqi Xu authored
This reverts commit af86957c. Close https://github.com/llvm/llvm-project/issues/61733. Previously I banned the eagerly loading for declarations from named modules to speedup the process of reading modules. But I didn't think about special decls like PragmaCommentDecl and PragmaDetectMismatchDecl. So here is the issue https://github.com/llvm/llvm-project/issues/61733. Note that the current behavior is still incorrect. Given: ``` // mod.cppm module; export module mod; ``` and ``` // user.cpp import mod; ``` Now the IR of `user.cpp` will contain the metadata '!0 = !{!"msvcprt.lib"}' incorrectly. The root cause of the problem is that `EagerlyDeserializedDecls` is designed for headers and it didn't take care for named modules. We need to redesign a new mechanism for named modules.
-
Anubhab Ghosh authored
-
Anubhab Ghosh authored
This commit adds the %lib <file> command to load a dynamic library to be used by the currently running interpreted code. For example `%lib libSDL2.so`. Differential Revision: https://reviews.llvm.org/D141824
-
Ben Shi authored
Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D147100
-
Brad Smith authored
Invert the logic and have the default being true. Disable the few spots where it looks like IAS is currently not used. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D147030
-
Fangrui Song authored
The .init_array code is ELF specific. For ELF platforms, `__USER_LABEL_PREFIX__` is defined as "". Make the simplification so that downstream ELF targets can build this file even if `__USER_LABEL_PREFIX__` is undefined. Reviewed By: barannikov88 Differential Revision: https://reviews.llvm.org/D147093
-
Phoebe Wang authored
This reverts commit db6a979a. Reland D102817 without any change. The previous revert was a mistake. Differential Revision: https://reviews.llvm.org/D102817
-
Roy Sundahl authored
This test has to be limited to darwin due to multiple failures on other platforms for multple reasons. (Timeout, puts() limit, etc.). This commit modifies D146189. Reviewed By: NoQ Differential Revision: https://reviews.llvm.org/D147094
-
Thomas Raoux authored
Without this bufferization cannot track operations removed during bufferization. Unfortunately there is currently no way to enforce that ops need to be erased through the rewriter and this causes sporadic errors when tracking pointers in Bufferization pass. Therefore there is no easy way to test that the pattern is doing the right thing. Reviewed By: mravishankar Differential Revision: https://reviews.llvm.org/D147095
-
wren romano authored
These warnings were introduced by D146561. Reviewed By: aartbik, Peiming Differential Revision: https://reviews.llvm.org/D147090
-
Aaron Siddhartha Mondal authored
Originally added in D128465. Used by `llvm:Support` and `lld:ELF`. Enabled by default. Disable with `--@llvm_zstd//:llvm_enable_zstd=false`. Reviewed By: MaskRay, GMNGeoffrey Differential Revision: https://reviews.llvm.org/D143344
-
Hongtao Yu authored
[CSSPGO][Preinliner] Trim cold call edges of the profiled call graph for a more stable profile generation. I've noticed that for some services CSSPGO profile is less stable than non-CS AutoFDO profile from profiling to profiling without source changes. This is manifested by comparing profile similarities. For example in my experiments, AutoFDO profiles are always 99+% similar over same binary but different inputs (very close dynamic traffics) while CSSPGO profile similarity is around 90%. The main source of the profile stability is the top-down order computed on the profiled call graph in the llvm-profgen CS preinliner. The top-down order is used to guide the CS preinliner to pre-compute an inline decision that is later on fulfilled by the compiler. A subtle change in the top-down order from run to run could cause a different inline decision computed. A deeper look in the diversion of the top-down order revealed that: - The topological sorting inside one SCC isn't quite right. This is fixed by {D130717}. - The profiled call graphs of the two sides of the A/B run isn't 100% the same. The call edges in the two runs do not subsume each other, and edges appear in both graphs may not have exactly the same weight. This is due to the nature that the graphs are dynamic. However, I saw that the graphs can be made more close by removing the cold edges from them and this bumped up the CSSPGO profile stableness to the same level of the AutoFDO profile. Removing cold call edges from the dynamic call graph may have an impact on cold inlining, but so far I haven't seen any performance issues since the CS preinliner mainly targets hot callsites, and cold inlining can always be done by the compiler CGSCC inliner. Also fixing an issue where the largest weight instead of the accumulated weight for a call edge is used in the profiled call graph. Reviewed By: wenlei Differential Revision: https://reviews.llvm.org/D147013 -
Jonas Devlieghere authored
Support universal Mach-O binaries with a fat64 header. After 4d683f7f, dsymutil can now generate such binaries when the offsets would otherwise overflow the 32-bit offsets in the regular fat header. rdar://107289570 Differential revision: https://reviews.llvm.org/D147012
-
Anshil Gandhi authored
Change target feature of __builtin_amdgcn_global_atomic_fadd_f32 to atomic-fadd-rtn-insts. Enable atomic-fadd-rtn-insts for gfx90a, gfx940 and gfx1100 as they all support the return variant of `global_atomic_add_f32`. Fixes https://github.com/llvm/llvm-project/issues/61331. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D146840
-
Aaron Siddhartha Mondal authored
Reviewed By: GMNGeoffrey Differential Revision: https://reviews.llvm.org/D147088
-