- May 14, 2022
-
-
Rafael Auler authored
When a profile is collected in a BOLTed binary, the generated profile is tagged with a header string "boltedcollection" in the first line of the fdata file. Fix merge-fdata to recognize this header string and preserve it into the output. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D125591
-
Nico Weber authored
-
Med Ismail Bennani authored
This patch renames the `SBCompileUnit::GetIndexForLineEntry` api to be an overload of `SBCompileUnit::FindLineEntryIndex` Differential Revision: https://reviews.llvm.org/D125594 Signed-off-by:
Med Ismail Bennani <medismail.bennani@gmail.com>
-
Richard authored
Add a recursive descent parser to match macro expansion tokens against fully formed valid expressions of integral literals. Partial expressions will not be matched -- they can't be valid initializing expressions for an enum. Differential Revision: https://reviews.llvm.org/D124500 Fixes #55055
-
Mogball authored
-
Mogball authored
Rename ZeroResult -> ZeroResults ZeroSuccessor -> ZeroSuccessors ZeroRegion -> ZeroRegions to be in line with ZeroOperands and grammatically correct.
-
owenca authored
This patch is the result of running clang-format version 753fe330 in clang/unittests/Format/: clang-format -style="{InsertBraces: true, RemoveBracesLLVM: true}" -i *.cpp *.h Differential Revision: https://reviews.llvm.org/D125510
-
Wolfgang Pieb authored
This patch is refactoring the allocation, initialization and deletion of MDNodes. It is intended as a preparatory patch for the upcoming addition of dynamic resizability of MDNodes. It is fundamentally NFC, but removes the necessity for suppressing the memory sanitizer for MDNode's operator delete. Reviewers: dexonsmith Differential Revision: https://reviews.llvm.org/D125489
-
Ben Dunbobbin authored
The CREATE/CREATETHIN commands should overwrite the output file: https://sourceware.org/binutils/docs/binutils/ar-scripts.html. This fixes a regression for MRI scripts introduced in: https://reviews.llvm.org/D123142 which put logic into performWriteOperation. performWriteOperation is called for all MRI commands that write an archive out (one's with a SAVE command). performWriteOperation is unaware of MRI semantics and loads an existing archive if present. If an existing archive is loaded, llvm-ar checks the properties of the existing archive for decisions about the output archive (for example making the output archive thin if the existing one was). https://reviews.llvm.org/D123142 adds the following logic... if (OldArchive) { if (Thin && !OldArchive->isThin()) fail("cannot convert a regular archive to a thin one"); if (OldArchive->isThin()) Thin = true; } ... which errors for a script with CREATETHIN in effect if there is an existing regular archive, and causes CREATE to output a thin archive if there is an existing thin archive. Differential Revision: https://reviews.llvm.org/D125439
-
bzcheeseman authored
Addresses use cases in Clang/MLIR that need pointer-to-pointer, reference-to-reference, and value-to-value casts from/to the same types. This should reduce boilerplate by allowing the user to simply specify the pointer cast and forward the reference cast directly to the pointer cast. This cast trait DOES NOT implement `castFailed` and `doCastIfPossible` because in the general case doing so could result in a nullptr dereference. Users can use `NullableValueCastFailed` and `DefaultDoCastIfPossible` as desired for those cases where `nullptr` is acceptable. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125576
-
bzcheeseman authored
Since cast_convert_val now has pointer specializations, we don't need the pointer partial specialization for CastInfo. We want to trim these down when possible to avoid future ambiguous partial specialization errors. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125578
-
Chris Lattner authored
The warning caused build errors on a couple flang testers that are building with -Werror. The diagnostic change makes the generated error correct. This is a followup to https://reviews.llvm.org/D125549 Differential Revision: https://reviews.llvm.org/D125587
-
Roger Ferrer Ibanez authored
When building the final merged node, we were using the original chain rather than the output chain of the new operation. After some collapsing of the chain this could cause the loads be incorrectly scheduled respect to later stores. This was uncovered by SingleSource/Regression/C/gcc-c-torture/execute/pr36038.c of the llvm testsuite. https://reviews.llvm.org/D125560
-
Roger Ferrer Ibanez authored
After a memmove is expanded, we load the higher addresses after we have stored onto the lower ones, as if we could do a forward copy. Differential Revision: https://reviews.llvm.org/D125553
-
Alexander Shaposhnikov authored
Adjust `optimizeGlobalCtorsList` to handle the case of different priorities. This addresses the issue https://github.com/llvm/llvm-project/issues/55083. Test plan: ninja check-all Differential revision: https://reviews.llvm.org/D125278
-
Joseph Huber authored
Summary: We should use the last argument so this flag can be overridden properly.
-
Alexey Bataev authored
If alternate node has only 2 instructions and the tree is already big enough, better to skip the vectorization of such nodes, they are not very profitable (the resulting code cotains 3 instructions instead of original 2 scalars). SLP can try to vectorize the buildvector sequence in the next attempt, if it is profitable. Metric: SLP.NumVectorInstructions Program SLP.NumVectorInstructions results results0 diff test-suite :: MultiSource/Benchmarks/DOE-ProxyApps-C/miniAMR/miniAMR.test 72.00 73.00 1.4% test-suite :: MultiSource/Benchmarks/Prolangs-C/TimberWolfMC/timberwolfmc.test 1186.00 1198.00 1.0% test-suite :: MultiSource/Benchmarks/DOE-ProxyApps-C++/miniFE/miniFE.test 241.00 242.00 0.4% test-suite :: MultiSource/Applications/JM/lencod/lencod.test 2131.00 2139.00 0.4% test-suite :: External/SPEC/CINT2017rate/523.xalancbmk_r/523.xalancbmk_r.test 6377.00 6384.00 0.1% test-suite :: External/SPEC/CINT2017speed/623.xalancbmk_s/623.xalancbmk_s.test 6377.00 6384.00 0.1% test-suite :: External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test 12650.00 12658.00 0.1% test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 26169.00 26147.00 -0.1% test-suite :: MultiSource/Benchmarks/Trimaran/enc-3des/enc-3des.test 99.00 86.00 -13.1% Gains: 526.blender_r - more vectorized trees. enc-3des - same. Others: 510.parest_r - no changes. miniFE - same 623.xalancbmk_s - some (non-profitable) parts of the trees are not vectorized. 523.xalancbmk_r - same lencod - same timberwolfmc - same miniAMR - same Differential Revision: https://reviews.llvm.org/D125571 -
Alan Zhao authored
EXTERN PROC isn't really well documented in MSVC, so after poking around it seems as if it's just a regular extern symbol. Interestingly enough, under MSVC the following is allowed: extern foo:proc mov eax, foo MSVC will output: mov eax, 0 while llvm-ml will currently output: mov eax, dword ptr [foo] (since foo is an extern) Arguably, llvm-ml's output makes more sense, even though it's inconsistent with MSVC ml. However, since moving an extern proc symbol to a register doesn't really make sense in the first place, we'll treat it as undefined behavior for now. Reviewed By: epastor Differential Revision: https://reviews.llvm.org/D125582
-
Eli Friedman authored
When GlobalISel fails, we need to report the error, and we need to set the FailedISel property. We skipped those steps if stack protector insertion failed, which led to a very strange miscompile. Differential Revision: https://reviews.llvm.org/D125584
-
Egor Zhdan authored
Some new DriverKit tests were added in https://reviews.llvm.org/D121911, and unfortunately they fail on Linux build bots.
-
Alexey Bataev authored
If the insert indes was used already or is not constant, we should stop looking for unique buildvector sequence, it mustbe splitted to 2 different buildvectors.
-
Joseph Huber authored
Summary: This patch allows users to compile the static library without CUDA installed on the system. This requires the new flag `--cuda-feature` to indicate that we need `+ptx61` in order to compile the runtime.
-
Joseph Huber authored
Summary: Normally we parse through the CUDA installation to disover the needed features. However, we may want to build libraries on targets that do not currently have CUDA installed but still need to know which features to make use of when creating the PTX or bitcode. This flag is a simple way to specify this so we can compile certain codes withotu a valid CUDA installation. Ideally this could be done via an -Xarch or simimlar flag but currently they cannot handle this. We would need to support using an -Xarch flag that takes multiple arguments that then pass them to the -Xclang functionality.
-
Amir Ayupov authored
Move BOLT libraries out of `LLVM_LINK_COMPONENTS` to `target_link_libraries`. Addresses issue #55432. Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D125568
-
River Riddle authored
There isn't really a good pre-existing syntax highlighter for tablegen, so this commit adds a textmate version that covers nearly everything in the current spec. Differential Revision: https://reviews.llvm.org/D125427
-
Amir Ayupov authored
- Fix common (arch-independent) tests to explicitly target -linux triple. - Override the triple inside arch-specific tests. - Add cflags to common tests. - Update individual tests. - Expand pipe stderr `|&` shorthand. Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D125548
-
Egor Zhdan authored
This is the second patch that upstreams the support for Apple's DriverKit. The first patch: https://reviews.llvm.org/D118046. Differential Revision: https://reviews.llvm.org/D121911
-
Jonas Devlieghere authored
When using dsymForUUID, the majority of time symbolication a crashlog with crashlog.py is spent waiting for it to complete. Currently, we're calling dsymForUUID sequentially when iterating over the modules. We can drastically cut down this time by calling dsymForUUID in parallel. This patch uses Python's ThreadPoolExecutor (introduced in Python 3.2) to parallelize this IO-bound operation. The performance improvement is hard to benchmark, because even with an empty local cache, consecutive calls to dsymForUUID for the same UUID complete faster. With warm caches, I'm seeing a ~30% performance improvement (~90s -> ~60s). I suspect the gains will be much bigger for a cold cache. dsymForUUID supports batching up multiple UUIDs. I considered going that route, but that would require more intrusive changes. It would require hoisting the logic out of locate_module_and_debug_symbols which we explicitly document [1] as a feature of Symbolication.py to locate symbol files. [1] https://lldb.llvm.org/use/symbolication.html Differential reviison: https://reviews.llvm.org/D125107
-
Amara Emerson authored
Differential Revision: https://reviews.llvm.org/D125041
-
Amir Ayupov authored
Addresses warnings when built with Apple Clang. Reviewed By: yota9 Differential Revision: https://reviews.llvm.org/D125483
-
Amir Ayupov authored
Address warnings in Release build without assertions. Tip @tschuett for reporting the issue #55404. Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D125475
-
Fangrui Song authored
-
Fangrui Song authored
-
Amir Ayupov authored
Fix build with Apple Clang. Tip @tschuett for reporting the issue #55404. Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D125480
-
Louis Dionne authored
For example, we used to trigger CI even for commits that touched a file whose path contained 'cmake', even if it's not the root cmake directory. Fix that.
-
Joseph Huber authored
The previous patches allowed us to create a static library containing all the device code. This patch uses that library to perform the device runtime linking late when performing LTO. This in addition to simplifying the libraries, allows us to transparently handle the runtime library as-needed without needing Clang to manually pass the necessary library in the linker wrapper job. Depends on D125315 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D125333
-
Joseph Huber authored
This patch adds the necessary CMake configuration to build a static library version of the device runtime, `libomptarget.devicertl.a`. Various improvements in how we handle static libraries and generating offloading code should allow us to treat the device library as a regular project without needing to invoke the clang front-end directly. Here we generate a job for each offloading architecture supported. Each offloading architecture will be embedded into the static library and used as-needed by the host. This library will primarily be used to replace the bitcode library when performing LTO. Currently, we need to manually pass in the bitcode library which requires foreknowledge of the offloading architecture. This approach lets us handle that in the linker wrapper instead. Furthermore this should improve our interface to the device runtime. We can now build it fully under a release build and have all the expected entry points, as well as supporting debug builds. Depends on D125265 D125256 D125260 D125314 D125563 Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D125315
-
Joseph Huber authored
We used to globally include the libomptarget include directory for all projects. This caused some conflicts with the other files named "Debug.h". This patch changes the cmake to include these files via the target include instead. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D125563
-
Joseph Huber authored
We use globals to configure debugging at compile-time for the device runtime. Because these are only used by the OpenMP runtime we shouldn't define them if we aren't using the device runtime. When a user passes in '-nogpulib' this indicates that we are not using the device runtime, so we should check for the precense of this flag and not emit these globals if used. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D125314
-
Joseph Huber authored
OpenMP uses several wrapper hearders to provide the definitions of needed symbols contained in the host. However, some users may use the `-nostdinc` option to override these definitions themselves. The OpenMP wrapper headers are stored in the same location as the clang install. If the user passes `-nostdinc` then this include directory is never looked at by default which means that including these wrappers will always fail. These headers should instead be included manually if they are needed with a `-nostdinc` build. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D125265
-