- Jun 19, 2023
-
-
Nikita Popov authored
-
Job Noorman authored
BOLT currently assumes (and asserts) that no two relocations can share the same offset. Although this is true in most cases, ELF has a feature called (not sure if this is an official term) composed relocations [1] where multiple relocations at the same offset are combined to produce a single value. For example, to support label subtraction (a - b) on RISC-V, two relocations are emitted at the same offset: - R_RISCV_ADD32 a + 0 - R_RISCV_SUB32 b + 0 which, when combined, will produce the value of (a - b). To support this in BOLT, first, RelocationSetType in BinarySection is changed to be a multiset in order to allow it to store multiple relocations at the same offset. Next, Relocation::emit() is changed to receive an iterator pair of relocations. In most cases, these will point to a single relocation in which case its behavior is unaltered by this patch. For composed relocations, they should point to all relocations at the same offset and the following happens: - A new method Relocation::createExpr() is called for every relocation. This method is essentially the same as the original emit() except that it returns the MCExpr without emitting it. - The MCExprs of relocations i and i+1 are combined using the opcode returned by the new method Relocation::getComposeOpcodeFor(). - After combining all MCExprs, the last one is emitted. Note that in the current patch, getComposeOpcodeFor() simply calls llvm_unreachable() since none of the current targets use composed relocations. This will change once the RISC-V target lands. Finally, BinarySection::emitAsData() is updated to group relocations by offset and emit them all at once. Note that this means composed relocations are only supported in data sections. Since this is the only place they seem to be used in RISC-V, I believe it's reasonable to only support them there for now to avoid further code complexity. [1]: https://www.sco.com/developers/gabi/latest/ch4.reloc.html Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D146546
-
Mark de Wever authored
Like yocto, zepto, zetta, and yotta. The new prefixes quecto, ronto, ronna, and quetta can't be implemented in a intmax_t. So their implementation does nothing. Implements - P2734R0 Adding the new SI prefixes Depends on D153192 Reviewed By: #libc, philnik Differential Revision: https://reviews.llvm.org/D153200
-
Mark de Wever authored
Implements part of: - P1614R2 The Mothership has Landed Reviewed By: #libc, H-G-Hristov, philnik Differential Revision: https://reviews.llvm.org/D153222
-
Mark de Wever authored
This FTM was introduced in P0553R4 Bit operations Which has been implemented since libc++ 9. This was noticed while working on D153192. Reviewed By: #libc, philnik Differential Revision: https://reviews.llvm.org/D153225
-
Nico Weber authored
lld/test/Unit/lit.site.cfg.py.in got cleaned up in the reland.
-
Mark de Wever authored
This updates: - The status tables - Feature test macros - New headers for modules The latter avoids forgetting about modules when implementing the feature in a new header. Reviewed By: #libc, philnik Differential Revision: https://reviews.llvm.org/D153192
-
Nicolas Vasilache authored
-
Nicolas Vasilache authored
This reverts commit aabea3d3. That commit had mistakenly squashed spurious changes in.
-
Nico Weber authored
-
Nico Weber authored
-
Nico Weber authored
The lld CL relanded in 6f2e92c1. This reverts commit d76b37e6.
-
David Green authored
For both CodeGen and CostModelling, this adds extran testing for the new lvm.vector.reduce.fmaximum and lvm.vector.reduce.fminimum intrinsics, as well as making sure there is test coverage for all the various cases.
-
Joe Nash authored
Patch eece6ba2 changed the src1 type of v_ldexp_f16 from i32 to i16. Though semantically src1 is an i16, the hardware reads this operand as an f16 type, which primarily enables floating point inline constants. Therefore this patch changes the operand type to f16. It maintains the current behavior where floating point source modifiers are not allowed on src1. SDWA sext modifier continues to be allowed. The test asm and disasm test changes in eece6ba2 are reverted, because the floating point inline constants are allowed. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153169
-
Vladislav Dzhidzhoev authored
RFC https://discourse.llvm.org/t/rfc-dwarfdebug-fix-and-improve-handling-imported-entities-types-and-static-local-in-subprogram-and-lexical-block-scopes/68544 Similar to imported declarations, the patch tracks function-local types in DISubprogram's 'retainedNodes' field. DwarfDebug is adjusted in accordance with the aforementioned metadata change and provided a support of function-local types scoped within a lexical block. The patch assumes that DICompileUnit's 'enums field' no longer tracks local types and DwarfDebug would assert if any locally-scoped types get placed there. Authored-by:
Kristina Bessonova <kbessonova@accesssoftek.com> Differential Revision: https://reviews.llvm.org/D144006 Depends on D144005
-
Roger Ferrer Ibanez authored
In D152070 we added many new intrinsic types required for the RISC-V Vector Extension. This was crashing when loading the AST as those types are intrinsically added to the AST (they don't come from the disk). The total number required now by clang exceeds 400 so increasing the value to 500 solves the problem. This value was already increased in D92715 but I assume this has some impact on the on-disk format. Also add a static assert to avoid this happening again in the future. Differential Revision: https://reviews.llvm.org/D153111
-
Florian Hahn authored
Add unit test triggering an assertion with abfeda5a.
-
Nicolas Vasilache authored
-
Nicolas Vasilache authored
Restriction to ModuleOp is ancient and unnecessarily restrictive.
-
Louis Dionne authored
This makes it such that new.cpp contains only the definitions of operator new and operator delete, like its libc++abi counterpart. Differential Revision: https://reviews.llvm.org/D153136
-
Nikita Popov authored
Avoid inserting the freeze if not necessary, as this allows LVI to continue reasoning about the expression.
-
Nikita Popov authored
-
Alexandros Lamprineas authored
After each iteration of the function specializer, constant stack values are promoted to constant globals in order to enable recursive function specialization. This should also be done once before running the specializer. Enables specialization of _QMbrute_forcePdigits_2 from SPEC2017:548.exchange2_r. Differential Revision: https://reviews.llvm.org/D152799
-
Adrian Munera authored
This patch implements the "__kmp_print_tdg_dot" function, that prints a task dependency graph into a dot file containing the tasks and their dependencies. It is activated through a new environment variable "KMP_TDG_DOT" Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D150962
-
Dmitry Makogon authored
SplitBlockAndInsertIfThen utility creates two new blocks, they're called ThenBlock and Tail (true and false destinations of a conditional branch correspondingly). The function has a bool parameter Unreachable, and if it's set, then ThenBlock is terminated with an unreachable. At the end of the function the new blocks are added to the loop of the split block. However, in case ThenBlock is terminated with an unreachable, it cannot belong to any loop. Differential Revision: https://reviews.llvm.org/D152434
-
David Truby authored
The LLVM comdat operation specifies how to deduplicate globals with the same key in two different object files. This is necessary on Windows where e.g. two object files with linkonce globals will not link unless a comdat for those globals is specified. It is also supported in the ELF format. Differential Revision: https://reviews.llvm.org/D150796
-
melonedo authored
Implement XCVbitmanip intrinsics for CV32E40P according to the specification. This commit is part of a patch-set to upstream the 7 vendor specific extensions of CV32E40P. Contributors: @CharKeaney, @jeremybennett, @lewis-revill, @liaolucy, @simoncook, @xmj. Spec: https://github.com/openhwgroup/cv32e40p/blob/62bec66b36182215e18c9cf10f723567e23878e9/docs/source/instruction_set_extensions.rst Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152915
-
Louis Dionne authored
The operations.cpp file contained the implementation of a ton of functionality unrelated to just the filesystem operations, and filesystem_common.h contained a lot of unrelated functionality as well. Splitting this up into more files will make it possible in the future to support parts of <filesystem> (e.g. path) on systems where there is no notion of a filesystem. Differential Revision: https://reviews.llvm.org/D152377
-
Louis Dionne authored
A few tests were also straightforward to translate to SFINAE tests instead, so in a few cases I did that and removed the .fail.cpp test entirely. Differential Revision: https://reviews.llvm.org/D153149
-
Louis Dionne authored
For consistency with other algorithms. Differential Revision: https://reviews.llvm.org/D153141
-
Matt Arsenault authored
These are leftover hacks from using asm declaratios to access intrinsics.
-
Matt Arsenault authored
-
Nikita Popov authored
-
Akash Banerjee authored
Key changes: - Refactor the createTargetData function to make use of the emitOffloadingArrays and emitOffloadingArraysArgument functions to generate code. - Added a new emitIfClause helper function to allow handling if clauses in a similar fashion to Clang. - Updated the MLIR side of code to account for changes to createTargetData. Depends on D149872 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D146557
-
David Spickett authored
This register is used as the pointer to the current thread local storage block and is read from NT_ARM_TLS on Linux. Though tpidr will be present on all AArch64 Linux, I am soon going to add a second register tpidr2 to this set. tpidr is only present when SME is implemented, therefore the NT_ARM_TLS set will change size. This is why I've added this as a dynamic register set to save changes later. Reviewed By: omjavaid Differential Revision: https://reviews.llvm.org/D152516
-
Alexandre Ganea authored
This reverts commit aa495214. As discussed in https://github.com/llvm/llvm-project/issues/53475 this patch allows for using LLD-as-a-lib. It also lets clients link only the drivers that they want (see unit tests). This also adds the unit test infra as in the other LLVM projects. Among the test coverage, I've added the original issue from @krzysz00, see: https://github.com/ROCmSoftwarePlatform/D108850-lld-bug-reproduction Important note: this doesn't allow (yet) linking in parallel. This will come a bit later hopefully, in subsequent patches, for COFF at least. Differential revision: https://reviews.llvm.org/D119049
-
Mehdi Amini authored
The wrapper, as most of compiler-generated functions, are intended to serve the IR for the current module. The safest linkage is to keep these private to avoid any possible collision with other modules. Differential Revision: https://reviews.llvm.org/D153255
-
Mehdi Amini authored
This was a spurious closing parenthese.
-
Nikita Popov authored
The ConstantRange specifies the range of the scalar elements in the vector. When converting into a Constant, we need to create a vector splat with the correct type. For that purpose, pass in the expected type for the constant. Fixes https://github.com/llvm/llvm-project/issues/63380.
-
Russell Greene authored
https://github.com/llvm/llvm-project/issues/62750 I setup a simple test with a large .so (~100MiB) that was only present on the target machine but not present on the local machine, and ran a lldb server on the target and connectd to it. LLDB properly downloads the file from the remote, but it does so at a very slow speed, even over a hardwired 1Gbps connection! Increasing the buffer size for downloading these helps quite a bit. Test setup: ``` $ cat gen.py print('const char* hugeglobal = ') for _ in range(1000*500): print(' "' + '1234'*50 + '"') print(';') print('const char* mystring() { return hugeglobal; }') $ gen.py > huge.c $ mkdir libdir $ gcc -fPIC huge.c -Wl,-soname,libhuge.so -o libdir/libhuge.so -shared $ cat test.c #include <string.h> #include <stdio.h> extern const char* mystring(); int main() { printf("%d\n", strlen(mystring())); } $ gcc test.c -L libdir -l huge -Wl,-rpath='$ORIGIN' -o test $ rsync -a libdir remote:~/ $ ssh remote bash -c "cd ~/libdir && /llvm/buildr/bin/lldb-server platform --server --listen '*:1234'" ``` in another terminal ``` $ rm -rf ~/.lldb # clear cache $ cat connect.lldb platform select remote-linux platform connect connect://10.0.0.14:1234 file test b main r image list c q $ time /llvm/buildr/bin/lldb --source connect.lldb ``` Times with various buffer sizes: 1kiB (current): ~22s 8kiB: ~8s 16kiB: ~4s 32kiB: ~3.5s 64kiB: ~2.8s 128kiB: ~2.6s 256kiB: ~2.1s 512kiB: ~2.1s 1MiB: ~2.1s 2MiB: ~2.1s I choose 512kiB from this list as it seems to be the place where the returns start diminishing and still isn't that much memory My understanding of how this makes such a difference is ReadFile issues a request for each call, and larger buffer means less round trip times. The "ideal" situation is ReadFile() being async and being able to issue multiple of these, but that is much more work for probably little gains. NOTE: this is my first contribution, so wasn't sure who to choose as a reviewer. Greg Clayton seems to be the most appropriate of those in CODE_OWNERS.txt Reviewed By: clayborg, jasonmolenda Differential Revision: https://reviews.llvm.org/D153060
-