- Jun 19, 2023
-
-
Akash Banerjee authored
Key changes: - Refactor the createTargetData function to make use of the emitOffloadingArrays and emitOffloadingArraysArgument functions to generate code. - Added a new emitIfClause helper function to allow handling if clauses in a similar fashion to Clang. - Updated the MLIR side of code to account for changes to createTargetData. Depends on D149872 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D146557
-
David Spickett authored
This register is used as the pointer to the current thread local storage block and is read from NT_ARM_TLS on Linux. Though tpidr will be present on all AArch64 Linux, I am soon going to add a second register tpidr2 to this set. tpidr is only present when SME is implemented, therefore the NT_ARM_TLS set will change size. This is why I've added this as a dynamic register set to save changes later. Reviewed By: omjavaid Differential Revision: https://reviews.llvm.org/D152516
-
Alexandre Ganea authored
This reverts commit aa495214. As discussed in https://github.com/llvm/llvm-project/issues/53475 this patch allows for using LLD-as-a-lib. It also lets clients link only the drivers that they want (see unit tests). This also adds the unit test infra as in the other LLVM projects. Among the test coverage, I've added the original issue from @krzysz00, see: https://github.com/ROCmSoftwarePlatform/D108850-lld-bug-reproduction Important note: this doesn't allow (yet) linking in parallel. This will come a bit later hopefully, in subsequent patches, for COFF at least. Differential revision: https://reviews.llvm.org/D119049
-
Mehdi Amini authored
The wrapper, as most of compiler-generated functions, are intended to serve the IR for the current module. The safest linkage is to keep these private to avoid any possible collision with other modules. Differential Revision: https://reviews.llvm.org/D153255
-
Mehdi Amini authored
This was a spurious closing parenthese.
-
Nikita Popov authored
The ConstantRange specifies the range of the scalar elements in the vector. When converting into a Constant, we need to create a vector splat with the correct type. For that purpose, pass in the expected type for the constant. Fixes https://github.com/llvm/llvm-project/issues/63380.
-
Russell Greene authored
https://github.com/llvm/llvm-project/issues/62750 I setup a simple test with a large .so (~100MiB) that was only present on the target machine but not present on the local machine, and ran a lldb server on the target and connectd to it. LLDB properly downloads the file from the remote, but it does so at a very slow speed, even over a hardwired 1Gbps connection! Increasing the buffer size for downloading these helps quite a bit. Test setup: ``` $ cat gen.py print('const char* hugeglobal = ') for _ in range(1000*500): print(' "' + '1234'*50 + '"') print(';') print('const char* mystring() { return hugeglobal; }') $ gen.py > huge.c $ mkdir libdir $ gcc -fPIC huge.c -Wl,-soname,libhuge.so -o libdir/libhuge.so -shared $ cat test.c #include <string.h> #include <stdio.h> extern const char* mystring(); int main() { printf("%d\n", strlen(mystring())); } $ gcc test.c -L libdir -l huge -Wl,-rpath='$ORIGIN' -o test $ rsync -a libdir remote:~/ $ ssh remote bash -c "cd ~/libdir && /llvm/buildr/bin/lldb-server platform --server --listen '*:1234'" ``` in another terminal ``` $ rm -rf ~/.lldb # clear cache $ cat connect.lldb platform select remote-linux platform connect connect://10.0.0.14:1234 file test b main r image list c q $ time /llvm/buildr/bin/lldb --source connect.lldb ``` Times with various buffer sizes: 1kiB (current): ~22s 8kiB: ~8s 16kiB: ~4s 32kiB: ~3.5s 64kiB: ~2.8s 128kiB: ~2.6s 256kiB: ~2.1s 512kiB: ~2.1s 1MiB: ~2.1s 2MiB: ~2.1s I choose 512kiB from this list as it seems to be the place where the returns start diminishing and still isn't that much memory My understanding of how this makes such a difference is ReadFile issues a request for each call, and larger buffer means less round trip times. The "ideal" situation is ReadFile() being async and being able to issue multiple of these, but that is much more work for probably little gains. NOTE: this is my first contribution, so wasn't sure who to choose as a reviewer. Greg Clayton seems to be the most appropriate of those in CODE_OWNERS.txt Reviewed By: clayborg, jasonmolenda Differential Revision: https://reviews.llvm.org/D153060
-
Alexandros Lamprineas authored
Allows constant folding of such instructions when estimating user bonus. Differential Revision: https://reviews.llvm.org/D153036
-
Nikita Popov authored
Fold uadd.sat(X, Y) uge X and usub.sat(X, Y) ule X to true. Proof: https://alive2.llvm.org/ce/z/596m9X Fixes https://github.com/llvm/llvm-project/issues/63381.
-
Nikita Popov authored
-
Jacob Crawley authored
Adds a new HLFIR operation for the COUNT intrinsic according to the design set out in flang/docs/HighLevel.md. This patch includes all the necessary changes to create a new HLFIR operation and lower it into the fir runtime call. Author was @jacob-crawley. Minor adjustments by @tblah Differential Revision: https://reviews.llvm.org/D152521
-
Tom Eccles authored
Codegen only supports conversions between logicals and integers. The verifier should reflect this. Differential Revision: https://reviews.llvm.org/D152935
-
Tom Eccles authored
When the ENTRY statement is used, the same source can return different types depending on the entry point. These different return values are storage associated (share the same storage). Previously, this led to the declaration of the results to all have the largest type. This patch adds a convert between the stack allocation and the declaration so that the hlfir.decl gets the right type. I haven't managed to generate code where this convert converted a reference to an allocation for a smaller type into an allocation for a larger one, but I have added an assert just in case. This is a different solution to https://reviews.llvm.org/D152725, see discussion there. Differential Revision: https://reviews.llvm.org/D152931
-
Dmitry Makogon authored
-
Hristo Hristov authored
-
Alexandros Lamprineas authored
The specialization bonus is zero in some unittests because the basic blocks containing the users of the constant arguments are executed less frequently than the entry block. Sinking them into loops solves that. Differential Revision: https://reviews.llvm.org/D153230
-
Momchil Velikov authored
Some transformation in CodeGenPrepare pass may create and/or delete basic block, but they don't update the LoopInfo, so the LoopInfo may end up containing dangling pointers and sometimes reused basic blocks, which leads to "interesting" non-deterministic behaviour. These transformations do not seem to alter the loop structure of the function, and updating the loop info is quite straighforward. Reviewed By: efriedma Differential Revision: https://reviews.llvm.org/D150384 Change-Id: If8ab3905749ea6be94fbbacd54c5cfab5bc1fba1
-
Martin Braenne authored
This patch includes a test that fails without the fix. I discovered that we weren't creating `Value`s for integer literals when, in a different patch, I tried to overwrite the value of a struct field with a literal for the purposes of a test and was surprised to find that the struct compared the same before and after the assignment. This functionality therefore seems useful at least for tests, but is probably also useful for actual analysis of code. Reviewed By: ymandel, xazax.hun, gribozavr2 Differential Revision: https://reviews.llvm.org/D152813
-
Joseph Huber authored
Recently the AMDGPU backend automatically enables a pass to optimize atomics. This results in the LTO build taking about 10x longer in all cases. For now we disable this by default as was the case before the patch in D152649. Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D153232
-
Ingo Müller authored
In https://reviews.llvm.org/D153029, I moved the loading/unloading mechanisms of shared libraries from the JIT runner to the execution engine in order to make that mechanism available in the latter (including its Python bindings). However, I realized that I introduced a small change in semantic: previously, the JIT runner checked for the presence of init/destroy functions and only loaded the library as JITDyLib if they were not present. After I moved the code, all libraries were loaded as JITDyLib, even if they registered their symbols explicitly in their init function. I am not sure if this is really a problem but (1) the previous behavior was different and (2) I guess it could cause a problem if some symbols are exported through the init function *and* have public visibility. This patch reestablishes the original behaviour in the new place of the code. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D153249
-
Christian Sigg authored
-
Felix authored
Fixes: https://github.com/llvm/llvm-project/issues/59119 Reviewed By: PiotrZSL Differential Revision: https://reviews.llvm.org/D152764
-
Bing1 Yu authored
Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D152819
-
Matthias Springer authored
As a convenience to the user, top-level sequence ops can optionally be used as matchers: the op type is specified by the type of the block argument. This is similar to how pass pipeline targets can be specified on the command line (`-pass-pipeline='builtin.module(func.func(...))`). Differential Revision: https://reviews.llvm.org/D153121
-
Matthias Springer authored
Add an extra check to make sure that transform IR is not getting modified by this op while it is being interpreted. This generally dangerous and we may want to enforce this for all transform ops that modify the payload in the future. Users should generally try to apply patterns only to the piece of IR where it is needed (e.g., a matched function) and not the entire module (which may contain the transform IR). This revision is in response to a crash in a downstream compiler that was caused by a dead `transform.structured.match` op that was removed by the GreedyPatternRewriteDriver's DCE while the enclosing sequence was being interpreted. Differential Revision: https://reviews.llvm.org/D153113
-
David Green authored
Similar to D152245, this adds integer addp patterns, using the larger v4i32 addp from addp extractlow, extracthi.
-
David Green authored
This adds some simple tablegen patterns for converting `faddp v2f32 extractlow(Rn), v2f32 extracthigh(Rn)` to `faddp v4f32 Rn, v4f32 Rn` using the q variants of the instructions, avoiding the extra ext needed to extract the high lanes. Only the bottom lanes of the new faddp are used, the second Rn operand is used as a placeholder. It uses Rn to prevent any false dependencies, but could equally by undef. Differential Revision: https://reviews.llvm.org/D152245
-
luxufan authored
-
luxufan authored
-
Yevgeny Rouban authored
Relax condition on runtime trip count unrolling loops with 1 non-latch exit that leads to a deop block. There are cases when the deopt blocks are common exits for different loops. LoopSimplify pass splits such edges to the common deopting blocks to make sure that all exit nodes of the loop only have predecessors that are inside of the loop (See simplifyOneLoop()). This breaks the current condition for unrolling. This patch allows the split transitive blocks that still lead to the deopting blocks. Differential Revision: https://reviews.llvm.org/D152639
-
Chuanqi Xu authored
Close https://github.com/llvm/llvm-project/issues/61940. The root cause is that clang will generate vtable as strong symbol now even if the corresponding class is defined in other module units. After I check the wording in Itanium ABI, I find this is not inconsistent. Itanium ABI 5.2.3 (https://itanium-cxx-abi.github.io/cxx-abi/abi.html#vague-vtable) says: > The virtual table for a class is emitted in the same object containing > the definition of its key function, i.e. the first non-pure virtual > function that is not inline at the point of class definition. So the current behavior is incorrect. This patch tries to address this. Also I think we need to do a similar change for MSVC ABI. But I don't find the formal wording. So I don't address this in this patch. Reviewed By: rjmccall, iains, dblaikie Differential Revision: https://reviews.llvm.org/D150023
-
Fangrui Song authored
Add the `S_ATTR_LIVE_SUPPORT` attribute to the sections so that `ld -dead_strip` will retain subsections that reference live functions, once we we add linker private "l" symbols as atoms.
-
Jianjian GUAN authored
Since we use match shl (v, splat 1) to vadd, we could also expand to widening add. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D153112
-
Fangrui Song authored
Fixes: c26c5e47 (essentially a no-op) The newly created MCDataFragment should inherit Atom (see MCMachOStreamer::finishImpl). To the best of my knowledge, this change cannot be tested at present, but this is important to ensure MCExpr.cpp:AttemptToFoldSymbolOffsetDifference gives the same result in case we evaluate the expression again with a MCAsmLayout. In the following case, ``` .section __DATA,xray_instr_map lxray_sleds_start1: .space 16 Lxray_sleds_end1: .section __DATA,xray_fn_idx .quad (Lxray_sleds_end1-lxray_sleds_start1)>>4 // can be folded without a MCAsmLayout ``` When we have a MCAsmLayout, without this change, evaluating (Lxray_sleds_end1-lxray_sleds_start1)>>4 again will fail due to `FA->getAtom() == nullptr && FB.getAtom() != nullptr` in MachObjectWriter::isSymbolRefDifferenceFullyResolvedImpl, called by AttemptToFoldSymbolOffsetDifference.
-
Fangrui Song authored
When the MCAssembler is non-null and the MCAsmLayout is null, we can fold A-B in these additional cases: * when A is a pending label (will be reassigned to a real fragment in flushPendingLabels()) * A and B are separated by a MCFillFragment with a constant size
-
Fangrui Song authored
If FA == FB, we can use SA.getOffset() - SB.getOffset() even if FA is not a MCDataFragment, as the only case this can be problematic (different offsets for a variable-size fragment) is invalid/unreachable. If FA != FB, the `if (FI->getKind() != MCFragment::FT_Data)` check below can bail out correctly. This change will help Mach-O fold more expressions. For ELF this is NFC, unless evaluateFixup has a bug that would evaluate an expression differently.
-
Fangrui Song authored
The newly created MCDataFragment should inherit Atom (see MCMachOStreamer::finishImpl). I cannot think of a case to test the behavior, but this is one step towards folding the Mach-O label difference below and making Mach-O more similar to ELF. ``` .section __DATA,xray_instr_map lxray_sleds_start1: .space 16 Lxray_sleds_end1: .section __DATA,xray_fn_idx .quad (Lxray_sleds_end1-lxray_sleds_start1)>>4 // error: expected relocatable expression // Mach-O ```
-
Alfred Persson Forsberg authored
Differential Revision: https://reviews.llvm.org/D153231
-
Fangrui Song authored
-
Kazu Hirata authored
-