- Oct 26, 2020
-
-
George Mitenkov authored
This patch introduces a SPIR-V runner. The aim is to run a gpu kernel on a CPU via GPU -> SPIRV -> LLVM conversions. This is a first prototype, so more features will be added in due time. - Overview The runner follows similar flow as the other runners in-tree. However, having converted the kernel to SPIR-V, we encode the bind attributes of global variables that represent kernel arguments. Then SPIR-V module is converted to LLVM. On the host side, we emulate passing the data to device by creating in main module globals with the same symbolic name as in kernel module. These global variables are later linked with ones from the nested module. We copy data from kernel arguments to globals, call the kernel function from nested module and then copy the data back. - Current state At the moment, the runner is capable of running 2 modules, nested one in another. The kernel module must contain exactly one kernel function. Also, the runner supports rank 1 integer memref types as arguments (to be scaled). - Enhancement of JitRunner and ExecutionEngine To translate nested modules to LLVM IR, JitRunner and ExecutionEngine were altered to take an optional (default to `nullptr`) function reference that is a custom LLVM IR module builder. This allows to customize LLVM IR module creation from MLIR modules. Reviewed By: ftynse, mravishankar Differential Revision: https://reviews.llvm.org/D86108
-
Fraser Cormack authored
This optmization produces incorrect results when the vector element type is not byte-sized. Related to D78568.
-
Andrew Ng authored
The code to detect the requirement for 64-bit offsets in the archive symbol table was not correctly accounting for the archive file signature and the size of all the contents of the symbol table itself, e.g. the symbol table's header and string table. Also was not considering the variation in symbol table formats. This could result in the creation of large archives with a corrupt symbol table. Change the testing environment variable SYM64_THRESHOLD to be an absolute value rather than a power of 2 in order to enable precise testing of this detection code. Differential Revision: https://reviews.llvm.org/D89891
-
Simon Pilgrim authored
We were using the unique_ptr M to determine the triple after it had been moved in the EngineBuilder constructor.
-
George Mitenkov authored
This patch introduces a pass for running `mlir-spirv-cpu-runner` - LowerHostCodeToLLVMPass. This pass emulates `gpu.launch_func` call in LLVM dialect and lowers the host module code to LLVM. It removes the `gpu.module`, creates a sequence of global variables that are later linked to the varables in the kernel module, as well as a series of copies to/from them to emulate the memory transfer to/from the host or to/from the device sides. It also converts the remaining Standard dialect into LLVM dialect, emitting C wrappers. Reviewed By: mravishankar Differential Revision: https://reviews.llvm.org/D86112
-
Haojian Wu authored
Because of typo-correction, the AST can be transformed, and the transformed AST is marginally useful for diagnostics purpose, the following diagnostics usually do harm than good (easily cause confusions). Given the following code: ``` void abcc(); void test() { if (abc()); // diagnostic 1 (for the typo-correction): the typo is correct to `abcc()`, so the code is treate as `if (abcc())` in AST perspective; // diagnostic 2 (for mismatch type): we perform an type-analysis on `if`, discover the type is not match } ``` The secondary diagnostic "convertable to bool" is likely bogus to users. The idea is to use RecoveryExpr (clang's dependent mechanism) to preserve the recovery behavior but suppress all follow-up diagnostics. Differential Revision: https://reviews.llvm.org/D89946 -
Simon Pilgrim authored
Alive2: https://alive2.llvm.org/ce/z/bCvvHd
-
Simon Pilgrim authored
-
Dmitry Vyukov authored
Enable mips64 support in buildgo.sh. Author: mzh (Meng Zhuo) Reviewed-in: https://reviews.llvm.org/D90130
-
Evgeny Leviant authored
-
Djordje Todorovic authored
-
Pavel Labath authored
-
Pavel Labath authored
Displaying large packed bitfields did not work if one was accessing them through a pointer, and he used the "->" notation ("[0]." notation is fine). The reason for that is that implicit dereference in -> is plumbed all the way down to ValueObjectChild::UpdateValue, where the process of fetching the child value was forked for this flag. The bitfield "sliding" code was implemented only for the branch which did not require dereferencing. This patch restructures the function to avoid this mistake. Processing now happens in two stages. - first the parent is dereferenced (if needed) - then the child value is computed (this step includes sliding and is common for both branches) Differential Revision: https://reviews.llvm.org/D89236 -
Michał Górny authored
Differential Revision: https://reviews.llvm.org/D90119
-
Michał Górny authored
Ensure that xs_xstate_bv is set correctly before calling WriteRegisterSet(). The bit can be clear if the relevant registers were at their initial state when they were read, and it needs to be set in order to apply changes from the XState structure. Differential Revision: https://reviews.llvm.org/D90105
-
Michał Górny authored
Reset registers to their 'initial' state instead of a semi-random pattern in write tests. While the latter might have been helpful while debugging failures (i.e. to distinguish unmodified registers from mistakenly written zeroes), the former makes it possible to test whether xstate_bv field is written correctly when using XSAVE. With this change, the four relevant tests start failing on NetBSD without D90105. Differential Revision: https://reviews.llvm.org/D90114
-
Michał Górny authored
Include <x86/fpu.h> rather than <machine/fpu.h>, as the latter is not present on i386. Differential Revision: https://reviews.llvm.org/D90128
-
Tyker authored
-
Jean Perier authored
2 Bug fixes: - Do not resolve procedure as intrinsic if they appeared in an EXTERNAL attribute statement (one path was not considering this flag) - Emit an error if a procedure resolved to be an intrinsic function (resp. subroutine) is used as a subroutine (resp. function). Lowering was attempted while the evaluate::Expression for the call was missing without any errors. 1 behavior change: - Do not implicitly resolve subroutines (resp. functions) as intrinsics because their name is the name of an intrinsic function (resp. subroutine). Add justification in documentation. Reviewed By: klausler, tskeith Differential Revision: https://reviews.llvm.org/D90049
-
Tyker authored
This allows using annotation in a much more contexts than it currently has. especially when annotation with template or constexpr. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D88645
-
Kazushi (Jam) Marukawa authored
Add VCMP/VCPS/VCPX/VCMS/VCMX vector instructions. Also add regression tests. Reviewed By: simoll Differential Revision: https://reviews.llvm.org/D89643
-
Kazushi (Jam) Marukawa authored
Add VADD/VADS/VADX/VSUB/VSBS/VSBX/VMPY/VMPS/VMPX/VMPD/VDIV/VDVS/VDVX instructions. Also add regression tests. Reviewed By: simoll Differential Revision: https://reviews.llvm.org/D89642
-
Florian Hahn authored
This patch adds a remarks that provides counts for each opcode per basic block. An snippet of the generated information can be seen below. The current implementation uses the target specific opcode for the counts. For example, on AArch64 this means we currently get 2 entries for `add` instructions if the block contains 32 and 64 bit adds. Similarly, immediate version are treated differently. Unfortunately there seems to be no convenient way to get only the mnemonic part of the instruction as a string AFAIK. This could be improved in the future. ``` --- !Analysis Pass: asm-printer Name: InstructionMix DebugLoc: { File: arm64-instruction-mix-remarks.ll, Line: 30, Column: 30 } Function: foo Args: - String: 'BasicBlock: ' - BasicBlock: else - String: "\n" - String: INST_MADDWrrr - String: ': ' - INST_MADDWrrr: '2' - String: "\n" - String: INST_MOVZWi - String: ': ' - INST_MOVZWi: '1' ``` Reviewed By: anemet, thegameg, paquette Differential Revision: https://reviews.llvm.org/D89892 -
Sebastian Neubauer authored
If no pal metadata is given, default to the msgpack format instead of the legacy metadata. This makes tests better readable. Differential Revision: https://reviews.llvm.org/D90035
-
Evgeny Leviant authored
-
Kai Luo authored
-
Kazushi (Jam) Marukawa authored
Support atomic load instruction and add a regression test. VE uses release consitency, so need to insert fence around atomic instructions. This patch enable AtomicExpandPass and use emitLeadingFence and emitTrailingFence mechanism for such purpose. Reviewed By: simoll Differential Revision: https://reviews.llvm.org/D90135
-
Evgeny Leviant authored
Differential revision: https://reviews.llvm.org/D90024
-
Evgeny Leviant authored
Differential revision: https://reviews.llvm.org/D90029
-
Evgeny Leviant authored
Differential revision: https://reviews.llvm.org/D90045
-
LLVM GN Syncbot authored
-
David Green authored
This adds a MultiHazardRecognizer and starts to make use of it in the ARM backend. The idea of the class is to allow multiple independent hazard recognizers to be added to a single base MultiHazardRecognizer, allowing them to all work in parallel without requiring them to be chained into subclasses. They can then be added or not based on cpu or subtarget features, which will become useful in the ARM backend once more hazard recognizers are being used for various things. This also renames ARMHazardRecognizer to ARMHazardRecognizerFPMLx in the process, to more clearly explain what that recognizer is designed for. Differential Revision: https://reviews.llvm.org/D72939
-
Kazushi (Jam) Marukawa authored
Support atomic fence instruction and add a regression test. Add MEMBARRIER pseudo insturction also to use it as a barrier against to the compiler optimizations. Reviewed By: simoll Differential Revision: https://reviews.llvm.org/D90112
-
Max Kazantsev authored
-
Max Kazantsev authored
-
Max Kazantsev authored
-
Max Kazantsev authored
No exact example where it would help, but it's a generally a more powerful way to prove predicates.
-
Kirill Bobyrev authored
It requires Index.proto to be built first. Failed builds: https://github.com/clangd/clangd/runs/1305985916
-
Christudasan Devadasan authored
We use an absolute address for stack objects and it would be necessary to have a constant 0 for soffset field. Fixes: SWDEV-228562 Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D89234
-
Craig Topper authored
The 0xf3 prefix has been defined as wbnoinvd on Icelake Server. So the prefix isn't ignored by the CPU. AMD documentation suggests that wbnoinvd is treated as wbinvd on older processors. Intel documentation is not clear. Perhaps 0xf2 and 0x66 are treated the same, but its not documented. This patch changes TB to PS in the td file so 0xf2 and 0x66 will be treated as errors. This matches versions of objdump after wbnoinvd was added.
-