- Mar 23, 2021
-
-
Matt Arsenault authored
-
Philip Reames authored
-
Florian Hahn authored
This patch adds CHECK-LABEL lines to llvm/test/Transforms/LoopVectorize/vplan-printing.ll in order to make failures slightly easier to diagnose.
-
Matt Arsenault authored
-
serge-sans-paille authored
Better safe than sorry here, quoting Craig Topper: > Clang passes a pretty lengthy feature string.
-
Chris Lattner authored
This allows adding a C function pointer as a matchAndRewrite style pattern, which is a very common case. This adopts it in ExpandTanh to show how it reduces a level of nesting. We could allow C++ lambdas here, but that doesn't work as well with type inference in the common case. Instead of: patterns.insert(convertTanhOp); you need to specify: patterns.insert<math::TanhOp>(convertTanhOp); which is boilerplate'y. Capturing state like this is very uncommon, so we choose to require clients to define their own structs and use the non-convenience method when they need to do so. Differential Revision: https://reviews.llvm.org/D99039
-
Stefan Pintilie authored
There is a bug when initial exec is relaxed to local exec. In the following situation: InitExec.c ``` extern __thread unsigned TGlobal; unsigned getConst(unsigned*); unsigned addVal(unsigned, unsigned*); unsigned GetAddrT() { return addVal(getConst(&TGlobal), &TGlobal); } ``` Def.c ``` __thread unsigned TGlobal; unsigned getConst(unsigned* A) { return *A + 3; } unsigned addVal(unsigned A, unsigned* B) { return A + *B; } ``` The problem is in InitExec.c but Def.c is required if you want to link the example and see the problem. To compile everything: ``` clang -O3 -mcpu=pwr10 -c InitExec.c clang -O3 -mcpu=pwr10 -c Def.c ld.lld InitExec.o Def.o -o IeToLe ``` If you objdump the problem object file: ``` $ llvm-objdump -dr --mcpu=pwr10 InitExec.o ``` you will get the following assembly: ``` 0000000000000000 <GetAddrT>: 0: a6 02 08 7c mflr 0 4: f0 ff c1 fb std 30, -16(1) 8: 10 00 01 f8 std 0, 16(1) c: d1 ff 21 f8 stdu 1, -48(1) 10: 00 00 10 04 00 00 60 e4 pld 3, 0(0), 1 0000000000000010: R_PPC64_GOT_TPREL_PCREL34 TGlobal 18: 14 6a c3 7f add 30, 3, 13 0000000000000019: R_PPC64_TLS TGlobal 1c: 78 f3 c3 7f mr 3, 30 20: 01 00 00 48 bl 0x20 0000000000000020: R_PPC64_REL24_NOTOC getConst 24: 78 f3 c4 7f mr 4, 30 28: 30 00 21 38 addi 1, 1, 48 2c: 10 00 01 e8 ld 0, 16(1) 30: f0 ff c1 eb ld 30, -16(1) 34: a6 03 08 7c mtlr 0 38: 00 00 00 48 b 0x38 0000000000000038: R_PPC64_REL24_NOTOC addVal ``` The lines of interest are: ``` 10: 00 00 10 04 00 00 60 e4 pld 3, 0(0), 1 0000000000000010: R_PPC64_GOT_TPREL_PCREL34 TGlobal 18: 14 6a c3 7f add 30, 3, 13 0000000000000019: R_PPC64_TLS TGlobal 1c: 78 f3 c3 7f mr 3, 30 ``` Which once linked gets turned into: ``` 10010210: ff ff 03 06 00 90 6d 38 paddi 3, 13, -28672, 0 10010218: 00 00 00 60 nop 1001021c: 78 f3 c3 7f mr 3, 30 ``` The problem is that register 30 is never set after the optimization. Therefore it is not correct to relax the above instructions by replacing the add instruction with a nop. Instead the add instruction should be replaced with a copy (mr) instruction. If the add uses the same resgiter as input and as ouput then it is safe to continue to replace the add with a nop. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D95262 -
Chia-hung Duan authored
In the original structure, it will try to match CHECK-LABEL first then see if the subsequent doesn't have the target strings. This is not what we are expected. We are expecting the two functions which will be deleted should be matched before CHECK-LABEL. Also fixed the function names. Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D99060
-
Matt Morehouse authored
-
Philip Reames authored
-
Philip Reames authored
-
Rob Suderman authored
Multiply-shift requires wider compute types or CPU specific code to avoid premature truncation, apply_shift fixes this issue Also, Tosa's mul op supports different input / output types. Added path that sign-extends input values to int-32 values before multiplying. Differential Revision: https://reviews.llvm.org/D99011
-
Peter Steinfeld authored
If you specify a specific procedure of a generic interface that has the same name as both the generic interface and a preceding derived type, the compiler would fail an internal call to CHECK(). I fixed this by testing for this situation when processing specific procedures. I also added a test that will cause the call to CHECK() to fail without this new code. Differential Revision: https://reviews.llvm.org/D99085
-
Philip Reames authored
-
Lang Hames authored
-
Philip Reames authored
-
Craig Topper authored
[LegalizeDAG] Add asserts to verify the types of custom legalized operation matches the original node. We've messed this up a few times recently on RISCV. Experiments with these asserts found a couple issues on other targets as well. They've all been cleaned up now so we can put in these asserts to catch future issues I had to waive Glue because ADDC/ADDE/etc legalization replaces Glue with i32 on at least AArch64. X86 used to do the same before we switched to ADDCARRY. So I guess that's just how that works. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D98979
-
Philip Reames authored
-
Craig Topper authored
I've split the gather/scatter custom handler to avoid complicating it with even more differences between gather/scatter. Tests are the scalable vector tests with the vscale removed and dropped the tests that used vector.insert. We're probably not as thorough on the splitting cases since we use 128 for VLEN here but scalable vector use a known min size of 64. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D98991
-
Arthur Eubanks authored
This reverts commits 56700e93 and c2f9086b. Breaks multiple Android bots, e.g. https://lab.llvm.org/buildbot/#/builders/77/builds/4777.
-
LLVM GN Syncbot authored
-
Frank Derry Wanye authored
This lint check is a part of the FLOCL (FPGA Linters for OpenCL) project out of the Synergy Lab at Virginia Tech. FLOCL is a set of lint checks aimed at FPGA developers who write code in OpenCL. The altera unroll loops check finds inner loops that have not been unrolled, as well as fully-unrolled loops that should be partially unrolled due to unknown loop bounds or a large number of loop iterations. Based on the Altera SDK for OpenCL: Best Practices Guide.
-
Raphael Isemann authored
Objective-C apparently allows name conflicts between instance and class properties, so this is valid code: ``` @protocol DupProp @property (class, readonly) int prop; @property (readonly) int prop; @end ``` The ASTImporter however isn't aware of this and will consider the two properties as if they are the same property because it just compares their name and types. This causes that when importing both properties we only end up with one property (whatever is imported first from what I can see). Beside generating a different AST this also leads to a bunch of asserts and crashes as we still correctly import the two different getters for both properties (the import code for methods does the correct check where it differentiated between instance and class methods). As one of the setters will not have its associated ObjCPropertyDecl imported, any call to `ObjCMethodDecl::findPropertyDecl` will just lead to an assert or crash. Fixes rdar://74322659 Reviewed By: shafik, kastiglione Differential Revision: https://reviews.llvm.org/D99077
-
Siva Chandra authored
-
Stefan Gränitz authored
The `callB()` template function always moved errors on return, because in the majority of cases its return type is an `Expected<T>` and the error must be moved into the implicit ctor. For the special case of a `void` result, however, the `ResultTraits` class is specialized and the return type is a raw `Error`. Some build bots complain, that in favor of NRVO errors should not be moved in this case. ``` llvm/include/llvm/ExecutionEngine/Orc/Shared/RPCUtils.h:1513:27: llvm/include/llvm/ExecutionEngine/Orc/Shared/RPCUtils.h:1519:27: llvm/include/llvm/ExecutionEngine/Orc/Shared/RPCUtils.h:1526:29: warning: moving a local object in a return statement prevents copy elision [-Wpessimizing-move] ``` The warning is reasonable from a type-system point of view. For performance it's entirely insignificant. Differential Revision: https://reviews.llvm.org/D98947
-
Stefan Gränitz authored
Don't leak ResourceKeys from MaterializationResponsibility::withResourceKeyDo() in notifyEmitted(). Also make some improvements in the overall implementation. Differential Revision: https://reviews.llvm.org/D98863
-
Stefan Gränitz authored
There can be multiple MaterializationResponsibilitys in-flight for a single ResourceKey. Hence, pending debug objects must be tracked by MaterializationResponsibility and not by ResourceKey. Differential Revision: https://reviews.llvm.org/D98785
-
Philip Reames authored
This patch exploits the knowledge that we may be running many fewer than bitwidth iterations of the loop, and may be able to disallow the overflow case. This patch specifically implements only the shl case, but this can be generalized to ashr and lshr without difficulty. Differential Revision: https://reviews.llvm.org/D98222
-
Bjorn Pettersson authored
Make sure we use PowerOf2Floor instead of PowerOf2Ceil when calculating max number of elements that fits inside a vector register (otherwise we could end up creating vectors larger than the maximum vector register size). Also make sure we honor the min/max VF (as given by TTI or cmd line parameters) when doing vectorizeStores. Reviewed By: anton-afanasyev Differential Revision: https://reviews.llvm.org/D97691
-
Bjorn Pettersson authored
-
Philip Reames authored
Triggered by discussion on D98222. The case where we have a loop variant step is suprising, and doesn't match the behavior of SCEV's recurrences. As such, make sure we call that out explicitly.
-
- Mar 22, 2021
-
-
Wenlei He authored
Switch to use cold threshold from profile summary for cold context merging and trimming, instead of relying on hard coded values. Minor refactoring included for switch names, etc. Differential Revision: https://reviews.llvm.org/D98921
-
Pavel Labath authored
The fix in 10d54e2f did not work.
-
Arthur O'Dwyer authored
The container headers don't need to include <functional> for any other reason (or at least, they wouldn't if we moved `less` and `equal_to` out of <functional>), so let's put `__libcpp_erase_if_container` somewhere that's common to the containers but outside of <functional>. Also, calling `std::erase_if(c, pred)` should not trigger ADL. Differential Revision: https://reviews.llvm.org/D99043
-
Matt Morehouse authored
Subsequent patches will implement page-aliasing mode for x86_64, which will initially only work for the primary heap allocator. We force callback instrumentation to simplify the initial aliasing implementation. Reviewed By: vitalybuka, eugenis Differential Revision: https://reviews.llvm.org/D98069
-
Matt Arsenault authored
-
Pavel Labath authored
The file contained bogus input - the DIE list was not properly terminated. This should not cause a crash, but it seems it was crashing at least on linux arm and x86 windows.
-
Stefan Pintilie authored
Do not try to materialize a constant using prefix instructions if the selection using non prefix instructions was able to do it using a single non prefix instruction. Reviewed By: nemanjai, #powerpc Differential Revision: https://reviews.llvm.org/D98791
-
Pavel Labath authored
lit has grown a feature where it stores the runtimes of all tests. Normally, these times should be stored in the build directory, but because our API tests have set test_exec_root to point to the source tree, it has ended up polluting our checkout and led to the .lit_test_times.txt being committed to the repository. Delete this file, and adjust the exec root of API tests. I've also needed to adjust the root of Shell tests, in order to avoid the two overlapping.
-
Joe Ellis authored
Previously only the i32 type was tested. Now, the {i,f}{16,32,64} types are tested. The v8{i,f}16 cases lower differently to the other cases, which is worth defending. The lowering for the other cases is currently identical, but probably worth having for the better coverage. Differential Revision: https://reviews.llvm.org/D98690
-