- Feb 11, 2020
-
-
John Regehr authored
-
Craig Topper authored
[X86] Custom lower ISD::FP16_TO_FP and ISD::FP_TO_FP16 on f16c targets instead of using isel patterns. We need to use vector instructions for these operations. Previously we handled this with isel patterns that used extra instructions and copies to handle the the conversions. Now we use custom lowering to emit the conversions. This allows them to be pattern matched and optimized on their own. For example we can now emit vpextrw to store the result if its going directly to memory. I've forced the upper elements to VCVTPHS2PS to zero to keep some code similar. Zeroes will be needed for strictfp. I've added a DAG combine for (fp16_to_fp (fp_to_fp16 X)) to avoid extra instructions in between to be closer to the previous codegen. This is a step towards strictfp support for f16 conversions.
-
Kai Luo authored
-
Fangrui Song authored
https://github.com/riscv/riscv-elf-psabi-doc/pull/131 assigned 58 to R_RISCV_IRELATIVE. Differential Revision: https://reviews.llvm.org/D74022
-
Johannes Doerfert authored
The existing wording leaves it unclear if C++ standard library data structures should be preferred over custom LLVM ones, e.g., SmallVector, even though common practice seems clear on the issue. This change makes the wording more explicit and aligns it better with the code base. Some motivating statistics: ``` ag SmallVector llvm/lib/ | wc 8846 40306 901421 ag 'std::vector' llvm/lib/ | wc 2123 8990 214482 ag SmallVector clang/lib/ | wc 3023 13824 281691 ag 'std::vector' clang/lib/ | wc 719 2914 72817 ``` Differential Revision: https://reviews.llvm.org/D74340
-
Evgenii Stepanov authored
The interceptor uses thread-local variables, which (until very recently) are emu-tls. An access to such variable may call malloc which can deadlock the runtime library.
-
Yuanfang Chen authored
With fix (somehow one hunk is missed).
-
Jason Molenda authored
-
Michael Kruse authored
Thanks Justin Paston-Cooper for the report.
-
Yuanfang Chen authored
This reverts commit 8a29cb44. fuzzer-linux bot has failure because of this.
-
River Riddle authored
Summary: There are a few field init values that are concrete but not complete/foldable (e.g. `?`). This allows for using those values as initializers without erroring out. Example: ``` class A { string value = ?; } class B<A impl> : A { let value = impl.value; // This currently emits an error. let value = ?; // This doesn't emit an error. } ``` Differential Revision: https://reviews.llvm.org/D74360 -
Nathan James authored
-
Michael Kruse authored
ISL changed some return types from unsigned to isl_size (typedef of int), which results in such warnings.
-
Michael Kruse authored
The primary motivation is to fix an assertion failure in isl_basic_map_alloc_equality: isl_assert(ctx, room_for_con(bmap, 1), return -1); Although the assertion does not occur anymore, I could not identify which of ISL's commits fixed it. Compared to the previous ISL version, Polly requires some changes for this update * Since ISL commit 20d3574 "perform parameter alignment by modifying both arguments to function" isl_*_gist_* and similar functions do not always align the paramter list anymore. This caused the parameter lists in JScop files to become out-of-sync. Since many regression tests use JScop files with a fixed parameter list and order, we explicitly call align_params to ensure a predictable parameter list. * ISL changed some return types to isl_size, a typedef of (signed) int. This caused some issues where the return type was unsigned int before: - No overload for std::max(unsigned,isl_size) - It cause additional 'mixed signed/unsigned comparison' warnings. Since they do not break compilation, and sizes larger than 2^31 were never supported, I am going to fix it separately. * With the change to isl_size, commit 57d547 "isl_*_list_size: return isl_size" also changed the return value in case of an error from 0 to -1. This caused undefined looping over isl_iterator since the 'end iterator' got index -1, never reached from the 'begin iterator' with index 0. * Some internal changes in ISL caused the number of operations to increase when determining access ranges to determine aliasing overlaps. In one test, this caused exceeding the default limit of 800000. The operations-limit was disabled for this test. -
Peter Collingbourne authored
Differential Revision: https://reviews.llvm.org/D74366
-
Yuanfang Chen authored
-
Yuanfang Chen authored
For CleanseCrashInput, discards stdout output anyway since it is not used. These changes are to defend against aggressive PID recycle on windows to reduce the chance of contention on files. Using pipe instead of file also workaround the problem that when the process is spawned by llvm-lit, the aborted process keeps a handle to the output file such that the output file can not be removed. This will cause random test failures. https://devblogs.microsoft.com/oldnewthing/20110107-00/?p=11803 Reviewers: kcc, vitalybuka Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D73329
-
diggerlin authored
SUMMARY: refator the std::tuple<uint64_t, StringRef, uint8_t> to structor Reviewers: daltenty Subscribers: wuzish, nemanjai, hiraditya Differential Revision: https://reviews.llvm.org/D74240
-
David Blaikie authored
I'm /guessing/ this isn't terribly testable without a very large input file. Even generated from a more compact assembly file, it's probably best not to generate a giant temporary test file - if I'm wrong about that/anyone has good suggestions for testing, I'm all ears! Based on post-commit review feedback from Igor Kudrin on eed02423
-
Amara Emerson authored
A downstream test exposed a simple logic bug with the manual pointer stripping code, fix that by just using stripPointerCasts() on the value. I don't think there's a way to expose this issue upstream.
-
Davide Italiano authored
-
Eric Christopher authored
-
Peter Collingbourne authored
This will be useful for optimizing the size class map. Differential Revision: https://reviews.llvm.org/D74098
-
Peter Collingbourne authored
Add an optional table lookup after the existing logarithm computation for MidSize < Size <= MaxSize during size -> class lookups. The lookup is O(1) due to indexing a precomputed (via constexpr) table based on a size table. Switch to this approach for the Android size class maps. Other approaches considered: - Binary search was found to have an unacceptable (~30%) performance cost. - An approach using NEON instructions (see older version of D73824) was found to be slightly slower than this approach on newer SoCs but significantly slower on older ones. By selecting the values in the size tables to minimize wastage (for example, by passing the malloc_info output of a target program to the included compute_size_class_config program), we can increase the density of allocations at a small (~0.5% on bionic malloc_sql_trace as measured using an identity table) performance cost. Reduces RSS on specific Android processes as follows (KB): Before After zygote (median of 50 runs) 26836 26792 (-0.2%) zygote64 (median of 50 runs) 30384 30076 (-1.0%) dex2oat (median of 3 runs) 375792 372952 (-0.8%) I also measured the amount of whole-system idle dirty heap on Android by rebooting the system and then running the following script repeatedly until the results were stable: for i in $(seq 1 50); do grep -A5 scudo: /proc/*/smaps | grep Pss: | cut -d: -f2 | awk '{s+=$1} END {print s}' ; sleep 1; done I did this 3 times both before and after this change and the results were: Before: 365650, 356795, 372663 After: 344521, 356328, 342589 These results are noisy so it is hard to make a definite conclusion, but there does appear to be a significant effect. On other platforms, increase the sizes of all size classes by a fixed offset equal to the size of the allocation header. This has also been found to improve density, since it is likely for allocation sizes to be a power of 2, which would otherwise waste space by pushing the allocation into the next size class. Differential Revision: https://reviews.llvm.org/D73824 -
Peter Collingbourne authored
This lets us remove two pointer indirections (one by removing the pointer, and another by making the AllocatorPtr declaration hidden) in the C++ wrappers. Differential Revision: https://reviews.llvm.org/D74356
-
Dimitry Andric authored
Summary: Instead of hand-crafting an offset into the structure returned by dlopen(3) to get at the link map, use the documented API. This is described in dlinfo(3): by calling it with `RTLD_DI_LINKMAP`, the dynamic linker ensures the right address is returned. This is a recommit of 92e267a9, with dlinfo(3) expliclity being referenced only for FreeBSD, non-Android Linux, NetBSD and Solaris. Other OSes will have to add their own implementation. Reviewers: devnexen, emaste, MaskRay, krytarowski Reviewed By: krytarowski Subscribers: krytarowski, vitalybuka, #sanitizers, llvm-commits Tags: #sanitizers, #llvm Differential Revision: https://reviews.llvm.org/D73990
-
Vedant Kumar authored
This breaks macOS, because TARGET_OS_EMBEDDED is always defined. Thanks to Jason Molenda for pointing this out. Revert "Do not define AcceptPIDFromInferior when it will not be used" This reverts commit d23c15a6. This reverts commit 936d1427.
-
Dimitry Andric authored
This reverts commit 92e267a9, as it appears Android is missing dlinfo(3).
-
Sanjay Patel authored
As discussed in PR41083: https://bugs.llvm.org/show_bug.cgi?id=41083 ...we can assert/crash in EarlyCSE using the current hashing scheme and instructions with flags. ValueTracking's matchSelectPattern() may rely on overflow (nsw, etc) or other flags when detecting patterns such as min/max/abs composed of compare+select. But the value numbering / hashing mechanism used by EarlyCSE intersects those flags to allow more CSE. Several alternatives to solve this are discussed in the bug report. This patch avoids the issue by doing simple matching of min/max/abs patterns that never requires instruction flags. We give up some CSE power because of that, but that is not expected to result in much actual performance difference because InstCombine will canonicalize these patterns when possible. It even has this comment for abs/nabs: /// Canonicalize all these variants to 1 pattern. /// This makes CSE more likely. (And this patch adds PhaseOrdering tests to verify that the expected transforms are still happening in the standard optimization pipelines. I left this code to use ValueTracking's "flavor" enum values, so we don't have to change the callers' code. If we decide to go back to using the ValueTracking call (by changing the hashing algorithm instead), it should be obvious how to replace this chunk. Differential Revision: https://reviews.llvm.org/D74285
-
Vedant Kumar authored
Null-check and adjut a TypeLoc before casting it to a FunctionTypeLoc. This fixes a crash in -fsanitize=nullability-return, and also makes the location of the nonnull type available when the return type is adjusted. rdar://59263039 Differential Revision: https://reviews.llvm.org/D74355
-
Ted Woodward authored
Summary: The lit feature object-emission was added because Hexagon did not support the integrated assembler, so some tests needed to be turned off with a Hexagon target. Hexagon now supports the integrated assembler, so this feature can be removed. Reviewers: bcain, kparzysz, jverma, whitequark, JDevlieghere Reviewed By: JDevlieghere Subscribers: mehdi_amini, hiraditya, steven_wu, dexonsmith, arphaman, llvm-commits Tags: #llvm Differential Revision: https://reviews.llvm.org/D73568
-
LLVM GN Syncbot authored
-
Hiroshi Yamauchi authored
Summary: It attempts to devirtualize a call on alloca through vtable loads. Reviewers: davidxl Subscribers: mgorny, Prazek, hiraditya, llvm-commits Tags: #llvm Differential Revision: https://reviews.llvm.org/D71308
-
Davide Italiano authored
This reverts commit 1a39f1b9 as it breaks macOS.
-
Lei Zhang authored
We have spv.entry_point_abi for specifying the local workgroup size. It should be decorated onto input gpu.func ops to drive the SPIR-V CodeGen to generate the proper SPIR-V module execution mode. Compared to using command-line options for specifying the configuration, using attributes also has the benefits that 1) we are now able to use different local workgroup for different entry points and 2) the tests contains the configuration directly. Differential Revision: https://reviews.llvm.org/D74012
-
Alexey Bataev authored
Added full support for 'release' clause in flush|atomic directives.
-
Hanhan Wang authored
Summary: After D72555 has been landed, `linalg.indexed_generic` also accepts ranked tensor as input and output. Add a test for it. Differential Revision: https://reviews.llvm.org/D74267
-
Martin Storsjö authored
The plugin expects to have undefined references to symbols exported by the loading process, which isn't supported by shared libraries on windows. Differential Revision: https://reviews.llvm.org/D74042
-
Nico Weber authored
-
Xiangling Liao authored
This patch: - enable frame pointer for AIX; - update some of red zone comments; - add/update testcases; Differential Revision: https://reviews.llvm.org/D72454
-