- Sep 20, 2022
-
-
Michał Górny authored
Make `--config`, `--no-default-config` and `--config-*-dir` CoreOptions to enable their availability to all clang driver modes. This improves consistency given that the default set of configuration files is processed independently of mode anyway. Differential Revision: https://reviews.llvm.org/D134191
-
Sebastian Peryt authored
This is the first patch in a series intended for removing flag -enable-new-pm=0 from lit tests. This is part of a bigger effort of completely removing legacy code related to legacy pass manager in favor of currently default new pass manager. In this patch flag has been removed only from tests where no significant change has been required because checks has been duplicated for both PMs. Reviewed By: fhahn Differential Revision: https://reviews.llvm.org/D134150
-
Felipe de Azevedo Piovezan authored
A common debugging pattern is to set a breakpoint that only stops after a number of hits is recorded. The current implementation never resets the hit count of breakpoints; as such, if a user re-`run`s their program, the debugger will never stop on such a breakpoint again. This behavior is arguably undesirable, as it renders such breakpoints ineffective on all but the first run. This commit changes the implementation of the `Will{Launch, Attach}` methods so that they reset the _target's_ breakpoint hitcounts. Differential Revision: https://reviews.llvm.org/D133858 -
mydeveloperday authored
Working in a mixed environment of both vscode/vim with a team configured prettier configuration, this can leave clang-format and prettier fighting each other over the formatting of arrays, both simple arrays of elements. This review aims to add some "control knobs" to the Json formatting in clang-format to help align the two tools so they can be used interchangeably. This will allow simply arrays `[1, 2, 3]` to remain on a single line but will break those arrays based on context within that array. Happy to change the name of the option (this is the third name I tried) Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D133589
-
Keith Smiley authored
I don't think __obj_selrefs is a thing, but __objc_selrefs definitely is. Differential Revision: https://reviews.llvm.org/D130221
-
Rahul Joshi authored
Use isa<> instead of dyn_cast END_PUBLIC Differential Revision: https://reviews.llvm.org/D134092
-
Krzysztof Parzyszek authored
Enable creating an idiom: V -> opJoin(SplitVectorOp(V))
-
Simon Pilgrim authored
This was achieved with the 'cost-tables vs llvm-mca' script D103695
-
Xiang Li authored
Remove check which disable BitInt as element type for ext_vector. Enabling it for HLSL to use _BitInt(16) as 16bit int at https://reviews.llvm.org/D133668 Reviewed By: erichkeane Differential Revision: https://reviews.llvm.org/D133634
-
zhijian authored
https://lab.llvm.org/buildbot/#/builders/174/builds/13432 XCOFFObjectFile.cpp:805:12: error: reinterpret_cast from 'unsigned long' to 'uintptr_t' (aka 'unsigned int') is not allowed return reinterpret_cast<uintptr_t>(0ul);
-
Florian Hahn authored
The test requires the AArch64 backend, so move it to the right subdir.
-
Florian Hahn authored
Includes a test for the miscompile in #57712.
-
Krzysztof Parzyszek authored
-
Mingming Liu authored
Use opaqueptr for test case llvm/test/Transforms/SimplifyCFG/preserve-llvm-loop-metadata.ll. - Adjust variable number accordingly since bitcast between different pointer types are not necessary. Differential Revision: https://reviews.llvm.org/D134159
-
- Sep 19, 2022
-
-
zhijian authored
https://lab.llvm.org/buildbot/#/builders/216/builds/9977 XCOFFOtFile.cpp: error C3487: 'unsigned long': all return expressions must deduce to the same type: previously it was 'uintptr_t'
-
Katherine Rasmussen authored
Write a semantics test for the atomic intrinsic subroutine, atomic_and. Reviewed By: rouson Differential Revision: https://reviews.llvm.org/D133727
-
Simon Pilgrim authored
This was achieved with the 'cost-tables vs llvm-mca' script D103695
-
spupyrev authored
After BOLT's merge to LLVM, there are two (almost identical) versions of the code layout algorithm. The diff unifies the implementations by keeping the one in LLVM. There are mild changes in the resulting block orders. I tested the changes extensively both on the clang binary and on prod services. Didn't see stat sig differences on average. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D129895
-
zhijian authored
Summary: according nm in AIX OS , https://www.ibm.com/docs/en/aix/7.2?topic=n-nm-command In AIX OS, The default is to process 32-bit object files (ignore 64-bit objects). The mode can also be set with the OBJECT_MODE environment variable. For example, OBJECT_MODE=64 causes nm to process any 64-bit objects and ignore 32-bit objects. The -X flag overrides the OBJECT_MODE variable. In non AIX OS. The default is to process all support object files. and not support the OBJECT_MODE environment variable. Reviewers: James Henderson Differential Revision: https://reviews.llvm.org/D132494
-
Tue Ly authored
-
Louis Dionne authored
In https://llvm.org/D56913, we added an emulation for the __atomic_always_lock_free compiler builtin when compiling in Freestanding mode. However, the emulation did (and could not) give exactly the same answer as the compiler builtin, which led to a potential ABI break for e.g. enum classes. After speaking to the original author of D56913, we agree that the correct behavior is to instead always use the compiler builtin, since that provides a more accurate answer, and __atomic_always_lock_free is a purely front-end builtin which doesn't require any runtime support. Furthermore, it is available regardless of the Standard mode (see https://godbolt.org/z/cazf3ssYY). However, this patch does constitute an ABI break. As shown by https://godbolt.org/z/1eoex6zdK: - In LLVM <= 11.0.1, an atomic<enum class with 1 byte> would not contain a lock byte. - In LLVM >= 12.0.0, an atomic<enum class with 1 byte> would contain a lock byte. This patch breaks the ABI again to bring it back to 1 byte, which seems like the correct thing to do. Fixes #57440 Differential Revision: https://reviews.llvm.org/D133377
-
Simon Pilgrim authored
-
Sam McCall authored
This feature relies on Relations in the index being complete. An out-of-tree index implementation is missing some override relations, so such renames end up breaking the code. We plan to fix it, but this flag is a cheap band-aid for now. Differential Revision: https://reviews.llvm.org/D133440
-
Simon Pilgrim authored
This was achieved with the 'cost-tables vs llvm-mca' script D103695
-
zhijian authored
Summary: llvm-readobj support a new option --exception-section for xcoff object file. https://www.ibm.com/docs/en/aix/7.2?topic=formats-xcoff-object-file-format#XCOFF__iua3i23ajbau Reviewers: James Henderson,Paul Scoropan Differential Revision: https://reviews.llvm.org/D133030
-
Sam McCall authored
When we aim a hint at some expanded tokens, we're only willing to attach it to spelled tokens that exactly corresponde. e.g. int zoom(int x, int y, int z); int dummy = zoom(NUMBERS); Here we want to place a hint "x:" on the expanded "1", but we shouldn't be willing to place it on NUMBERS, because it doesn't *exactly* correspond (it has more tokens). Fortunately we don't even have to implement this algorithm from scratch, TokenBuffer has it. Fixes https://github.com/clangd/clangd/issues/1289 Fixes https://github.com/clangd/clangd/issues/1118 Fixes https://github.com/clangd/clangd/issues/1018 Differential Revision: https://reviews.llvm.org/D133982
-
Benjamin Kramer authored
-
Guray Ozen authored
This revision adds a new op `map_nested_foreach_thread_to_gpu_threads` to transform dialect. The op searches `scf.foreach_threads` inside the `gpu_launch` and distributes them with `gpu.thread_id` attribute. Loop mapping is explicit and given by the `map_nested_foreach_thread_to_gpu_threads` op. Mapping is done one-to-one, therefore the loops dissappear. The dynamic trip count or trip count that are larger than thread size are not supported for the time being. However, we can indeed support them by generating a loop inside with cyclic scheduling. For the time being, trip counts that are dynamic or bigger than thread sizes are not supported. However, in the future the compiler can indeed generate a loop with static cyclic scheduling to support these cases. Current mechanism allows `scf.foreach_threads` to be siblings or nested. There cannot be interleaving code between the loops when they are nested. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D133950
-
Tue Ly authored
-
Tue Ly authored
Implement exp10f function correctly rounded to all rounding modes. Algorithm: perform range reduction to reduce ``` 10^x = 2^(hi + mid) * 10^lo ``` where: ``` hi is an integer, 0 <= mid * 2^5 < 2^5 -log10(2) / 2^6 <= lo <= log10(2) / 2^6 ``` Then `2^mid` is stored in a table of 32 entries and the product `2^hi * 2^mid` is performed by adding `hi` into the exponent field of `2^mid`. `10^lo` is then approximated by a degree-5 minimax polynomials generated by Sollya with: ``` > P = fpminimax((10^x - 1)/x, 4, [|D...|], [-log10(2)/64. log10(2)/64]); ``` Performance benchmark using perf tool from the CORE-MATH project on Ryzen 1700: ``` $ CORE_MATH_PERF_MODE="rdtsc" ./perf.sh exp10f GNU libc version: 2.35 GNU libc release: stable CORE-MATH reciprocal throughput : 10.215 System LIBC reciprocal throughput : 7.944 LIBC reciprocal throughput : 38.538 LIBC reciprocal throughput : 12.175 (with `-msse4.2` flag) LIBC reciprocal throughput : 9.862 (with `-mfma` flag) $ CORE_MATH_PERF_MODE="rdtsc" ./perf.sh exp10f --latency GNU libc version: 2.35 GNU libc release: stable CORE-MATH latency : 40.744 System LIBC latency : 37.546 BEFORE LIBC latency : 48.989 LIBC latency : 44.486 (with `-msse4.2` flag) LIBC latency : 40.221 (with `-mfma` flag) ``` This patch relies on https://reviews.llvm.org/D134002 Reviewed By: orex, zimmermann6 Differential Revision: https://reviews.llvm.org/D134104
-
Tue Ly authored
-
Nikita Popov authored
This should fix the expensive checks build. Ideally we would not have invalid loops in LoopDispositions.
-
Simon Pilgrim authored
Without LZCNT/BMI, the *_ZERO_UNDEF costs are cheaper as they can avoid the zero handling.
-
Nikita Popov authored
This reverts commit e5581df6. This causes major compile-time regressions, about 2-3% end-to-end on CTMark.
-
Tue Ly authored
Optimize the core part of `tanhf` implementation that is to compute `e^x` similar to https://reviews.llvm.org/D133870. Factor the constants and polynomial approximation out so that it can be used for `exp10f` Performance benchmark using perf tool from the CORE-MATH project on Ryzen 1700: ``` $ CORE_MATH_PERF_MODE="rdtsc" ./perf.sh tanhf GNU libc version: 2.35 GNU libc release: stable CORE-MATH reciprocal throughput : 13.377 System LIBC reciprocal throughput : 55.046 BEFORE: LIBC reciprocal throughput : 75.674 LIBC reciprocal throughput : 33.242 (with `-msse4.2` flag) LIBC reciprocal throughput : 25.927 (with `-mfma` flag) AFTER: LIBC reciprocal throughput : 26.359 LIBC reciprocal throughput : 18.888 (with `-msse4.2` flag) LIBC reciprocal throughput : 14.243 (with `-mfma` flag) $ CORE_MATH_PERF_MODE="rdtsc" ./perf.sh tanhf --latency GNU libc version: 2.35 GNU libc release: stable CORE-MATH latency : 43.365 System LIBC latency : 123.499 BEFORE LIBC latency : 112.968 LIBC latency : 104.908 (with `-msse4.2` flag) LIBC latency : 92.310 (with `-mfma` flag) AFTER LIBC latency : 69.828 LIBC latency : 63.874 (with `-msse4.2` flag) LIBC latency : 57.427 (with `-mfma` flag) ``` Reviewed By: orex, zimmermann6 Differential Revision: https://reviews.llvm.org/D134002
-
Simon Pilgrim authored
Only AVX512 has decent CTTZ/CTLZ vector ops, add tests to ensure we definitely vectorize these
-
Aaron Ballman authored
This spotted a mistake with the original patch, so it puts the status back to "partial" in the C status tracking page. This amends 51038362.
-
Simon Pilgrim authored
Allows to determine known zero elements, which particularly helps simplification of DIV/REM by constant patterns
-
Aaron Ballman authored
-
Nicolas Vasilache authored
Given an opOperand uniquely determined by the operation `%op` and the operand number `num`, the `transform.get_producer_of_operand %op[num]` returns the handle to the unique operation that produced the SSA value used as opOperand. The transform fails if the operand is a block argument. Differential Revision: https://reviews.llvm.org/D134171
-