- May 24, 2023
-
-
Aaron Ballman authored
-
Markus Böck authored
C++s iterator concept requires operator* to return the same type as is specified by the iterators reference type. This functionality is especially important for older generic code that did not yet make use of auto. An example from within LLVM is iterator_adaptor_base which uses the reference type of the iterator it is wrapping as its return type for operator* (this class is used as base for a lot of other functionality like filter iterators and so on). Using any of the graph traversal iterators listed above with it would previously fail to compile due to reference being non-const while operator* returned a const reference. This patch fixes that by correctly specifying reference and using it as the return type of operator* explicitly to prevent further issues in the future. Differential Revision: https://reviews.llvm.org/D151198
-
Azat Khuzhin authored
This will avoid hardcoding all unsupported targets, since even after one more follow up fix [1], there is one more failure. [1]: https://reviews.llvm.org/D150886 Plus, if you want to run it locally on some target that CI does not covers, it could also false-positively fail, which is not good. Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D151046
-
Alex Langford authored
[lldb][NFCI] Merge implementations of ObjectFileMachO::GetMinimumOSVersion and ObjectFileMachO::GetSDKVersion These functions do the exact same thing (even if they look slightly different). I yanked the common implementation, cleaned it up, and shoved it into its own function. Differential Revision: https://reviews.llvm.org/D151120
-
Diego Caballero authored
This patch adds support to shape cast a vector<1x1x1...1xElemenType> to a vector<ElementType> and the other way around. Differential Revision: https://reviews.llvm.org/D151169
-
Joseph Huber authored
Summry: This was accidentally dropped from a previous patch following a rebase. Fix it to where it's consistent. Differential Revision: https://reviews.llvm.org/D151232
-
Aaron Ballman authored
We documented -fmsc-version as defaulting to 1300 and -fms-compatibility-version as defaulting to 1800, neither of which were accurate. We currently default to 1920. See MSVCToolChain::computeMSVCVersion() for details.
-
Jin Xin Ng authored
Ensures a subsequent call (via an external caller) to __sanitizer_get_allocated_size via hooks will return a valid size. This allows a faster version of __sanitizer_get_allocated_size to be implemented, which can skip checks. Test to ensure RunFreeHooks' call order will come with __sanitizer_get_allocated_size_fast Differential Revision: https://reviews.llvm.org/D151151
-
Mark de Wever authored
This is a followup of the review comments in D144499. Reviewed By: ldionne, philnik, #libc Differential Revision: https://reviews.llvm.org/D150585
-
Mark de Wever authored
This member has been added in D148641 so it can be used in the formatter to avoid creating a "temporary" string. Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D150791
-
Mark de Wever authored
This adds the cppm files of D144994. These files by themselves will do nothing. The goal is to reduce the size of D144994 and making it easier to review the real changes of the patch. Implements parts of - P2465R3 Standard Library Modules std and std.compat Reviewed By: ldionne, ChuanqiXu, aaronmondal, #libc Differential Revision: https://reviews.llvm.org/D151030
-
Fangrui Song authored
There are two motivations. `-fno-pic -fstack-protector -mstack-protector-guard=global` created `__stack_chk_guard` is referenced directly on all ELF OSes except FreeBSD. This patch allows referencing the symbol indirectly with -fno-direct-access-external-data. Some Linux kernel folks want `-fno-pic -fstack-protector -mstack-protector-guard-reg=gs -mstack-protector-guard-symbol=__stack_chk_guard` created `__stack_chk_guard` to be referenced directly, avoiding R_X86_64_REX_GOTPCRELX (even if the relocation may be optimized out by the linker). https://github.com/llvm/llvm-project/issues/60116 Why they need this isn't so clear to me. --- Add module flag "direct-access-external-data" and set the dso_local property of the stack protector symbol. The module flag can benefit other LLVMCodeGen synthesized symbols that are not represented in LLVM IR. Nowadays, with `-fno-pic` being uncommon, ideally we should set "direct-access-external-data" when it is true. However, doing so would require ~90 clang/test tests to be updated, which are too much. As a compromise, we set "direct-access-external-data" only when it's different from the implied default value. Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D150841
-
Mark de Wever authored
During the ISO C++ Committee meeting plenary session the C++23 Standard has been voted as technical complete. This updates the reference to c++2b to c++23 and updates the __cplusplus macro. Note since we use clang-tidy 16 a small work-around is needed. Clang knows -std=c++23 but clang-tidy not so for now force the lit compiler flag to use -std=c++2b instead of -std=c++23. Reviewed By: #libc, philnik, jloser, ldionne Differential Revision: https://reviews.llvm.org/D150795
-
Jin Xin Ng authored
Previously lsan would not invoke hooks on reallocations. An accompanying regression test is included in sanitizer_common. This change also moves hook calls to a location where subsequent calls (via an external caller) to __sanitizer_get_allocated_size via hooks will return a valid size. This allows a faster version of __sanitizer_get_allocated_size to be implemented, which can skip checks. Test to ensure RunFreeHooks' call order will come with __sanitizer_get_allocated_size_fast Differential Revision: https://reviews.llvm.org/D151175
-
Slava Zakharin authored
There are several observations regarding the copy-in/copy-out: * Actual argument associated with INTENT(OUT) dummy argument that requires finalization (7.5.6.3 p. 7) may be read by the finalization function, so a copy-in is required. * A temporary created for the copy-in/copy-out must be destroyed without finalization after the call (or after the corresponding copy-out), otherwise, memory leaks may occur. * The copy-out assignment must not perform finalization for the LHS. * The copy-out assignment from the temporary to the actual argument may or may not need to initialize the LHS. This change-set introduces new runtime methods: CopyOutAssign and DestroyWithoutFinalization. They are called by the compiler generated code to match the behavior described above. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D151135 -
max authored
Differential Revision: https://reviews.llvm.org/D151167
-
Pavel Iliin authored
On AArch64 for function multiversioning target_version/target_clones attributes should be used. The patch fixes the defect allowing target attribute to cause multiversioning. Differential Revision: https://reviews.llvm.org/D150867
-
Craig Topper authored
[LegalizeTypes][ARM][AArch6][RISCV][VE][WebAssembly] Add special case for smin(X, -1) and smax(X, 0) to ExpandIntRes_MINMAX. We can compute a simpler expression for Lo for these cases. This is an alternative for the test cases in D151180 that works for more targets. This is similar to some of the special cases we have for expanding setcc operands. Differential Revision: https://reviews.llvm.org/D151182
-
Joseph Huber authored
These files aren't fully formatted. I'm guessing this was a holdover from when `clang-format` was totally broken for OpenMP offloading. Format the files to be more consistent. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D151226
-
- May 23, 2023
-
-
Joseph Huber authored
Currently we have the `send_n` and `recv_n` routines to stream data, such as a string to print, to the other side. The first operation is to send the size so the other side knows the number of bytes to recieve. However, this wasted 56 bytes that could've been sent. This meant that small values, like the arguments to a function to call on the host for example, needed to perform an extra send. This patch sends the first 56 bytes in the first packet and continues if necessary. Depends on D150992 Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D151041
-
Joseph Huber authored
We provide the `send_n` and `recv_n` utilities as a generic way to stream data between both sides of the process. This was previously tested and performed as expected when using a string of constant size. However, when the size was allowed to diverge between the threads in the warp or wavefront this could deadlock. This did not occur on NVPTX because of the use of the explicit warp sync. However, on AMD one of the work items in the wavefront could continue executing and hit the next `recv` call before the other threads, then we would deadlock as we violated the RPC invariants. This patch replaces the for loop with a thread ballot. This will cause every thread in the warp or wavefront to continue executing the loop until all of them can exit. This acts as a more explicit wavefront sync. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D150992
-
Nikolas Klauser authored
This also removes some tests which we have grouped together into robust_from_*.pass.cpp tests. Specifically, checking that - `ranges::dangling` is returned is done in `libcxx/test/std/algorithms/ranges_robust_against_dangling.pass.cpp` - `std::invoke` is used is done in `libcxx/test/std/algorithms/ranges_robust_against_omitting_invoke.pass.cpp`. - implicit conversion to bool works is done in `libcxx/test/std/algorithms/ranges_robust_against_nonbool_predicates.pass.cpp` Checking the comparison order is invalid because the `operator==` isn't symmetric. Checking what the exact type of `operator==` is, is invalid because comparing the same object has to yield the same results if the objects are not modified. Reviewed By: ldionne, #libc Spies: EricWF, libcxx-commits Differential Revision: https://reviews.llvm.org/D150588
-
Nikolas Klauser authored
`basic_ios` is defined in `<ios>`, so it seems weird that we declare the explicit instantiation for it i `<streambuf>`, which is technically unrelated. Reviewed By: #libc, EricWF, ldionne Spies: ldionne, EricWF, libcxx-commits Differential Revision: https://reviews.llvm.org/D150912
-
Yaxun (Sam) Liu authored
D106463 caused a regression that prevents std::malloc to be called in the device function, which is allowed with nvcc. Basically the standard C++ header introducing malloc in std namespace by using ::malloc. The device ::malloc function needs to be declared before using ::malloc to be introduced into std namespace. Revert D106463 and add a test. Reviewed by: Artem Belevich Differential Revision: https://reviews.llvm.org/D150965
-
Nikolas Klauser authored
Reviewed By: ldionne, #libc, Mordante Spies: arichardson, Mordante, libcxx-commits Differential Revision: https://reviews.llvm.org/D151119
-
Felipe de Azevedo Piovezan authored
In an effort to unify the different dwarf parsers available in the codebase, this commit removes LLDB's custom parsing for the `.debug_ranges` DWARF section, instead calling into LLVM's parser. Subsequent work should look into unifying `llvm::DWARDebugRangeList` (whose entries are pairs of (start, end) addresses) with `lldb::DWARFRangeList` (whose entries are pairs of (start, length)). The lists themselves are also different data structures, but functionally equivalent. Depends on D150363 Differential Revision: https://reviews.llvm.org/D150366
-
Tue Ly authored
Implement double precision log1p function correctly rounded to all rounding modes. **Performance** - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.93%. - Benchmarks with `./perf.sh` tool from the CORE-MATH project, unit is (CPU clocks / call). - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log1p GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 39.792 + 1.011 clc/call; Median-Min = 0.940 clc/call; Max = 41.373 clc/call; -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 87.285 + 1.135 clc/call; Median-Min = 1.299 clc/call; Max = 89.715 clc/call; -- System LIBC reciprocal throughput -- [####################] 100 % Ntrial = 20 ; Min = 20.666 + 0.123 clc/call; Median-Min = 0.125 clc/call; Max = 20.828 clc/call; -- LIBC reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 20.928 + 0.771 clc/call; Median-Min = 0.725 clc/call; Max = 22.767 clc/call; -- LIBC reciprocal throughput -- without FMA [####################] 100 % Ntrial = 20 ; Min = 31.461 + 0.528 clc/call; Median-Min = 0.602 clc/call; Max = 36.809 clc/call; ``` - Latency from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log1p --latency GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 77.875 + 0.062 clc/call; Median-Min = 0.051 clc/call; Max = 78.003 clc/call; -- CORE-MATH latency -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 101.958 + 1.202 clc/call; Median-Min = 1.325 clc/call; Max = 104.452 clc/call; -- System LIBC latency -- [####################] 100 % Ntrial = 20 ; Min = 60.581 + 1.443 clc/call; Median-Min = 1.611 clc/call; Max = 62.285 clc/call; -- LIBC latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 48.817 + 1.108 clc/call; Median-Min = 1.300 clc/call; Max = 50.282 clc/call; -- LIBC latency -- without FMA [####################] 100 % Ntrial = 20 ; Min = 61.121 + 0.599 clc/call; Median-Min = 0.761 clc/call; Max = 62.020 clc/call; ``` - Accurate pass latency: ``` $ ./perf.sh log1p --latency --simple_stat GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA 760.444 -- CORE-MATH latency -- without FMA (-march=x86-64-v2) 827.880 -- LIBC latency -- with FMA 711.837 -- LIBC latency -- without FMA 764.317 ``` Reviewed By: zimmermann6 Differential Revision: https://reviews.llvm.org/D151049
-
Nikita Popov authored
When sinking and users are dropped, add the using instructions to the worklist, as they can likely be removed as well. This should be NFC apart from worklist order effects.
-
Jean Perier authored
This patch moves the counter and storage management part of the array constructor inlined temporary strategy into its own utility so that it can be reused for the simple cases of temporary creations inside WHERE and FORALL. It actually fixes a bug where the counter first value used for addressing was "2" leading to read/write after the allocated storage... It seems I ran the tests end-to-end without the HLFIR flag when previously testing this. So this may clear some segfaults. Differential Revision: https://reviews.llvm.org/D151106
-
Tom Eccles authored
In upstream mlir, the dialect conversion infrastructure is used for lowering from one dialect to another: the passes are of the form XToYPass. Whereas, transformations within the same dialect tend to use applyPatternsAndFoldGreedily. In this case, the full complexity of applyPatternsAndFoldGreedily isn't needed so we can get away with the simpler applyOpPatternsAndFold. This change was suggested by @jeanPerier Differential Revision: https://reviews.llvm.org/D150853
-
Tue Ly authored
Implement double precision log2 function correctly rounded to all rounding modes. See https://reviews.llvm.org/D150014 for a more detail description of the algorithm. **Performance** - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.91%. - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log2 GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 15.458 + 0.204 clc/call; Median-Min = 0.224 clc/call; Max = 15.867 clc/call; -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 23.711 + 0.524 clc/call; Median-Min = 0.443 clc/call; Max = 25.307 clc/call; -- System LIBC reciprocal throughput -- [####################] 100 % Ntrial = 20 ; Min = 14.807 + 0.199 clc/call; Median-Min = 0.211 clc/call; Max = 15.137 clc/call; -- LIBC reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 17.666 + 0.274 clc/call; Median-Min = 0.298 clc/call; Max = 18.531 clc/call; -- LIBC reciprocal throughput -- without FMA [####################] 100 % Ntrial = 20 ; Min = 26.534 + 0.418 clc/call; Median-Min = 0.462 clc/call; Max = 27.327 clc/call; ``` - Latency from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log2 --latency GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 46.048 + 1.643 clc/call; Median-Min = 1.694 clc/call; Max = 48.018 clc/call; -- CORE-MATH latency -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 62.333 + 0.138 clc/call; Median-Min = 0.119 clc/call; Max = 62.583 clc/call; -- System LIBC latency -- [####################] 100 % Ntrial = 20 ; Min = 45.206 + 1.503 clc/call; Median-Min = 1.467 clc/call; Max = 47.229 clc/call; -- LIBC latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 43.042 + 0.454 clc/call; Median-Min = 0.484 clc/call; Max = 43.912 clc/call; -- LIBC latency -- without FMA [####################] 100 % Ntrial = 20 ; Min = 57.016 + 1.636 clc/call; Median-Min = 1.655 clc/call; Max = 58.816 clc/call; ``` - Accurate pass latency: ``` $ ./perf.sh log2 --latency --simple_stat GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA 177.632 -- CORE-MATH latency -- without FMA (-march=x86-64-v2) 231.332 -- LIBC latency -- with FMA 459.751 -- LIBC latency -- without FMA 463.850 ``` Reviewed By: zimmermann6 Differential Revision: https://reviews.llvm.org/D150374
-
Krzysztof Parzyszek authored
A prior commit accidentally affected a safety check allowing aliased memory instructions to be moved across one another.
-
Manna, Soumi authored
Reported by Static Code Analyzer Tool, Coverity: Dereference null return value Inside "ExprConstant.cpp" file, in <unnamed>::RecordExprEvaluator::VisitCXXStdInitializerListExpr(clang::CXXStdInitializerListExpr const *): Return value of function which returns null is dereferenced without checking. bool RecordExprEvaluator::VisitCXXStdInitializerListExpr( const CXXStdInitializerListExpr *E) { // returned_null: getAsConstantArrayType returns nullptr (checked 81 out of 93 times). //var_assigned: Assigning: ArrayType = nullptr return value from getAsConstantArrayType. const ConstantArrayType *ArrayType = Info.Ctx.getAsConstantArrayType(E->getSubExpr()->getType()); LValue Array; //Condition !EvaluateLValue(E->getSubExpr(), Array, this->Info, false), taking false branch. if (!EvaluateLValue(E->getSubExpr(), Array, Info)) return false; // Get a pointer to the first element of the array. //Dereference null return value (NULL_RETURNS) //dereference: Dereferencing a pointer that might be nullptr ArrayType when calling addArray. Array.addArray(Info, E, ArrayType); This patch adds an assert for unexpected type for array initializer. Reviewed By: erichkeane Differential Revision: https://reviews.llvm.org/D151040 -
Kadir Cetinkaya authored
Differential Revision: https://reviews.llvm.org/D149948
-
Nikita Popov authored
Requeue the modified instruction. This should be NFC apart from worklist order effects.
-
Tue Ly authored
Implement double precision log function correctly rounded to all rounding modes. See https://reviews.llvm.org/D150014 for a more detail description of the algorithm. **Performance** - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.93%. - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 17.465 + 0.596 clc/call; Median-Min = 0.602 clc/call; Max = 18.389 clc/call; -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 54.961 + 2.606 clc/call; Median-Min = 2.180 clc/call; Max = 59.583 clc/call; -- System LIBC reciprocal throughput -- [####################] 100 % Ntrial = 20 ; Min = 12.608 + 0.276 clc/call; Median-Min = 0.359 clc/call; Max = 13.147 clc/call; -- LIBC reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 20.952 + 0.468 clc/call; Median-Min = 0.602 clc/call; Max = 21.881 clc/call; -- LIBC reciprocal throughput -- without FMA [####################] 100 % Ntrial = 20 ; Min = 18.569 + 0.552 clc/call; Median-Min = 0.601 clc/call; Max = 19.259 clc/call; ``` - Latency from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log --latency GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 48.431 + 0.699 clc/call; Median-Min = 0.073 clc/call; Max = 51.269 clc/call; -- CORE-MATH latency -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 64.865 + 3.235 clc/call; Median-Min = 3.475 clc/call; Max = 71.788 clc/call; -- System LIBC latency -- [####################] 100 % Ntrial = 20 ; Min = 42.151 + 2.090 clc/call; Median-Min = 2.270 clc/call; Max = 44.773 clc/call; -- LIBC latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 35.266 + 0.479 clc/call; Median-Min = 0.373 clc/call; Max = 36.798 clc/call; -- LIBC latency -- without FMA [####################] 100 % Ntrial = 20 ; Min = 48.518 + 0.484 clc/call; Median-Min = 0.500 clc/call; Max = 49.896 clc/call; ``` - Accurate pass latency: ``` $ ./perf.sh log --latency --simple_stat GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA 598.306 -- CORE-MATH latency -- without FMA (-march=x86-64-v2) 632.925 -- LIBC latency -- with FMA 455.632 -- LIBC latency -- without FMA 488.564 ``` Reviewed By: zimmermann6 Differential Revision: https://reviews.llvm.org/D150131
-
Manna, Soumi authored
[NFC][Clang] Fix Coverity bug with dereference null return value in clang::CodeGen::CodeGenFunction::EmitOMPArraySectionExpr() Reported by Coverity: Inside "CGExpr.cpp" file, in clang::CodeGen::CodeGenFunction::EmitOMPArraySectionExpr(clang::OMPArraySectionExpr const *, bool): Return value of function which returns null is dereferenced without checking. } else { //returned_null: getAsConstantArrayType returns nullptr (checked 83 out of 95 times). // var_assigned: Assigning: CAT = nullptr return value from getAsConstantArrayType. auto *CAT = C.getAsConstantArrayType(ArrayTy); //identity_transfer: Member function call CAT->getSize() returns an offset off CAT (this). // Dereference null return value (NULL_RETURNS) //dereference: Dereferencing a pointer that might be nullptr CAT->getSize() when calling APInt. ConstLength = CAT->getSize(); } This patch adds an assert to resolve the bug. Reviewed By: erichkeane Differential Revision: https://reviews.llvm.org/D151137 -
Nikita Popov authored
-
Nikita Popov authored
Make sure the old load/store operand is queued for DCE. This should be NFC apart from worklist order effects.
-
Joseph Huber authored
The AMDGPU backend has a built-in pass to lower constructors. We do this manually in the `start.cpp` implementation so we can disable this to keep the binaries smaller. Differential Revision: https://reviews.llvm.org/D151213
-