- Jul 28, 2023
-
-
Slava Zakharin authored
It looks like a regression after D151737: shape of the elemental call became rank-0. Reviewed By: klausler Differential Revision: https://reviews.llvm.org/D156386
-
Slava Zakharin authored
When creating a temporary for conflicting LHS and RHS we have to deep copy the dynamic (allocatable, automatic) components from RHS to the temp. Otherwise, the conflict may still be present between LHS and temp. gfortran/regression/alloc_comp_assign_1.f90 is an example where the current runtime code produces wrong result: https://github.com/llvm/llvm-test-suite/blob/7b5b5dcbf9bdde729a14722eb67f9c3ab01647c7/Fortran/gfortran/regression/alloc_comp_assign_1.f90#L50 Reviewed By: klausler, tblah Differential Revision: https://reviews.llvm.org/D156364
-
Douglas Yung authored
This reverts commit 96ff464d. The test in this change was failing on many buildbots: https://lab.llvm.org/buildbot/#/builders/164/builds/41292 https://lab.llvm.org/buildbot/#/builders/258/builds/4491 https://lab.llvm.org/buildbot/#/builders/192/builds/3566 https://lab.llvm.org/buildbot/#/builders/123/builds/20411 https://lab.llvm.org/buildbot/#/builders/58/builds/42553 https://lab.llvm.org/buildbot/#/builders/247/builds/7037 https://lab.llvm.org/buildbot/#/builders/139/builds/46259 https://lab.llvm.org/buildbot/#/builders/216/builds/24650 https://lab.llvm.org/buildbot/#/builders/234/builds/12571 https://lab.llvm.org/buildbot/#/builders/232/builds/12574 https://lab.llvm.org/buildbot/#/builders/235/builds/975
-
Reid Kleckner authored
These are small include-only changes in the AArch64 and ARM backends that seem sufficiently small to commit separately without review. See issue #64166 for more information about layering.
-
Reid Kleckner authored
These are small include-only changes in the X86, Mips, and SystemZ backend that seem sufficiently small to commit separately without review. See issue #64166 for more information about layering.
-
-
-
Yaxun (Sam) Liu authored
By default, clang assumes HIP kernels are launched with uniform block size, which is the case for kernels launched through triple chevron or hipLaunchKernelGGL. Clang adds uniform-work-group-size function attribute to HIP kernels to allow the backend to do optimizations on that. However, in some rare cases, HIP kernels can be launched through hipExtModuleLaunchKernel where global work size is specified, which may result in non-uniform block size. To be able to support non-uniform block size for HIP kernels, an option `-f[no-]offload-uniform-block is added. This option is generic for offloading languages. Its default value is on for CUDA/HIP and off otherwise. Make -cl-uniform-work-group-size an alias to -foffload-uniform-block. Reviewed by: Siu Chi Chan, Matt Arsenault, Fangrui Song, Johannes Doerfert Differential Revision: https://reviews.llvm.org/D155213 Fixes: SWDEV-406592
-
Brian Cain authored
Before applying this fix, clang would not include the specified library path arguments: $ ./bin/clang --target=hexagon-unknown-linux-musl -o tprog tprog.o -L/tmp -### ... clang: warning: argument unused during compilation: '-L/tmp' [-Wunused-command-line-argument] "/local/mnt/workspace/install/clang-latest/bin/ld.lld" "-z" "relro" "-o" "tprog" "-dynamic-linker=/lib/ld-musl-hexagon.so.1" "/usr/lib/crt1.o" "-L/usr/lib" "tprog.o" "-lclang_rt.builtins-hexagon" "-lc" Differential Revision: https://reviews.llvm.org/D156330 -
Piotr Zegar authored
Sort printed options in --dump-config output. Fixes: #64153 Reviewed By: carlosgalvezp Differential Revision: https://reviews.llvm.org/D156452
-
Valentin Clement authored
This patch adds support to lower the link clause on OpenACC declare construct in module declaration. Depends on D156463 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D156464
-
Valentin Clement authored
This patch adds support to lower the device_resident clause on OpenACC declare construct in module declaration. Depends on D156457 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D156463
-
Valentin Clement authored
This patch adds support to lower the copyin clause on OpenACC declare construct in module declaration. Depends on D156353 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D156457
-
Valentin Clement authored
Add the acc.global_dtor when lowering the OpenACC declare construct. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D156353
-
Piotr Zegar authored
Looks like directory Inputs/config-files/2 does not exist in git because it didnt had any files, this caused some clang-tidy test to silently fail.
-
Saleem Abdulrasool authored
This relocation type is often used for debug information on Windows. We would previously abort due to the unreachable for the unhandled relocation type. Add support for this to prevent LLDB from aborting if it encounters this relocation type.
-
Hiroshi Yamauchi authored
This fixes a code gen issue where savings the swift async context register (x22) accidentally overwrites the saved value of another callee-saved register, corrupts its value and causes a crash. Differential Revision: https://reviews.llvm.org/D156391
-
Craig Topper authored
-
Piotr Zegar authored
This reverts commit 2cdb8437.
-
Craig Topper authored
This makes the assembly more readable. Reviewed By: luke Differential Revision: https://reviews.llvm.org/D156348
-
Yitzhak Mandelbaum authored
This patch adds support for building a weak topological ordering (WTO) of the CFG blocks, based on a limit flow graph constructed via (repeated) interval partitioning of the CFG. This patch is part 2 of 2 for adding WTO support. Differential Revision: https://reviews.llvm.org/D153058
-
Kelvin Li authored
This patch includes the a subset of MMA intrinsics that are included in the mma intrinsic module: mma_assemble_acc mma_assemble_pair mma_build_acc mma_disassemble_acc mma_disassemble_pair Submit on behalf of Daniel Chen <cdchen@ca.ibm.com> Differential Revision: https://reviews.llvm.org/D155725
-
Alexandros Lamprineas authored
This patch allows constant folding of PHIs when estimating the user bonus. Phi nodes are a special case since some of their inputs may remain unresolved until all the specialization arguments have been processed by the InstCostVisitor. Therefore, we keep a list of dead basic blocks and then lazily visit the Phi nodes once the user bonus has been computed for all the specialization arguments. In addition to the last revision this one fixes the bug reported on Phabricator. Differential Revision: https://reviews.llvm.org/D154852
-
Quinn Dawkins authored
Introduces a generic parameter type intended for transform dialect use cases that uses/manipulates the underlying attribute (e.g. op annotation). Includes a disclaimer that mixing parameter based and attribute based control is discouraged. Differential Revision: https://reviews.llvm.org/D155980
-
Teresa Johnson authored
This avoids the need to regenerate the PGO raw profile on version changes. Modify the update script to autogenerate new PGO proftext inputs. Differential Revision: https://reviews.llvm.org/D156460
-
Piotr Zegar authored
Sort printed options in --dump-config output. Fixes: #64153 Reviewed By: carlosgalvezp Differential Revision: https://reviews.llvm.org/D156452
-
Wanyi Ye authored
Recently we've observed lldb crashes caused by missing object file linked to a thin archive (.a) files. The crash is due to a missing NULL check in the code when looking for child object file referred by the thin archive. Malformed archive file should not crash LLDB. Instead, it should report the error and continue. New error message will look like the following ``` error: libfoo.a(__objects__/foo/barAppDelegate.mm.o) failed to load objfile for path/to/libfoo.a. Debugging will be degraded for this module. ``` Test Plan: llvm-lit test ``` ./bin/llvm-lit -sv ../llvm-project/lldb/test/API/functionalities/archives/TestBSDArchives.py ``` Test without code change will error out with LLDB crash ``` -- Command Output (stderr): -- PASS: LLDB (~/llvm-upstream/Debug/bin/clang-arm64) :: test (TestBSDArchives.BSDArchivesTestCase) PASS: LLDB (~/llvm-upstream/Debug/bin/clang-arm64) :: test_frame_var_errors_when_archive_missing (TestBSDArchives.BSDArchivesTestCase) FAIL: LLDB (~/llvm-upstream/Debug/bin/clang-arm64) :: test_frame_var_errors_when_mtime_mistmatch_for_object_in_archive (TestBSDArchives.BSDArchivesTestCase) PLEASE submit a bug report to https://github.com/llvm/llvm-project/issues/ and include the crash backtrace. Stack dump: 0. HandleCommand(command = "b a") 1. HandleCommand(command = "breakpoint set --name 'a'") Fatal Python error: Segmentation fault Current thread 0x00000001f7b99e00 (most recent call first): File "~/llvm-upstream/Debug/bin/LLDB.framework/Resources/Python/lldb/__init__.py", line 3270 in HandleCommand File "~/llvm-upstream/llvm-project/lldb/packages/Python/lldbsuite/test/lldbtest.py", line 2070 in runCmd File "~/llvm-upstream/llvm-project/lldb/packages/Python/lldbsuite/test/lldbtest.py", line 2421 in expect File "~/llvm-upstream/llvm-project/lldb/test/API/functionalities/archives/TestBSDArchives.py", line 156 in test_frame_var_errors_when_thin_archive_malformed ... ``` Differential Revision: https://reviews.llvm.org/D156367
-
Sam McCall authored
This updates clangd to take advantage of the APIs added in D155819. The main difficulties here are around path normalization. For layering and performance reasons Includes compares paths lexically, and so we should have consistent paths that can be compared across addSearchPath() and add(): symlinks resolved or not, relative or absolute. This patch spells out that requirement, for most tools consistent use of FileManager/HeaderSearch is enough. For clangd this does not work: IncludeStructure doesn't hold FileEntrys due to the preamble/main-file split. It records paths, but canonicalizes them first. We choose to use this canonical form as our common representation, so we have to canonicalize the directory entries too. This is done in preamble-build and recorded in IncludeStructure, as canonicalization is quite expensive. Differential Revision: https://reviews.llvm.org/D155878
-
Matt Arsenault authored
-
Daniel Hoekwater authored
PATCHABLE_TYPED_EVENT_CALL and PATCHABLE_EVENT_CALL are pseudo instructions that expand to XRay sleds, so getInstSizeInBytes should reflect the size of the sleds, not the pseudo-instructions. Differential Revision: https://reviews.llvm.org/D156272
-
Nico Weber authored
57329ca9 added a .inc include within an unnamed namespace, creating ::(unnamed namespace)::llvm in addition to ::llvm. This confuses things on Windows.
-
spupyrev authored
We are bringing a new algorithm for function layout (reordering) based on the call graph (extracted from a profile data). The algorithm is an improvement of top of a known heuristic, C^3. It tries to co-locate hot and frequently executed together functions in the resulting ordering. Unlike C^3, it explores a larger search space and have an objective closely tied to the performance of instruction and i-TLB caches. Hence, the name CDS = Cache-Directed Sort. The algorithm can be used at the linking or post-linking (e.g., BOLT) stage. The algorithm shares some similarities with C^3 and an approach for basic block reordering (ext-tsp). It works with chains (ordered lists) of functions. Initially all chains are isolated functions. On every iteration, we pick a pair of chains whose merging yields the biggest increase in the objective, which is a weighted combination of frequency-based and distance-based locality. That is, we try to co-locate hot functions together (so they can share the cache lines) and functions frequently executed together. The merging process stops when there is only one chain left, or when merging does not improve the objective. In the latter case, the remaining chains are sorted by density in the decreasing order. **Complexity** We regularly apply the algorithm for large data-center binaries containing 10K+ (hot) functions, and the algorithm takes only a few seconds. For some extreme cases with 100K-1M nodes, the runtime is within minutes. **Perf-impact** We extensively tested the implementation extensively on a benchmark of isolated binaries and prod services. The impact is measurable for "larger" binaries that are front-end bound: the cpu time improvement (on top of C^3) is in the range of [0% .. 1%], which is a result of a reduced i-TLB miss rate (by up to 20%) and i-cache miss rate (up to 5%). Reviewed By: rahmanl Differential Revision: https://reviews.llvm.org/D152834
-
Sam McCall authored
A verbatim header usually corresponds to a symbol from a header with a pragma "IWYU pragma: private, include <foo.h>". Currently this is only satisfied if the main file contains exactly #include <foo.h> In practice this is too strict, we also want to allow #include "path/to/foo.h" so long as they resolve to the same file. We cannot be 100% sure without doing IO, and we're not willing to do that, but we can detect the common cases based on paths. Differential Revision: https://reviews.llvm.org/D155819
-
Vitaly Buka authored
Breaks window bot. This reverts commit e3f935c7.
-
Craig Topper authored
We only want to be able to parse a missing offset. We don't want to print with no offset. This matches the non-compressed form of these aliases.
-
- Jul 27, 2023
-
-
LLVM GN Syncbot authored
-
spupyrev authored
Adding some logs related to stale profile matching. The new data can be helpful to understand how "stale" the input profile is and how well the inference is able to utilize the stale data. Example of outputs on clang-10 built with LTO (profile collected on a year-old release): ``` BOLT-INFO: inferred profile for 2101 (18.52% of profiled, 100.00% of stale) functions responsible for 30.95% samples (14754697 out of 47670654) BOLT-INFO: stale inference matched 89.42% of basic blocks (79052 out of 88402 stale) responsible for 76.99% samples (645737 out of 838719 stale) ``` LTO+AutoFDO: ``` BOLT-INFO: inferred profile for 6146 (57.57% of profiled, 100.00% of stale) functions responsible for 90.34% samples (50891403 out of 56330313) BOLT-INFO: stale inference matched 74.55% of basic blocks (191295 out of 256589 stale) responsible for 57.30% samples (1288632 out of 2248799 stale) ``` Reviewed By: Amir, maksfb Differential Revision: https://reviews.llvm.org/D154737
-
Simon Pilgrim authored
As noted on #63980 rotate by immediate amounts is much cheaper than variable amounts. This still needs to be expanded to vector rotate cases, and we need to add reasonable funnel-shift costs as well (very tricky as there's a huge range in CPU behaviour for these).
-
Piotr Zegar authored
Detects implicit conversions between pointers of different levels of indirection. Reviewed By: xgupta Differential Revision: https://reviews.llvm.org/D149084
-
Maksim Kita authored
Fold strcmp for short string literals with size 2. Depends D155742. Differential Revision: https://reviews.llvm.org/D155743
-