- Sep 10, 2023
-
-
Jan Svoboda authored
After 98e6deb6 the 'HeadersForSymbolTest.IWYUTransitiveExportWithPrivate' test in 'ClangIncludeCleanerTest' started failing. This is most likely because `FileEntryRef::getName()` now starts with ".\" on Windows, whereas `FileEntry::getName()` did not. This commit fixes assumption of forward slash separators.
-
Joseph Huber authored
Summary: This patch implements fwrite, putc, putchar, and fputc on the GPU. These are very straightforward, the main difference for the GPU implementation is that we are currently ignoring `errno`. This patch also introduces a minimal smoke test for `putc` that is an exact copy of the `puts` test except we print the string char by char. This also modifies the `fopen` test to use `fwrite` to mirror its use of `fread` so that it is tested as well.
-
Jonas Devlieghere authored
Add a CODE_OF_CONDUCT.md file to the root of the repository. The file itself references the LLVM Community Code of Conduct. GitHub will recognize this file and put a link to it to the right of the repository, similar to the license and security policy, making the CoC easier to discover.
-
Tyler Lanphear authored
Fix crash on RAUW due to locals and globals having different address spaces. This is the intent of the original code, but it assumes the alloca address space is 0. This patch fixes the code to check that the global's address space matches `DL.getAllocaAddrSpace()` instead. Fixes #65155
-
Jan Svoboda authored
-
Jan Svoboda authored
-
Jan Svoboda authored
-
Jan Svoboda authored
-
- Sep 09, 2023
-
-
Fabian Mora authored
The revert happened due to a build bot failure that threw 'CUDA_ERROR_UNSUPPORTED_PTX_VERSION'. The failure's root cause was a pass using "+ptx76" for compilation and an old CUDA driver on the bot. This commit relands the patch with "+ptx60". Original Gh PR: #65768 Original commit message: Migrate tests referencing `gpu-to-cubin` to the new compilation workflow using `TargetAttrs`. The `test-lower-to-nvvm` pass pipeline was modified to use the new compilation workflow to simplify the introduction of future tests. The `createLowerGpuOpsToNVVMOpsPass` function was removed, as it didn't allow for passing all options available in the `ConvertGpuOpsToNVVMOp` pass. -
Mark de Wever authored
-
Sergei Barannikov authored
According to the manual, cas, casl, casx and casxl are synthetic instructions. They map to casa and casxa with certain ASI tags.
-
Fabian Mora authored
Revert "[mlir][test][gpu] Migrate CUDA tests to the TargetAttr compilation workflow (#65768) (#65848) This reverts commit d21b6729.
-
Fabian Mora authored
Migrate tests referencing `gpu-to-cubin` to the new compilation workflow using `TargetAttrs`. The `test-lower-to-nvvm` pass pipeline was modified to use the new compilation workflow to simplify the introduction of future tests. The `createLowerGpuOpsToNVVMOpsPass` function was removed, as it didn't allow for passing all options available in the `ConvertGpuOpsToNVVMOp` pass.
-
Job Noorman authored
Linker relaxation may change relocations (offsets and types). However, when --emit-relocs is used, relocations are simply copied from the input section causing a mismatch with the corresponding (relaxed) code section. This patch fixes this as follows: for non-relocatable RISC-V binaries, `InputSection::copyRelocations` reads relocations from the relocated section's `relocations` array (since this gets updated by the relaxation code). For all other cases, relocations are read from the input section directly as before. In order to reuse as much code as possible, and to keep the diff small, the original `InputSection::copyRelocations` is changed to accept the relocations as a range of `Relocation` objects. This means that, in the general case when reading from the input section, raw relocations need to be converted to `Relocation`s first, which introduces quite a bit of boiler plate. It also means there's a slight code size increase due to the extra instantiations of `copyRelocations` (for both range types). Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D159082
-
Amy Wang authored
This enables canonicalization to fold away unnecessary tensor.dim ops which in turn enables folding away of other operations, as can be seen in conv_tensors_dynamic where affine.min operations were folded away.
-
Timm Bäder authored
Now what we have delegate() we can use it here.
-
Timm Bäder authored
-
Timm Bäder authored
-
Timm Bäder authored
-
liqin.weng authored
Spill/reload instructions are artificially generated by the compiler and have no relation to the original source code. So the best thing to do is not attach any debug location to them (instead of just taking the next debug location we find on following instructions). Refered to https://reviews.llvm.org/rG3e081703c349dd00b8ef6991c2d15964915dd8f4 Reviewed By: asb, kito-cheng, benshi001 Differential Revision: https://reviews.llvm.org/D129173
-
Amara Emerson authored
-
Job Noorman authored
Relocation used for store instructions.
-
liqin.weng authored
Add basic handling for VP ops that can expand to cast intrinsics Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D159478
-
Kazu Hirata authored
-
Nuno Lopes authored
Continuing the discussion in https://discourse.llvm.org/t/codegen-layout-of-si-class-type-info-doesnt-match-the-actual-size/73274 Before we had this code: @_ZTVN10__cxxabiv117__class_type_infoE = external global ptr now we'll produce: @_ZTVN10__cxxabiv117__class_type_infoE = external global [0 x ptr] This is because we may not know the exact size of this data, and clang issues gep inbounds with idx=2. Before, that gep would always result in poison.
-
Douglas Yung authored
Should fix https://lab.llvm.org/buildbot/#/builders/216/builds/27001.
-
Mehdi Amini authored
Improve the pull-request subcription notification format by adding the description and files statistics (#65828)
-
Johannes Doerfert authored
-
Dhruv Chawla authored
This patch adds a hidden CLI option "--sroa-max-alloca-slices", which is an integer that controls the maximum number of alloca slices SROA can consider before bailing out. This is useful because it may not be profitable to split memcpys into (possibly tens of) thousands of loads/stores. This also prevents an issue with exponential compile time explosion in passes like DSE and MemCpyOpt caused by excessive alloca splitting. Fixes https://github.com/rust-lang/rust/issues/88580. Differential Revision: https://reviews.llvm.org/D159354
-
Timm Bäder authored
... when initializing. Fixes a problem pointed out in https://reviews.llvm.org/D156045/
-
Tom Stellard authored
We cannot use the default github token for labeling PRs, because this will not trigger the PR Subscriber job. However, we weren't allowed to use a different token via a secret, because secrets aren't allowed in PR workflows. The solution is to create two workflows, the first accepts the pull_request_taget event extracts the PR number and then starts the second workflow which adds the labels to the PRs. This separation ensures that nothing malicious in the first workflow is able to access the secret we use in the second workflow.
-
Jan Svoboda authored
-
Fangrui Song authored
-
Brad Smith authored
-
Siva Chandra authored
The options added via COMPILE_OPTIONS will be treated as INTERFACE options. This will help in setting compile options based on libc config options in future patches.
-
Jan Svoboda authored
-
Shilei Tian authored
-
Jan Svoboda authored
This reapplies ddbcc10b, except for a tiny part that was reverted separately: 65331da0. That will be reapplied later on, since it turned out to be more involved. This commit is enabled by 5523fefb and f0f548a6, specifically the part that makes 'clang-tidy/checkers/misc/header-include-cycle.cpp' separator agnostic.
-
Brad Smith authored
Some fixes for the header / library paths.. - Use concat macro for all paths - Correct the C++ header paths - Add library paths Differential Revision: https://reviews.llvm.org/D159414
-
Fabian Mora authored
Migrate tests referencing `gpu-to-hsaco` to the new compilation workflow using TargetAttrs.
-