- Oct 30, 2020
-
-
Joachim Meyer authored
This is very similar to 7f1e6fcf, just fixing a left-over. With this, it should be possible to use both, -x cuda and -fopenmp in the same invocation, enabling to use both OpenMP, targeting CPU, and CUDA, targeting the GPU. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D90415
-
Jim Ingham authored
-
Thomas Raoux authored
Fix semantic in the distribute integration test based on offline feedback. This exposed a bug in block distribution, we need to make sure the id is multiplied by the stride of the vector. Fix the transformation and unit test. Differential Revision: https://reviews.llvm.org/D89291
-
Craig Topper authored
This combine makes two calls to SimplifyDemandedBits, one for the LHS and one for the RHS. If the LHS call returns true, we don't make the RHS call. When SimplifyDemandedBits makes a change, it will add the nodes around the change to the DAG combiner worklist. If the simplification happens on the first recursion step, the N will get added to the worklist. But if the simplification happens deeper in the recursion, then N will not be revisited until the next time the DAG combiner runs. This patch explicitly addes N to the worklist anytime a Simplification is made. Without this we might miss additional simplifications on the LHS or never simplify the RHS. Special care also needs to be taken to not add N if it has been CSEd by the simplification. There are similar examples in DAGCombiner and the X86 target, but I don't have a test for it for RISC-V. I've also returned SDValue(N, 0) instead of SDValue() so DAGCombiner knows a change was made and will update its Statistic variable. The test here was constructed so that 2 simplifications happen to the LHS. Without this fix one happens in the post type legalization DAG combine and the other happens after LegalizeDAG. This prevents the RHS from ever being simplified causing the left and right shift to clear the upper 32 bits of the RHS to be left behind. Differential Revision: https://reviews.llvm.org/D90339
-
Craig Topper authored
-
Jim Ingham authored
we should be exporting one by one. Differential Revision: https://reviews.llvm.org/D78972
-
Jim Ingham authored
The intention is not to allow stop-hook commands to query the user, so this is correct. It also works around a deadlock in switching to the Python Session to execute python based commands in the stop hook when the Debugger stdin is backed by a FILE *. Differential Revision: https://reviews.llvm.org/D90332
-
Christian Sigg authored
Do not use the pass yet, except in a test. Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D89937
-
Christian Sigg authored
For the synchronous case, destroy the stream after synchronization. Sneak in a unrelated change to report why the gpu.wait conversion pattern didn't match. Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D89933
-
Nikita Popov authored
Make the existing VECREDUCE based code more generic, but expressing it in terms of the neutral value of the base opcode instead.
-
Stefanos Baziotis authored
Differential Revision: https://reviews.llvm.org/D89739
-
Christian Sigg authored
This is a roll-forward of rGec7780eb, now that the remaining gpu.launch_func have been converted to custom form in rGb22f1110. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D90420
-
Nikita Popov authored
Use -0.0 instead of 0.0 as the start value. The previous use of 0.0 was fine for all existing uses of this function though, as it is always generated with fast flags right now, and thus nsz.
-
Ilya Bukonkin authored
-
Florian Hahn authored
Some architectures do not have general vector select instructions (e.g. AArch64). But some cmp/select patterns can be vectorized using other instructions/intrinsics. One example is using min/max instructions for certain patterns. This patch updates the cost calculations for selects in the SLP vectorizer to consider using min/max intrinsics. This patch does not change SLP vectorizer's codegen itself to actually generate those intrinsics, but relies on the backends to lower the vector cmps & selects. This keeps things simple on the SLP side and works well in practice for AArch64. This exposes additional SLP vectorization opportunities in some benchmarks on AArch64 (-O3 -flto). Metric: SLP.NumVectorInstructions Program base slp diff test-suite...ications/JM/ldecod/ldecod.test 502.00 697.00 38.8% test-suite...ications/JM/lencod/lencod.test 1023.00 1414.00 38.2% test-suite...-typeset/consumer-typeset.test 56.00 65.00 16.1% test-suite...6/464.h264ref/464.h264ref.test 804.00 822.00 2.2% test-suite...006/453.povray/453.povray.test 3335.00 3357.00 0.7% test-suite...CFP2000/177.mesa/177.mesa.test 2110.00 2121.00 0.5% test-suite...:: External/Povray/povray.test 2378.00 2382.00 0.2% Reviewed By: RKSimon, samparker Differential Revision: https://reviews.llvm.org/D89969
-
Thomas Lively authored
This commit removes unused FileCheck prefixes from WebAssembly test files to avoid causing test failures once FileCheck disallows unused prefixes by default. See D90281 and the corresponding llvm-dev thread for context. Reviewed By: aardappel Differential Revision: https://reviews.llvm.org/D90416
-
Nikita Popov authored
The neutral value for FADD is -0.0, not 0.0, so this is what we need to pad vectors with.
-
Nikita Popov authored
The neutral value is -0.0, not 0.0. This doesn't matter for "fast" reductions due to nsz, but does matter for reassoc-only and seq reductions. Change tests to mostly use -0.0 where the neutral value was intended, and add some additional test coverage in some places. Also update LangRef to use the right value.
-
Christian Sigg authored
This should fix the reason for the failures after ec7780eb. I will roll forward in a separate change. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D90410
-
Tony authored
- AMDGPUUsage.rst: Correct AMD GPU DWARF address space table address sizes which are in bits and not bytes. - clang/.../Options.td: Improve description of AMD GPU options. - Re-generate ClangComamndLineReference.rst from clang/.../Options.td . Differential Revision: https://reviews.llvm.org/D90364
-
Alex Orlov authored
Use LLVM/utils/remote-exec.py to run compiler-rt tests remotely on the target. Reviewed By: vvereschaka Differential Revision: https://reviews.llvm.org/D90054
-
Marcel Hlopko authored
This preprocessor define was meant to be used to conditionally include VCSVersion.inc. However, the define was always set, and it was the content of the header that was conditionally generated. Therefore HAVE_VCS_VERSION_INC should be cleaned up. Reviewed By: gribozavr2, MaskRay Differential Revision: https://reviews.llvm.org/D84623
-
Nikita Popov authored
-
Adhemerval Zanella authored
On aarch64 with kernel 4.12.13 the test sporadically fails with RSS at start: 1564, after mmap: 103964, after mmap+set label: 308768, \ after fixed map: 206368, after another mmap+set label: 308768, after \ munmap: 206368 release_shadow_space.c.tmp: [...]/release_shadow_space.c:80: int \ main(int, char **): Assertion `after_fixed_mmap <= before + delta' failed. It seems on some executions the memory is not fully released, even after munmap. And it also seems that ASLR is hurting it by adding some fragmentation, by disabling it I could not reproduce the issue in multiple runs.
-
Peyton, Jonathan L authored
Patch by Nawrin Sultana Differential Revision: https://reviews.llvm.org/D90403
-
Dávid Bolvanský authored
One step closer to fix PR47644. Differential Revision: https://reviews.llvm.org/D89645
-
Paul-Antoine Arras authored
This diff adds support for LLVM bitcode objects to llvm-libtool-darwin. Test plan: make check-all Differential revision: https://reviews.llvm.org/D88722
-
Utkarsh Saxena authored
With every incremental change, one needs to check-in new model upstream. This also significantly increases the size of the git repo with every new model. Testing and comparing the old and previous model is also not possible as we run only a single model at any point. One solution is to have a "staging" decision forest which can be injected into clangd without pushing it to upstream. Compare the performance of the staging model with the live model. After a couple of enhancements have been done to staging model, we can then replace the live model upstream with the staging model. This reduces upstream churn and also allows us to compare models with current baseline model. This is done by having a callback in CodeCompleteOptions which is called only when we want to use a decision forest ranking model. This allows us to inject different completion model internally. Differential Revision: https://reviews.llvm.org/D90014
-
Craig Topper authored
RISCVRegisterInfo.h is part of the CodeGen layer. The Utils library is intended to be shared with the MC layer so shouldn't use files from the CodeGen layer. The register enum names are already available from RISCVMCTargetDesc.h. It appears what was coming from this include was a transitive include of the Register class which I've replaced with MCRegister. Register has a constructor from MCRegister so it should be convertible.
-
Teresa Johnson authored
I finally see why this test is failing (on now 2 bots). Somehow the path name is getting messed up, and the "linux" converted to "1". I suspect there is something in the environment causing the macro expansion in the test to get messed up: http://lab.llvm.org:8011/#/builders/112/builds/555/steps/5/logs/FAIL__MemProfiler-x86_64-linux__log_path_test_cpp http://lab.llvm.org:8011/#/builders/37/builds/275/steps/31/logs/stdio On the avr bot: -DPROFILE_NAME_VAR="/home/buildbot/llvm-avr-linux/llvm-avr-linux/stage1/projects/compiler-rt/test/memprof/X86_64LinuxConfig/TestCases/Output/log_path_test.cpp.tmp.log2" after macros expansions becomes: /home/buildbot/llvm-avr-1/llvm-avr-1/stage1/projects/compiler-rt/test/memprof/X86_64LinuxConfig/TestCases/Output/log_path_test.cpp.tmp.log2 Similar (s/linux/1/) on the other bot. Disable it while I investigate
-
Sylvestre Ledru authored
-
Thomas Lively authored
As proposed in https://github.com/WebAssembly/simd/pull/124, using the opcodes adopted by V8 in https://chromium-review.googlesource.com/c/v8/v8/+/2486235/2/src/wasm/wasm-opcodes.h. Uses new builtin functions and a new target intrinsic exclusively to ensure that the new instructions are only emitted when a user explicitly opts in to using them since they are still in the prototyping and evaluation phase. Differential Revision: https://reviews.llvm.org/D90357
-
Louis Dionne authored
-
Jody Sankey authored
The zx_clock_get syscall on Fuchsia is deprecated - ref https://fuchsia.dev/fuchsia-src/reference/syscalls/clock_get This changes to the recommended replacement; calling zx_clock_read on the userspace UTC clock. Reviewed By: mcgrathr, phosek Differential Revision: https://reviews.llvm.org/D90169
-
Roland McGrath authored
Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D90279
-
Jay Foad authored
By setting up the AsmStrings correctly we can remove some special cases from AMDGPUInstPrinter::printOffset. Differential Revision: https://reviews.llvm.org/D90307
-
Mehdi Amini authored
This reverts commit ec7780eb. One of the bot is crashing in a test related to this change.
-
Teresa Johnson authored
After 81f7b96e, I can see that the reason this test is failing on llvm-avr-linux is that it doesn't think the directory exists (error comes during file open for write command). Not sure why since this is the main test Output directory and we created a different file there earlier in the test from the same file open invocation. Print directory contents in an attempt to debug.
-
Jan Kratochvil authored
-
Mircea Trofin authored
When passing -lto-embed-bitcode=post-merge-pre-opt, we were getting empty .llvmcmd sections. It turns out that is because the CodeGenOptions::CmdArgs field was only populated when clang saw -fembed-bitcode={all|marker}. This patch always populates the CodeGenOptions::CmdArgs. The overhead of carrying through in memory in all cases is likely negligible in the grand schema of things, and it keeps the using code simple. Differential Revision: https://reviews.llvm.org/D90366
-